Relocate process-wide global variables from legacy parser that ended up
there because it used to be conf.c, but which is now now frozen at the
4.x feature set. Each variable is moved to their respective "owner".
Give cgroup_current[] and cgroup_settings_current[] named bounds. Their
extern declarations were unsized, so sizeof() on them stopped compiling
once the definitions moved to another translation unit.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The one-liner format has grown crowded and very wide, and every new
service option makes it worse.
Add a second, block-based format, parsed with libconfuse:
service sshd {
description = "OpenSSH daemon"
runlevel = "2345"
command = "/usr/sbin/sshd -D $SSHD_OPTS"
}
Both formats keep the .conf extension and are detected per file by
content. Try-parse strictly with libconfuse; on a parse error,
re-parse leniently to tell a block file with a typo from a one-liner
file. Only a one-liner file reaches the legacy parser, a typo is
reported with its file and line.
Each block is translated to the canonical one-liner and registered
through the existing entry points, so the two formats cannot drift.
The one-liner parser is frozen at the 4.x feature set, new options
land only in the block schema. libconfuse 3.3 or later is required,
CFGF_KEYSTRVAL does not exist before it.
Covers service, task, run, sysv and tty blocks, the static directives,
and the cgroup, rlimit, set and log blocks. Templating and the
documentation rewrite are still to come.
The regression test covers translation of a service block to the
one-liner, a block-format /etc/finit.conf booting with set {} applied
at bootstrap, both formats side by side, and rejection of a typo at
block and at root level.
A rejected file must not fall through to the legacy parser, which
registers a bogus unstartable service per line. assert_num_children
cannot see that, the bogus service has no children either, so the
check is assert_num_services.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The one-liner parser is about to be joined by a second, block-based
format. Give it a name that says which of the two it implements,
before any content changes make the diff hard to follow.
No functional change.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Every test left a stray `sleep 300` behind, reparented to PID 1, where
it lingered for up to five minutes after the test had finished.
wdstart() runs the watchdog in a subshell, so $! is the pid of the
subshell, not of the sleep it forks. wdkill() killed the subshell and
orphaned the sleep.
Kill the child first, killing the subshell puts the sleep beyond the
reach of pkill -P. Neither kill is sure to match, and wdkill() runs
from the EXIT trap under set -e, so both must tolerate failure. Also
return early when wdpid is unset, for failures before wdstart() runs.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Finit 5.0 changes the .conf syntax, which has been essentially
unchanged since 1.x. The published docs track master, so when 5.x
lands, 4.x users lose their reference.
Publish the site under a per-major directory, /4.x/ for now, with
the Material version selector to switch between them. The selector
only needs mike's file layout -- a versions.json at the site root --
which the deploy job now generates from the version directories in
the pages repo, so mike itself is not needed.
The major comes from AC_INIT and the future 4.x maintenance branch
is already in the workflow triggers, so once 5.0 is on master, doc
fixes on the 4.x branch keep /4.x/ updated. A root index.html
redirects to the newest version, and a 404.html rewrites
pre-versioned deep links so old bookmarks and search hits land in
the right place.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Both the Finit project and Infix use the same MkDocs Material setup, and
in the latter the User Guide has picked up a lot of polish that never
made it back here: a single sidebar with section indexes instead of
tabs, footnote tooltips, more pymdownx markup, image zoom tuning, and no
generator advert in the footer.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
* src/pid.c: note the stale-pidfile-cleanup exception to the
documented "Finit does not touch pid:! pidfiles" rule.
* doc/config/services.md: add a user-facing paragraph on the same.
* doc/ChangeLog.md: add Unreleased section covering this PR --
stale pidfile cleanup, restart log with signal name and core
dump flag, and the SIGUNKOWN typo fix.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Cover the scenario fixed in "service: clean stale pidfile after
unclean daemon exit": a daemon with a pid:!/path config dies via
SIGKILL, leaving its pidfile behind, and the next instance must
still come up.
Add a 'serv -x' flag (refuse to start when the pidfile already
exists, dbus-style) so the test actually exercises the cleanup --
without it, plain 'serv' would happily overwrite the file and the
test would pass with or without the fix.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The fallback for unknown signal numbers in sig_name() returned the
misspelled "SIGUNKOWN". Now that this string surfaces in user-
facing logs ("killed by …"), fix the typo.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Replace the bare signal number ("by signal: 9") with the symbolic
name ("killed by SIGKILL") and annotate when the kernel wrote a
core:("killed by SIGSEGV, core dumped"). Makes the restart line
self-explanatory and gives operators a strong breadcrumb when a
daemon dies unexpectedly.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
With `pid:!/path` Finit does not manage the file -- the daemon
creates it on start and removes it on graceful exit. If the daemon
dies before cleanup (SIGKILL, OOM, segfault, exit during startup)
the file lingers and can block the next instance from starting,
e.g. dbus-daemon refuses with EEXIST and the restart loop fails.
Remove the file when it still names the just-reaped PID and that
PID is no longer alive (the liveness check guards against reuse).
Called from service_cleanup(), and from service_monitor()'s
forking+starting branch where cleanup was previously skipped.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
status() returns a pointer to a single static buffer, so calling it twice
in the same cprintf() argument list — status(3) and status(rc) — causes
one to overwrite the other before the format string is rendered. When
status(3) wins, the line shows [ ⋯ ] instead of [ OK ]. Fix by copying
status(rc) into a local buffer before calling status(3).
Also drop the delline() calls added to print() — that macro writes \033[2K
to buffered stdout while cprintf() writes unbuffered to stderr, so the
erase sequences can arrive out of order. The \r\e[K already present in
the cprintf format strings makes them redundant anyway.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Check return value of remove() in delete_cb() and log failures via
dbg(), CID 909395
Replace stat() calls with open(O_DIRECTORY)+ fstat() for newroot and "/"
checks. Eliminates the check-then-use race and lets O_DIRECTORY do the
isdir validation atomically, CID 909394
Drop the explicit close(0/1/2) before opening /dev/console. dup2()
closes the old targets itself, so open() returns a fd > STDERR_FILENO
that can always be closed unconditionally, removing the conditional
guard, CID 909393
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Refactor print() to emit description + final status in a single call to
cprintf(), preventing kernel messages from splitting the two parts.
For two-phase print(-1,...) + print_result() sequences used by, e.g.,
run_interactive, save the last description and re-print it before the
[ OK ] / [FAIL] output so the status is never left stranded on a blank
line when command output or kernel messages have scrolled away the
original description.
Finally, add print_exit() which drains the console output buffer with
tcdrain(2) and resets ANSI SGR attributes + cursor visibility before the
kernel takes back the console on reboot/halt, preventing escape code
leakage into bootloader or early-kernel output.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
When a service without SIGHUP reload support (noreload) is touched and
'initctl reload' is called, service_update_rdeps() correctly identifies
its reverse dependencies but only marks them dirty. It does not clear
the service's condition, so when service_step_all() runs:
- rdeps supporting SIGHUP hit the sm_in_reload() guard and break early,
left running while their dependency is being killed.
- rdeps without SIGHUP support may receive SIGTERM too late, after the
dependency has already died and broken their connection, causing them
to exit from RUNNING state and have their restart counter incremented.
Fix by calling cond_clear() on the service's condition immediately in
service_update_rdeps(), before service_step_all() runs. cond_clear()
calls cond_update() which calls service_step() inline on all affected
services, which see COND_OFF and transition to STOPPING_STATE — all
before SIGTERM is ever sent to the dependency itself.
This mirrors the pattern already used in api.c:do_reload() for direct
'initctl reload <svc>' calls.
Fixes: avahi/mdns stop causing mdns-alias restart counter increment
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Add support for the --exclude-prefix=PATH option to skip rules whose
path starts with the specified prefix. The option can be specified
multiple times to exclude multiple path prefixes.
The -E flag is a shortcut for:
--exclude-prefix=/dev --exclude-prefix=/proc \
--exclude-prefix=/run --exclude-prefix=/sys
This is useful to avoid creating files below virtual or memory-backed
file system mount points.
devmon: assert condition immediately if device already exists
1. udev fires early, creates device nodes in /dev/
2. Config is parsed later, calling devmon_add_cond() for each dev/ condition
3. At this point the device already exists but the inotify event was missed
4. The PR's fexist() check catches this — new node is added to the TAILQ and condition is immediately asserted
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The config parsing happens after udev triggers the initial event,
make sure to assert the condition if the device node exists when adding
it from configuration.
Signed-off-by: Mattias Walström <lazzer@gmail.com>
A service that is reloaded should not trigger dependants to be reloaded
unless the new <~cond> is used. Which is reserve for tightly coupled
services.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
For SysV services with pid:!/path, the pidfile belongs to the service
itself and Finit shouldn't delete it.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
This fixes a real bug where `initctl reload syslogd` unconditionally
clears syslogd's pid condition, causing all dependent services (dbus,
dnsmasq, etc.) to be stopped even though syslogd handles SIGHUP
gracefully and its PID/pidfile persist.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
A reload: script, like 'frrinit.sh reload' could potentially take a
while to finish, during which Finit would be blocked. This change
reuses the service_script_add(), used for ready: scripts, to track
these background helpers.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
When the kernel manages to reap a child process before we've moved it to
its proper cgroup it will return ESRCH (No such process), we can safely
ignore such errors for short-lived processes.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Conditions in Finit are dependencies: if A is asserted, service B is
allowed to run. When A goes through FLUX (e.g., upstream reloads),
dependents are PAUSED and then simply resumed when the condition is
reasserted -- this is the correct behavior for barrier-style deps
like <pid/syslogd>.
However, some setups have tightly coupled services where dependents
must be reloaded/restarted when an upstream service reloads, not just
resumed. E.g., the FRR routing stack on Infix OS:
netd <pid/mgmtd> ← zebra <!pid/netd> ← {staticd,ripd} <!pid/zebra>
When netd reloads (SIGHUP), zebra and its dependents must be restarted
to pick up the new configuration.
The new '~' condition prefix marks a dependency as flux-sensitive:
service <!~pid/netd> name:zebra ...
When the upstream condition goes FLUX and returns to ON, the dependent
is reloaded (SIGHUP) or restarted (noreload '!') instead of merely
resumed. Transitivity follows naturally through the condition chain.
Closes#416Closes#476
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>