A service whose condition goes into flux during a reload is paused
with its reload still pending. If another reload was requested in
the meantime, re-parsing its unchanged .conf file cleared the pending
mark, so the service was resumed without ever being reloaded. Seen
with sshd <pid/syslogd> on Infix, where a configuration change that
touched both landed as two reloads in a row and sshd kept its old
listen addresses.
The mark is only ever cleared once the change has been applied, so a
mark that is still set when the file is parsed again means exactly
that: not applied yet. Leave it alone.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Covers the longer user and group names, the tmpfiles mode and owner
work, the shutdown remount order, and the script timeout crash.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Same treatment as d/D got: e went through fisdir() and a path based
chmod/chown, f/F did the chmod/chown by path after closing the file.
Both now work on the open fd, so the mode and owner end up on the
directory or file that was just checked or written. f/F use open(2)
directly, which lets a plain f rely on O_EXCL for create-if-missing
instead of the earlier stat().
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
All workflow runs now warn:
Node.js 20 is deprecated. The following actions target Node.js 20
but are being forced to run on Node.js 24.
Update the actions/* dependencies to their current major versions,
all of which target Node.js 24: checkout v7, upload-artifact v7,
download-artifact v8, setup-python v7, and cache v6. The release
action tracks the floating v1 tag and updates on its own.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
With 5.0 on master, maintenance releases move to a 4.x branch: run
push builds, docs deploy, and weekly distcheck for any N.x branch,
and mark only the highest stable tag as the latest release.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
'make check' refreshes the sysroot through the setup-chroot rule, but
running a test script by hand does not, so the test exercises whichever
finit was installed last and reports on code that is no longer there.
Both a passing and a failing run are then meaningless, and nothing says
so.
Compare the built binary against the installed one at startup and fail
with the command that fixes it.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
A stop: or reload: script written with a timeout killed Finit at
config load:
service stop:5,/bin/true service.sh -- Boom
parse_script() takes the timeout as a pointer and the caller decides
whether it wants one. However, both stop: and reload: scripts so far
have no timeout, i.e., NULL. Guard the branch that reads a leading
number.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Every test left a stray `sleep 300` behind, reparented to PID 1, where
it lingered for up to five minutes after the test had finished.
wdstart() runs the watchdog in a subshell, so $! is the pid of the
subshell, not of the sleep it forks. wdkill() killed the subshell and
orphaned the sleep.
Kill the child first, killing the subshell puts the sleep beyond the
reach of pkill -P. Neither kill is sure to match, and wdkill() runs
from the EXIT trap under set -e, so both must tolerate failure. Also
return early when wdpid is unset, for failures before wdstart() runs.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Finit 5.0 changes the .conf syntax, which has been essentially
unchanged since 1.x. The published docs track master, so when 5.x
lands, 4.x users lose their reference.
Publish the site under a per-major directory, /4.x/ for now, with
the Material version selector to switch between them. The selector
only needs mike's file layout -- a versions.json at the site root --
which the deploy job now generates from the version directories in
the pages repo, so mike itself is not needed.
The major comes from AC_INIT and the future 4.x maintenance branch
is already in the workflow triggers, so once 5.0 is on master, doc
fixes on the 4.x branch keep /4.x/ updated. A root index.html
redirects to the newest version, and a 404.html rewrites
pre-versioned deep links so old bookmarks and search hits land in
the right place.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Both the Finit project and Infix use the same MkDocs Material setup, and
in the latter the User Guide has picked up a lot of polish that never
made it back here: a single sidebar with section indexes instead of
tabs, footnote tooltips, more pymdownx markup, image zoom tuning, and no
generator advert in the footer.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
* src/pid.c: note the stale-pidfile-cleanup exception to the
documented "Finit does not touch pid:! pidfiles" rule.
* doc/config/services.md: add a user-facing paragraph on the same.
* doc/ChangeLog.md: add Unreleased section covering this PR --
stale pidfile cleanup, restart log with signal name and core
dump flag, and the SIGUNKOWN typo fix.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Cover the scenario fixed in "service: clean stale pidfile after
unclean daemon exit": a daemon with a pid:!/path config dies via
SIGKILL, leaving its pidfile behind, and the next instance must
still come up.
Add a 'serv -x' flag (refuse to start when the pidfile already
exists, dbus-style) so the test actually exercises the cleanup --
without it, plain 'serv' would happily overwrite the file and the
test would pass with or without the fix.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The fallback for unknown signal numbers in sig_name() returned the
misspelled "SIGUNKOWN". Now that this string surfaces in user-
facing logs ("killed by …"), fix the typo.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Replace the bare signal number ("by signal: 9") with the symbolic
name ("killed by SIGKILL") and annotate when the kernel wrote a
core:("killed by SIGSEGV, core dumped"). Makes the restart line
self-explanatory and gives operators a strong breadcrumb when a
daemon dies unexpectedly.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
With `pid:!/path` Finit does not manage the file -- the daemon
creates it on start and removes it on graceful exit. If the daemon
dies before cleanup (SIGKILL, OOM, segfault, exit during startup)
the file lingers and can block the next instance from starting,
e.g. dbus-daemon refuses with EEXIST and the restart loop fails.
Remove the file when it still names the just-reaped PID and that
PID is no longer alive (the liveness check guards against reuse).
Called from service_cleanup(), and from service_monitor()'s
forking+starting branch where cleanup was previously skipped.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
status() returns a pointer to a single static buffer, so calling it twice
in the same cprintf() argument list — status(3) and status(rc) — causes
one to overwrite the other before the format string is rendered. When
status(3) wins, the line shows [ ⋯ ] instead of [ OK ]. Fix by copying
status(rc) into a local buffer before calling status(3).
Also drop the delline() calls added to print() — that macro writes \033[2K
to buffered stdout while cprintf() writes unbuffered to stderr, so the
erase sequences can arrive out of order. The \r\e[K already present in
the cprintf format strings makes them redundant anyway.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Check return value of remove() in delete_cb() and log failures via
dbg(), CID 909395
Replace stat() calls with open(O_DIRECTORY)+ fstat() for newroot and "/"
checks. Eliminates the check-then-use race and lets O_DIRECTORY do the
isdir validation atomically, CID 909394
Drop the explicit close(0/1/2) before opening /dev/console. dup2()
closes the old targets itself, so open() returns a fd > STDERR_FILENO
that can always be closed unconditionally, removing the conditional
guard, CID 909393
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Refactor print() to emit description + final status in a single call to
cprintf(), preventing kernel messages from splitting the two parts.
For two-phase print(-1,...) + print_result() sequences used by, e.g.,
run_interactive, save the last description and re-print it before the
[ OK ] / [FAIL] output so the status is never left stranded on a blank
line when command output or kernel messages have scrolled away the
original description.
Finally, add print_exit() which drains the console output buffer with
tcdrain(2) and resets ANSI SGR attributes + cursor visibility before the
kernel takes back the console on reboot/halt, preventing escape code
leakage into bootloader or early-kernel output.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
When a service without SIGHUP reload support (noreload) is touched and
'initctl reload' is called, service_update_rdeps() correctly identifies
its reverse dependencies but only marks them dirty. It does not clear
the service's condition, so when service_step_all() runs:
- rdeps supporting SIGHUP hit the sm_in_reload() guard and break early,
left running while their dependency is being killed.
- rdeps without SIGHUP support may receive SIGTERM too late, after the
dependency has already died and broken their connection, causing them
to exit from RUNNING state and have their restart counter incremented.
Fix by calling cond_clear() on the service's condition immediately in
service_update_rdeps(), before service_step_all() runs. cond_clear()
calls cond_update() which calls service_step() inline on all affected
services, which see COND_OFF and transition to STOPPING_STATE — all
before SIGTERM is ever sent to the dependency itself.
This mirrors the pattern already used in api.c:do_reload() for direct
'initctl reload <svc>' calls.
Fixes: avahi/mdns stop causing mdns-alias restart counter increment
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
devmon: assert condition immediately if device already exists
1. udev fires early, creates device nodes in /dev/
2. Config is parsed later, calling devmon_add_cond() for each dev/ condition
3. At this point the device already exists but the inotify event was missed
4. The PR's fexist() check catches this — new node is added to the TAILQ and condition is immediately asserted
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
A service that is reloaded should not trigger dependants to be reloaded
unless the new <~cond> is used. Which is reserve for tightly coupled
services.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
For SysV services with pid:!/path, the pidfile belongs to the service
itself and Finit shouldn't delete it.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
This fixes a real bug where `initctl reload syslogd` unconditionally
clears syslogd's pid condition, causing all dependent services (dbus,
dnsmasq, etc.) to be stopped even though syslogd handles SIGHUP
gracefully and its PID/pidfile persist.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
A reload: script, like 'frrinit.sh reload' could potentially take a
while to finish, during which Finit would be blocked. This change
reuses the service_script_add(), used for ready: scripts, to track
these background helpers.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
When the kernel manages to reap a child process before we've moved it to
its proper cgroup it will return ESRCH (No such process), we can safely
ignore such errors for short-lived processes.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Conditions in Finit are dependencies: if A is asserted, service B is
allowed to run. When A goes through FLUX (e.g., upstream reloads),
dependents are PAUSED and then simply resumed when the condition is
reasserted -- this is the correct behavior for barrier-style deps
like <pid/syslogd>.
However, some setups have tightly coupled services where dependents
must be reloaded/restarted when an upstream service reloads, not just
resumed. E.g., the FRR routing stack on Infix OS:
netd <pid/mgmtd> ← zebra <!pid/netd> ← {staticd,ripd} <!pid/zebra>
When netd reloads (SIGHUP), zebra and its dependents must be restarted
to pick up the new configuration.
The new '~' condition prefix marks a dependency as flux-sensitive:
service <!~pid/netd> name:zebra ...
When the upstream condition goes FLUX and returns to ON, the dependent
is reloaded (SIGHUP) or restarted (noreload '!') instead of merely
resumed. Transitivity follows naturally through the condition chain.
Closes#416Closes#476
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
When 'initctl reload' is called after marking a service in a dependency
chain dirty, Finit fails to restart (unfreeze) affected services.
This patch updates the pidfile plugin to watch for IN_ATTRIB changes,
e.g. when a process uses utimensat() to update its pidfile, and adds
service_step_all() at end of reload cycle to guarantee convergence
after conditions are reasserted.
Issue #476
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
In a setup like this, when 'netd' is marked dirty and subsequently is
reloaded, e.g., using 'initctl reload', zebra is properly restarted,
but staticd isn't:
mgmtd <!> ← netd <pid/mgmtd> ← zebra <!pid/netd> ← staticd <!pid/zebra>
Finit must invalidate the condition of zebra to trigger a restart also
of staticd. This to guard against daemons like zebra that may fail to
clean up their pidfiles.
Fixes#475
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Verify that 'initctl reload foo' properly triggers dependent
services by checking that bar gets a new PID after the reload.
Also change the second test case from service/foo/running to
service/foo/ready which is the actual condition set by pidfile.so.
Fix a race in slay where the target process could exit between
the PID lookup and kill -9, causing spurious test failures in
tight kill loops (e.g., start-kill-service.sh).
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
When reloading a specific service with 'initctl reload foo', the
pid/foo and service/foo/ready conditions were never cleared, so
dependent services were not notified of the reload.
Clear the service's pid condition and, for pid/none notify types,
the ready condition before reloading. The conditions are then
reasserted by the pidfile inotify handler when the service touches
its PID file after processing SIGHUP.
For s6/systemd services the ready condition is left intact since
their readiness notification may not re-trigger on SIGHUP.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Users starting Finit based systems using U-Boot or Barebox may otherwise
not get a visible cursor at their prompt.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Define __NR_clone3 (435) ourselves when not provided by the toolchain
headers. The syscall number is stable kernel ABI and the same on all
architectures since Linux 5.3.
The existing runtime fallback to fork() handles older kernels that don't
support the syscall.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Strings like command, description, and environment may contain characters
that need escaping for valid JSON, e.g., embedded quotes in command line
arguments like -V "NanoPi R2S".
Add json_escape() helper to handle quotes, backslashes, and control chars.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
When a TTY exited with non-zero code (e.g., user with shell=/sbin/false),
it would enter restart state but never recover, requiring manual restart.
The throttling logic from commit f0032ab had two issues:
1. Duplicate exit code check in service_retry() created infinite timer loop
2. TTYs lacked default restart_tmo, causing timer to never start
Fix by removing duplicate check and ensuring TTYs get a 2-second default
restart_tmo for proper throttling.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
At least the sysvpart.sh regression test cannot run in parallel yet with
other tests (probably runparts.sh), so we must ensure the tests never
run in parallel, in particular at release.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Finit now requires being able to query at least for the root user and
group before starting any services.
Also, add support for using libraries installed in /usr/local
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Finit has support for "Please press Enter to activate this console."
which means there's no getty yet running. However, when profiling
systems with Finit, and embedded systems in general, a common metric
is the time from power-on to getty has started.
This commit makes sure to rename the process so that BusyBox pidof is
capable of detecting that "getty" has started. This is mostly for the
bootchart2 project's bootchartd, the native BusyBox bootchartd does not
have this issue.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
This is a refactor of getuser() and getgroup() so that they always
return a valid user, and group, for all normal use-cases. When an
error occurs we now handle it properly in service_fork() so as to
not attempt to start services with an invalid user/group setting
as root.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
When running Finit under boothcartd (bootchart2 project) the PATH is
lost due to a bug. This was a wakeup, so set critical variables in
main() early, before calling fs_init().
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Try out navigation.tabs, if we don't like it we can revert.
Reorg for stricter sections, what information actually belongs where?
E.g., introduction is now in a Getting Started section.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Some 'respawn' type services, like gettys, may hog the CPU in error
states if the service immediately exits. E.g., due to missing dev.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
During /bin/login phase the TTY device node is chowned and chmodded to
the authenticated user. It will remain in this state until the next
call to getty.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
A user reports inability to re-start getty on the console after logging
out from a serial console. After some digging it was found that the tty
was owned by the last user logged in and 600. Even thougn getty runs as
root, it did not have permission to re-open the device node.
Turns out there was a minor bug in the new capability code that cleared
all capabilities from the root user. A surprising amount of programs
worked just fine, but restarting getty gave it away.
The fix is to only call cap_setuid() when capabilities are set for the
service, otherwise we just fall back to setuid().
Also, refactor service_register() wrt. capabilities a bit so that we can
give users an early warning if the configuration is invalid, by adding a
parse_caps() helper function that calls cap_iab_from_text() to verify.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>