A service that drops privileges cannot create its own PID file in
/run, root owns it. Finit can create the file with pidfile-create,
but the daemon still cannot touch it to confirm a SIGHUP.
Five new settings, block format only: runtime-dir, state-dir,
cache-dir, logs-dir, and config-dir. The value is a directory name,
resolved under /run, /var/lib, /var/cache, /var/log, and /etc,
respectively. The directory is created before the service starts,
mode 0755 owned by user/group, and the full path is exported to the
process as RUNTIME_DIRECTORY, STATE_DIRECTORY, CACHE_DIRECTORY,
LOGS_DIRECTORY, and CONFIGURATION_DIRECTORY. Mode and ownership are
asserted at creation only, a daemon may tighten them afterwards.
The runtime directory is removed when the unit stops, after any
exec-stop-post script, like systemd with RuntimeDirectoryPreserve=no.
A completed run/task counts as stopped unless remain-after-exit keeps
it up. The other four persist across restarts.
These are the first settings with no legacy token: they are validated
by service_set_dir() and stored on the svc that service_register()
now returns. systemd accepts a list of directories per setting; this
is a single name for now, widening later is compatible since
libconfuse accepts a bare value for a list option.
The test sysroot gains libnss_files.so.2, which ldd cannot see, glibc
dlopen()s it. Without it getpwnam() fails inside the chroot, so
user/group settings never resolved and directory ownership could not
be tested.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The block format spells conditions as bare strings everywhere else, so
requiring `if = "<usr/foo>"` left one sigil behind, carried over from
the line-based `if:` token. A namespace separator already tells the two
apart: a value with a '/' is a condition, anything else is a service
name.
svc_ifthen() picks its mode from the start of the statement and applies
it to the whole, so a statement naming both kinds cannot be evaluated.
That is now an error, as are the old angle brackets, and either one
skips the block:
/etc/finit.conf: mixed: if: cannot mix a service name with a
condition in 'anchor,usr/enable-me', a statement must be all of
one kind, skipping
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
parse_cgroup() takes two arguments that are not cgroupfs files: the
leaf directory to place the service in, and whether to hand the subtree
over to it. The block format could express neither. 'name' happened
to work, because a free-form key is emitted as name:VALUE and that is
what the parser looks for, but 'delegate' came out as delegate:true and
was filed as a cgroup setting, so it silently did nothing.
Declare both, and emit delegate as the bare flag the parser expects.
Neither means anything on a top-level group definition, so say so there
rather than emitting something that would be written to cgroupfs.
service podman {
cgroup containers { name = "podman" delegate = true }
...
}
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Both keys prompt the question they should be answering.
'remain' decides whether a finished run or task keeps existing: without
it the entry is pruned, so the work re-runs on every runlevel entry,
initctl cannot see it, and its post script never fires. With it the
entry stays, is not re-run, and gets a teardown when stopped or when it
leaves its runlevels. That is systemd's RemainAfterExit, and 'remain'
is that name with the informative half cut off.
'manual' says how a service is started but not that it is about
starting at all.
remain -> remain-after-exit
manual -> manual-start
Both keep their old spelling as an alias, which they qualify for twice
over, as abbreviations of the canonical name and as the legacy
spellings.
While here, give sec_getbool() the alias argument its string and list
counterparts already take.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
A condition list could be led by '!', which is not a condition and not
a negation. It is a flag on the block, and it means two unrelated
things depending on which block it sits in: a service or sysv does not
handle SIGHUP and must be restarted to reload, while a run or task
must not hold up bootstrap. Writing '<!>' with no condition at all is
legal, which gives away that it was never an operator.
Give each meaning its own key, valid only where it applies:
service foo { reload-signal = "none" } # restart to reload
task bar { required = false } # do not hold up bootstrap
Using either on a block type it does not apply to warns, as does a '!'
left in a conditions list. Both still translate to that same '!',
which is all a legacy line can carry, so reload-signal takes SIGHUP or
none for now; str2sig() already accepts any case and an optional SIG
prefix.
This also clears the way for the conditions list to grow real
operators, '+' and '-' for asserted and deasserted, without '!'
sitting among them meaning something else entirely.
The '~' prefix stays. It belongs to the list: it marks a dependency
whose reload should propagate here.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The initial implementation was done using a naive translator of the
legacy one-liner format key by key, so it inherited encodings that the
block format exists to remove: a timeout packed into a script path, a
small comma-and-colon language inside the log string, sigils standing in
for booleans, and a log key (services) carrying three (!) types.
Settled naming against systemd, OpenRC, FreeBSD rc.subr, s6 and SMF.
Match systemd's semantics, not its naming.
pid -> pidfile, plus pidfile-create for the rare case
where Finit writes the file rather than the daemon
environment -> envfile, since it names a file to source, and the
top-level environment {} block sets variables
pre/post/... -> exec-start-pre, exec-start-ready, exec-stop,
exec-stop-post, exec-reload, exec-cleanup, each
with its own -timeout instead of "SEC,script"
halt, kill -> stop-signal, stop-timeout
restart -> restart for the policy, restart-max for the count
log -> a block with file, priority and identity, where
/dev/null and /dev/console are spelled as paths
group -> group and extra-groups, no longer positional
nowarn -> a leading - on command, as on envfile
List-valued keys take plural names. Aliases are desc, cond, mod,
caps, env, halt and kill; an alias may abbreviate the canonical name
or preserve a legacy spelling, nothing else.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
'make check' refreshes the sysroot through the setup-chroot rule, but
running a test script by hand does not, so the test exercises whichever
finit was installed last and reports on code that is no longer there.
Both a passing and a failing run are then meaningless, and nothing says
so.
Compare the built binary against the installed one at startup and fail
with the command that fixes it.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
A stop: or reload: script written with a timeout killed Finit at
config load:
service stop:5,/bin/true service.sh -- Boom
parse_script() takes the timeout as a pointer and the caller decides
whether it wants one. However, both stop: and reload: scripts so far
have no timeout, i.e., NULL. Guard the branch that reads a leading
number.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
A template in the block format registered garbage. conf_parse_file()
routed every file with an '@' in its name straight to the legacy
parser, which read the block line by line: the section header became a
service whose command was the section title, and each key = value line
below it became an environment variable.
service serv:%i { ... } -> service 'serv:eth0' with argument '{'
Substitute %i over the whole file before parsing instead, so format
detection and both parsers see finished text. A bare name@.conf is
still skipped, it is the template rather than an instance of one.
The legacy parser no longer opens the file or substitutes per line, it
is handed the instantiated buffer, so the template convention now has
one implementation instead of two. conf_is_template() applies
basenm(), a directory with an '@' in its name is not a template.
libconfuse cannot name a buffer it parses before 3.4, so a typo in a
template would be reported against "[buf]". Parse through fmemopen()
with the file name preset until the floor moves.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The one-liner format has grown crowded and very wide, and every new
service option makes it worse.
Add a second, block-based format, parsed with libconfuse:
service sshd {
description = "OpenSSH daemon"
runlevel = "2345"
command = "/usr/sbin/sshd -D $SSHD_OPTS"
}
Both formats keep the .conf extension and are detected per file by
content. Try-parse strictly with libconfuse; on a parse error,
re-parse leniently to tell a block file with a typo from a one-liner
file. Only a one-liner file reaches the legacy parser, a typo is
reported with its file and line.
Each block is translated to the canonical one-liner and registered
through the existing entry points, so the two formats cannot drift.
The one-liner parser is frozen at the 4.x feature set, new options
land only in the block schema. libconfuse 3.3 or later is required,
CFGF_KEYSTRVAL does not exist before it.
Covers service, task, run, sysv and tty blocks, the static directives,
and the cgroup, rlimit, set and log blocks. Templating and the
documentation rewrite are still to come.
The regression test covers translation of a service block to the
one-liner, a block-format /etc/finit.conf booting with set {} applied
at bootstrap, both formats side by side, and rejection of a typo at
block and at root level.
A rejected file must not fall through to the legacy parser, which
registers a bogus unstartable service per line. assert_num_children
cannot see that, the bogus service has no children either, so the
check is assert_num_services.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Every test left a stray `sleep 300` behind, reparented to PID 1, where
it lingered for up to five minutes after the test had finished.
wdstart() runs the watchdog in a subshell, so $! is the pid of the
subshell, not of the sleep it forks. wdkill() killed the subshell and
orphaned the sleep.
Kill the child first, killing the subshell puts the sleep beyond the
reach of pkill -P. Neither kill is sure to match, and wdkill() runs
from the EXIT trap under set -e, so both must tolerate failure. Also
return early when wdpid is unset, for failures before wdstart() runs.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Cover the scenario fixed in "service: clean stale pidfile after
unclean daemon exit": a daemon with a pid:!/path config dies via
SIGKILL, leaving its pidfile behind, and the next instance must
still come up.
Add a 'serv -x' flag (refuse to start when the pidfile already
exists, dbus-style) so the test actually exercises the cleanup --
without it, plain 'serv' would happily overwrite the file and the
test would pass with or without the fix.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
A service that is reloaded should not trigger dependants to be reloaded
unless the new <~cond> is used. Which is reserve for tightly coupled
services.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Conditions in Finit are dependencies: if A is asserted, service B is
allowed to run. When A goes through FLUX (e.g., upstream reloads),
dependents are PAUSED and then simply resumed when the condition is
reasserted -- this is the correct behavior for barrier-style deps
like <pid/syslogd>.
However, some setups have tightly coupled services where dependents
must be reloaded/restarted when an upstream service reloads, not just
resumed. E.g., the FRR routing stack on Infix OS:
netd <pid/mgmtd> ← zebra <!pid/netd> ← {staticd,ripd} <!pid/zebra>
When netd reloads (SIGHUP), zebra and its dependents must be restarted
to pick up the new configuration.
The new '~' condition prefix marks a dependency as flux-sensitive:
service <!~pid/netd> name:zebra ...
When the upstream condition goes FLUX and returns to ON, the dependent
is reloaded (SIGHUP) or restarted (noreload '!') instead of merely
resumed. Transitivity follows naturally through the condition chain.
Closes#416Closes#476
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
When 'initctl reload' is called after marking a service in a dependency
chain dirty, Finit fails to restart (unfreeze) affected services.
This patch updates the pidfile plugin to watch for IN_ATTRIB changes,
e.g. when a process uses utimensat() to update its pidfile, and adds
service_step_all() at end of reload cycle to guarantee convergence
after conditions are reasserted.
Issue #476
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Device conditions tracked by devmon were lost on `initctl reload`
because the reconf path did not re-assert them. Add devmon_reconf()
to iterate all tracked device nodes and set or clear their conditions
based on current device presence.
Signed-off-by: Mattias Walström <lazzer@gmail.com>
Verify that 'initctl reload foo' properly triggers dependent
services by checking that bar gets a new PID after the reload.
Also change the second test case from service/foo/running to
service/foo/ready which is the actual condition set by pidfile.so.
Fix a race in slay where the target process could exit between
the PID lookup and kill -9, causing spurious test failures in
tight kill loops (e.g., start-kill-service.sh).
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Finit now requires being able to query at least for the root user and
group before starting any services.
Also, add support for using libraries installed in /usr/local
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
A following compilation error was observed:
| libsystemd/sd-daemon.c:64: undefined reference to `strlcpy'
fix it by include the required libite dependency.
Signed-off-by: Ming Liu <liu.ming50@gmail.com>
This commit introduces a bare-bones replacement for libsystemd:
- Build .so file and add --with-libsystemd to configure
- Add capabilities support to test/src/serv.c
- Update tests to account for a Finit built w/o libsystemd support
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
A systemd service should only use NOTIFY_SOCKET and an s6 style
service reads its notify descript from the command line. For
details, see notify.sh
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
This commit introduces a stripped down sd_notify(), taken from the
systemd man page example, which is used by the serv daemon in lieu
of the previous broken implementation.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
New test adds /bin/fail.sh to verify that a failing pre:script (that
also takes too long to run) is detected: exit code and timeout.
Ensure existing test pass full path to /sbin/fail.sh script.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
As of Finit v4.6 we no longer assert the PID condition for services
declaring themselves as notify != pid. We replace D with a forking
service to catch any future regressions in the pidfile plugin.
No need to check reload PID of D, it is enough to check PID of C.
Also, reduce the number of retries at startup. If we haven't gone
up within 10 sec with this tiny config something is really wrong.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
This fixes an issue when finding a global environment variable with
spaces in the variable name:
set COLORTERM=yes
Literally, 'set COLORTERM' was the name of the variable.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>