The block title is the service identity, and libconfuse merges two
sections that share one, without a word. Two blocks titled the same
in one file therefore loaded as a single service holding a mix of both
declarations, with scalars taken from the last block and lists reset
by it.
system/10-hotplug.conf.in is written this way: two udevd blocks, one
per candidate binary, the way the line-based format spelled a
fallback. Only the second survived the merge, so a system that has
/lib/systemd/systemd-udevd but no udevd got no udevd service at all,
and the whole `if = "udevd"` chain behind it went with it.
CFGF_NO_TITLE_DUPES turns the merge into a parse error naming the file
and the title, and the file is then rejected as a whole. The same
title in another file is untouched, that is how an administrator
overrides a system .conf.
Format detection had to stop agreeing with it. is_new_format() probes
by parsing, so a duplicate title made it answer "not block format" and
conf_parse_file() handed the file to the legacy parser, whose errors
buried the real message. The probe now clears the flag on its own
copy of the option array, keeping the verdict syntactic.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The reload-signal translation calls str2sig() and compares against
SIGHUP, but conf.c never included signal.h. glibc pulls it in
transitively, so the omission went unnoticed until a cross-compile
against uClibc-ng in Buildroot failed to build.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The runparts directory and the dbus pidfile and daemon paths are
embedded in double-quoted values of the generated block files. A
literal quote in either ends the value early and libconfuse rejects
the whole file, and a backslash is read as an escape sequence,
silently mangling the path. The legacy one-liners had no quoting, so
neither failure existed before the block conversion.
conf_escape() doubles backslashes and escapes quotes; verified by
round-tripping hostile paths through cfg_parse_buf().
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Finit's own generated services -- watchdogd, keventd, runparts, and
the dbus plugin -- still went through conf_save_service() as legacy
one-liners, so `initctl show keventd` taught the old format on a
system otherwise converted to the new one.
conf_save_service() now takes the block title and a printf-style body
and writes the file itself:
# Generated by finit:conf_save_service()
service keventd {
description = "Finit kernel event daemon"
runlevel = "S12345789"
notify = "none"
cgroup init {}
command = "/libexec/finit/keventd"
}
vfprintf() into the file also removes the fixed-size staging buffers
in the callers, where a long dbus pidfile path could truncate inside
a quoted string and take the whole generated file with it.
Semantics preserved: watch-only pid:! maps to pidfile without
pidfile-create, the watchdog keeps its watchdog:finit identity, and
log:console becomes log { file = "/dev/console" }.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The README.md symlinks exist for GitHub browsing and collide with
index.md when mkdocs renders both; exclude them like TODO.md.
Two links pointed at anchors that never existed: features.md has bold
captions rather than headings, so "Automatic Reload" gets an explicit
attr_list anchor for the link from the front page, and the TTY link
now spells the actual heading, controlling-tty-for-services.
mkdocs build is silent after this.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The new-format reference describes the block format on its own terms,
which is the wrong lookup direction for someone holding a legacy
one-liner. Aaron migrated Finix OS from the PR description, proving
the need for a token-in, key-out mapping in the user guide.
One table per part of a stanza, worked conversions for the shapes
that changed structurally -- cgroup selection and the three tty
variants -- and the dropped tokens listed with their replacements.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The line-based format accepts `service :80 ...`, deriving the name
from the command basename. The block format has no counterpart, the
title carries both name and ID. Implied by the format description,
but anyone converting such a line deserves to find it written down.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The workflows install libuev and libite from source and everything
else from apt, but never libconfuse, so every build job on this
branch dies in configure:
checking for libconfuse >= 3.3... no
Ubuntu ships libconfuse 3.3 with the static library included, which
covers both the static and regular builds. Staying on 3.3 in CI is
deliberate: it exercises the fallback paths marked
"XXX: Workaround for libConfuse <3.4" that a from-source 3.4 would
leave untested.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Aaron Andersen points out in the #492 discussion that the *Directory
settings carry more contract than create-and-chown: per-directory
modes, specific ownership rules, and cleanup toggles. Without them
config-dir was chowned to the service user, which systemd never does,
an existing directory with drifted ownership was left wrong, and the
runtime directory could not survive a restart.
Now matching systemd.exec(5), and where the man page is vague, the
code in setup_exec_directory():
- each directory takes a matching -mode key, octal with the leading
zero, default 0755. The mode of the named directory is locked
down again on every start, also when it already exists
- config-dir is created but never chowned
- the contents of an existing directory are left alone as long as
the owner is right; on drift everything under it is chowned back
- runtime-dir-preserve = no | restart | yes maps
RuntimeDirectoryPreserve=. A service still qualified to run when
the runtime directory would be removed is restarting, not
stopping, which is what svc_enabled() answers
The dir mechanics move to mksubsysd(), taking resolved ids, with
mksubsys() reduced to a name-resolving wrapper for the dbus plugin.
The child resolves uid/gid once for both directory setup and
privilege drop.
The symlink form, RuntimeDirectory=foo:bar, is not adopted.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
A service that drops privileges cannot create its own PID file in
/run, root owns it. Finit can create the file with pidfile-create,
but the daemon still cannot touch it to confirm a SIGHUP.
Five new settings, block format only: runtime-dir, state-dir,
cache-dir, logs-dir, and config-dir. The value is a directory name,
resolved under /run, /var/lib, /var/cache, /var/log, and /etc,
respectively. The directory is created before the service starts,
mode 0755 owned by user/group, and the full path is exported to the
process as RUNTIME_DIRECTORY, STATE_DIRECTORY, CACHE_DIRECTORY,
LOGS_DIRECTORY, and CONFIGURATION_DIRECTORY. Mode and ownership are
asserted at creation only, a daemon may tighten them afterwards.
The runtime directory is removed when the unit stops, after any
exec-stop-post script, like systemd with RuntimeDirectoryPreserve=no.
A completed run/task counts as stopped unless remain-after-exit keeps
it up. The other four persist across restarts.
These are the first settings with no legacy token: they are validated
by service_set_dir() and stored on the svc that service_register()
now returns. systemd accepts a list of directories per setting; this
is a single name for now, widening later is compatible since
libconfuse accepts a bare value for a list option.
The test sysroot gains libnss_files.so.2, which ldd cannot see, glibc
dlopen()s it. Without it getpwnam() fails inside the chroot, so
user/group settings never resolved and directory ownership could not
be tested.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
rmrf() is needed outside tmpfiles.c. The move also deduplicates the
nftw callback: the contents-only removal used by tmpfiles 'D' entries
is now rmcontents(), sharing the callback with rmrf().
mksubsys() did nothing at all when the user could not be resolved, no
directory and no message, and callers had no way to tell. Now the
directory is always created, ownership is best effort, and an unknown
user is warned about.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Settings that exist only in the block format have nowhere to go: the
legacy line cannot carry them, and service_register() returned an errno
that no caller ever read, so conf.c had no handle on the service it just
created. Return the svc instead, NULL with errno set on failure, errno
zero when a block is skipped on purpose.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The block format spells conditions as bare strings everywhere else, so
requiring `if = "<usr/foo>"` left one sigil behind, carried over from
the line-based `if:` token. A namespace separator already tells the two
apart: a value with a '/' is a condition, anything else is a service
name.
svc_ifthen() picks its mode from the start of the statement and applies
it to the whole, so a statement naming both kinds cannot be evaluated.
That is now an error, as are the old angle brackets, and either one
skips the block:
/etc/finit.conf: mixed: if: cannot mix a service name with a
condition in 'anchor,usr/enable-me', a statement must be all of
one kind, skipping
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Reference sections kept pointing at the line-based format they no longer
document. `sysv` and `task` sent the reader to Services for "<COND>",
the cgroups chapter opened by listing three legacy directives and then
explained further down that only two of them exist here, and the logging
chapter still gave "log:prio:facility.level,tag:ident" as the full
syntax.
Some claims were wrong independent of the format:
- a sysv is a supervised daemon, grouped with service in
SVC_TYPE_DAEMON, not a variation on task
- restart-max has no upper bound of 255, or any other
- the built-in rescue fallback runs in 12345789, not 12345
- conditional loading quotes system/10-hotplug.conf, not
system/hotplug.conf
- the key spells conflicts, not conflict
- the built-in getty no longer wants TERM last, it is a key
`if` takes either a service name or, in angle brackets, a condition,
decided in svc_ifthen(). Only the examples showed this, so it is now
said.
Terminology follows the split index.md already draws: a block is the new
format, a stanza the line-based one.
src/rescue.conf was still line-based, missed because it sits in src/
rather than system/ or contrib/.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
An ambient capability only reaches the effective set when euid is
non-zero, so a service that pairs `capabilities = { "^cap_..." }` with a
root user gets none of the restriction it asks for, and keeps the full
root set instead. Finit read the list, applied it, and said nothing. A
build without libcap dropped the list on the floor just as quietly.
Both now warn, naming the service:
nginx: ambient capabilities ('^') have no effect as root, use a
non-root user, or '%' and '!' entries
The ambient entries are read back from the parsed IAB value rather than
matched in the text, so inheritable ('%') and bounding ('!') entries stay
silent -- those work fine as root.
The warning repeats when the .conf files are re-read on runlevel change,
as parse warnings here already do.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The block conversion changed the bodies of the reference sections but
left every "**Syntax:**" header spelling the line-based format, so each
page opened by teaching the format it then stopped using. Six files
were missed entirely: runparts, files, capabilities, requirements,
runlevels, and switchroot.
runparts had no block spelling written down anywhere, though the parser
has read `runparts`, `runparts-progress`, and `runparts-sysv` all along.
tty gains a table per variant. Its three syntax lines carried nine
positional fields between them, which no longer describes anything the
parser accepts.
Fixes#148
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The examples people copy from were still written in the line-based
format, so the block format was documented but nowhere demonstrated.
Two names in contrib were accidents of the old format, where the
service name falls out of the command basename: the Alpine and Void
keymap task was called zcat, and Debian's console/keyboard setup tasks
carried a .sh suffix. They now carry the name their file implies.
Nothing referenced the old names.
The mdevd coldplug path keeps the name it has always had. Its legacy
line spelled the name inside the cgroup argument, where it names the
cgroup leaf and not the service, so the barrier condition really is
<run/mdevd-coldplug/success> and not the <run/coldplug/success> the
comment above it promises. Converted as-is so boot ordering does not
change; the discrepancy is now written down where it happens.
A list may not contain comments, the lexer sees the entries after the
'#' regardless:
modules = {
# "fbcon",
"softdog"
}
so the commented-out module candidates sit above the list instead.
setup-sysroot.sh removes 10-hotplug.conf from the test sysroot, so that
file is covered by parsing only, not by make check.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The syntax overview no longer describes a line-based format, since that
is not what the rest of the documentation shows. It now covers the
grammar, the two naming conventions, the nine aliases, and the leading
'-' on a path, and it says plainly that both formats are still read and
told apart per file by content. Without that, a reader with an
existing configuration is left wondering what happened to it.
service-opts.md was a list of modifiers to place between a directive
and its command, so it needed rewriting rather than translating: there
are no positions left to describe. It is now grouped by what the
settings do.
conditions.md needed correcting. It presented '!' as a condition
prefix alongside '~'. It is neither a condition nor a negation, it is
a flag on the block that means one thing on a service and another on a
run or task, so it is spelled reload-signal and required here, and the
page maps the old form to both.
Two things the pages claimed are not true. The kill delay range is
1-300, not 1-60, and stop and reload scripts are no longer run without
a timeout.
ChangeLog.md keeps its line-based examples. Those sit in historical
release entries, and rewriting them in a syntax that did not exist at
the time would misdate the format.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
parse_cgroup() takes two arguments that are not cgroupfs files: the
leaf directory to place the service in, and whether to hand the subtree
over to it. The block format could express neither. 'name' happened
to work, because a free-form key is emitted as name:VALUE and that is
what the parser looks for, but 'delegate' came out as delegate:true and
was filed as a cgroup setting, so it silently did nothing.
Declare both, and emit delegate as the bare flag the parser expects.
Neither means anything on a top-level group definition, so say so there
rather than emitting something that would be written to cgroupfs.
service podman {
cgroup containers { name = "podman" delegate = true }
...
}
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
This is what 'initctl create' and 'initctl edit -c' put in front of a
user writing their first .conf file, so it is also the whole of the
"initctl emits the new format" work: neither command generates syntax,
they copy this file and open an editor on it.
The ASCII diagram naming eight positional fields goes with it. A
block has no positions to explain.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Both keys prompt the question they should be answering.
'remain' decides whether a finished run or task keeps existing: without
it the entry is pruned, so the work re-runs on every runlevel entry,
initctl cannot see it, and its post script never fires. With it the
entry stays, is not re-run, and gets a teardown when stopped or when it
leaves its runlevels. That is systemd's RemainAfterExit, and 'remain'
is that name with the informative half cut off.
'manual' says how a service is started but not that it is about
starting at all.
remain -> remain-after-exit
manual -> manual-start
Both keep their old spelling as an alias, which they qualify for twice
over, as abbreviations of the canonical name and as the legacy
spellings.
While here, give sec_getbool() the alias argument its string and list
counterparts already take.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
A condition list could be led by '!', which is not a condition and not
a negation. It is a flag on the block, and it means two unrelated
things depending on which block it sits in: a service or sysv does not
handle SIGHUP and must be restarted to reload, while a run or task
must not hold up bootstrap. Writing '<!>' with no condition at all is
legal, which gives away that it was never an operator.
Give each meaning its own key, valid only where it applies:
service foo { reload-signal = "none" } # restart to reload
task bar { required = false } # do not hold up bootstrap
Using either on a block type it does not apply to warns, as does a '!'
left in a conditions list. Both still translate to that same '!',
which is all a legacy line can carry, so reload-signal takes SIGHUP or
none for now; str2sig() already accepts any case and an optional SIG
prefix.
This also clears the way for the conditions list to grow real
operators, '+' and '-' for asserted and deasserted, without '!'
sitting among them meaning something else entirely.
The '~' prefix stays. It belongs to the list: it marks a dependency
whose reload should propagate here.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The initial implementation was done using a naive translator of the
legacy one-liner format key by key, so it inherited encodings that the
block format exists to remove: a timeout packed into a script path, a
small comma-and-colon language inside the log string, sigils standing in
for booleans, and a log key (services) carrying three (!) types.
Settled naming against systemd, OpenRC, FreeBSD rc.subr, s6 and SMF.
Match systemd's semantics, not its naming.
pid -> pidfile, plus pidfile-create for the rare case
where Finit writes the file rather than the daemon
environment -> envfile, since it names a file to source, and the
top-level environment {} block sets variables
pre/post/... -> exec-start-pre, exec-start-ready, exec-stop,
exec-stop-post, exec-reload, exec-cleanup, each
with its own -timeout instead of "SEC,script"
halt, kill -> stop-signal, stop-timeout
restart -> restart for the policy, restart-max for the count
log -> a block with file, priority and identity, where
/dev/null and /dev/console are spelled as paths
group -> group and extra-groups, no longer positional
nowarn -> a leading - on command, as on envfile
List-valued keys take plural names. Aliases are desc, cond, mod,
caps, env, halt and kill; an alias may abbreviate the canonical name
or preserve a legacy spelling, nothing else.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
'make check' refreshes the sysroot through the setup-chroot rule, but
running a test script by hand does not, so the test exercises whichever
finit was installed last and reports on code that is no longer there.
Both a passing and a failing run are then meaningless, and nothing says
so.
Compare the built binary against the installed one at startup and fail
with the command that fixes it.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Both were bounded by killdelay, the delay between the stop signal and
SIGKILL, because service_run_script() had nothing else to reach for.
That conflates two things: how long the daemon may take to die, and
how long its stop script may run.
Give each hook a timeout of its own, defaulting to killdelay when
unset, so the existing behaviour is what you get until you ask for
something else. parse_script() already falls back that way for the
hooks that had one.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
A stop: or reload: script written with a timeout killed Finit at
config load:
service stop:5,/bin/true service.sh -- Boom
parse_script() takes the timeout as a pointer and the caller decides
whether it wants one. However, both stop: and reload: scripts so far
have no timeout, i.e., NULL. Guard the branch that reads a leading
number.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Three private ones had grown: fnread() in util.c, flen() behind
pid_cmdline()/pid_cgroup() in cgutil.c, and conf_read_template() in
conf.c. Two of them were also wrong in ways the others were not.
fnread() formatted the path into a char[256] and stat()ed it before
opening, so a longer path was silently truncated and then read from
whichever file the truncation happened to name, and the size could
change between the look and the read. flen() existed because neither
of those approaches works on procfs at all, where stat() reports zero
and the only way to learn the size is to read to EOF.
Add fslurp() to util.[ch], which every tool already links. It opens
first and sizes the fd it holds, treats st_size as a hint, and reads
until EOF, so procfs and regular files take the same path. Paths are
formatted by libite's vfopenf(), which allocates to fit. Callers that
need the byte count, /proc/PID/cmdline embeds NUL, ask for it.
fnread() keeps its signature and becomes a bounded copy out of the
result, so its one caller is unaffected.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
A template in the block format registered garbage. conf_parse_file()
routed every file with an '@' in its name straight to the legacy
parser, which read the block line by line: the section header became a
service whose command was the section title, and each key = value line
below it became an environment variable.
service serv:%i { ... } -> service 'serv:eth0' with argument '{'
Substitute %i over the whole file before parsing instead, so format
detection and both parsers see finished text. A bare name@.conf is
still skipped, it is the template rather than an instance of one.
The legacy parser no longer opens the file or substitutes per line, it
is handed the instantiated buffer, so the template convention now has
one implementation instead of two. conf_is_template() applies
basenm(), a directory with an '@' in its name is not a template.
libconfuse cannot name a buffer it parses before 3.4, so a typo in a
template would be reported against "[buf]". Parse through fmemopen()
with the file name preset until the floor moves.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Relocate process-wide global variables from legacy parser that ended up
there because it used to be conf.c, but which is now now frozen at the
4.x feature set. Each variable is moved to their respective "owner".
Give cgroup_current[] and cgroup_settings_current[] named bounds. Their
extern declarations were unsized, so sizeof() on them stopped compiling
once the definitions moved to another translation unit.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The one-liner format has grown crowded and very wide, and every new
service option makes it worse.
Add a second, block-based format, parsed with libconfuse:
service sshd {
description = "OpenSSH daemon"
runlevel = "2345"
command = "/usr/sbin/sshd -D $SSHD_OPTS"
}
Both formats keep the .conf extension and are detected per file by
content. Try-parse strictly with libconfuse; on a parse error,
re-parse leniently to tell a block file with a typo from a one-liner
file. Only a one-liner file reaches the legacy parser, a typo is
reported with its file and line.
Each block is translated to the canonical one-liner and registered
through the existing entry points, so the two formats cannot drift.
The one-liner parser is frozen at the 4.x feature set, new options
land only in the block schema. libconfuse 3.3 or later is required,
CFGF_KEYSTRVAL does not exist before it.
Covers service, task, run, sysv and tty blocks, the static directives,
and the cgroup, rlimit, set and log blocks. Templating and the
documentation rewrite are still to come.
The regression test covers translation of a service block to the
one-liner, a block-format /etc/finit.conf booting with set {} applied
at bootstrap, both formats side by side, and rejection of a typo at
block and at root level.
A rejected file must not fall through to the legacy parser, which
registers a bogus unstartable service per line. assert_num_children
cannot see that, the bogus service has no children either, so the
check is assert_num_services.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The one-liner parser is about to be joined by a second, block-based
format. Give it a name that says which of the two it implements,
before any content changes make the diff hard to follow.
No functional change.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Every test left a stray `sleep 300` behind, reparented to PID 1, where
it lingered for up to five minutes after the test had finished.
wdstart() runs the watchdog in a subshell, so $! is the pid of the
subshell, not of the sleep it forks. wdkill() killed the subshell and
orphaned the sleep.
Kill the child first, killing the subshell puts the sleep beyond the
reach of pkill -P. Neither kill is sure to match, and wdkill() runs
from the EXIT trap under set -e, so both must tolerate failure. Also
return early when wdpid is unset, for failures before wdstart() runs.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Finit 5.0 changes the .conf syntax, which has been essentially
unchanged since 1.x. The published docs track master, so when 5.x
lands, 4.x users lose their reference.
Publish the site under a per-major directory, /4.x/ for now, with
the Material version selector to switch between them. The selector
only needs mike's file layout -- a versions.json at the site root --
which the deploy job now generates from the version directories in
the pages repo, so mike itself is not needed.
The major comes from AC_INIT and the future 4.x maintenance branch
is already in the workflow triggers, so once 5.0 is on master, doc
fixes on the 4.x branch keep /4.x/ updated. A root index.html
redirects to the newest version, and a 404.html rewrites
pre-versioned deep links so old bookmarks and search hits land in
the right place.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Both the Finit project and Infix use the same MkDocs Material setup, and
in the latter the User Guide has picked up a lot of polish that never
made it back here: a single sidebar with section indexes instead of
tabs, footnote tooltips, more pymdownx markup, image zoom tuning, and no
generator advert in the footer.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
* src/pid.c: note the stale-pidfile-cleanup exception to the
documented "Finit does not touch pid:! pidfiles" rule.
* doc/config/services.md: add a user-facing paragraph on the same.
* doc/ChangeLog.md: add Unreleased section covering this PR --
stale pidfile cleanup, restart log with signal name and core
dump flag, and the SIGUNKOWN typo fix.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Cover the scenario fixed in "service: clean stale pidfile after
unclean daemon exit": a daemon with a pid:!/path config dies via
SIGKILL, leaving its pidfile behind, and the next instance must
still come up.
Add a 'serv -x' flag (refuse to start when the pidfile already
exists, dbus-style) so the test actually exercises the cleanup --
without it, plain 'serv' would happily overwrite the file and the
test would pass with or without the fix.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The fallback for unknown signal numbers in sig_name() returned the
misspelled "SIGUNKOWN". Now that this string surfaces in user-
facing logs ("killed by …"), fix the typo.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Replace the bare signal number ("by signal: 9") with the symbolic
name ("killed by SIGKILL") and annotate when the kernel wrote a
core:("killed by SIGSEGV, core dumped"). Makes the restart line
self-explanatory and gives operators a strong breadcrumb when a
daemon dies unexpectedly.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
With `pid:!/path` Finit does not manage the file -- the daemon
creates it on start and removes it on graceful exit. If the daemon
dies before cleanup (SIGKILL, OOM, segfault, exit during startup)
the file lingers and can block the next instance from starting,
e.g. dbus-daemon refuses with EEXIST and the restart loop fails.
Remove the file when it still names the just-reaped PID and that
PID is no longer alive (the liveness check guards against reuse).
Called from service_cleanup(), and from service_monitor()'s
forking+starting branch where cleanup was previously skipped.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
status() returns a pointer to a single static buffer, so calling it twice
in the same cprintf() argument list — status(3) and status(rc) — causes
one to overwrite the other before the format string is rendered. When
status(3) wins, the line shows [ ⋯ ] instead of [ OK ]. Fix by copying
status(rc) into a local buffer before calling status(3).
Also drop the delline() calls added to print() — that macro writes \033[2K
to buffered stdout while cprintf() writes unbuffered to stderr, so the
erase sequences can arrive out of order. The \r\e[K already present in
the cprintf format strings makes them redundant anyway.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Check return value of remove() in delete_cb() and log failures via
dbg(), CID 909395
Replace stat() calls with open(O_DIRECTORY)+ fstat() for newroot and "/"
checks. Eliminates the check-then-use race and lets O_DIRECTORY do the
isdir validation atomically, CID 909394
Drop the explicit close(0/1/2) before opening /dev/console. dup2()
closes the old targets itself, so open() returns a fd > STDERR_FILENO
that can always be closed unconditionally, removing the conditional
guard, CID 909393
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Refactor print() to emit description + final status in a single call to
cprintf(), preventing kernel messages from splitting the two parts.
For two-phase print(-1,...) + print_result() sequences used by, e.g.,
run_interactive, save the last description and re-print it before the
[ OK ] / [FAIL] output so the status is never left stranded on a blank
line when command output or kernel messages have scrolled away the
original description.
Finally, add print_exit() which drains the console output buffer with
tcdrain(2) and resets ANSI SGR attributes + cursor visibility before the
kernel takes back the console on reboot/halt, preventing escape code
leakage into bootloader or early-kernel output.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
When a service without SIGHUP reload support (noreload) is touched and
'initctl reload' is called, service_update_rdeps() correctly identifies
its reverse dependencies but only marks them dirty. It does not clear
the service's condition, so when service_step_all() runs:
- rdeps supporting SIGHUP hit the sm_in_reload() guard and break early,
left running while their dependency is being killed.
- rdeps without SIGHUP support may receive SIGTERM too late, after the
dependency has already died and broken their connection, causing them
to exit from RUNNING state and have their restart counter incremented.
Fix by calling cond_clear() on the service's condition immediately in
service_update_rdeps(), before service_step_all() runs. cond_clear()
calls cond_update() which calls service_step() inline on all affected
services, which see COND_OFF and transition to STOPPING_STATE — all
before SIGTERM is ever sent to the dependency itself.
This mirrors the pattern already used in api.c:do_reload() for direct
'initctl reload <svc>' calls.
Fixes: avahi/mdns stop causing mdns-alias restart counter increment
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
devmon: assert condition immediately if device already exists
1. udev fires early, creates device nodes in /dev/
2. Config is parsed later, calling devmon_add_cond() for each dev/ condition
3. At this point the device already exists but the inotify event was missed
4. The PR's fexist() check catches this — new node is added to the TAILQ and condition is immediately asserted
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
A service that is reloaded should not trigger dependants to be reloaded
unless the new <~cond> is used. Which is reserve for tightly coupled
services.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
For SysV services with pid:!/path, the pidfile belongs to the service
itself and Finit shouldn't delete it.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
This fixes a real bug where `initctl reload syslogd` unconditionally
clears syslogd's pid condition, causing all dependent services (dbus,
dnsmasq, etc.) to be stopped even though syslogd handles SIGHUP
gracefully and its PID/pidfile persist.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
A reload: script, like 'frrinit.sh reload' could potentially take a
while to finish, during which Finit would be blocked. This change
reuses the service_script_add(), used for ready: scripts, to track
these background helpers.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
When the kernel manages to reap a child process before we've moved it to
its proper cgroup it will return ESRCH (No such process), we can safely
ignore such errors for short-lived processes.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Conditions in Finit are dependencies: if A is asserted, service B is
allowed to run. When A goes through FLUX (e.g., upstream reloads),
dependents are PAUSED and then simply resumed when the condition is
reasserted -- this is the correct behavior for barrier-style deps
like <pid/syslogd>.
However, some setups have tightly coupled services where dependents
must be reloaded/restarted when an upstream service reloads, not just
resumed. E.g., the FRR routing stack on Infix OS:
netd <pid/mgmtd> ← zebra <!pid/netd> ← {staticd,ripd} <!pid/zebra>
When netd reloads (SIGHUP), zebra and its dependents must be restarted
to pick up the new configuration.
The new '~' condition prefix marks a dependency as flux-sensitive:
service <!~pid/netd> name:zebra ...
When the upstream condition goes FLUX and returns to ON, the dependent
is reloaded (SIGHUP) or restarted (noreload '!') instead of merely
resumed. Transitivity follows naturally through the condition chain.
Closes#416Closes#476
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
When 'initctl reload' is called after marking a service in a dependency
chain dirty, Finit fails to restart (unfreeze) affected services.
This patch updates the pidfile plugin to watch for IN_ATTRIB changes,
e.g. when a process uses utimensat() to update its pidfile, and adds
service_step_all() at end of reload cycle to guarantee convergence
after conditions are reasserted.
Issue #476
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
In a setup like this, when 'netd' is marked dirty and subsequently is
reloaded, e.g., using 'initctl reload', zebra is properly restarted,
but staticd isn't:
mgmtd <!> ← netd <pid/mgmtd> ← zebra <!pid/netd> ← staticd <!pid/zebra>
Finit must invalidate the condition of zebra to trigger a restart also
of staticd. This to guard against daemons like zebra that may fail to
clean up their pidfiles.
Fixes#475
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Verify that 'initctl reload foo' properly triggers dependent
services by checking that bar gets a new PID after the reload.
Also change the second test case from service/foo/running to
service/foo/ready which is the actual condition set by pidfile.so.
Fix a race in slay where the target process could exit between
the PID lookup and kill -9, causing spurious test failures in
tight kill loops (e.g., start-kill-service.sh).
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
When reloading a specific service with 'initctl reload foo', the
pid/foo and service/foo/ready conditions were never cleared, so
dependent services were not notified of the reload.
Clear the service's pid condition and, for pid/none notify types,
the ready condition before reloading. The conditions are then
reasserted by the pidfile inotify handler when the service touches
its PID file after processing SIGHUP.
For s6/systemd services the ready condition is left intact since
their readiness notification may not re-trigger on SIGHUP.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Users starting Finit based systems using U-Boot or Barebox may otherwise
not get a visible cursor at their prompt.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Define __NR_clone3 (435) ourselves when not provided by the toolchain
headers. The syscall number is stable kernel ABI and the same on all
architectures since Linux 5.3.
The existing runtime fallback to fork() handles older kernels that don't
support the syscall.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Strings like command, description, and environment may contain characters
that need escaping for valid JSON, e.g., embedded quotes in command line
arguments like -V "NanoPi R2S".
Add json_escape() helper to handle quotes, backslashes, and control chars.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
When a TTY exited with non-zero code (e.g., user with shell=/sbin/false),
it would enter restart state but never recover, requiring manual restart.
The throttling logic from commit f0032ab had two issues:
1. Duplicate exit code check in service_retry() created infinite timer loop
2. TTYs lacked default restart_tmo, causing timer to never start
Fix by removing duplicate check and ensuring TTYs get a 2-second default
restart_tmo for proper throttling.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
At least the sysvpart.sh regression test cannot run in parallel yet with
other tests (probably runparts.sh), so we must ensure the tests never
run in parallel, in particular at release.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>