Rewrite doc/keventd.md from a 14-line stub into comprehensive
documentation covering all features of the new unified keventd:
device node creation, persistent symlinks, firmware loading,
module loading, coldplug, conditions, and command-line usage.
Update doc/conditions.md to list keventd as the primary provider
of dev/* and sys/pwr/* conditions, with devmon as fallback when
an external device manager is used instead.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
A pass over the whole branch before merge, mostly in libink since
that is the new code and the part exposed to the wire. Grouped here
rather than scattered so the review is easy to read in one place.
libink parser and dispatch:
- Bound reader lengths so a 32-bit size_t can't wrap a wire length
past the guard and read out of bounds. Reachable pre-auth on any
bus, so it matters on the 32-bit targets Finit runs on.
- Drop a peer when a reply send fails instead of limping on with a
half-written frame; a built-in whose send failed used to fall
through and put a second frame on the wire.
initctl:
- Copy a D-Bus error name out of the reply before closing the client;
the reply points into memory the close frees. Both error paths now
share one helper so this can't creep back.
Authorization:
- Take the caller's groups from the kernel (SO_PEERCRED plus
SO_PEERGROUPS) rather than getpwuid()/getgrouplist(), which go
through NSS and can block PID 1 on a slow LDAP or SSSD backend.
The check is now a lookup against the group resolved once at init,
with no NSS and no 256 KiB array on the stack. A caller reaching
us through a broker carries no group set, so system-bus privileged
methods are root-only; the local bus keeps group support. See
libink/README.md for the note on lifting that.
Shutdown:
- Call dbus_exit() from the shutdown path so the server, its peers,
and the socket are let go cleanly. The teardown existed but nobody
called it.
Tests, CI, docs:
- A fuzz target for the message parser, run as a quick sweep in the
suite and properly under libFuzzer in CI, with the corpus carried
between runs. The -as-uid tests drop groups the way a login does
so SO_PEERGROUPS sees the right set, and widen the test socket to
reach the per-method check behind the 0660 gate. Bring the GitHub
actions up to versions that run on Node 24, and tidy a few small
things a /simplify pass turned up.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
__msg_parse() turns bytes off a socket into pointers, before anything
has vouched for the peer, and it is the only place in libink that
does. It had no test of its own beyond whatever the other tests
happened to send it, all of it well-formed.
The target checks the parser's contract, not merely that it survived.
A header field must point into the header field array, and terminate
inside it, and the parse must never claim more bytes than it was
handed. Crash-only would pass a parser that walked into the body and
returned fields from there, since those bytes were handed over too.
The expected bounds are derived from the raw header rather than from
the parser, so the two have to agree independently.
Every input is copied into an allocation sized to it first. Reading
past the end of a roomy buffer stays inside the allocation and the
sanitizer never sees it; against an exact one the same read is a
fault, which is where the sharpest findings come from.
Under libFuzzer it is an ordinary fuzz target and named files replay,
which is how a find gets reproduced. With no arguments it runs a
fixed sweep -- every truncation, every single-byte corruption, every
value of the length that decides where the header ends, and seeded
garbage -- so the suite covers the same contract on every build,
without clang or a corpus in the tree. It takes 40 ms.
CI fuzzes it properly on every pull request, keeps the crashers, and
carries the corpus between runs so it reaches deeper over time than
any single run can. Note that clang links the fuzzer runtime against
the newest GCC tree it finds, so the libstdc++ headers have to match
that one and not the default compiler, which is worth saying since
installing the obvious package leaves you exactly where you started.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
"Set not yet implemented" reads as a promise. Finit exposes no
writable property and has no use for one: everything a caller might
want to change is a Manager1 method, where the authorization lives.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Finit speaks D-Bus itself now and claims org.finit on the system bus
when it finds one, but nothing in a default build ever brings that bus
up. The plugin that does was opt-in, so the built-in support sat idle
unless the integrator knew to ask for both halves.
Defaulting it on is only reasonable if the result stays the admin's to
change, and a service registered from C through conf_save_service() is
not: it lands in the run path where it cannot be overridden or emptied
out. So the daemon moves to 20-dbus.conf and its directories to
tmpfiles.d/dbus.conf, the same way hotplug and every other daemon we
ship them for. The plugin keeps only what has to look at the running
system, the stale pidfile and the machine UUID.
Those directories are no longer chowned to messagebus. tmpfiles.d
skips a line whose user does not exist rather than falling back, so
the plugin's messagebus/dbus/root ladder has no equivalent there, and
dbus-daemon binds its socket before dropping privileges anyway.
The plugin already bows out where there is no dbus-daemon installed,
so systems that never wanted a bus are unaffected, and
--disable-dbus-plugin is there for those that have one and would still
rather init left it alone.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
On the local bus SO_PEERCRED says who is calling and the kernel is the
one saying it. Behind a broker one connection carries every caller,
so that credential describes dbus-daemon and nothing else, and every
privileged method was refused there, root included.
Ask the bus driver instead. libink parks the call and hands us the
sender; we ask GetConnectionUnixUser and answer when the reply lands,
through the same event loop as everything else. Nothing blocks:
blocking in PID 1 is why libuEv exists. That needs calls libink can
make on a connection it already has, so it gained those too.
Answers are cached, since a bus never reuses a unique name while it
runs. Not across a restart though: a new dbus-daemon numbers from
scratch and :1.7 becomes somebody else, so the cache goes when the
broker does. A sender name too long to key on is refused rather than
truncated, two callers sharing a truncated key would share an
identity.
Privilege is no longer uid 0 alone. The socket is already owned by
the --with-group group, so refusing its members every method that
changes anything left a wheel user able to open the bus and unable to
reboot. Both gates now say the same thing.
Group membership needs NSS, which the C library loads with dlopen(),
so the lookup is compiled out where Finit is built to link statically.
That leaves such a build root-only, which is worth saying out loud
rather than leaving to be discovered.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The D-Bus socket was bound world read/write, on the reasoning that
SO_PEERCRED authorizes each method anyway. That leaves the read-only
surface open to every local user, and it quietly ignores --with-group:
a system that restricts initctl to the wheel group still handed the
same service state to anyone who asked over the bus.
Bind it 0660 and chown it to the configured group, the same gate the
fallback socket has always had. libink takes the mode as an argument
rather than assuming one, since who may connect is the embedder's
policy, not the library's.
The mode is applied at bind(), so there is no window where the socket
is more permissive than intended.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
libink was written against the only bus it had, its own, where the
peer on the other end is the client. A broker is not: it routes for
senders it names itself, expects a DESTINATION on anything addressed
through it, and answers on its own schedule rather than next.
Runlevels go on the wire as S and N rather than the digits Finit
keeps internally, since that is what a caller outside Finit means by
one.
The library stays a convenience library, linked into finit and
initctl and installed nowhere: the ABI promise waits until libink is
its own project.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The summary table, the per-service detail, JSON and the quiet and
ident forms all read state Finit already publishes, so they read it
from the bus like everything else rather than through a second path
that has to be kept in step.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Runlevel and version are state, not actions, so they belong behind
org.freedesktop.DBus.Properties rather than another method each.
Finit also claims org.finit on the system bus when it finds one, so
ordinary D-Bus clients can reach it without knowing about
/run/finit/bus. Opportunistic on purpose: no dbus-daemon is a normal
state for the systems Finit runs on, not an error to report.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
A cgroup holding processes cannot enable controllers for its children,
so init/ had to stay a leaf. The hotplug helpers 10-hotplug.conf.in
places there ended up in groups where cpu.weight and friends could
never be set.
Keeping PID 1 in the root cgroup makes init/ a domain like the others.
It also unbreaks lxc-based runtimes: liblxc bases the container tree on
PID 1's cgroup and only special-cases systemd's init.scope/, so under
Finit it landed containers in init/, with no controllers available.
Issue #497
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The line-based format has had the flag since v4.4 (issue #286), where
it prepends -p to the built-in getty, which turns it into login -p and
passes the environment on. The block format was written from the three
documented tty variants and the flags listed in the tty documentation,
and passenv was in neither, so it was left out. Converting a tty line
that used it therefore lost it, with nothing said.
It only reaches the built-in getty. An external getty is handed its
arguments through command, so there is nowhere to put a -p, and the
setting is refused with a warning rather than quietly ignored.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The migration guide told anyone holding the repeated-stanza idiom for a
per-platform service to split the variants across files or stay on the
line-based format, because a block title is an identity and the
variants have to share one barrier. provides is the answer, so the
guide converts that shape now instead of routing around it, and the
header of 10-hotplug.conf.in no longer points at the workaround.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The migration guide covered a stanza at a time, which is the wrong
shape for the two idioms that repeated a whole stanza. One of them,
several candidate binaries for one service, is now a command list.
The other, one service gated differently per platform, has no block
equivalent: those blocks share an identity because they share the
barrier condition downstream services wait for, so they cannot be
given separate titles. For that one the guide says to split the
variants across files, or leave that file in the line-based format,
which Finit still reads.
The udevd example in services.md taught the merge-broken form, and
system/10-hotplug.conf.in pointed readers at it for their syslogd.
Also lists libConfuse among the build dependencies. It has been
mandatory since the new .conf format landed, and build.md still said
two libraries. And corrects the note on variable expansion: it is
${VAR} that libconfuse expands when the file is read, with
${VAR:-default} supported. A plain $VAR reaches the service, which is
what makes `command = "syslogd -F $SYSLOGD_ARGS"` work with envfile.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The README.md symlinks exist for GitHub browsing and collide with
index.md when mkdocs renders both; exclude them like TODO.md.
Two links pointed at anchors that never existed: features.md has bold
captions rather than headings, so "Automatic Reload" gets an explicit
attr_list anchor for the link from the front page, and the TTY link
now spells the actual heading, controlling-tty-for-services.
mkdocs build is silent after this.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The new-format reference describes the block format on its own terms,
which is the wrong lookup direction for someone holding a legacy
one-liner. Aaron migrated Finix OS from the PR description, proving
the need for a token-in, key-out mapping in the user guide.
One table per part of a stanza, worked conversions for the shapes
that changed structurally -- cgroup selection and the three tty
variants -- and the dropped tokens listed with their replacements.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The line-based format accepts `service :80 ...`, deriving the name
from the command basename. The block format has no counterpart, the
title carries both name and ID. Implied by the format description,
but anyone converting such a line deserves to find it written down.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Aaron Andersen points out in the #492 discussion that the *Directory
settings carry more contract than create-and-chown: per-directory
modes, specific ownership rules, and cleanup toggles. Without them
config-dir was chowned to the service user, which systemd never does,
an existing directory with drifted ownership was left wrong, and the
runtime directory could not survive a restart.
Now matching systemd.exec(5), and where the man page is vague, the
code in setup_exec_directory():
- each directory takes a matching -mode key, octal with the leading
zero, default 0755. The mode of the named directory is locked
down again on every start, also when it already exists
- config-dir is created but never chowned
- the contents of an existing directory are left alone as long as
the owner is right; on drift everything under it is chowned back
- runtime-dir-preserve = no | restart | yes maps
RuntimeDirectoryPreserve=. A service still qualified to run when
the runtime directory would be removed is restarting, not
stopping, which is what svc_enabled() answers
The dir mechanics move to mksubsysd(), taking resolved ids, with
mksubsys() reduced to a name-resolving wrapper for the dbus plugin.
The child resolves uid/gid once for both directory setup and
privilege drop.
The symlink form, RuntimeDirectory=foo:bar, is not adopted.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
A service that drops privileges cannot create its own PID file in
/run, root owns it. Finit can create the file with pidfile-create,
but the daemon still cannot touch it to confirm a SIGHUP.
Five new settings, block format only: runtime-dir, state-dir,
cache-dir, logs-dir, and config-dir. The value is a directory name,
resolved under /run, /var/lib, /var/cache, /var/log, and /etc,
respectively. The directory is created before the service starts,
mode 0755 owned by user/group, and the full path is exported to the
process as RUNTIME_DIRECTORY, STATE_DIRECTORY, CACHE_DIRECTORY,
LOGS_DIRECTORY, and CONFIGURATION_DIRECTORY. Mode and ownership are
asserted at creation only, a daemon may tighten them afterwards.
The runtime directory is removed when the unit stops, after any
exec-stop-post script, like systemd with RuntimeDirectoryPreserve=no.
A completed run/task counts as stopped unless remain-after-exit keeps
it up. The other four persist across restarts.
These are the first settings with no legacy token: they are validated
by service_set_dir() and stored on the svc that service_register()
now returns. systemd accepts a list of directories per setting; this
is a single name for now, widening later is compatible since
libconfuse accepts a bare value for a list option.
The test sysroot gains libnss_files.so.2, which ldd cannot see, glibc
dlopen()s it. Without it getpwnam() fails inside the chroot, so
user/group settings never resolved and directory ownership could not
be tested.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The block format spells conditions as bare strings everywhere else, so
requiring `if = "<usr/foo>"` left one sigil behind, carried over from
the line-based `if:` token. A namespace separator already tells the two
apart: a value with a '/' is a condition, anything else is a service
name.
svc_ifthen() picks its mode from the start of the statement and applies
it to the whole, so a statement naming both kinds cannot be evaluated.
That is now an error, as are the old angle brackets, and either one
skips the block:
/etc/finit.conf: mixed: if: cannot mix a service name with a
condition in 'anchor,usr/enable-me', a statement must be all of
one kind, skipping
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Reference sections kept pointing at the line-based format they no longer
document. `sysv` and `task` sent the reader to Services for "<COND>",
the cgroups chapter opened by listing three legacy directives and then
explained further down that only two of them exist here, and the logging
chapter still gave "log:prio:facility.level,tag:ident" as the full
syntax.
Some claims were wrong independent of the format:
- a sysv is a supervised daemon, grouped with service in
SVC_TYPE_DAEMON, not a variation on task
- restart-max has no upper bound of 255, or any other
- the built-in rescue fallback runs in 12345789, not 12345
- conditional loading quotes system/10-hotplug.conf, not
system/hotplug.conf
- the key spells conflicts, not conflict
- the built-in getty no longer wants TERM last, it is a key
`if` takes either a service name or, in angle brackets, a condition,
decided in svc_ifthen(). Only the examples showed this, so it is now
said.
Terminology follows the split index.md already draws: a block is the new
format, a stanza the line-based one.
src/rescue.conf was still line-based, missed because it sits in src/
rather than system/ or contrib/.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
An ambient capability only reaches the effective set when euid is
non-zero, so a service that pairs `capabilities = { "^cap_..." }` with a
root user gets none of the restriction it asks for, and keeps the full
root set instead. Finit read the list, applied it, and said nothing. A
build without libcap dropped the list on the floor just as quietly.
Both now warn, naming the service:
nginx: ambient capabilities ('^') have no effect as root, use a
non-root user, or '%' and '!' entries
The ambient entries are read back from the parsed IAB value rather than
matched in the text, so inheritable ('%') and bounding ('!') entries stay
silent -- those work fine as root.
The warning repeats when the .conf files are re-read on runlevel change,
as parse warnings here already do.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The block conversion changed the bodies of the reference sections but
left every "**Syntax:**" header spelling the line-based format, so each
page opened by teaching the format it then stopped using. Six files
were missed entirely: runparts, files, capabilities, requirements,
runlevels, and switchroot.
runparts had no block spelling written down anywhere, though the parser
has read `runparts`, `runparts-progress`, and `runparts-sysv` all along.
tty gains a table per variant. Its three syntax lines carried nine
positional fields between them, which no longer describes anything the
parser accepts.
Fixes#148
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The syntax overview no longer describes a line-based format, since that
is not what the rest of the documentation shows. It now covers the
grammar, the two naming conventions, the nine aliases, and the leading
'-' on a path, and it says plainly that both formats are still read and
told apart per file by content. Without that, a reader with an
existing configuration is left wondering what happened to it.
service-opts.md was a list of modifiers to place between a directive
and its command, so it needed rewriting rather than translating: there
are no positions left to describe. It is now grouped by what the
settings do.
conditions.md needed correcting. It presented '!' as a condition
prefix alongside '~'. It is neither a condition nor a negation, it is
a flag on the block that means one thing on a service and another on a
run or task, so it is spelled reload-signal and required here, and the
page maps the old form to both.
Two things the pages claimed are not true. The kill delay range is
1-300, not 1-60, and stop and reload scripts are no longer run without
a timeout.
ChangeLog.md keeps its line-based examples. Those sit in historical
release entries, and rewriting them in a syntax that did not exist at
the time would misdate the format.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
* src/pid.c: note the stale-pidfile-cleanup exception to the
documented "Finit does not touch pid:! pidfiles" rule.
* doc/config/services.md: add a user-facing paragraph on the same.
* doc/ChangeLog.md: add Unreleased section covering this PR --
stale pidfile cleanup, restart log with signal name and core
dump flag, and the SIGUNKOWN typo fix.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Conditions in Finit are dependencies: if A is asserted, service B is
allowed to run. When A goes through FLUX (e.g., upstream reloads),
dependents are PAUSED and then simply resumed when the condition is
reasserted -- this is the correct behavior for barrier-style deps
like <pid/syslogd>.
However, some setups have tightly coupled services where dependents
must be reloaded/restarted when an upstream service reloads, not just
resumed. E.g., the FRR routing stack on Infix OS:
netd <pid/mgmtd> ← zebra <!pid/netd> ← {staticd,ripd} <!pid/zebra>
When netd reloads (SIGHUP), zebra and its dependents must be restarted
to pick up the new configuration.
The new '~' condition prefix marks a dependency as flux-sensitive:
service <!~pid/netd> name:zebra ...
When the upstream condition goes FLUX and returns to ON, the dependent
is reloaded (SIGHUP) or restarted (noreload '!') instead of merely
resumed. Transitivity follows naturally through the condition chain.
Closes#416Closes#476
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Similar to systemd's RemainAfterExit=yes. Prevents the task from
re-running on runlevel re-entry and ensures the post: script runs
when explicitly stopped or when leaving valid runlevels.
Useful for tasks that set up persistent state like firewall rules:
task [2345] remain:yes \
post:/usr/sbin/teardown-firewall \
/usr/sbin/setup-firewall -- Firewall setup
Not supported for bootstrap-only tasks (runlevel S only) since these
are deleted immediately after completion.
in an initramfs, then transition to the real root filesystem. Useful
for systems requiring early boot tasks like LUKS unlock, LVM activation,
or network boot before mounting the real root.
Adds INIT_CMD_SWITCH_ROOT API command, `initctl switch-root` subcommand,
and HOOK_SWITCH_ROOT plugin hook point. The implementation gracefully
stops services, moves virtual filesystems (/dev, /proc, /sys, /run) to
the new root, deletes initramfs contents to free memory, then execs the
new init as PID 1.
See GitHub Discussion #292 for background.
Some 'respawn' type services, like gettys, may hog the CPU in error
states if the service immediately exits. E.g., due to missing dev.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Implement supplementary group support for services, allowing them to
access resources owned by multiple groups. Uses the @user:group,sup1,sup2
syntax to explicitly specify supplementary groups, in addition to now
reading group membership from /etc/group.
Cgroups v2 limits are hierarchical - a process is constrained by the
most restrictive limit in its ancestor chain, not just its immediate
cgroup. This patch updates cg_conf() to walk up the hierarchy and
report effective limits by comparing values at each level.
This fixes incorrect "max" (unlimited) reporting in 'initctl --json
status', 'initctl cgroup', and 'initctl top' when child cgroups have
no explicit limits but parents do.
For memory.max and cpu.max: take minimum (most restrictive)
For memory.min: take maximum (most protection)
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>