Files
Joachim Wiberg 61b0e0f3e6 Fix #420: run services inside a PAM session
Apply a PAM session to run/task/sysv/services Finit starts, pam_limits
above all, so a service running as a given user picks up that user's
limits the way a login does.

Add a new `pam` setting for the new block format (only), like the
per-service directories, naming a file in /etc/pam.d:

    service weston {
        user    = "weston"
        pam     = "weston-autologin"
        command = "/usr/bin/weston --continue-without-input"
    }

pam_close_session() has to be called by a process still holding the
handle, and the handle does not survive exec().  Hence the keeper: it
holds the handle, drops to the service's credentials, and waits for a
parent-death signal before closing the session.  Same shape as
systemd's (sd-pam), for the same reason, and one per fork, so the
script hooks open and close their own.

The keeper closes the descriptors it inherited from Finit and only
those.  Closing everything would also take out what pam_open_session()
opened for itself, a keyring fd or a lock file, and leave the modules
to close a session with those pulled out from under them.  Closing
nothing, as (sd-pam) does, would leave it holding the write end of the
notify pipe for the service's whole lifetime and starve notify = "s6"
services of their ready signal.  So the fds open before pam_start()
are snapshotted and exactly those are closed, while the ones PAM opens
after are marked close-on-exec so the daemon does not inherit them
either.

A refused value, a denied account stack, an uninstalled pam.d file,
and a build without PAM support all keep the service from starting
rather than running it with the stacks skipped: one that quietly loses
pam_limits and its private /tmp, with nothing said.  Capabilities a
module like pam_cap.so granted are merged into the IAB Finit applies
instead of being replaced by it, which only helps a service that also
sets capabilities, the other arm being a plain setuid() with nothing
left to restore once permitted is empty.

The test sysroot gains pam_permit.so, pam_deny.so and pam_limits.so,
which ldd cannot see, libpam dlopen()s them, and the test skips when
the host has none to stage.  The negative cases pin the exit status
rather than only asserting crashed, which serv reports for any early
exit, so a bad command or an unwritable pidfile cannot pass for a
rejected session.

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2026-09-23 16:30:14 +02:00

204 lines
7.0 KiB
Markdown

Runlevels
=========
Finit supports runlevels, but unlike other init systems runlevels are
declared per service/run/task/sysv command. When booting up a system
Finit pass through three phases:
1. Setting up the console, parsing any command line options, and other
housekeeping tasks like mounting all filesystems, and calling `fsck`
2. Starting all run/task/services in runlevel S, then waiting for all
services to have started, and all run/tasks to have completed
3. Go to runlevel 2, or whatever the user has set in the configuration
Available runlevels:
- ` S`: bootStrap
- ` 1`: Single user mode
- `2-5`: traditional multi-user mode
- ` 6`: reboot
- `7-9`: multi-user mode (extra)
- ` 0`: shutdown
Runlevel S (bootStrap), is for tasks supposed to run once at boot, and
services like `syslogd`, which need to start early and run throughout
the lifetime of your system.
Example:
task console-setup {
runlevel = "S"
command = "/lib/console-setup/console-setup.sh"
}
service rsyslogd {
runlevel = "S12345"
envfile = "-/etc/default/rsyslog"
command = "rsyslogd -n $RSYSLOGD_ARGS"
}
When bootstrap has completed, Finit moves to runlevel 2. This can be
changed in `/etc/finit.conf` with `runlevel = N`, or by a script running
in runlevel S that calls, e.g., `initctl runlevel 9`.
The latter is useful if startup scripts detect problems outside of
Finit's control, e.g., critical services/devices missing or hardware
problems.
Each runlevel must be allowed to "complete". Meaning, all services in
runlevel S must have started and all run/tasks have been started and
collected (exited). Finit waits 120 seconds for all run/tasks in S to
complete before proceeding to 2.
Finit first stops everything that is not allowed to run in 2, and then
brings up networking. Networking is expected to be available in all
runlevels except: S, 1 (single user level), 6, and 0. Networking is
enabled either by `network = "script"`, or if you have an
`/etc/network/interfaces` file, Finit calls `ifup -a` -- at the very
least the loopback interface is brought up.
> [!NOTE]
> When moving from runlevel S to 2, all run/task/services that were
> constrained to runlevel S only are dropped from bookkeeping. So when
> reaching the prompt, `initctl` will not show these run/tasks. This is
> a safety mechanism to prevent bootstrap-only tasks from accidentally
> being run again. E.g., `console-setup.sh` above.
Runlevel Configuration
----------------------
**Syntax:** `runlevel = N`
The system runlevel to go to after bootstrap (S) has completed. `N` is
the runlevel number 0-9, where 6 is reserved for reboot and 0 for halt.
Completed in this context means all services have been started and all
run/tasks have been started and collected.
It is recommended to keep runlevel 1 as single-user mode, because
Finit disables networking in this mode.
*Default:* 2
> [!NOTE]
> Only read and executed in runlevel S (bootstrap).
Networking
----------
**Syntax:** `network = "PATH"`
Script or program to bring up networking, with optional arguments.
Deprecated. We recommend using dedicated task/run blocks per runlevel,
or `/etc/network/interfaces` if you have a system with `ifupdown`, like
Debian, Ubuntu, Linux Mint, or an embedded BusyBox system.
> [!NOTE]
> Only read and executed in runlevel S (bootstrap).
System Hostname
---------------
**Syntax:** `hostname = "NAME"`
Set system hostname to NAME, unless `/etc/hostname` exists in which case
the contents of that file is used.
Deprecated. We recommend using `/etc/hostname` instead.
> [!NOTE]
> Only read and executed in runlevel S (bootstrap).
Kernel Modules
--------------
**Syntax:** `modules = { "MODULE [ARGS]", ... }`, alias `mod`
Load kernel modules, each with optional arguments. Similar to the
`insmod` command line tool.
modules = { "button", "evdev", "softdog" }
> [!NOTE]
> A list cannot hold comments; the lexer reads the entries after a `#`
> regardless. Put commented-out candidates above the list.
Deprecated, there is both a `modules-load.so` and a `modprobe.so` plugin
that can handle module loading better. The former supports loading from
`/etc/modules-load.d/`, the latter uses kernel modinfo to automatically
load (or coldplug) every required module. Hotplug module loading is
handled by [keventd](../keventd.md), the built-in device manager. On
systems using the hotplug plugin with BusyBox mdev instead, add to
`/etc/mdev.conf`:
$MODALIAS=.* root:root 0660 @modprobe -b "$MODALIAS"
> [!NOTE]
> Only read and executed in runlevel S (bootstrap).
Resource Limits
---------------
**Syntax:** `rlimit { RESOURCE = LIMIT }`, with `soft.` or `hard.` prefix
Set the hard or soft limit for a resource, or both if the prefix is
omitted. `RESOURCE` is the lower-case `RLIMIT_` string constants from
`setrlimit(2)`, without prefix. E.g. to set `RLIMIT_CPU`, use `cpu`.
LIMIT is an integer that depends on the resource being modified, see
[setrlimit(2)](https://man7.org/linux/man-pages/man2/setrlimit.2.html),
or the kernel `/proc/PID/limits` file, for details. Finit versions
before v3.1 used `infinity` for `unlimited`, which is still supported,
albeit deprecated.
rlimit {
hard.as = 8388608 # no more than 8MB of address space
soft.core = unlimited # core dumps may be arbitrarily large
cpu = 10 # soft & hard = 10 sec
}
`rlimit` can be set globally, in `/etc/finit.conf`, or locally per
each `/etc/finit.d/*.conf` read. I.e., a set of task/run/service
blocks can share the same rlimits if they are in the same .conf.
For a service with [`pam`](pam.md), `pam_limits` runs after these
limits are applied, so `limits.conf` has the last word.
Miscellaneous Settings
----------------------
**Syntax:** `reboot-delay = 0-60`
Optional delay at reboot (or shutdown or halt) to allow kernel
filesystem threads to complete after calling `sync(2)` before
rebooting. This applies primarily to filesystems that do not
have a reboot notifier implemented. At the point of writing,
the only known filesystems affected are: ubifs, jffs2.
*Default:* 0 (disabled)
When enabled (non-zero), this delay runs after file systems have been
unmounted and the root filesystem has been remounted read-only, and
sync(2) has been called, twice.
> "On Linux, sync is only guaranteed to schedule the dirty blocks for
> writing; it can actually take a short time before all the blocks are
> finally written.
**Syntax:** `reboot-watchdog = true|false`
Controls whether the system should reboot via the watchdog timer (WDT)
or directly via the SoC/kernel. When enabled, Finit will:
1. Send `SIGPWR` to the registered watchdog daemon before shutdown
2. Send `SIGTERM` to the watchdog daemon and wait up to 10 seconds
for the watchdog to trigger a hardware reset
When disabled (default), Finit skips the watchdog reboot logic and
calls the kernel's `reboot(2)` syscall directly for a clean SoC reboot.
*Default:* off (reboot via SoC)
> [!NOTE]
> This setting only affects reboots. The watchdog daemon will still
> run and monitor the system during normal operation.