When a SysV init script starts a daemon, Finit knows nothing of the PID
it should monitor. The PID is written, by start-stop-daemon or the
daemon itself, to the PID file. Finit monitors for new PID files and
can match the PID in such files with an svc_t.
For the regular use-case, we prefer first looking up the matching svc_t
based on the PID -- assuming we start and monitor the service. As a
fallback we resort to mathching the svc_t's declared PID file with the
new file we just discovered.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
We want to find the PID of the daemon the init script starts, so we need
a way to declare this to Finit.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
SysV init scripts should not go to "done" state but "halted" so we can
do: `initctl stop foo; initctl start foo`, like we do for our regular
monitored services.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
When running in a container we still want to use any syslog daemon
available for our logging needs. However, the time between the first
logit() in Finit and any such daemon having started can be long. In a
normal (non-containerized) setup we log to the kernel ring buffer, but
that's not available in a container scenario. At least not for
unprivileged containers. So we need to detect all these cases and be
prepared to fall back to log to the console, either using these LOG_CONS
flag to openlog(), or by simply calling vfprintf() to stderr.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
This trick can also be used by others who want to run Finit in an
unshare. Set the container environment variable to 'unshare',
like lxc and docker do for their products.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Turns out the kill(2) syscall returns ENOENT, not ESRCH, in our test
suite. Don't know why, the man page never mentions ENOENT, only the
ESRCH code. Let's check for both, either way ithe PID is not there.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
A service may have unexpectedly died, and we never got the signal, so
when stopping services we must set the new state after we've tried to
stop the service. Otherwise the svc_set_state() function starts a
background timer for the SIGKILL job, which may block a reboot.
The kill() syscall tells us if the service was there or not, if not we
must clean up and go to HALTED state.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
At startup (and reconf) of systems with lots of services there is a risk
of losing inotify events, e.g., PID file creation/delete events. This
patch increase the receive buffer (doubles it).
On Linux the getsockopt() for SO_RCVBUF returns double the set size, due
to housekeeping in the kernel. So we don't have to do any adjustments
when setting it.
Issue #226
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
In some (error) cases the PID known to Finit may no longer exist, or may
not have been added to a cgroup (yet). Handle this case by skipping the
output of cgroup info in such conditions.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
On, e.g., a container based system logs may not be available so check
also that the messages fallback exists instead of causing confusing
error in the status output for the service.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
To silence warnings at startup/shutdown, check for existance of swapon
and swapoff before calling them.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
When a run task is started with svc->started = 1, it should be
considered started successfully or failed on the other hand.
Signed-off-by: Sergio Morlans <sergio.morlans@atlascopco.com>
Signed-off-by: Ming Liu <liu.ming50@gmail.com>
This patch adds support for optional logging of output from all run()
commands. For run_interactive() we've opted to log instead of just
redirect, meaning output on error is till on console but also in log.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
We've been discussing, over the years, that we'd like to have an easier
way to debug bringup with Finit. Network bringup is one such case where
it's hard to debug if/what you've misspelled in /etc/network/interfaces
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
This fixes a seemingly long-running bug; we only brought down networking
in shutdown/halt and for runlevel 1, single-user mode -- not reboot.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
In commit f0f358a13:
[ service.c: set/clear condition 'done' for run tasks ]
a runtask done condition would be set/cleared when entering DONE/HALTED
states, but it did not cover all the user cases, for instance, sometimes
an end user may want to know if a runtask has finished sucessfully or
to decide what to do on its failures.
So we now change the conditions to: tsktype/tskname/success and
tsktype/tskname/failure.
And this change not only applies to run/task types, but also applies to
sysv type, in case it fails, a sysv/tskname/failure condition would be
set.
Signed-off-by: Ming Liu <liu.ming50@gmail.com>
Refactor new kill/shutdown implementation from 3e0063e to fix the
regression in compiler output:
sig.c: In function ‘kill_callback’:
sig.c:222:12: warning: cast from pointer to integer of different size [-Wpointer-to-int-cast]
222 | kill(pid, (int)context);
| ^
Also, some minor renames and simplificactions to match project style.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
- Check all pointers
- Declaratons always at top of func/scope
- Use established variable nomenclature
- Skip useless if() stmt
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
- Comments preferably at beginning of func/sect
- Reorder code slightly, add whitespace for readability
- Drop useless comment
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The implementation looks up the named service by using
`svc_parse_jobstr`. The callbacks for `svc_parse_jobstr` has been
augmented to accept a user data parameter. For this use case,
a carrier for the actual signal was needed. The address of the
signal parameter is taken and passed on as a `void *`. The
callback then simply deferences it as an int - the signal number.
Blocking SIGTERM means Finit will wait another two seconds before it
sends SIGKILL, which casues a delayed [WARN] Killing ... before we can
proceed to shutdown/reboot -- not OK. Don't know what I was smoking
back in 2017, but it must've been good.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
When running in a container we might not have the necessary privileges
to mount cgroups. Don't cause error in this case, just log it.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
A regression was introduced by commit f0f358a1:
[ service.c: set/clear condition 'done' for run tasks ]
svc->type is a integer but mistakenly being used as string, which will
cause crash.
Signed-off-by: Ming Liu <liu.ming50@gmail.com>
Instead of unconditionally waiting 2 seconds for processes to die,
check continuously for remaining processes, and break the loop when
none remain.
Turn PID 1 to a RT process with highest priority 99 during shutdown,
this ensures it would not be preempted by other RT processes.
Signed-off-by: Robert Andersson <robert.m.andersson@atlascopco.com>
Signed-off-by: Mathias Thore <mathias.thore@atlascopco.com>
Signed-off-by: Ming Liu <liu.ming50@gmail.com>