Files
Joachim Wiberg e5040385f5 test: fix flaky crashing.sh
Two races, both around the post:script Finit runs when it gives up on a
service.

Finit marks the service crashed before it forks post:script, so the file
the test greps for lands a moment later.  Waiting for the state is not
enough.

slay then takes the PID from initctl status, and service_post_script()
sets svc->pid to the script's PID, so once Finit has given up the PID
reported for the service is the post:script.  Killing that takes out the
script instead of the service and the file never arrives at all.  The
guard for this was already there, with a comment describing it, but
inside the loop that waits for a PID to appear, so it only covered the
case where there was none.  A PID that was already there went straight
to kill -9.  Hence the gcc leg killing once more at lap 13, after Finit
had stopped restarting, where clang stopped at 12.

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2026-08-22 10:25:28 +02:00

73 lines
1.7 KiB
Bash
Executable File

#!/bin/sh
nm=$1
getpid()
{
initctl status "$1" |awk '/PID :/{ print $3; }'
}
pid=$(getpid "$nm")
if [ -f oldpid ]; then
oldpid=$(cat oldpid)
if [ "$oldpid" = "$pid" ]; then
echo "Looks bad, wait for it ..."
sleep 3
pid=$(getpid "$nm")
if [ "$oldpid" = "$pid" ]; then
echo "Finit did not deregister old PID $oldpid vs $pid"
initctl status "$nm"
ps
echo "Reloading finit ..."
initctl reload
sleep 1
initctl status "$nm"
exit 1
fi
fi
fi
timeout=50
while [ "$pid" -le 1 ]; do
sleep 0.1
# Once Finit gives up on a service it forks the post:script and
# reports that PID as the service's, so a slay still waiting for
# the service to come back would kill the script instead.
st=$(initctl -p status "$nm" 2>/dev/null | awk '/Status/{print $3}')
if [ "$st" = "crashed" ]; then
exit 1
fi
pid=$(getpid "$nm")
if [ "$pid" -le 1 ]; then
timeout=$((timeout - 1))
if [ "$timeout" -gt 0 ]; then
continue
fi
if [ $pid -ne 0 ]; then
echo "Got a bad PID: $pid, aborting ..."
ps
sleep 1
initctl status "$nm"
ps
fi
exit 1
fi
break
done
echo "$pid" > /tmp/oldpid
# The check above only runs while waiting for a PID to appear. Finit
# reports the post:script's PID as the service's once it has given up
# restarting, so a PID that was already there needs the same guard:
# killing it would take out the script instead of the service, and the
# side effects the test is waiting for never happen.
st=$(initctl -p status "$nm" 2>/dev/null | awk '/Status/{print $3}')
if [ "$st" = "crashed" ]; then
exit 1
fi
#echo "PID $pid, kill -9 ..."
kill -9 "$pid" 2>/dev/null || true