What
In a full pytest tests/ -m tier_a run, the process dies at
tests/test_0053_hang_watchdog.py::test_report_carries_the_main_thread_stack with no pytest summary. macOS crash reports show the same signature every time:
EXC_BAD_ACCESS (SIGSEGV) KERN_INVALID_ADDRESS
thread 0 (main):
python dump_traceback
python _Py_DumpTracebackThreads
python faulthandler_user
libsystem_platform.dylib _sigtramp
libsystem_c.dylib nanosleep
python time_sleep
... (pytest frames)
other threads: PMIx progress_engine (libevent kevent), a Python lock wait
The test arms uw.mpi.watch(seconds=0.3) and sleeps. The watchdog fires faulthandler from a signal handler, and CPython's asynchronous traceback walker reads an invalid address while walking the main thread's frames. No Underworld code is on the faulting stack. Python 3.12.12 (conda-forge), macOS.
When it happens
| build |
full tier_a runs |
crashed here |
development @ 1b4f3b5 |
1 |
0 |
| bugfix/picard-warmup-791, first commit |
1 |
0 |
| bugfix/picard-warmup-791, revised commit (520b7d1) |
3 |
3 |
It does NOT reproduce, on either build, when:
test_0053 runs alone (14/14 pass);
- the full suite is collected but only
-k hang_watchdog runs;
- exactly the preceding set runs (
tests/analytic_full tests/parallel tests/test_00[0-5]*.py -m tier_a, 453 passed on both builds).
So the trigger is state that builds up over a full-collection run, and it is sensitive to process layout rather than to any code path in the change that exposed it. #791's diff touches only the Stokes solve warm-up and continuation stages; nothing in it runs in test_0053 or touches signals, threads or faulthandler.
Why it matters
- A hard crash kills the run with no summary. A background runner can then report success: in our case the shell exit code of a trailing
tail was 0.
- The test drives faulthandler from a timer signal while other threads (PMIx progress engine) are live. Async traceback dumping is known to be fragile in exactly that setting, so this can surface on unrelated changes.
Suggestions
- Run
test_0053 in a subprocess (e.g. pytest-forked / --forked, or subprocess + a small driver). The watchdog's faulthandler dump then runs in a fresh process with no accumulated suite state, and a crash fails one test instead of killing the run.
- Separately worth confirming: whether any earlier test leaves a
faulthandler.register / dump_traceback_later or an interval timer armed (the watchdog restores a pre-existing SIGALRM handler, per its own tests, but stale faulthandler registrations are not covered).
What
In a full
pytest tests/ -m tier_arun, the process dies attests/test_0053_hang_watchdog.py::test_report_carries_the_main_thread_stackwith no pytest summary. macOS crash reports show the same signature every time:The test arms
uw.mpi.watch(seconds=0.3)and sleeps. The watchdog fires faulthandler from a signal handler, and CPython's asynchronous traceback walker reads an invalid address while walking the main thread's frames. No Underworld code is on the faulting stack. Python 3.12.12 (conda-forge), macOS.When it happens
development@ 1b4f3b5It does NOT reproduce, on either build, when:
test_0053runs alone (14/14 pass);-k hang_watchdogruns;tests/analytic_full tests/parallel tests/test_00[0-5]*.py -m tier_a, 453 passed on both builds).So the trigger is state that builds up over a full-collection run, and it is sensitive to process layout rather than to any code path in the change that exposed it. #791's diff touches only the Stokes solve warm-up and continuation stages; nothing in it runs in
test_0053or touches signals, threads or faulthandler.Why it matters
tailwas 0.Suggestions
test_0053in a subprocess (e.g.pytest-forked/--forked, orsubprocess+ a small driver). The watchdog's faulthandler dump then runs in a fresh process with no accumulated suite state, and a crash fails one test instead of killing the run.faulthandler.register/dump_traceback_lateror an interval timer armed (the watchdog restores a pre-existing SIGALRM handler, per its own tests, but stale faulthandler registrations are not covered).