Skip to content

test_0053 hang_watchdog: SIGSEGV inside CPython faulthandler dump_traceback during full tier_a runs (state-dependent) #793

Description

@lmoresi

What

In a full pytest tests/ -m tier_a run, the process dies at
tests/test_0053_hang_watchdog.py::test_report_carries_the_main_thread_stack with no pytest summary. macOS crash reports show the same signature every time:

EXC_BAD_ACCESS (SIGSEGV)  KERN_INVALID_ADDRESS
thread 0 (main):
  python  dump_traceback
  python  _Py_DumpTracebackThreads
  python  faulthandler_user
  libsystem_platform.dylib  _sigtramp
  libsystem_c.dylib  nanosleep
  python  time_sleep
  ... (pytest frames)
other threads: PMIx progress_engine (libevent kevent), a Python lock wait

The test arms uw.mpi.watch(seconds=0.3) and sleeps. The watchdog fires faulthandler from a signal handler, and CPython's asynchronous traceback walker reads an invalid address while walking the main thread's frames. No Underworld code is on the faulting stack. Python 3.12.12 (conda-forge), macOS.

When it happens

build full tier_a runs crashed here
development @ 1b4f3b5 1 0
bugfix/picard-warmup-791, first commit 1 0
bugfix/picard-warmup-791, revised commit (520b7d1) 3 3

It does NOT reproduce, on either build, when:

  • test_0053 runs alone (14/14 pass);
  • the full suite is collected but only -k hang_watchdog runs;
  • exactly the preceding set runs (tests/analytic_full tests/parallel tests/test_00[0-5]*.py -m tier_a, 453 passed on both builds).

So the trigger is state that builds up over a full-collection run, and it is sensitive to process layout rather than to any code path in the change that exposed it. #791's diff touches only the Stokes solve warm-up and continuation stages; nothing in it runs in test_0053 or touches signals, threads or faulthandler.

Why it matters

  • A hard crash kills the run with no summary. A background runner can then report success: in our case the shell exit code of a trailing tail was 0.
  • The test drives faulthandler from a timer signal while other threads (PMIx progress engine) are live. Async traceback dumping is known to be fragile in exactly that setting, so this can surface on unrelated changes.

Suggestions

  • Run test_0053 in a subprocess (e.g. pytest-forked / --forked, or subprocess + a small driver). The watchdog's faulthandler dump then runs in a fresh process with no accumulated suite state, and a crash fails one test instead of killing the run.
  • Separately worth confirming: whether any earlier test leaves a faulthandler.register / dump_traceback_later or an interval timer armed (the watchdog restores a pre-existing SIGALRM handler, per its own tests, but stale faulthandler registrations are not covered).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions