Skip to content

Read spin file coverage from the file at ingest and require it before jobs run #1663

Description

@sapols

Problem

Spin tables are delivered in batches, on median every 2 days, with gaps of up to 8 days (2026-07-27 to 08-04 and 2026-09-10 to 09-18 since May). The last file in a batch ends partway through the delivery day. The indexer records only the day-of-year range from the filename, and the orchestration hands every overlapping spin file to a job without checking that they cover its window: get_spin_files_inputs has passed require_coverage=False since the Dagster port, even though verify_spin_coverage has existed since #1041.

So the arrival of a spin file wakes the partition of the day it ends in, and that partition fails in processing with Query times ... are outside of the spin data range. It regenerates on its own when the next delivery lands, because spin is a triggering input. For daily consumers that is one failed Batch run per delivery. Verified for MAG L1D on 2026-09-10: file imap_2026_252_2026_253_01.spin (ends 13:47 on the 10th) arrived 09-10 19:50 and triggered the failing run in IMAP-Science-Operations-Center/imap_processing#3471; the next file arrived 09-18 15:05 and the product was produced at 15:54. SWE L2 hit the same thing in #1155.

Jobs whose spin dependency is trigger_job: false (Hi L1B, Hi L1C, Ultra L1C) do not get the retry.

What a proper fix needs

Worked out in #1657 and its review:

  • Index real coverage. Open the spin file at ingest and record the first row's spin_start_utc and the last row's spin_start_utc plus its spin_period_sec. The spin_files columns are already DateTime, so no schema change. Filename dates cannot serve pointing partitions: spin files do not follow pointing boundaries (imap_2026_252_2026_253_01.spin runs 09-09 21:23 to 09-10 13:47, pointing 367 runs 09-09 10:03 to 09-10 10:03), so a day-granular check would make every pointing job wait one delivery longer than it needs.
  • Require coverage for required spin dependencies, at spin precision, with a seam tolerance: adjacent files overlap or gap by up to 20 ms, and a real gap is one spin, about 15 s.
  • Fix verify_spin_coverage to track the running end of coverage. The current pairwise loop misreads a longer file that spans shorter ones; the production case is 2026-07-24 (imap_2026_204_2026_206_02.spin next to single-day pieces), the only day in the table it gets wrong, and with coverage required it would skip forever.
  • Pick the latest version by filename, before the overlap filter. Versions of a file differ in real coverage by milliseconds (imap_2025_316_2025_317_01 and _02 start 20 ms apart), so grouping on (start_date, end_date) no longer collapses them.
  • Let non-triggering inputs unblock partitions that never produced output. A run that skips for missing dependencies ends in SUCCESS, and the kickoff sensor treats any successful run as done. The narrowest change is "done = successful run and an output materialized", which retries skips and leaves failed-with-partial-output behavior as it is today.
  • Backfill the existing rows (669 today) before enforcement, in each environment. Re-indexing must keep ingestion_date so sensors do not treat refreshed rows as new files.
  • Drop the midnight floor on the partition start in get_spin_files_inputs, a workaround for day-granular dates.

Starting points

Reviewed drafts with tests, not merged:

#1657 is being closed in favor of this issue.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

  • Status
    Todo

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions