Skip to content

fix(text): aligned tables, cover-only form detection, 8-Ks read with their exhibits - #96

Merged
jfrench9 merged 4 commits into
mainfrom
bugfix/eight-k-exhibits-and-table-columns
Oct 4, 2026
Merged

jfrench9 merged 4 commits into
mainfrom
bugfix/eight-k-exhibits-and-table-columns

Conversation

@jfrench9

@jfrench9 jfrench9 commented Oct 4, 2026 •

Copy link
Copy Markdown
Member

Summary

Comparing the two servers on Workiva's latest earnings 8-K showed three problems in xbrlkit's text tools. search_text on the 8-K found nothing in the release, because an 8-K loaded as its cover page alone. The financial tables put values in the wrong column. And the release, loaded by its own URL, came back labelled a 10-K. This PR fixes all three. The table fix also reaches the RoboSystems sec server, which renders filing text through xbrlkit.

Changes

text/tables.py (table → markdown)

  • When a $, ( or ($ merged into its number, the cell it sat in was dropped, so every value after it in the row moved one column left. The same happened when a trailing ) or % was attached. Cells now keep their positions. Each column a symbol stood in is joined to the next column when no row has values in both. That also lines up Workiva's layout, where a value without a $ spans the symbol and number columns.
  • A self-closing <td ... /> (Apple's filing agent writes spacers this way) was read as an open tag and swallowed the cell after it. It now reads as an empty cell.

serve/session.py (cover form)

  • _form_on_the_cover searched the first 400,000 characters for "Form 10-K", so a press release's safe-harbor citation made it a 10-K. Only the top 2,000 characters are searched now. Real covers name their form within about 200.

serve/ (8-K exhibits)

  • An 8-K or 8-K/A loaded from EDGAR, on both the XBRL path and the document-only path, now fetches its EX-99 exhibits through other_documents / read_other_document. They are joined into one text: ## Form 8-K, then ## EX-99.1 and so on, each a section with its offset. This is the layout the RoboSystems sec server reads an 8-K in. If the exhibits can't be reached, the 8-K loads on its own as before.
  • join_texts and FORM_SECTION_ID are added to the declared xbrlkit.serve surface, so a host that builds the text itself can use the same assembly.
  • The items_note, next steps, server instructions, the documents tool description and serve/README.md now point to read_text / search_text when the exhibits are in the text and the text tools read it whole. Under pure, or when only the tagged text is read, they still point to documents, because there the exhibits aren't in the text the tools search.
  • If one exhibit can't be read, the others still go into the text. Two exhibits both typed plain EX-99 each get their own section id.

Output Impact

CHANGED OUTPUT in the text tools (search_text, read_text, describe_filing sections), not in the serializations. Holon, TAVI and model.json don't go through the table renderer.

  • Tables: every filing's markdown tables change shape. Spacer columns are gone and values sit under their headers. Character offsets into the text change with them, so an offset saved from an earlier version no longer points at the same place.
  • 8-Ks from EDGAR: the text now includes the EX-99 exhibits, and sections.items gains form_8k / ex_99_1. A load costs the index fetch plus one fetch per exhibit.
  • Documents loaded by URL or path: a document whose only "Form X" mention is below its first 2,000 characters now reports form: null instead of a guess.
  • RoboSystems: it renders through build_text, so it picks up the table change when its pin moves. Text it has already cached or indexed keeps the old rendering until rebuilt.

Testing

  • The pre-commit hook ran ruff, basedpyright and the suite on each commit; the last run had 696 passed, 2 skipped. New tests cover the $ position shift, the colspan-2 layout, a symbol column that holds real values, self-closing cells, cover vs. citation, join_texts, an 8-K read with its release, an 8-K whose exhibits can't be reached, one exhibit failing while the others load, the exhibit guidance in pure and tagged-text modes, and a 10-K that doesn't reach for exhibits.

  • Tables, old vs. new renderer: five documents, counting rows with values outside the columns most rows use. The total fell from 1,574 to 329:

    • Workiva Q2 2026 EX-99.1 (0001445305-26-000059): 47 → 0
    • Apple 10-K (0000320193-25-000079): 79 → 5
    • Microsoft 10-K (0001193125-26-323660): 270 → 27
    • NVIDIA 10-K (0001045810-26-000021): 111 → 11
    • JPMorgan 10-K (0001628280-26-008131): 1,067 → 286

    I read the six tables this count flagged as worse; each is now correct. For example, Apple's buyback total sits under "Approximate Dollar Value", and JPMorgan's loan maturities sit under their buckets.

  • Cover form: the first "Form X" match is at 125–215 characters in the four 10-Ks and the Workiva 8-K, and at 13,455 in the Workiva release.

  • 8-Ks live from EDGAR:

    • WK 8-K (0001445305-26-000059, inline XBRL cover): EX-99.1 is read in, and "free cash flow|guidance" returns 20 hits, all labelled EX-99.1.
    • A 2018 Workiva 8-K with no XBRL (0001445305-18-000131, document-only path): EX-99.1 is read in, with 7 hits.

🤖 Generated with Claude Code

A "$" merged into its number dropped the symbol's cell, so every value
after it sat one column left of the rows without one; a value spanning the
symbol and number columns (colspan 2) landed apart from one beside a "$".
A self-closing <td /> spacer swallowed the cell after it. Merges now keep
each cell's position, a symbol column joins the number's where no row fills
both, and a self-closing cell reads as empty.
A document opened on its own took the first "Form 10-K" anywhere in its
first 400,000 characters as its form, so an earnings release whose safe
harbor cites the annual report loaded as a 10-K. A cover names its form in
its first few hundred characters; only those are read now.
An 8-K loaded from EDGAR was its cover page alone, so search_text found
nothing the release says and told the caller to reword. Its EX-99 exhibits
are now read into the same text, each a section under its type with its
offset — the layout the RoboSystems sec server reads an 8-K in, so the two
answer alike. join_texts is the assembly, exported for a host that builds
the text itself. The items note and next steps point at the text when the
exhibits are in it; an 8-K whose exhibits cannot be reached loads alone.
…read

Under pure, or reading the tagged text alone, the exhibits are not in the
text the tools search and sections.items is not shown, yet the note and next
steps sent the caller there; they now say documents in those modes. One
exhibit out of reach no longer drops the rest, and two exhibits typed plain
EX-99 each keep a section id of their own.
@jfrench9
jfrench9 merged commit 3ba12d3 into main Oct 4, 2026
4 checks passed
@jfrench9
jfrench9 deleted the bugfix/eight-k-exhibits-and-table-columns branch October 4, 2026 22:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant