Skip to content

Prepare filing PDFs, analyze originals, and require previews - #272

Merged
nonprofittechy merged 8 commits into
mainfrom
feature/document-preparation-preview
Sep 30, 2026
Merged

nonprofittechy merged 8 commits into
mainfrom
feature/document-preparation-preview

Conversation

@nonprofittechy

@nonprofittechy nonprofittechy commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Uploading an interactive PDF or Word document now produces a filing PDF before case review. Filers see the actual PDF.js preview, can download the preserved original, and are asked to check each filing PDF before pressing Continue. PDF preview has no confirmation checkboxes. Additional previews are available on Upload, Confirm filing, Check documents, Organize documents, Fees and interview handoff screens.

Closes #113.

Preparation behavior

  • Use Gotenberg for DOC/DOCX conversion and flattening, without adding PDFtk to the runtime. Request tagged PDF output and lossless images for Word conversion.
  • Keep static PDFs byte for byte. Repair missing or stale text appearance streams using the existing pypdf dependency before flattening; preserve usable existing appearances.
  • Retain private originals separately. Preview and submission use the filing PDF. Analysis reads the original PDF with its stored form values, or locally extracted DOCX text from docx2python; binary DOC uses the converted PDF. Preserve recovery and review metadata through legacy document rebuilds and clean up both copies when appropriate.
  • Reject unreadable/encrypted/XFA inputs, certificate signatures that flattening would invalidate, lost filled text, changed page counts, remaining form fields, oversized responses and service failures with recovery instructions.
  • Add a draft-scoped, authenticated, uncached content endpoint and self-host the pinned PDF.js legacy build, worker, fonts and character maps. Generate those assets during npm ci and in the Docker build.
  • Record preview-step completion on Continue and tie it to the current document list and storage keys. A later upload requires another preview; direct submission cannot bypass it.
  • Add jurisdiction-aware preparation guidance and policy. Vermont explicitly enables flattening and links to its verified File & Serve instructions.

Regression hardening

  • Reset original/preparation/approval metadata whenever a legacy lead key changes. Prepare legacy stored PDF/Word attachments before Continue and enforce the preview step on direct submission.
  • Resolve indirect PDF annotation arrays. Apply literal text checks only to ordinary visible text fields, preserving valid comb/password/formatted/hidden appearances.
  • Return 422 for permanent handoff preparation failures and 503 for temporary service/storage failures. Word handoffs use the same pipeline as app uploads.
  • Queue extraction after commit; clean staged uploads on database and storage exceptions; remove replaced blobs after commit while preserving cross-draft references.
  • Add regression coverage for those cases and make downstream-flow fixtures explicitly represent files that completed the preview step.

Preview interaction and copy

  • Remove PDF preview checkboxes. Ask filers to check each PDF, then Continue records the step for the current files. Preserve stale-file and preparation checks. The existing case-details review checkbox remains.
  • Shorten the upload guidance, preview instructions, download labels and errors. Remove repeated filenames and conversion details; use short sentences and plain words.
  • Re-run the real browser flow, refresh desktop/mobile screenshots, and verify Continue works without preview checkboxes. The focused preparation/preview/opt-out/copy suite passed 89 tests; the browser run had zero Axe violations and page errors.

Original document analysis

  • Analyze the unflattened PDF and supply stored form answers alongside its original bytes. Preserve the AcroForm when limiting pages and exclude fields from omitted pages.
  • Extract DOCX text locally with docx2python, including tables, Unicode, headers and notes, then use text evidence extraction. Retain AI opt-out for every format.
  • Bound Word text with DOCUMENT_EXTRACTION_MAX_TEXT_CHARS (100,000 by default), identify truncation on review, and prevent source replacement from committing stale analysis.
  • A real synthetic DOCX analysis smoke test extracted its docket correctly in 15.52 seconds through the configured AI and live taxonomy service. Local synthetic text extraction took a median 2.85 ms over 20 runs. These are smoke timings, not a controlled PDF comparison.

Engine comparison

The corpus scan found 11 distinct filled PDFs / 22 pages in six local Docassemble repositories. LITEFile and raw Gotenberg accepted 11/11; PDFtk Java 3.3.3 accepted 10/11. Raster inspection exposed a limitation that text extraction missed: raw Gotenberg/QPDF renders multiline answers on one line when appearance streams are missing. The pypdf repair preserves their line breaks before Gotenberg flattens them. A Unicode stress case that could not be safely flattened is rejected with a recovery path.

Also filled the official Vermont Small Claims Answer with synthetic names, a checked checkbox, a three-line answer and typed signature. Checked both prepared pages visually, then checked a four-page packet containing that answer and synthetic exhibits. No court submission was made. These results support the engine choice for this tested corpus; previews remain necessary for font/layout/accessibility changes.

Validation

  • 1,240 backend tests passed, with the opt-in browser test skipped in the ordinary suite; 64 JavaScript tests passed.
  • 76 focused original-source, extraction, worker-claim, opt-out and preview tests passed, including nine new regression cases.
  • Separate real Gotenberg + Django live server + Chromium test passed: multi-file PDF/DOCX upload, real PDF.js rendering, page navigation, zoom, Continue without preview checkboxes, organize and fees previews, mobile, invalid upload, preview failure and successful retry.
  • GitHub accessibility workflow passed. Mocked conversion tests now configure their service explicitly; Conversion/preview tests and the full suite pass with Gotenberg environment variables cleared.
  • Zero Axe violations in the preview screen and zero browser page errors. S3 is an in-memory test double for the browser run; conversion and rendering are real.
  • Ruff, formatting, ty, ESLint, Prettier, template checks, Bandit and migration consistency checks passed. Stylelint has zero errors and 19 existing warnings elsewhere.
  • Docker image and Docusaurus production builds passed.
  • The initial GitHub dependency audit caught four advisories in the existing virtualenv development dependency. Updated its lockfile entry to 21.7.13 and compatible dependencies; the local Python dependency audit and virtual environment smoke check now pass.

Screenshots and the Axe result are stored only in the validation gist; screenshot files have been removed from the feature branch history.

Full evidence and reproduction commands: validation record.

Gist: https://gist.github.com/nonprofittechy/54857d2ed0b841a486dbf7a30b1ad915

Filing PDF preview

Vermont answer retains checkbox and multiline text

Deployment and limits

Apply migration 0029, configure GOTENBERG_URL and optional basic auth credentials, and provide fonts used by court documents on Gotenberg. The Docker build includes preview assets; local development uses npm ci in efile_app. Existing editable drafts without preparation metadata must pass preparation and preview; missing stored files must be replaced.

Flattening can affect accessibility, links or annotations; tagged Word output is not a PDF/UA certification. The original stays available and the filer must check the filing copy. The corpus includes partially populated templates as well as synthetic stress cases, so continued validation against additional completed filings is useful. Existing npm audit findings in development/test dependencies are documented in the validation record.

@nonprofittechy nonprofittechy changed the title Prepare filing PDFs and require document previews Prepare filing PDFs, analyze originals, and require previews Sep 30, 2026
@nonprofittechy
nonprofittechy force-pushed the feature/document-preparation-preview branch from 51a93d5 to 50a7cc8 Compare September 30, 2026 20:44
…dant code

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@nonprofittechy
nonprofittechy merged commit 53f6df3 into main Sep 30, 2026
8 checks passed
@nonprofittechy
nonprofittechy deleted the feature/document-preparation-preview branch September 30, 2026 22:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Flatten PDFs and DOCX on upload

1 participant