You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Read support implemented on branch claude/security-bugs-performance-review-qdmmco in commit a9f46d4 — and since formats 113–115 share one layout, this covers 113 (Stata 8/9) and 114 (Stata 10/11) as well as 115 (Stata 12).
The reader detects the format from the first byte (< → XML-tagged 117+ parser, 113/114/115 → new legacy parser for the flat binary layout: fixed header, descriptors, expansion fields). Legacy one-byte type codes are translated to the modern ones at parse time, so everything downstream — missing-value handling, %td date conversion, Latin-1→UTF-8 transcoding, value labels, projection/filter pushdown, parallel scans — works identically on legacy files. Value-label tables (same layout as 117 minus the tags) parse through a now-shared code path.
Fixtures and tests: format_114.dta and value_labels_114.dta written by pandas; format_115.dta is a byte-identical 114-layout file with the format byte set to 115 (pandas can't write 115) carrying Latin-1 content (café, naïve) to exercise transcoding. Suite is green at 625 assertions.
Scope note: this is read-only, as requested — COPY ... (FORMAT dta) still writes format 119. If write support for older formats is wanted, that's the same work item as the 117-write half of #3 (version-parameterized writer metadata). Also not covered: formats older than 113 (pre-Stata 8), which use different descriptor/value-label layouts.
https://www.stata.com/help.cgi?dta_115