doc-batch-gradients measures how much gradient noise whole-document long-context batches add, on open Pythia and OLMo-2 checkpoints. Within-document gradient correlation reaches out to 16K tokens. On code repositories this makes a 16K-token batch behave like roughly 3.5–5× fewer independent samples than packed web text (about 2× for arXiv, small for books). The work is preregistered, the results are released, and everything runs on CPU.
Contact: hanyu.yang.92@gmail.com



