Repository navigation
Run the local arm for real: local_logits and generative have never left the test suite #3
Description
Activity
- addedenhancementNew capability or a measurement the tool cannot make yetNew capability or a measurement the tool cannot make yetadapterAdapter transports and vendor configurationAdapter transports and vendor configurationmethodologyHow a number is computed and what it does or does not meanHow a number is computed and what it does or does not meanv0.2Planned for the v0.2 releasePlanned for the v0.2 release
on Sep 22, 2026 - added a commit that references this issue
on Sep 25, 2026 The local half is done (#96).
local_logitsran end to end againstQwen/Qwen2.5-1.5B-Instruct, pinned to commit989aa7980e4cf806f80c7fef2b1adb7bc71aa306, on the vendored fixture. It used a GPU (RTX 4060, CUDA torch 2.14, transformers 5.17, Windows). Thelocalextra installs and imports, the revision pin is enforced, the artifact is written, every figure is read against its null, and the report renders. Two runs gave identical figures.- 39 of 105 rows scored: the 38 yes/no rows and 1 choice row. The other 66 choice rows were refused, each by name, because their options are multi-token identifiers.
- Accuracy 0.4615 against a chance null of 0.4957: INCONCLUSIVE. ECE 0.1257 against a floor of 0.1016 (p95 0.1869): INCONCLUSIVE. Brier 0.2762 against a floor of 0.2386 (p95 0.2662): distinguishable from noise.
- Latency p50 71 ms, p99 488 ms.
The run found four faults, all fixed in #96:
- The workers loaded the checkpoint once each.
- There was no
--deviceoption. - Kernel setup was timed as a case's latency.
- Two report lines were misworded for arms with failures.
Still open here:
- The generative half. It needs a real Anthropic key and a paid run. The goal is to observe the refusal path, and to confirm the report leaves the arm out of calibration rather than just meaning to.
- A methodology question the real run raised. On this fixture the local arm can score almost only the yes/no rows, because the choice options are multi-token. Should the null score multi-token options by summed sequence log-probability instead of refusing them? That would change what
restricted_softmaxmeans, so it should be decided before Per-label temperature scaling, so a refused global fit has a remedy rather than only a diagnosis #4 builds on this arm. - A local-machine note: loading took anywhere from about 100 s to 731 s with warm caches. It's no part of any figure, but a first run can look hung.
- added a commit that references this issue
on Sep 25, 2026 The multi-token decision is implemented (#102): letter labels, opt-in.
--option-style letterasks the options as A, B, C and reads the letter tokens. On the public fixture it scores all 105 rows, where the default scores 39. The pinned Qwen2.5-1.5B landed at chance accuracy (0.343 against a chance null of 0.342) with an ECE of 0.349 against a floor of 0.087, which is distinguishable from noise: a small model that is confidently wrong. The default still reads each option's own token.The generative half is still open, and it needs a paid run.
Why this is the highest-value item in the backlog
local_logitsandgenerativehave never run outside the test suite. Both areexercised against fakes; neither has been pointed at a real checkpoint or a
real chat model on a real dataset.
That matters more for
local_logitsthan it sounds, because therestricted_softmaxclass is the null hypothesis of the entire tool.plumbline's argument is that a vendor's calibration claim has to be measured
rather than believed, and the thing it should be measured against is a plain
softmax over the declared option tokens, read straight out of an open model
with no calibration claim attached. METHODOLOGY states this: the local arm "is
the null hypothesis the calibrated claims have to beat."
Right now that is an assertion. Nobody has run it.
What done looks like
local_logitsruns against a pinned open checkpoint on the vendored publicfixture, and the run completes end to end: artifact written, metrics computed
against their nulls, report rendered.
generativeruns against a real chat model and its refusal path is observedrather than simulated. It declares
nonesemantics, so it contributesaccuracy only; confirm the report actually excludes it from calibration
rather than merely intending to.
localextra installs and imports on a machine that has a GPU, and therevision-pinning guard is exercised against a real HuggingFace revision.
restricted_softmaxsection is updated with what was measured,with its n and its null, replacing the current assertion.
docs/PLAN.mdrecords the checkpoint, the revision, and the date.The outcome worth wanting
A result showing the local readout is well calibrated would weaken a claim
this project makes. It would mean a restricted softmax over option tokens,
with no calibration training and no vendor claim, is already doing the job the
calibrated claim is sold on, and that the gap plumbline exists to detect is
smaller than implied.
That is the correct outcome to want. A measurement tool that only produces
results flattering to its own framing is not a measurement tool. If the null is
strong, the honest move is to say so in METHODOLOGY and let readers draw the
conclusion, not to quietly stop reporting it.
Either result is publishable. Only not running it is not.