Skip to content

Run the local arm for real: local_logits and generative have never left the test suite #3

Description

@TMHSDigital

Why this is the highest-value item in the backlog

local_logits and generative have never run outside the test suite. Both are
exercised against fakes; neither has been pointed at a real checkpoint or a
real chat model on a real dataset.

That matters more for local_logits than it sounds, because the
restricted_softmax class is the null hypothesis of the entire tool.

plumbline's argument is that a vendor's calibration claim has to be measured
rather than believed, and the thing it should be measured against is a plain
softmax over the declared option tokens, read straight out of an open model
with no calibration claim attached. METHODOLOGY states this: the local arm "is
the null hypothesis the calibrated claims have to beat."

Right now that is an assertion. Nobody has run it.

What done looks like

  • local_logits runs against a pinned open checkpoint on the vendored public
    fixture, and the run completes end to end: artifact written, metrics computed
    against their nulls, report rendered.
  • generative runs against a real chat model and its refusal path is observed
    rather than simulated. It declares none semantics, so it contributes
    accuracy only; confirm the report actually excludes it from calibration
    rather than merely intending to.
  • The local extra installs and imports on a machine that has a GPU, and the
    revision-pinning guard is exercised against a real HuggingFace revision.
  • METHODOLOGY's restricted_softmax section is updated with what was measured,
    with its n and its null, replacing the current assertion.
  • docs/PLAN.md records the checkpoint, the revision, and the date.

The outcome worth wanting

A result showing the local readout is well calibrated would weaken a claim
this project makes.
It would mean a restricted softmax over option tokens,
with no calibration training and no vendor claim, is already doing the job the
calibrated claim is sold on, and that the gap plumbline exists to detect is
smaller than implied.

That is the correct outcome to want. A measurement tool that only produces
results flattering to its own framing is not a measurement tool. If the null is
strong, the honest move is to say so in METHODOLOGY and let readers draw the
conclusion, not to quietly stop reporting it.

Either result is publishable. Only not running it is not.

Activity

  1. added this to the v0.2 milestone on Sep 22, 2026
  2. added
    enhancementNew capability or a measurement the tool cannot make yet
    adapterAdapter transports and vendor configuration
    methodologyHow a number is computed and what it does or does not mean
    v0.2Planned for the v0.2 release
    on Sep 22, 2026
  3. TMHSDigital commented on Sep 25, 2026

    @TMHSDigital
    OwnerAuthor

    The local half is done (#96).

    local_logits ran end to end against Qwen/Qwen2.5-1.5B-Instruct, pinned to commit 989aa7980e4cf806f80c7fef2b1adb7bc71aa306, on the vendored fixture. It used a GPU (RTX 4060, CUDA torch 2.14, transformers 5.17, Windows). The local extra installs and imports, the revision pin is enforced, the artifact is written, every figure is read against its null, and the report renders. Two runs gave identical figures.

    • 39 of 105 rows scored: the 38 yes/no rows and 1 choice row. The other 66 choice rows were refused, each by name, because their options are multi-token identifiers.
    • Accuracy 0.4615 against a chance null of 0.4957: INCONCLUSIVE. ECE 0.1257 against a floor of 0.1016 (p95 0.1869): INCONCLUSIVE. Brier 0.2762 against a floor of 0.2386 (p95 0.2662): distinguishable from noise.
    • Latency p50 71 ms, p99 488 ms.

    The run found four faults, all fixed in #96:

    • The workers loaded the checkpoint once each.
    • There was no --device option.
    • Kernel setup was timed as a case's latency.
    • Two report lines were misworded for arms with failures.

    Still open here:

    1. The generative half. It needs a real Anthropic key and a paid run. The goal is to observe the refusal path, and to confirm the report leaves the arm out of calibration rather than just meaning to.
    2. A methodology question the real run raised. On this fixture the local arm can score almost only the yes/no rows, because the choice options are multi-token. Should the null score multi-token options by summed sequence log-probability instead of refusing them? That would change what restricted_softmax means, so it should be decided before Per-label temperature scaling, so a refused global fit has a remedy rather than only a diagnosis #4 builds on this arm.
    3. A local-machine note: loading took anywhere from about 100 s to 731 s with warm caches. It's no part of any figure, but a first run can look hung.
  4. TMHSDigital commented on Sep 25, 2026

    @TMHSDigital
    OwnerAuthor

    The multi-token decision is implemented (#102): letter labels, opt-in.

    --option-style letter asks the options as A, B, C and reads the letter tokens. On the public fixture it scores all 105 rows, where the default scores 39. The pinned Qwen2.5-1.5B landed at chance accuracy (0.343 against a chance null of 0.342) with an ECE of 0.349 against a floor of 0.087, which is distinguishable from noise: a small model that is confidently wrong. The default still reads each option's own token.

    The generative half is still open, and it needs a paid run.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    adapterAdapter transports and vendor configurationenhancementNew capability or a measurement the tool cannot make yetmethodologyHow a number is computed and what it does or does not meanv0.2Planned for the v0.2 release

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions