The build's memory, not a spec. What is left, what was already decided, and what is still unknown. Kept short on purpose: a cleared context should be able to read this page and pick up where the work stopped.
v0.1 is complete. Code, docs, CI, licence and the live validation all
landed on 2026-09-21, and v0.1.0 was released that day.
v0.1.1 is released, 2026-09-25: the review release. A repository review
filed 43 issues against v0.1.0 and every one is fixed in it, the local arm has
run against a real checkpoint, and it is the first release whose wheel and
sdist are built, checked, and attested by the release workflow. CI is green on
Ubuntu, Windows, and macOS across Python 3.12 to 3.14, and on the oldest
dependencies pyproject allows.
The repository is public as of 2026-09-21. The "Pre-public checklist" at the bottom of this page is kept as the record of what was verified before that happened, not as outstanding work.
The remaining work is the v0.2 milestone, and #3 comes first. Its local half is done; its generative half needs a paid run. Two decisions are open before v0.2 builds on the local arm: whether it should score multi-token options by sequence probability rather than refuse them (raised on #3), and whether MCE stays in the report at all (#7).
- Phase 6: adapters,
local_logitsandgenerative. Landed. The restricted softmax over a pinned checkpoint, and the text-generating control arm. - Phase 7: datasets. Landed. The JSONL loader that refuses unscoreable
rows, the JevBench public fixture and its translation, and the end-to-end
smoke run in
examples/smoke_public_dataset.py. - Phase 8: report and CLI. Landed. The markdown report groups arms by
probability_semantics, states every figure's row count and null, and demotes MCE to diagnostics.src/plumbline/cli.pyexists, so theplumblineconsole script pyproject declares now works:run,report,adapters,version. - Phase 9: recalibration, the cascade, and the methodology. Landed. The report fits a temperature on a held-out half and prints the verdict rather than the number when the verdict is a refusal; the cascade section ends in one sentence naming the threshold, the coverage, the expected cost, and the cost of escalating everything. METHODOLOGY has no pending sections left.
All three items are done. Kept here because the answers matter, not the list.
- Terms read, 2026-09-21. No clause mentions benchmarking or restricts publishing results. Three others constrain what a public README may carry, and all three are handled: 14.1 (pricing confidential) is why no figure ships, 16.4 (publicity) is why the example report uses the mock, and 2.3(c) (reverse engineering) turned out not to apply because the vendor publishes the confidence formula themselves.
- Live call made, 2026-09-21. See "Live validation" below. Both open questions are settled, and the call found two bugs that had never been exercised.
- Decision on going public: made. The repository is public, the v0.1.0 tag is pushed, and the checklist at the bottom of this page records what was done.
- Ordinal score questions. Landed (#6): mean absolute error against a
permutation null, the ranked probability score, and a cumulative calibration
error, each against a calibrated-model floor, in a score block of their own.
typesafe_wireasks a real Score with the row's rubric. Recalibration and the cascade for scores are not designed yet. - Batching. One request per case today, so cost and latency are both conservative relative to batched use. The vendor's own documentation describes packing many questions against one shared state in a single call, which is a materially different cost and latency profile and is the single largest measurement gap in v0.1.
- Per-label scaling. Landed (#4): after a global fit refused as the wrong shape or stopped short of the floor, a temperature per predicted label, on the same split and verdict rule, with a 100-row gate per label. Vector and matrix scaling stay out until per-label proves insufficient on real data.
- Option descriptions on the wire. Landed in v0.1.1 (#39):
typesafe_wiresends a Choice's descriptions as its criteria, they join the cache key and the dataset hash, and the report says when an adapter does not send them.
Proposals, not code. Each says what would be built, what it refuses, and what is still undecided, so the build can be judged against something written first.
Blocked on #3: the generative half of the local arm has not run live, and nothing here should be built on an arm that has only met fakes.
What a run would report. The same rows, measured twice: once one request per case (today's behaviour) and once batched. Cost and latency are reported for each, so the saving is measured on the user's data and never quoted from a vendor page. How much state the rows share changes the saving a great deal, so no single "N times cheaper" figure is ever printed.
Who decides the batches: the caller. Batch composition is an input, given as a file mapping each batch to its case ids, the same way escalation cost and error cost are never defaulted. A documented helper can write that file by grouping rows on a shared field the user names, but the grouping is a choice the user can read and edit, not a heuristic that moves the result unseen. Automatic grouping is deliberately not offered. The cost is friendliness; the gain is that two runs of the same file mean the same measurement.
What the artifact records. Each case record gains a batch_id, and the
artifact gains the batch table (batch id, case ids in request order, a hash of
the whole request). Which cases shared a request changes the result, so it sits
beside prompt_hash, and it joins the cache key: the same case asked alone and
asked in a batch are different measurements and never share a cache entry.
Latency is two quantities, in two columns. An unbatched case has a per-case wall-clock time. A batch has one wall-clock time covering its cases. The report shows per-case latency for the unbatched pass and per-batch latency for the batched pass. It never divides a batch's time by its size and presents the result as a per-case figure, because that is a different quantity.
What is refused. A batch whose answer cannot be attributed to its case with certainty is refused whole, by name, and none of its cases are scored. That covers a response with the wrong number of answers, an answer that does not echo its case id, and a duplicated id. Cases in one batch never contribute to each other's score, and a refused batch counts toward the refusal tally like any other refusal.
What it needs from an adapter. A new capability flag beside the existing
cost one, so an adapter that cannot batch says so and the batched pass is skipped
with a stated reason instead of silently falling back to one request per case.
Only typesafe_wire is a candidate, and only if the wire format's batch call
returns one attributable answer per question.
Undecided. Whether the batched pass may use a different retry policy (one failed case would cost a whole batch a retry). Whether to report the saving as a ratio at all, or only the two absolute figures. When this lands, the README Limitations entry on conservative cost and latency is replaced with what was measured.
The question in the issue is answered in METHODOLOGY: on a scalar-only column isotonic regression escapes the shape temperature cannot fix, and loses to it at small sizes where temperature already fits. What is left is whether the report should choose between them, and that is a design problem before it is a coding one.
The trap. Picking the correction with the lower held-out ECE and then reporting that held-out ECE counts the choice twice, so the reported figure is optimistic by an amount nobody has measured.
Proposed shape. Three disjoint parts of the rows, never two. The first fits both candidates. The second chooses between them. The third is judged once, for the chosen correction only, and that is the figure reported. The third part is never used to choose.
Rows it needs. Isotonic spends rows finding a shape, and the existing table shows it does not beat temperature until thousands of rows even where it should. So the gate is higher than the 200-row gate for temperature alone: isotonic is considered only when each of the three parts clears a stated floor, and below that the report fits temperature only, as it does today, and says why isotonic was not tried.
Refusal stays. If neither correction reaches the floor on the third part, the report refuses to recommend either, as it does now. A fit that chooses between two corrections is not allowed to turn a refusal into an answer.
Undecided. The row floor per part, which needs the same kind of sweep the existing table came from. Whether the choice rule is lowest ECE on the second part or lowest ECE as a multiple of the floor, since the table shows the two disagree across sizes. Neither should be settled without that sweep.
Settled during the build. Reopen one only with a reason, not from scratch.
-
ECE and MCE are always reported against their calibrated-null floor.
-
MCE is demoted to a diagnostics block, never beside ECE, because it cannot detect gross overconfidence at 500 rows.
-
Recalibration has three verdicts (recommended, partial, refused) and emits no temperature on refusal.
-
Latency percentiles use nearest rank, not interpolation.
-
Adapters are organized by transport; a new vendor is config, not code.
-
plumbline is not a leaderboard. JevBench is.
-
Every pricing entry carries its source and the date it was read; an entry older than
DEFAULT_PRICING_MAX_AGE_DAYSis used but flagged rather than presented as current. -
A blank cost column names its reason (
CostBasis): an adapter that reports no tokens at all is a different finding from an API that reported none on this run. -
local_logitsis fixed atrestricted_softmaxandgenerativeatnone; neither is configurable, because the report groups on that field. -
A row whose gold label is not one of its options is refused by the loader, never scored: it would mark every system wrong and read as a model failure.
-
Every calibration figure carries its row count, and the artifact records
dataset_rowsbesidedataset_hash, because the floor depends on n. -
The JevBench public rows are vendored as a fixture under MIT with attribution. Each is asked as the question type it states, through plumbline's own harness, so results from them are never comparable with JevBench's published numbers.
-
A yes/no row is asked as a Noul where the transport has one, and every record carries both what the row asks and how it was asked. A noul figure is never compared with a two-option-choice figure without that line between them.
-
Ordinal score rows are read by rank in a block of their own and left out of every choice figure; before v0.2 they were excluded from every figure.
-
Artifacts never overwrite each other: the timestamp is only accurate to the second, so a repeated name gets a suffix rather than replacing user records.
-
A refused recalibration prints no number: the verdict, the split sizes, and nothing that could be lifted into production code.
-
The cascade threshold is chosen on the fit half of the rows and its cost and coverage are reported on the held-out half, which neither the threshold nor any temperature has seen; a threshold chosen and scored on the same rows would report its best case. Both halves are on the recalibrated scale when a temperature was recommended, because a threshold set against a raw overconfident probability sits in the wrong place.
-
Escalation cost and error cost are supplied by the caller and never defaulted. No benchmark can know them, and a made-up default would decide the threshold.
-
Below 200 held-out rows no threshold is printed at all.
-
generativeconfigures no fallback model. This is deliberate and deviates from the SDK's default advice for this model family: server-side fallback would re-run a declined case on another model inside the same call, so one dataset would be answered partly by two models and the run would no longer be about the system under test. A decline is recorded as a refusal with its category and counted. The tradeoff is stated at the call site inadapters/generative.py; it is not an oversight. -
CodeQL default setup is kept for its
actionscoverage, not its Python coverage. Workflow script injection is a real class of bug andci.ymlis where it would hide. The Python queries are expected to be low yield on this codebase: it is a CLI with no attacker in its threat model, run by an operator on their own data with their own key. The first full scan produced exactly one alert, a false positive on a test assertion (py/incomplete-url-substring-sanitization, a substring check that makes no security decision), dismissed with that reasoning. Turn it off if a second false positive appears on ordinary work; at that point it is costing review attention it is not repaying. It is a setting, not a workflow file, so disabling it is one API call and leaves no trace in the tree. -
Secret scanning and push protection guard credential formats and nothing else. The disclosure risk this project actually has is contractual, and the control for it is the
--pricingdesign plus a test. SECURITY.md says so in full, because a green scanning badge invites the wrong assumption. -
The site's calculator (
site/floor.js) is a JavaScript port of the floor that reproduces numpy's random stream draw for draw, not a statistical approximation of it. It is held to golden values from the Python within 1e-9 by.github/workflows/site.yml, and a failing check blocks the deploy: a stale site is the better failure than a wrong one. -
The site's worked example is regenerated at deploy time from the mock adapter by the command
docs/example-report.mdrecords.scripts/build_site.pyrefuses any other adapter, so no vendor's rows can reach the site through it.
numpy does not promise that Generator streams are stable across versions. The
port reproduces numpy 2.5.3's PCG64, its ziggurat normal and exponential
samplers (with their tables copied from that release), random_standard_gamma,
and random_beta. If a numpy release changes any of those, the port and the
Python silently diverge.
- What breaks. The first draw that differs shifts every draw after it, so the floors move by around 1e-3, not by rounding error.
- How it shows up.
site.ymlruns on every pull request, with no path filter, so a Dependabot pull request that bumps numpy inuv.lockruns it.scripts/floor_golden.py --checkfails first if the Python's own floors moved ("fixture is stale"); after regenerating the fixture,check_floor_parity.mjsfails with every case listed. Either way it fails on the pull request, not onmain, and the site keeps serving the last good build. - The fix. Re-port the sampler that changed from the new numpy source
(
numpy/random/src/distributions/distributions.c,ziggurat_constants.h, andpcg64/pcg64.hat the new tag), update the version named infloor.js, regenerate the fixture, and let the check pass. Do not loosen the tolerance to make it pass: a tolerance wide enough to absorb a desynchronised stream is wide enough to absorb a wrong port. - Why numpy is not pinned tighter in
pyproject.toml. The site builds withuv sync --locked, so the version the port runs against isuv.lock's exact pin, which is already tighter than a major.minor bound. A bound inpyproject.tomlwould constrain everyone who installs plumbline, for the sake of a page they never run, and would protect the site from nothing the lock does not already cover. The existingnumpy>=2.1floor is also not looser than the port tolerates: on 2026-09-22,floor_golden.py --checkmatched all 27 cases under numpy 2.1.3, 2.2.6, 2.3, 2.4.6, and 2.5.3, so a user on any of those gets the floors the site shows. To recheck a version, runuv run --isolated --with 'numpy==X.Y.*' python scripts/floor_golden.py --check.
Considered and declined on 2026-09-21. Listed so they are not re-proposed as oversights. Each would be defensible later for a stated reason; none is defensible merely because projects usually have one.
.editorconfig. ruff already owns formatting here, CI enforcesruff format --check, and no tool in this repository reads an editorconfig. Adding one creates a second source of truth for line length and indentation that can silently disagree with the first..gitattributesalready pins the vendored fixture's bytes, which is the only line-ending rule that affects correctness. Revisit if a contributor's editor is actually fighting ruff.- pre-commit. It would catch exactly what CI already catches, in exchange
for a setup step in CONTRIBUTING, a pinned-hook config to keep current, and a
second place where the lint versions live.
uv run ruff check . && uv run mypy --strictis already documented and is one command. Revisit if CI minutes or review round-trips become the bottleneck, which at this size they are not. - CODEOWNERS. One maintainer. The file would assign every path to the person who would be reviewing it anyway, and the ruleset already requires a pull request. Revisit on the second maintainer.
CITATION.cff. Nobody has cited this. A citation file asserting how to cite work nobody has referenced is a claim about its significance rather than a service to a reader. Revisit if someone references the METHODOLOGY results.- PyPI publishing. Premature, and the name is taken:
plumblineon PyPI is an unrelated project, so publishing needs a different distribution name first. It commits the project to a name and to a release cadence before the API has settled, and the API is explicitly not stable before v0.2. The README says cloning is the install path, which is honest and costs a reader one command. Revisit when the CLI flags stop moving. - CI dependency caching was not added because it was already there:
astral-sh/setup-uvruns withenable-cache: true.
Run on 2026-09-21. Model requested jev-latest; model that answered
jev-1.13.0. 40 choice rows from the vendored JevBench public fixture
(datasets/public/jevbench-hard.jsonl, choice rows only, first 40 of 67), one
request per case, 4 concurrent workers, no cache. Total billed input 66,458
tokens, output 2,555 tokens.
The artifact is not committed: it holds per-case records from a vendor call and
lives in gitignored results/.
No dollar figure from this run appears in this file or in any other committed
file. The vendor's MCA makes its pricing information confidential and overrides
the public-knowledge carve-out, and a spend total stated next to a token count
lets a reader divide one out. The figures exist in the local artifact under
results/, which is gitignored, and that is where they stay.
jev-1is not a model. The adapter's defaultmodel_requestedwasjev-1, which the API rejects with400 Unknown model: jev-1. The SDK's own default isjev-latest(typesafe_sdk/constants.py,DEFAULT_MODEL). Fixed inadapters/typesafe_wire.py. The first five-case probe spent nothing because every request 400'd.- The pricing table was keyed on aliases only.
pricing_forprefers the model the API reported, and a call tojev-latestreportsjev-1.13.0. With only alias keys, every live row priced asmodel_not_priced. The table is now keyed on the reported version as well as the aliases, with no prefix matching, so a futurejev-1.14.0will not silently inherit a superseded rate.
-
Does
response.usagepopulate the token counts? Yes. 40 of 40 rows returned non-Noneinput_tokensandoutput_tokens. The wire schema is right and the SDK'sint | Noneis a widening, not a description of behaviour. The None-handling path stays: it costs nothing and it is the difference between a blank cost and a false zero. But the report no longer needs to hedge about how oftentokens_not_reportedis taken on this vendor. Not observed once in 40 calls. -
What is
response.model, and does it match what was requested? It isjev-1.13.0on every row, and it does not match the requestedjev-latest. This is documented behaviour, not a fault: TypeSafe's models page states thatjev-latestandjev-previeware aliases that both currently resolve tojev-1.13.0, and that "the response'smodelfield reports the versioned ID that answered". Recording both strings in the artifact is what makes a result readable after an alias moves. -
Do
probabilitiessum to 1? Yes, to floating-point exactness. Maximum observed deviation across 40 rows is 1.11e-16, mean 2.78e-18, and 39 of 40 rows sum to exactly 1.0. Every one of the 165 probability values returned lies exactly on a two-decimal grid, so the distribution is quantized to 0.01 on the wire. Worth knowing for calibration work: with 10 equal-width ECE bins, a 0.01 grid is finer than the binning and does not bias it, but it does put a floor on how finely a threshold can be tuned. -
Is
probabilities[choice]always the maximum? The value always is: zero rows of 40 hadprob_selecteddiffering frommax(distribution.values()). But the reported choice is not always what a naive argmax picks, because ties happen. One row (hard-opus-c-temporal_numeric-04) returned two options tied at 0.23; the API selectedusd_2303_01and the gold label was the other member of the tie,usd_2335_00. That row scored wrong on a tie-break. At a 0.01 quantization over 3 to 6 options, ties are not rare events and code that re-derives the choice by argmax instead of readinganswer.choicewill disagree with the vendor. plumbline readsanswer.choice, which is correct. -
Does
confidenceequal(n * max_prob - 1) / (n - 1)? Effectively yes, to within wire rounding. Maximum absolute deviation over 40 rows is 0.0167, mean 0.0053, median 0.0050; 10 rows match exactly and all 40 are within 0.02. Both the probabilities and the confidences are returned on a two-decimal grid, and a +/-0.005 rounding ofmax_probpropagates to +/-0.005 * n/(n-1) in the confidence, which accounts for the spread. The residual is consistent with the vendor computing confidence from full-precision internal probabilities and rounding both for the wire.This is not a reverse-engineering finding. TypeSafe publishes the formula themselves on https://docs.typesafe.ai/confidence, where the interactive explainer computes
(count * peak - 1) / (count - 1)clamped to [0, 1], and the accompanying text describes it as an approximation of the production definition. The docs also confirm that Noul answers carry no confidence, which is what the adapter already assumes.Consequence for METHODOLOGY: confidence is a deterministic function of
max_probat fixed n, so AUROC parity between the probability column and the confidence column on fixed-width rows is structural rather than empirical, and the divergence on mixed-width rows follows from n varying. That is exactly what METHODOLOGY already claims from the seeded mock, so the existing hedged wording stands and now has a mechanism behind it. Per the decision recorded below, the relationship is not published. -
Latency percentiles, nearest rank, over 40 live calls (4 concurrent workers, so these include some self-inflicted queueing): min 135ms, p50 178ms, p90 235ms, p95 367ms, p99 469ms, max 469ms.
-
Did anything fail, rate limit, or refuse? No. 40 of 40 rows produced a prediction, zero refusals, zero rate limits. One row needed a second attempt and succeeded on it. No
429was seen; the published limits are 250,000 tokens per second and 1,200 requests per minute, far above this run.
TypeSafe publishes a tariff for jev-1.13.0 on
https://docs.typesafe.ai/models.md, covering both the input and the output
side. The figures are deliberately not reproduced here; read the page.
That settles the open question that had blocked cost entirely. The previous
shipped entry quoted an SDK field description for the output side and carried no
input price at all, so is_priced was False and the cost guard refused any run
with --max-cost-usd set: it cannot bound a run it cannot cost.
What plumbline ships is a separate decision from what the tariff says. MCA 14.1 makes the vendor's pricing information confidential and explicitly overrides the public-knowledge carve-out in 14.3, so a figure written into the default table would be published to everyone who clones this repository. The shipped entries therefore price nothing and name the page instead.
Of the two options considered, the number moves out of the committed table
into operator-supplied config, rather than the table keeping the number while
the source string omits it. The second does not work: a populated
input_usd_per_million field in a committed file is the per-token price,
whatever the adjacent string says. Only removing the number removes the
disclosure.
So:
config.pyshipsjev-1.13.0,jev-latestandjev-previewwith both prices None and a source naming the tariff page. None rather than 0.0 on the output side, because a zero there would be plumbline asserting those tokens are free.load_pricing_fileandmerged_pricing_tableread an operator's own JSON table and lay it over the shipped one, wired toplumbline run --pricing. Every supplied entry must carry its ownsourceandas_of, on the same reasoning that applies to a shipped one.docs/pricing.example.jsonis a template whose prices are null and whoseas_ofis the placeholderYYYY-MM-DD, which the loader refuses, so an unedited copy cannot price a run. An entry with both prices at 0 is refused unless it says"free": true.- A test asserts no shipped entry for this vendor carries a numeric price, so a figure cannot drift back in unnoticed.
Verified end to end with a local supplied table: estimate before the run, actual
after, cost_basis priced on all 40 rows, and the guard admitted the run on
its own arithmetic rather than being bypassed. With no supplied table the same
run reports model_not_priced and --max-cost-usd refuses, which is the honest
default.
The working tree was cleaned first: config.py no longer quotes the SDK field
description that stated the output price, METHODOLOGY.md no longer repeats
that quotation, and the old open-questions section that carried it is gone.
A scan of all 48 commits then found the claim in five files rather than the three first identified, plus one commit message:
src/plumbline/config.py, asJEV_OUTPUT_SOURCEMETHODOLOGY.md, in "Prices are dated, and so is every result"docs/PLAN.md, in the old open-questions sectionsrc/plumbline/metrics/cost.py, in the module docstring (not previously spotted)tests/test_cost_and_latency.py, in an assertion on the source string (not previously spotted)- the message of the commit that introduced dated pricing
The numeric input price never reached a commit, and neither did any spend total from the live run: both were removed before the first commit that would have carried them. The zero output price was also present in history as a populated price field, which is the same claim in another form, so it was included.
History was rewritten with git filter-repo on 2026-09-21, using a replace-text
rule over file contents and a replace-message rule over commit messages, rather
than by rebasing three commits by hand. The replacement is a neutral pointer to
the vendor's own models page, not a redaction marker: a marker would advertise
that something was taken out, which is the opposite of the point. The populated
zero price became None, which keeps the historical file valid Python and says
what the current entry says.
The replacement rules were checked against datasets/public/LICENSE-jevbench
and datasets/public/jevbench-hard.jsonl before running, because both contain
similar wording for unrelated reasons: MIT licence boilerplate in the first, and
warranty scenarios and a grant figure in the dataset rows of the second. No rule
matched either file, and both blobs are byte-identical before and after the
rewrite.
Accuracy 0.6500 over 40 rows against a chance null of 0.2571, better than chance. ECE 0.1070 against a calibrated-model floor of 0.1400, inconclusive: 40 rows cannot tell this apart from a perfectly calibrated model, which establishes nothing in either direction. Brier 0.1737 against a floor of 0.1512, also inconclusive. Confidence AUROC 0.7981 against a permutation null of 0.5013, separates correct from incorrect. Recalibration refused: it needs 200 held-out rows and this split has 20.
That is the intended behaviour at n = 40 and it is the argument the README makes.
-
Ask TypeSafe, in writing, for permission to name them in a public repository's example output. MCA 16.4 is the clause to ask about: it bars either party from publicly announcing that they entered the agreement, and a committed report from a keyed run arguably does that. Note the clause is one-sided in the other direction, since it expressly lets TypeSafe name its customers.
This is why
docs/example-report.mdcurrently shows the seeded mock arm rather than the live run, and says so at the top. Publicity clauses of this kind are boilerplate and routinely waived, and a vendor is usually glad to be named by a tool that measures it fairly and refuses to rank it. A one-line written yes swaps the mock report for a real vendor one, which is a strictly better page: a reader could then see the tool decline to draw a conclusion from an actual vendor at 40 rows.Worth asking about 14.1 in the same message, since a written yes on restating the published tariff would let the shipped pricing table carry figures again and remove the
--pricingstep for every user of this vendor.Follow-up, not a release blocker. Nothing about v0.1 depends on the answer.
A yes swaps the mock example report for a real one with no other work required. The live run already happened, its artifact is in gitignored
results/, andplumbline report <artifact> --out docs/example-report.mdregenerates the page from it. The header above the generated half is the only hand-written part, and only its provenance table and the sentence explaining why the arm is a mock would change. -
Whether the residual in the confidence relationship (max 0.0167) is purely wire rounding or a slightly different production formula. Not worth another spend to settle, and nothing in plumbline depends on the answer.
-
Resolved:datasets/private/.gitkeepis not tracked..gitignorenow ignoresdatasets/private/*with a negation for.gitkeep, so a fresh clone has the directory the README names and still never commits what goes in it.
For the GitHub About box. Paste as is.
Description (108 characters, as set):
Measure whether a decision model's probabilities are trustworthy on your own labeled data. Not a leaderboard.
Topics:
calibration
evaluation
llm
classification
machine-learning
benchmarking
uncertainty-quantification
model-evaluation
confidence-calibration
expected-calibration-error
python
cli
No vendor name is included. The tool is not about one vendor, the README says so in its first section, and a vendor topic would file it as a fan project.
Everything below is a manual step. Nothing in this repository does any of it for you.
Verified already, listed so you can re-check rather than re-derive:
- Full secret scan across all history, not just the working tree. The live
key appears in no commit and no blob;
.envhas never been committed;apik_,sk-ant-andBearermatch nothing anywhere in history. -
results/,cache/,models/anddatasets/private/are all ignored and none has ever been committed. No run artifact is in history. -
uv.lockis tracked and current (uv lock --checkpasses). - CI green on
ubuntu-latestandwindows-latest, Python 3.12 and 3.13. No cross-platform failure appeared; the suspected path handling was fine. - Fresh install from the built wheel: entry point resolves,
--help,adaptersandversionall work, the package imports with core dependencies only, and nothing writes into the working directory. -
LICENSEpresent, complete, Apache-2.0, dated 2026, holder TMHSDigital. - History rewritten to remove the vendor price claim, and re-scanned after. The MIT fixture and licence are byte-identical before and after.
Done since: the v0.1.0 tag is pushed, the repository is public, the About box carries the description above, and the CI badge renders for logged-out readers.
Still open: email TypeSafe about 16.4 and 14.1 (see "Open questions"). A written yes converts the mock example report into a real vendor one and would let the shipped pricing table carry figures again. Not a blocker for anything.