Senior software engineer in Dubai, with 15+ years of production experience. I build full-stack applications and applied AI tools, with a focus on evaluation, usable interfaces, and reproducible results.
Check whether a classifier's confidence matches how often it is correct. Import prediction files, keep calibration, policy-validation, and test rows separate, and export an HTML report, experiment JSON, or prediction CSV.
The recorded UCI digits example retains all 1,797 original test rows. Temperature scaling leaves accuracy at 94.88% and reduces test negative log-likelihood from 0.17595 to 0.15753. This is one reference result, not a guarantee for other data. Confidence-only CSV inputs support measurement; fitting needs logits and known labels.
Open the app · Website copy · Recorded evidence
A local companion for Codex and Claude Code: find sessions across projects, inspect their history, and resume the intended conversation. It has a browser interface and an optional Windows app. Preview releases are labelled; the optional Ask Ocelin feature sends a compact snapshot to the configured model.
See the workflow · Downloads and release notes
In one recorded comparison using Claude Haiku 4.5, domain-specific prompts scored 67.75% against 82.75% for a generalist prompt under strict first-token scoring, the rule registered before the run. For every question the domain prompts lost to the generalist under that rule, the reply showed its working first, answered "No." with a full stop, or was cut off at the 256-token limit; none was a finished wrong answer. Scored on final answers, a rule chosen after reading the replies, the domain prompts scored 91% and the generalist 88%, and the difference was not significant (exact McNemar p=0.14, or 0.40 over the 372 distinct prompts). One model, 400 questions (372 distinct) and one recording: this is not evidence about routing in general. Read the re-score analysis.
| Project | Evidence and reproduction |
|---|---|
| Calibration Explorer | Tagged source, setup, reference provenance, and replay instructions |
| calibrated | Numerical corrections and distribution status for v2.0.1 |
| Routing study | Recorded model transcript and replay under the registered scorer, a post-hoc final-answer re-score of the same transcript, and a separately labelled synthetic study |
| Careful Machine | Public reference implementation and interactive evidence demo |
The Careful Machine repository illustrates an approach with public reference code and synthetic examples. It is not my employer's production system. Reproducing recorded output establishes repeatability; correctness and generalisation need additional evidence.
Smaller libraries and engineering tools
These repositories contain focused utilities with their own installation instructions, tests, and limitations.
- Evidence and verification: grounded-claims, careful-verifier, u-pack.
- Evaluation and calibration: calibrated, ab-significance, probe-heads, frozen-eval, silent-zero.
- Runtime and process controls: careful-router, gpu-quiescence, training-forge, clean-room-guard.
- Accuracy stayed the same. Confidence changed.
- Building AI That Cites or Refuses
- How I Evaluate Production RAG Systems
- What 15 Years of Software Engineering Taught Me About AI Engineering
If you try a project, an issue describing the task, release, and point of confusion is useful feedback. Please omit private inputs and credentials.



