Skip to content
View m-sanchez's full-sized avatar

Block or report m-sanchez

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
m-sanchez/README.md

Miguel Sánchez Durán

Senior software engineer in Dubai, with 15+ years of production experience. I build full-stack applications and applied AI tools, with a focus on evaluation, usable interfaces, and reproducible results.

Website · LinkedIn · Writing

Featured work

Check whether a classifier's confidence matches how often it is correct. Import prediction files, keep calibration, policy-validation, and test rows separate, and export an HTML report, experiment JSON, or prediction CSV.

The recorded UCI digits example retains all 1,797 original test rows. Temperature scaling leaves accuracy at 94.88% and reduces test negative log-likelihood from 0.17595 to 0.15753. This is one reference result, not a guarantee for other data. Confidence-only CSV inputs support measurement; fitting needs logits and known labels.

Open the app · Website copy · Recorded evidence

A local companion for Codex and Claude Code: find sessions across projects, inspect their history, and resume the intended conversation. It has a browser interface and an optional Windows app. Preview releases are labelled; the optional Ask Ocelin feature sends a compact snapshot to the configured model.

See the workflow · Downloads and release notes

In one recorded comparison using Claude Haiku 4.5, domain-specific prompts scored 67.75% against 82.75% for a generalist prompt under strict first-token scoring, the rule registered before the run. For every question the domain prompts lost to the generalist under that rule, the reply showed its working first, answered "No." with a full stop, or was cut off at the 256-token limit; none was a finished wrong answer. Scored on final answers, a rule chosen after reading the replies, the domain prompts scored 91% and the generalist 88%, and the difference was not significant (exact McNemar p=0.14, or 0.40 over the 372 distinct prompts). One model, 400 questions (372 distinct) and one recording: this is not evidence about routing in general. Read the re-score analysis.

Inspect and reproduce

Project Evidence and reproduction
Calibration Explorer Tagged source, setup, reference provenance, and replay instructions
calibrated Numerical corrections and distribution status for v2.0.1
Routing study Recorded model transcript and replay under the registered scorer, a post-hoc final-answer re-score of the same transcript, and a separately labelled synthetic study
Careful Machine Public reference implementation and interactive evidence demo

The Careful Machine repository illustrates an approach with public reference code and synthetic examples. It is not my employer's production system. Reproducing recorded output establishes repeatability; correctness and generalisation need additional evidence.

Smaller libraries and engineering tools

These repositories contain focused utilities with their own installation instructions, tests, and limitations.

Writing

If you try a project, an issue describing the task, release, and point of confusion is useful feedback. Please omit private inputs and credentials.

miguelsanchez.co.uk · contact@miguelsanchez.co.uk · Dubai

Pinned Loading

  1. calibration-explorer calibration-explorer Public

    Compare classifier confidence with accuracy using your own predictions. Browser-based calibration analysis and reports.

    TypeScript

  2. ocelin ocelin Public

    Ocelin: a local Windows companion for Codex and Claude. Live sessions, app memory, conversation previews, searchable history and selectable taskbar integrations.

    JavaScript 11

  3. careful-machine-reference careful-machine-reference Public

    Reference implementation for The Careful Machine: models propose, deterministic components certify. One question, two machines, and each core test is a book claim made executable.

    TypeScript

  4. grounded-claims grounded-claims Public

    Cite-or-refuse for LLM answers, enforced in code: an ordered gate pipeline over the turn, a check chain over each claim, contained plug-ins, a composer that cannot introduce or forge, and replayabl…

    TypeScript

  5. calibrated calibrated Public

    Is your model's confidence honest? ECE, Brier decomposition, reliability diagrams and temperature scaling. Zero dependencies.

    TypeScript

  6. routing-study routing-study Public

    Does routing to specialists beat one generalist? A reproducible study composed from the toolkit: careful-router routes, frozen-eval holds the bars, calibrated checks confidence, ab-significance tes…

    TypeScript