OpenAI-compatible server with a live view into any HF model's residual stream. pip install brainscope
-
Updated
Sep 23, 2026 - Python
OpenAI-compatible server with a live view into any HF model's residual stream. pip install brainscope
A toolkit+visualizer that is designed for mech interp researchers working with chess models that use transformers tokenized by square.
Local GGUF workbench for quantization, evaluation, chat, and model interpretability.
A Jacobian-Lens (J-Lens) observer for vision-language models — read what a VLM is poised to say, before it says it. Multimodal J-Lens on Qwen3.5, concept-race, and a forward-only prompt helper.
In-browser LLM interpretability: from-scratch logit lens and activation steering on WebGPU
Open-source EU AI Act Annex IV documentation toolkit. Mechanistic interpretability + circuit discovery for transformers. One function call generates a structured, hash-chained evidence package.
OKI TRACE: Local LLM observability. See step-by-step, layer-by-layer what your AI thinks. Logit Lens & Attention for HuggingFace models.
InterpLens is an open-source mechanistic interpretability toolkit that accelerates neural network reverse-engineering with interactive circuit debugging and real-time telemetry. It bridges research prototyping and model transparency with zero-copy GPU memory attachment and universal transformer hooking.
Watch a language model's thoughts form before it speaks — interactive workbench + findings for Anthropic's Jacobian lens
Find the layer where a language model commits a decision — and steer it. Any open-weight HF model. (WANDERING arc paper #6)
Local Streamlit app for mechanistic interpretability of transformer models.
Hallucination detector for GigaChat3-10B: internal-state probing (257 features) + LLM-as-judge blend — PR-AUC 0.7955 with 5.7 ms overhead
White-box and behavioural audit of fine-tuned Qwen2.5-7B models with hidden loyalties (Apart Research hackathon, Jul 2026).
Decoding the black box of LLMs: A comparative analysis of Logit Lens vs. Tuned Lens to interpret intermediate Transformer layers in GPT-2.
OMTR — can LLM memorization be separated from predictability? Preregistered causal probing (activation patching) in Pythia & OLMo on consumer hardware; three honest non-separations, full corrections history
A J-space-inspired AI visual art and interpretability playground for watching hidden-state word candidates swarm and collapse into language.
Companion code for the paper "Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens" (arXiv:2609.01936). Pretrained dictionaries: hf.co/hematteo/sparse-readout-prism
🏛️ Champollion cracked hieroglyphs in 1822. I applied the same logic to LLM internals. 95% accuracy, $0 cost, fully reproducible. Contributors welcome.
Local web playground to look inside Qwen3-0.6B: logit lens, attention maps, activation steering, chat, image gen
A small, extensible mechanistic-interpretability lab — logit lens & activation patching on GPT-2 and Qwen3 behind a unified backend adapter. Config-driven, tested, laptop-friendly.
To associate your repository with the logit-lens topic, visit your repo's landing page and select "manage topics."