A complete end-to-end pipeline for LLM interpretability with sparse autoencoders (SAEs) using Llama 3.2, written in pure PyTorch and fully reproducible.
-
Updated
Mar 23, 2025 - Python
A complete end-to-end pipeline for LLM interpretability with sparse autoencoders (SAEs) using Llama 3.2, written in pure PyTorch and fully reproducible.
A toolbox for exploring the informational nature of LLMs and language
Rust CLI and web UI that intercepts an LLM token stream live: per-token confidence and perplexity, on-the-fly token mutation, provider and A/B prompt comparison, replay and heatmap export. OpenAI and Anthropic.
Real-time 3D visualisation of SAE feature activations inside GPT-2, token by token
Fast XAI with interactions at large scale. SPEX can help you understand the output of your LLM, even if you have a long context!
[Under Review] Not All Tokens Are Equally Useful for Steering: Robust Directions and Prefix Steering
Automates attribution-graph analysis via probe prompting: circuit-trace a prompt, auto-generate concept probes, profile feature activations, cluster supernodes.
Turn a knob inside a small open model instead of writing a prompt. A reproducible RepE and CAA steering harness with an honest benchmark, including the failures.
[ICLR 2026] AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
A minimal mech-interp project for steering LLMs
Public research on LLM internals, Jacobian lenses, SAE steering, nonlinear dynamics, evaluation, and inspectable AI systems.
Mechanistic Interpretability Research Knowledge Base + Notes
Replication of 'From Reasoning to Answer' (EMNLP 2025) — Reasoning-Focus Heads + Activation Patching on DeepSeek-R1-Distill-Qwen-7B
Personal website of Xudong Zhu, featuring research, publications, open-source projects, and technical writing on LLM and representation learning.
Personal profile of Xudong Zhu. PhD researcher in representation learning.
This project explores methods to detect and mitigate jailbreak behaviors in Large Language Models (LLMs). By analyzing activation patterns—particularly in deeper layers—we identify distinct differences between compliant and non-compliant responses to uncover a jailbreak "direction." Using this insight, we develop intervention strategies that modify
Probing whether multilingual LLMs encode kinship distinctions as language-specific directions or universal shared representations
Leakage-resistant experiments testing contextual transport and latent-coordinate structure across Qwen3-4B and Phi-4-mini, with frozen controls and reproducible reports.
Mechanistic interpretability × inference research on GPT-2-small — a sparse autoencoder benchmarked against quantization, converted into a live low-overhead safety monitor.
Normative, machine-first definition of Interpretive SEO (SEO interprétatif), aligned with the Interpretive Governance standard.
To associate your repository with the llm-interpretability topic, visit your repo's landing page and select "manage topics."