Pivotal Token Search
-
Updated
Sep 24, 2026 - Python
Pivotal Token Search
Adversarial Manipulation of CoT
Analysed determinism, faithfulness, reasoning patterns, & steering. Developed and tested methods to enhance control and fail-safes
mech-interp suite for Granite4 models that use Mamba-2 architecture
All code, stimuli, and results for a mechanistic interpretability study investigating how large language models internally represent emotional content
Implementation and analysis of Sparse Autoencoders for neural network interpretability research. Features interactive visualization dashboard and W&B integration.
Detect safety degradation during LLM fine-tuning before it becomes behavioral
Collection and learnings of my journey in Artificial Intelligence
Official implementation of the 'Uncovering Competency Gaps in Large Language Models and Their Benchmarks' paper
Investigated JEPA-WM predicted futures, specifically how activation steering can impact robotics task planning and outcomes. Performed a series of ablations to test hypotheses about CEM action planning in latent space for world models.
Local agent-driven mechanistic interpretability research platform for Apple Silicon
Unofficial implementation to reproduce the experiments from "Superposition as a Phase Change" of "Toy Models of Superposition".
A re-implementation of the "Refusal in Language Models is Meditated by a Single Direction" to understand mech interp
Cross-architecture mechanistic interpretability toolkit — first OSS Mamba SSM state extraction. Works on transformer + SSM + hybrid models with unified API.
To associate your repository with the mech-interp topic, visit your repo's landing page and select "manage topics."