DSD Harness is a multi-model research harness for DSD-style hidden-state interpretation experiments.
It packages the current shared runner for hidden-state verbalization and oracle-style comparisons across several open-weight model families, while keeping older exploratory scripts and synthetic-training experiments available in clearer subdirectories.
This repository currently focuses on:
- hidden-state verbalization (
interpret) - cross-sentence same-word vector-difference explanation (
INTERPRET_VECTOR_MODE=direct_sub) - projected/vector-diff explanation (
INTERPRET_VECTOR_MODE=projected_vdiff) - two-vector conceptual comparison (
INTERPRET_VECTOR_MODE=two_vector_compare) - layer-to-layer direct vector difference explanation (
INTERPRET_VECTOR_MODE=layer_direct_diff) - causal / oracle / drift / random comparisons (
frozen_probe) - probe fitting against oracle-style targets (
train_probe) - prototype / metaphor pair comparisons (
pair_qualitative) - runtime patching for row-wise oracle masking and model-specific compatibility
The primary public entrypoints are:
run_key_interpretation.pyuniversal_patcher.py
Practical status right now:
- Qwen — strongest current baseline
- Gemma — usable and often the next best qualitative path, but still prompt/example sensitive
- Llama — partially supported; quality varies and higher-layer outputs can collapse to generic assistant text
- GPT-OSS — partially supported; shell/runtime stability still needs work
- Gemma4 — engineering path is connected, but still below stable paper-grade quality
This means the repository is public and usable, but still an actively cleaned-up research harness rather than a finalized benchmark package.
See the usage docs for the full setup flow:
docs/usage/install.mddocs/usage/model-setup.mddocs/usage/experiment-modes.mddocs/usage/examples.mddocs/usage/status-and-limitations.mddocs/usage/repository-layout.md
Minimal example:
MODEL_TYPE=qwen \
PROMPT_MODE=summary \
EXPERIMENT_MODE=interpret \
TEST_LAYERS=0,10,20 \
PROMPT_TEXT='The key looked ordinary until the final clue revealed what it unlocked.' \
python run_key_interpretation.pyrun_key_interpretation.py currently dispatches the following modes:
interpret— verbalize a hidden state at selected layersfrozen_probe— comparecausal,oracle,drift, andrandomtargetstrain_probe— fit a probe from causal states to oracle-style targetspair_qualitative— compare prototype/metaphor pair behaviors from shared prefixes
The drift target uses a vdiff-style construction that removes the causal-direction projection from the oracle vector.
The repository is organized so the public root stays small:
run_key_interpretation.py— main harness entrypointuniversal_patcher.py— patching and masking utilitiesdocs/usage/— setup and usage docsexamples/— smoke tests, validation scripts, and prompt-comparison helpersexperiments/— synthetic-training, corpus-building, and sweep scriptslegacy/— historical or one-off scripts kept for reference but not treated as the main supported path
This code was originally developed against a private server layout, and some environment/path assumptions are still being generalized.
Before treating this as a polished external release, expect continued cleanup in:
- dependency packaging
- model-loading defaults
- path/config parameterization
- curated examples
- head-level analysis interfaces
License is not finalized yet.