Skip to content

Repository files navigation

DSD Harness

DSD Harness is a multi-model research harness for DSD-style hidden-state interpretation experiments.

It packages the current shared runner for hidden-state verbalization and oracle-style comparisons across several open-weight model families, while keeping older exploratory scripts and synthetic-training experiments available in clearer subdirectories.

What this repository is for

This repository currently focuses on:

  • hidden-state verbalization (interpret)
  • cross-sentence same-word vector-difference explanation (INTERPRET_VECTOR_MODE=direct_sub)
  • projected/vector-diff explanation (INTERPRET_VECTOR_MODE=projected_vdiff)
  • two-vector conceptual comparison (INTERPRET_VECTOR_MODE=two_vector_compare)
  • layer-to-layer direct vector difference explanation (INTERPRET_VECTOR_MODE=layer_direct_diff)
  • causal / oracle / drift / random comparisons (frozen_probe)
  • probe fitting against oracle-style targets (train_probe)
  • prototype / metaphor pair comparisons (pair_qualitative)
  • runtime patching for row-wise oracle masking and model-specific compatibility

The primary public entrypoints are:

  • run_key_interpretation.py
  • universal_patcher.py

Current model status

Practical status right now:

  • Qwen — strongest current baseline
  • Gemma — usable and often the next best qualitative path, but still prompt/example sensitive
  • Llama — partially supported; quality varies and higher-layer outputs can collapse to generic assistant text
  • GPT-OSS — partially supported; shell/runtime stability still needs work
  • Gemma4 — engineering path is connected, but still below stable paper-grade quality

This means the repository is public and usable, but still an actively cleaned-up research harness rather than a finalized benchmark package.

Quickstart

See the usage docs for the full setup flow:

Minimal example:

MODEL_TYPE=qwen \
PROMPT_MODE=summary \
EXPERIMENT_MODE=interpret \
TEST_LAYERS=0,10,20 \
PROMPT_TEXT='The key looked ordinary until the final clue revealed what it unlocked.' \
python run_key_interpretation.py

Main experiment modes

run_key_interpretation.py currently dispatches the following modes:

  • interpret — verbalize a hidden state at selected layers
  • frozen_probe — compare causal, oracle, drift, and random targets
  • train_probe — fit a probe from causal states to oracle-style targets
  • pair_qualitative — compare prototype/metaphor pair behaviors from shared prefixes

The drift target uses a vdiff-style construction that removes the causal-direction projection from the oracle vector.

Repository layout

The repository is organized so the public root stays small:

  • run_key_interpretation.py — main harness entrypoint
  • universal_patcher.py — patching and masking utilities
  • docs/usage/ — setup and usage docs
  • examples/ — smoke tests, validation scripts, and prompt-comparison helpers
  • experiments/ — synthetic-training, corpus-building, and sweep scripts
  • legacy/ — historical or one-off scripts kept for reference but not treated as the main supported path

Notes for open-source use

This code was originally developed against a private server layout, and some environment/path assumptions are still being generalized.

Before treating this as a polished external release, expect continued cleanup in:

  • dependency packaging
  • model-loading defaults
  • path/config parameterization
  • curated examples
  • head-level analysis interfaces

License

License is not finalized yet.

About

No description, website, or topics provided.

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages