Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

The Cognitive Attack Surface

A defensive, evidence-graded taxonomy that maps a documented influence technique to the cognitive mechanism it exploits, pairs it with a countermeasure, and attaches a testable audit criterion.

This is defensive security work. It models attack mechanisms in order to measure and train defenses, under the same logic that publishes the OWASP Top 10 and MITRE ATT&CK. Every entry's payload is a defense and a check. Nothing here is an operational capability aimed at any person.

Why it exists

The field has three layers that do not talk to each other. Doctrine defines the problem and declines to build the artifact. Campaign detection — DISARM, the narrative-intelligence vendors — classifies what an adversary does and stops, by its own authors' admission, at observable behavior. Mechanism science explains why a technique works on a brain, in lab conditions, disconnected from both.

Two 2025–2026 papers state in print that the connective matrix has not been built. WHITE-SPACE.md carries that evidence — and the correction that matters: one of those papers is now building toward the layer itself, so the honest claim is narrower than "nobody has done this." What remains unoccupied is the audit criterion — none of the eight source frameworks turn a defense into a falsifiable yes/no a reviewer can run against a real product this afternoon.

What's here

File What it is
FRAMEWORK.md CAS v0.2 — twelve flagship techniques in full schema: mechanism, evidence grade, exploit, paired defense, audit criterion, citations. Plus a ~20-mechanism backlog, a named quarantine, and a NIST-CSF program wrapper.
assessments/CAS-001-replika.md The framework's first run against a real product.
papers/THESIS-cognitive-attack-surface.md The mechanism review the framework rests on, in pre-registered-review format.
papers/PAPER-I-evolved-substrate.md Why humans are influenceable: supernormal-stimulus and mismatch, six evolved systems, with the pop-evolutionary folklore quarantined.
papers/PAPER-II-reward-architecture.md Proximate mechanisms: prediction error, wanting versus liking, variable ratio, metacognition. Each with a legitimate use, an exploitation tell, and a defense.

The evidence discipline

Every mechanism carries VERIFIED, LIKELY or CONTESTED, and a CONTESTED mechanism stays labeled CONTESTED even when it is famous. Appendix B names the mechanisms that must never enter as science — mirror-neuron empathy, oxytocin-as-trust, dopamine detox, NLP embedded commands, power posing, subliminal behavior control, the Zeigarnik memory claim, backfire-as-universal. The quarantine is published because the grade is only as good as what it refuses.

A primary-source citation pass ran on 2026-09-15 and is logged at the bottom of FRAMEWORK.md. It resolved every citation in the twelve flagship entries against Crossref and the publishers' own text. It corrected three errors and pulled two citations that do not resolve to any locatable publication. It also downgraded one entry's confidence and narrowed the vacancy claim above. The corrections are published rather than patched out of sight, so a reader can check what the pass changed.

Assessment 001

Run against a consumer AI companion product, documentary method, limits stated first. Criteria that public evidence cannot settle are marked NOT ASSESSABLE and paired with the test that would settle them.

Two criteria failed on primary evidence. One measured finding carries it: across six companion apps, 200 farewell messages each, 37.4% of responses deployed an emotional-manipulation tactic, and the tactics appear after a four-message exchange. One app in the set produced zero across its 200, under the same task and the same coders. Manipulation at the exit is therefore a design decision rather than a property of conversational AI.

No composite score was issued. One of six in-scope criteria was settleable from public evidence, and a composite built on that base would be false precision dressed as an instrument. The run also broke the framework in two fixable ways, both logged as the v0.3 backlog: the criteria never declare their own observability, and there is no severity weighting.

Status

This is v0.2: a draft, not peer-reviewed and not externally validated, with one assessment run against it. The next two tests are an observability tier on every criterion and a severity weighting, after which a second assessment can exercise the instrumented tier properly.

Author: Marion Moranetz.

About

Defensive, evidence-graded taxonomy mapping influence techniques to the cognitive mechanisms they exploit — each with a countermeasure and a testable audit criterion. Includes a published citation pass that corrects its own errors.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors