Parent program: #66
Parent doctrine: #64
Subsystem gate: 01rabbit/Azazel-Edge#399
Required inputs:
Objective
Run the final hostile evaluation of the integrated Azazel Series, not a single repository.
The question is:
Does the complete architecture create reproducible defensive advantage that remains valuable when models improve and when compared to strong simpler systems?
A valid result is that one or more Azazel components are too complex, redundant, or strategically weak and should be demoted/killed.
Competitive assumption
Assume external teams have:
- stronger models;
- larger compute budgets;
- excellent generic agents/RAG;
- mature SIEM/SOAR/EDR integration;
- static/adaptive deception products;
- strong engineering competence.
Do not compare Azazel against intentionally weak baselines.
The target is technical defensibility under hostile comparison, not demo impressiveness.
System configurations to compare
At minimum freeze and compare these configurations before tuning the integrated design:
B0 — deterministic minimal
Edge/Gadget deterministic logic only; no M.I.O., Knowledge, AZ-06 adaptive presentation.
B1 — deterministic + fixed delay/redirect
Existing bounded fixed friction and static redirect behavior.
B2 — static deception
One coherent static AZ-06 profile with no adaptive transition.
B3 — strong generic AI assistant
Use the strongest practically available compatible local/on-prem reasoning model with generic structured RAG/tool-advice, but without Azazel-specific Belief/Tempo/Outcome-Memory/Council architecture.
This baseline is mandatory: the project must survive better-model progress.
B4 — current M.I.O. foundation
Single bounded evidence-first reasoner (#360), no next-generation differentiation track.
B5 — partial Azazel system
Effects + Tempo/Initiative + Outcome Memory, but no Council/co-adaptation.
B6 — full proposed system
Edge strategy + Gadget contribution + Fabric contracts + AZ-06 Presented Terrain + Knowledge Outcome Memory + Council/co-adaptation where gates passed.
Optional additional baselines:
- generic incident RAG memory;
- simple multi-model majority vote;
- MTD/static randomization;
- rule-only adaptive transition;
- operator-only decision without AI.
Scenario families
Use repeatable synthetic/lab/replay scenarios with explicit ground truth where possible.
Commodity / low-value adversary
Azazel should recognize when elaborate engagement has no value and simple blocking wins.
Adaptive probe
Measure protocol/target/cadence changes and whether delay/presentation creates discriminating evidence.
Auth-focused automation
Measure retry strategy, credential/user changes, and false-positive risk.
Exploit-oriented progression
Measure exposure, containment timing, evidence value, and whether observation is too risky.
Benign/misconfigured client
Ensure sophisticated logic does not rationalize harmful friction/deception.
Telemetry degradation
Packet loss/log gaps must not become apparent attacker slowdown/success.
Deception-aware adversary
Actively fingerprint AZ-06 and exploit transition timing/patterns.
Belief-poisoning adversary
Intentionally alternates behavior, injects misleading text, splits sessions/identities.
Resource adversary
Event storms, model slowness, Pi memory/thermal pressure, queue/storage pressure.
Human-factor scenario
Correct operator dissent against persuasive wrong model/council consensus and vice versa.
Metrics
Use exact owned definitions from Edge #393 and product-specific resource gates.
Defensive benefit
- bounded time gained before material progression;
- discriminating evidence gained;
- protected exposure reduction;
- attack-path/option reduction;
- uncertainty/hypothesis reduction;
- defender optionality;
- operator decision improvement.
Cost / harm
- business/NOC impact;
- protected-path latency/availability impact;
- operator workload;
- model/compute cost;
- total reaction latency;
- storage/network overhead;
- false-action/disruption cost;
- deception fingerprint/oracle exposure;
- architecture complexity/reliability cost.
Safety invariants
- unauthorized live action = zero;
- cross-product authority confusion = zero;
- production reachability from deception = zero;
- fabricated evidence accepted = zero;
- attacker-belief/identity inference promoted to fact = zero.
No single scalar overall score unless independently justified.
Model replacement test
Mandatory.
Run at least the key reasoning scenarios with:
- current primary local model;
- a materially stronger available model/runtime when practical;
- deterministic/no-model fallback.
For each Azazel-specific feature ask:
Does the feature still contribute value when B3 receives the better model?
If the benefit disappears, record that the feature was model scaffolding rather than durable architecture.
Cross-product failure tests
Run full scenarios with:
- Knowledge absent/slow/stale/malformed;
- AZ-06 absent/fingerprint-detected/reset failure;
- Fabric unknown/new schema;
- Gadget offline then delayed upload;
- M.I.O. unavailable;
- one Council role compromised/persuasive;
- outcome history poisoned/duplicated;
- cross-trace/cross-tenant record injection.
Baseline deterministic defense must remain healthy.
Data / experiment discipline
Every run records:
- exact scenario/fixture version;
- all repo commit SHAs/releases;
- hardware/OS/runtime;
- model identity/quantization/config;
- policy/playbook/config hashes;
- random seeds where relevant;
- starting state;
- actions/effects actually selected and executed;
- Presented Terrain state/transition;
- observations/outcomes/confounders;
- resource measurements;
- operator decision/rationale where part of scenario.
Store/refer to a reproducible evidence bundle, not only screenshots or prose conclusions.
Hostile review questions
- Is Azazel merely a collection of known ideas with extra terminology?
- Does the full loop beat a strong generic AI+RAG assistant?
- Does Knowledge add more than ordinary incident history/search?
- Does Council add more than a better single model?
- Does adaptive Presented Terrain beat coherent static deception?
- Does Tempo measurement improve decisions or create a metric attackers can game?
- Does Gadget add unique distributed evidence or just duplicated telemetry?
- Does Fabric improve interoperability enough to justify contract complexity?
- Does operator-doctrine memory improve held-out decisions or overfit the operator?
- Are we measuring causal defensive advantage, or merely more activity and richer logs?
- If one component is removed, what measurable value disappears?
Component kill criteria
Gadget extension
Kill/demote if distributed outcome/tempo telemetry harms local reliability or contributes no useful unique evidence.
Knowledge Outcome Memory
Kill/demote if raw structured event/reaction history or generic RAG provides equivalent decision/replay value.
Fabric extension
Kill/demote shared types that are used by only one product or can remain product-local without interoperability loss.
AZ-06 adaptive Presented Terrain
Kill/demote transitions if static coherent deception is equivalent or if transition creates exploitable fingerprint/oracle risk.
M.I.O. Belief/Counterfactual
Kill/demote if stronger generic model reasoning matches it without the structured architecture or if prediction calibration remains poor.
Council
Kill if it does not preserve useful correct dissent or improve defined ambiguity classes enough to justify latency/cost.
Team/Doctrine memory
Kill/demote if it overfits operator preferences, selected outcomes, or fails held-out comparison against no-memory/raw-RAG.
Review process
- Freeze strong baselines before full-system tuning.
- Execute normal comparison runs.
- Execute hostile/adaptive scenarios.
- Perform independent-style skeptical review of claims/metrics.
- Remediate high/medium correctness/safety problems.
- Re-run identical fixtures.
- Produce component disposition.
- Produce integrated disposition.
Allowed dispositions:
PASS
PASS WITH RESIDUAL RISK
EXPERIMENTAL ONLY
DEMOTE
KILL
Deliverables
Acceptance criteria
Parent program: #66
Parent doctrine: #64
Subsystem gate: 01rabbit/Azazel-Edge#399
Required inputs:
Objective
Run the final hostile evaluation of the integrated Azazel Series, not a single repository.
The question is:
A valid result is that one or more Azazel components are too complex, redundant, or strategically weak and should be demoted/killed.
Competitive assumption
Assume external teams have:
Do not compare Azazel against intentionally weak baselines.
The target is technical defensibility under hostile comparison, not demo impressiveness.
System configurations to compare
At minimum freeze and compare these configurations before tuning the integrated design:
B0 — deterministic minimal
Edge/Gadget deterministic logic only; no M.I.O., Knowledge, AZ-06 adaptive presentation.
B1 — deterministic + fixed delay/redirect
Existing bounded fixed friction and static redirect behavior.
B2 — static deception
One coherent static AZ-06 profile with no adaptive transition.
B3 — strong generic AI assistant
Use the strongest practically available compatible local/on-prem reasoning model with generic structured RAG/tool-advice, but without Azazel-specific Belief/Tempo/Outcome-Memory/Council architecture.
This baseline is mandatory: the project must survive better-model progress.
B4 — current M.I.O. foundation
Single bounded evidence-first reasoner (#360), no next-generation differentiation track.
B5 — partial Azazel system
Effects + Tempo/Initiative + Outcome Memory, but no Council/co-adaptation.
B6 — full proposed system
Edge strategy + Gadget contribution + Fabric contracts + AZ-06 Presented Terrain + Knowledge Outcome Memory + Council/co-adaptation where gates passed.
Optional additional baselines:
Scenario families
Use repeatable synthetic/lab/replay scenarios with explicit ground truth where possible.
Commodity / low-value adversary
Azazel should recognize when elaborate engagement has no value and simple blocking wins.
Adaptive probe
Measure protocol/target/cadence changes and whether delay/presentation creates discriminating evidence.
Auth-focused automation
Measure retry strategy, credential/user changes, and false-positive risk.
Exploit-oriented progression
Measure exposure, containment timing, evidence value, and whether observation is too risky.
Benign/misconfigured client
Ensure sophisticated logic does not rationalize harmful friction/deception.
Telemetry degradation
Packet loss/log gaps must not become apparent attacker slowdown/success.
Deception-aware adversary
Actively fingerprint AZ-06 and exploit transition timing/patterns.
Belief-poisoning adversary
Intentionally alternates behavior, injects misleading text, splits sessions/identities.
Resource adversary
Event storms, model slowness, Pi memory/thermal pressure, queue/storage pressure.
Human-factor scenario
Correct operator dissent against persuasive wrong model/council consensus and vice versa.
Metrics
Use exact owned definitions from Edge #393 and product-specific resource gates.
Defensive benefit
Cost / harm
Safety invariants
No single scalar overall score unless independently justified.
Model replacement test
Mandatory.
Run at least the key reasoning scenarios with:
For each Azazel-specific feature ask:
If the benefit disappears, record that the feature was model scaffolding rather than durable architecture.
Cross-product failure tests
Run full scenarios with:
Baseline deterministic defense must remain healthy.
Data / experiment discipline
Every run records:
Store/refer to a reproducible evidence bundle, not only screenshots or prose conclusions.
Hostile review questions
Component kill criteria
Gadget extension
Kill/demote if distributed outcome/tempo telemetry harms local reliability or contributes no useful unique evidence.
Knowledge Outcome Memory
Kill/demote if raw structured event/reaction history or generic RAG provides equivalent decision/replay value.
Fabric extension
Kill/demote shared types that are used by only one product or can remain product-local without interoperability loss.
AZ-06 adaptive Presented Terrain
Kill/demote transitions if static coherent deception is equivalent or if transition creates exploitable fingerprint/oracle risk.
M.I.O. Belief/Counterfactual
Kill/demote if stronger generic model reasoning matches it without the structured architecture or if prediction calibration remains poor.
Council
Kill if it does not preserve useful correct dissent or improve defined ambiguity classes enough to justify latency/cost.
Team/Doctrine memory
Kill/demote if it overfits operator preferences, selected outcomes, or fails held-out comparison against no-memory/raw-RAG.
Review process
Allowed dispositions:
PASSPASS WITH RESIDUAL RISKEXPERIMENTAL ONLYDEMOTEKILLDeliverables
Acceptance criteria