compute_attacks does not finish within 180 seconds on a small generated ASPIC+ theory: 8 literals, 4 ordinary premises, 11 strict rules and 3 defeasible rules. build_arguments turns this theory into 18,850 arguments in about 0.5s, and 18,496 of those arguments conclude p. The property test tests/structured/aspic/test_aspic.py::TestAttackProperties::test_rebutting_targets_defeasible_conclusions generated the theory. Hypothesis saved it in a local .hypothesis database; after that, the test hit the 600s pytest timeout on every run, and pytest then raised a MemoryError while formatting the traceback. Main CI passes this test because CI starts with an empty database.
Theory (shrunk by Hypothesis)
- Language:
p, q, r, s and their negations, with each atom and its negation contradictory.
- Ordinary premises:
q, s, ~p, ~r. There are no axioms.
- Strict rules:
p -> ~s
q, ~r -> s
q, ~s -> r
s -> ~p
~p, ~q -> q
~q, ~p -> q
~q, ~q -> p
~r, q -> s
~r, ~s -> ~q
~s, q -> r
~s, ~r -> ~q
- Defeasible rules:
r, q => ~r (named d0)
s => q (named d3)
~p => ~s (named d1)
Reproduction
Run with the repository development environment (uv run python repro.py), with an external timeout as the bound. At ad3c01c plus the #96 branch:
import time
from argumentation.structured.aspic.aspic import (
ArgumentationSystem, ContrarinessFn, GroundAtom, KnowledgeBase, Literal, Rule,
build_arguments, compute_attacks,
)
P, Q, R, S = (Literal(GroundAtom(n)) for n in "pqrs")
NP, NQ, NR, NS = (a.contrary for a in (P, Q, R, S))
strict = frozenset(Rule(b, h, "strict") for b, h in [
((P,), NS), ((Q, NR), S), ((Q, NS), R), ((S,), NP), ((NP, NQ), Q), ((NQ, NP), Q),
((NQ, NQ), P), ((NR, Q), S), ((NR, NS), NQ), ((NS, Q), R), ((NS, NR), NQ)])
defeasible = frozenset({Rule((R, Q), NR, "defeasible", "d0"),
Rule((S,), Q, "defeasible", "d3"),
Rule((NP,), NS, "defeasible", "d1")})
system = ArgumentationSystem(frozenset({P, Q, R, S, NP, NQ, NR, NS}),
ContrarinessFn(frozenset((a, a.contrary) for a in (P, Q, R, S))), strict, defeasible)
kb = KnowledgeBase(axioms=frozenset(), premises=frozenset({Q, S, NP, NR}))
t = time.perf_counter(); args = build_arguments(system, kb)
print(len(args), "arguments", round(time.perf_counter() - t, 2), "s")
compute_attacks(args, system) # does not return within 180 s
Operational evidence (AGENTS.md)
build_arguments: 18,850 arguments in 0.40–0.54s.
- Arguments per conclusion:
p 18,496, r 144, ~q 136, s 31, q 18, ~r 17, ~p 4, ~s 4.
- Rebut attack points (a sub-argument whose top rule is defeasible): 98,849. Rebutting attack pairs: 8,687,712. This counts rebuttals only; undermining pairs are extra.
compute_attacks still had not returned when killed at 180s.
- Profile: py-spy, 30s at 50 Hz on the real interpreter rather than the uv launcher, 1,499 samples. The top leaf frames are
compute_attacks at aspic.py:1159–1164, the loop that adds one attack_keys entry per (attack point, candidate attacker) pair, with Argument.__hash__ (aspic.py:249) behind it. The cost is the size of the attack relation itself, not one slow step.
The underlying shape is argument multiplication. Cyclic strict rules, including ones with a repeated antecedent (~q, ~q -> p) and mirrored bodies (~p, ~q -> q / ~q, ~p -> q), give p 18,496 structurally distinct arguments. Each rebuttable sub-argument then pairs with every argument for the contrary conclusion.
Expected
The finite theory should be decided within a bounded budget, or the test generators should not produce theories whose argument sets explode. Which of those, if either, is a design decision for this issue. A useful first operational contract would be to bound len(build_arguments(...)) and the attack-pair count on this theory, so the regression fails fast instead of timing out.
No production change is included in this report.
compute_attacksdoes not finish within 180 seconds on a small generated ASPIC+ theory: 8 literals, 4 ordinary premises, 11 strict rules and 3 defeasible rules.build_argumentsturns this theory into 18,850 arguments in about 0.5s, and 18,496 of those arguments concludep. The property testtests/structured/aspic/test_aspic.py::TestAttackProperties::test_rebutting_targets_defeasible_conclusionsgenerated the theory. Hypothesis saved it in a local.hypothesisdatabase; after that, the test hit the 600s pytest timeout on every run, and pytest then raised a MemoryError while formatting the traceback. Main CI passes this test because CI starts with an empty database.Theory (shrunk by Hypothesis)
p, q, r, sand their negations, with each atom and its negation contradictory.q, s, ~p, ~r. There are no axioms.p -> ~sq, ~r -> sq, ~s -> rs -> ~p~p, ~q -> q~q, ~p -> q~q, ~q -> p~r, q -> s~r, ~s -> ~q~s, q -> r~s, ~r -> ~qr, q => ~r(namedd0)s => q(namedd3)~p => ~s(namedd1)Reproduction
Run with the repository development environment (
uv run python repro.py), with an external timeout as the bound. At ad3c01c plus the #96 branch:Operational evidence (AGENTS.md)
build_arguments: 18,850 arguments in 0.40–0.54s.p18,496,r144,~q136,s31,q18,~r17,~p4,~s4.compute_attacksstill had not returned when killed at 180s.compute_attacksat aspic.py:1159–1164, the loop that adds oneattack_keysentry per (attack point, candidate attacker) pair, withArgument.__hash__(aspic.py:249) behind it. The cost is the size of the attack relation itself, not one slow step.The underlying shape is argument multiplication. Cyclic strict rules, including ones with a repeated antecedent (
~q, ~q -> p) and mirrored bodies (~p, ~q -> q/~q, ~p -> q), givep18,496 structurally distinct arguments. Each rebuttable sub-argument then pairs with every argument for the contrary conclusion.Expected
The finite theory should be decided within a bounded budget, or the test generators should not produce theories whose argument sets explode. Which of those, if either, is a design decision for this issue. A useful first operational contract would be to bound
len(build_arguments(...))and the attack-pair count on this theory, so the regression fails fast instead of timing out.No production change is included in this report.