Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,7 @@ All notable changes to this project will be documented in this file.
- Grouped surrogate validation holds out whole designs or regimes alongside separate row-validation metrics. Prediction and recommendation expose observed-support diagnostics and warn on extrapolation for GP and RF (#152).
- `run_sequential()` allocates bounded extra replication to unresolved comparisons and feasibility boundaries, preserves raw replicate ids, and reports budgets, stopping reasons and simultaneous finite-horizon mean intervals under explicit bounded-score assumptions (#153).
- Adaptive sessions queue known configurations and import compatible completed observations with persistent identity/provenance and duplicate protection. Versioned session schemas validate reopened journals and refuse unverifiable legacy storage (#154).
- Opt-in `EvaluationCache` reuses grid evaluations by typed configuration, replicate namespace/id, model/scorer revision, fidelity, objective and annotation definitions; includes provenance, bypass, invalidation and conflicting-evidence checks (#154).

## [0.3.0] — 2026-10-02

Expand Down
2 changes: 2 additions & 0 deletions docs/api/io.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,4 +4,6 @@ Serialization and deserialization of study results.

::: trade_study.load_results

::: trade_study.EvaluationCache

::: trade_study.save_results
38 changes: 38 additions & 0 deletions docs/api/runner.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,44 @@ between an external side effect and saving its result may require repeating
that evaluation. Make side effects safe to repeat. Retry limits do not bound
the number of explicit resume invocations.

## Reusing identified evaluations

Reuse is opt-in through an explicit cache:

```python
cache = EvaluationCache(
"evaluations.sqlite",
revision="simulator-v2-scorer-v1-data-v3-annotations-v1",
replicate_namespace="experiment-2026-seed-17",
fidelity="high",
)
results = run_grid(world, scorer, grid, observables, n_reps=3, cache=cache)
```

The key includes a typed canonical configuration, replicate id, caller revision,
replicate namespace, fidelity, inspectable simulator/scorer class code,
objective definitions and annotation semantics. Grid position and total replicate
count are excluded: reordering designs or requesting more replicates can reuse
identical existing evaluations. Raw result rows still retain the current grid's
design-point ids. Different replicate ids and independent experiment namespaces
remain distinct even when their configurations match.

The caller must update the revision when instance settings, external data,
globals or opaque callable behavior changes. The replicate namespace must
identify the actual randomness/seed convention. Use a fresh namespace for an
independent experiment; matching configuration alone never establishes reuse.
Supported scalar/container configurations preserve types; opaque values are
refused. Cached evaluators must not mutate their configuration.

`cache_bypass=True` skips reads and writes, leaving stored evidence untouched.
`cache.clear()` invalidates all contexts in the cache file. Changing identity
creates a new context. Conflicting scores for one identity are refused.
Metadata exposes `cache_hit`, `cache_key`, and the revision/namespace/fidelity
context. Cached wall time is the original evaluation's time. Completed grid
checkpoints can populate the cache without re-evaluation; bypassing the cache
does not bypass checkpoints. Concurrent workers may evaluate an absent identity
more than once before either saves; reuse does not promise exactly-once effects.

::: trade_study.run_grid

::: trade_study.run_adaptive
Expand Down
2 changes: 2 additions & 0 deletions src/trade_study/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@
from ._pareto import extract_front, hypervolume, igd_plus, pareto_rank
from ._scoring import coverage_curve, score
from ._version import __version__
from .cache import EvaluationCache
from .design import (
Factor,
FactorConstraint,
Expand Down Expand Up @@ -57,6 +58,7 @@
"Annotation",
"Constraint",
"Direction",
"EvaluationCache",
"Factor",
"FactorConstraint",
"FactorType",
Expand Down
211 changes: 211 additions & 0 deletions src/trade_study/cache.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,211 @@
"""Opt-in evaluation reuse with explicit revision and replicate identity."""

from __future__ import annotations

import hashlib
import json
import sqlite3
from contextlib import closing
from dataclasses import dataclass
from enum import Enum
from pathlib import Path
from typing import TYPE_CHECKING, Any

import numpy as np

from ._checkpoint import _type_identity
from ._recovery import _grid_identity
from .protocols import TrialResult

if TYPE_CHECKING:
from .protocols import Annotation, Observable, Scorer, Simulator


class EvaluationCache:
"""Persistent immutable evaluation evidence for explicitly identified runs."""

def __init__(
self,
path: str | Path,
*,
revision: str,
replicate_namespace: str,
fidelity: str,
) -> None:
"""Create or reopen an evaluation cache.

Args:
path: SQLite cache file; compatible contexts may share one file.
revision: Caller-managed simulator/scorer/data/annotation revision.
Include every behavior input not identifiable from class code.
replicate_namespace: Explicit randomness/seed experiment identity.
Use a new namespace for genuinely independent replication.
fidelity: Explicit simulation/evaluation fidelity identity.

Raises:
ValueError: If an identity is empty or the cache format is incompatible.
"""
if any(
not value.strip() for value in (revision, replicate_namespace, fidelity)
):
msg = "revision, replicate_namespace and fidelity must be nonempty"
raise ValueError(msg)
self.path = Path(path)
self.revision = revision
self.replicate_namespace = replicate_namespace
self.fidelity = fidelity
self.path.parent.mkdir(parents=True, exist_ok=True)
with closing(sqlite3.connect(self.path, timeout=30)) as connection, connection:
connection.execute(
"CREATE TABLE IF NOT EXISTS cache_format (version INTEGER)"
)
stored = connection.execute("SELECT version FROM cache_format").fetchone()
if stored is None:
connection.execute("INSERT INTO cache_format VALUES (1)")
elif stored[0] != 1:
msg = "Incompatible evaluation cache format"
raise ValueError(msg)
connection.execute(
"CREATE TABLE IF NOT EXISTS evaluations "
"(key TEXT PRIMARY KEY, payload TEXT NOT NULL)"
)

def clear(self) -> None:
"""Invalidate all stored evaluations across every context in this file."""
with closing(sqlite3.connect(self.path, timeout=30)) as connection, connection:
connection.execute("DELETE FROM evaluations")


def _canonical(value: object) -> object:
"""Canonicalize configurations without collapsing opaque or container types.

Returns:
A typed JSON representation with stable dictionary ordering.

Raises:
ValueError: If a value has unsupported opaque behavior.
"""
if isinstance(value, Enum):
return {"enum": _type_identity(value), "value": _canonical(value.value)}
if isinstance(value, np.generic):
return {"numpy": str(value.dtype), "value": _canonical(value.item())}
if type(value) in {type(None), bool, int, float, str}:
return {"type": type(value).__name__, "value": value}
if isinstance(value, (list, tuple)) and type(value) in {list, tuple}:
return {"type": type(value).__name__, "value": [_canonical(v) for v in value]}
if isinstance(value, dict) and type(value) is dict:
return {
"dict": sorted(
(
json.dumps(_canonical(k), sort_keys=True, allow_nan=False),
_canonical(v),
)
for k, v in value.items()
)
}
msg = "Cached configurations require supported JSON/scalar values"
raise ValueError(msg)


@dataclass(frozen=True)
class _BoundCache:
cache: EvaluationCache
identity: str
context: dict[str, str]

def key(self, config: dict[str, Any], rep: int) -> str:
value = json.dumps(
{"context": self.identity, "config": _canonical(config), "rep": rep},
sort_keys=True,
allow_nan=False,
)
return hashlib.sha256(value.encode()).hexdigest()

def load(self, config: dict[str, Any], rep: int) -> TrialResult | None:
key = self.key(config, rep)
with closing(sqlite3.connect(self.cache.path, timeout=30)) as connection:
row = connection.execute(
"SELECT payload FROM evaluations WHERE key = ?", (key,)
).fetchone()
if row is None:
return None
payload = json.loads(row[0])
return TrialResult(
config,
payload["scores"],
payload["wall_seconds"],
{
"rep": rep,
"cache_hit": True,
"cache_key": key,
"cache_context": dict(self.context),
},
)

def save(self, config: dict[str, Any], rep: int, result: TrialResult) -> None:
key = self.key(config, rep)
payload = json.dumps(
{"scores": result.scores, "wall_seconds": result.wall_seconds},
sort_keys=True,
)
with (
closing(sqlite3.connect(self.cache.path, timeout=30)) as connection,
connection,
):
connection.execute(
"INSERT OR IGNORE INTO evaluations VALUES (?, ?)", (key, payload)
)
row = connection.execute(
"SELECT payload FROM evaluations WHERE key = ?", (key,)
).fetchone()
if json.dumps(json.loads(row[0])["scores"], sort_keys=True) != json.dumps(
result.scores, sort_keys=True
):
msg = (
"Conflicting scores for one cached evaluation identity; "
"change the replicate namespace/revision"
)
raise ValueError(msg)
result.metadata.update({
"cache_hit": False,
"cache_key": key,
"cache_context": dict(self.context),
})


def _bind_cache(
cache: EvaluationCache,
world: Simulator,
scorer: Scorer,
observables: list[Observable],
annotations: list[Annotation] | None,
) -> _BoundCache:
identity = json.dumps(
{
"definition": _grid_identity(
world, scorer, [], observables, annotations, 1, cache.revision
),
"replicate_namespace": cache.replicate_namespace,
"fidelity": cache.fidelity,
"annotation_lookups": [
_canonical(a.lookup)
if isinstance(a.lookup, dict)
else {
"module": getattr(a.lookup, "__module__", None),
"name": getattr(a.lookup, "__qualname__", None),
}
for a in (annotations or [])
],
},
sort_keys=True,
)
return _BoundCache(
cache,
identity,
{
"revision": cache.revision,
"replicate_namespace": cache.replicate_namespace,
"fidelity": cache.fidelity,
"schema_digest": hashlib.sha256(identity.encode()).hexdigest(),
},
)
Loading
Loading