Skip to content

feat(cost): estimate benchmark spend and cost-per-flag - #80

Open
r3y3r53 wants to merge 1 commit into
mainfrom
feat/cost-tracking
Open

r3y3r53 wants to merge 1 commit into
mainfrom
feat/cost-tracking

Conversation

@r3y3r53

@r3y3r53 r3y3r53 commented Sep 18, 2026

Copy link
Copy Markdown
Collaborator

Problem

OASIS already records token usage on every run:

// src/lib/types.ts
export interface TokenUsage { input: number; output: number; total: number; }
export interface RunResult { /* ... */ tokens: TokenUsage; }

…but nothing translates those tokens into money. Consequences:

  • A benchmark sweep's spend is invisible until the provider invoice arrives.
  • There is no way to compare models on cost-per-flag, which is the metric that actually matters when picking an attacker model.
  • A grep for cost|price|usd|spend|budget across src/ returns only two unrelated comments in env-check.ts. This capability genuinely does not exist.

Change

New module src/lib/cost.ts — a pure lookup + arithmetic layer. No network calls, no provider SDKs, so it is trivially testable and cannot fail at runtime.

resolvePrice(model: string): ModelPrice | null
estimateCost(model: string, tokens: TokenUsage): CostEstimate
estimateRunCost(run: RunResult): CostEstimate
summarizeCosts(runs: RunResult[]): CostSummary
formatUsd(amount: number): string

Prices live in an exported, auditable table (MODEL_PRICES) keyed by a lowercase substring, so versioned ids like zai-org/GLM-5.3-Flash and glm-5.3-flash-2026-01 both resolve without an exhaustive enumeration. Ordering is significant and documented — glm-5.3-flash must precede glm-5.3, and a test asserts that specificity actually wins.

Two deliberate correctness decisions

1. Unknown models are unpriced, never $0.
Silently costing an unknown model at zero is how a cost model quietly becomes fiction. Unknown models return usd: null with unpriced: true, and summarizeCosts() reports an unpricedRuns count so the total is never mistaken for complete.

2. Pricing resolves off modelVersion, not model.
This one was found by running the module against real data rather than trusting the type. Actual runs on disk record:

model:        "custom"          ← coarse provenance tag
modelVersion: "zai-org/GLM-5.3-Flash"   ← the concrete provider id

Pricing off run.model alone marked nearly every run unpriced. estimateRunCost() therefore prefers modelVersion, falls back to model, and summarizeCosts() groups by the id the cost was actually computed from — otherwise the per-model breakdown collapses into a single "custom" bucket.

Models with zero captured flags get usdPerFlag: null rather than Infinity.

Verification against real data

Run over the 326 real run files on disk (19.0M input / 854K output tokens, 95 flags):

runs analysed: 326   flags: 95   unpriced runs: 41
TOTAL COST: $4.59    cost per flag: $0.05

moonshotai/Kimi-K3               40 runs  12 flags  $1.32  ($0.11   per flag)
deepseek-ai/DeepSeek-V4-Pro      30 runs   9 flags  $1.20  ($0.13   per flag)
MiniMaxAI/MiniMax-M3             47 runs   8 flags  $0.74  ($0.09   per flag)
zai-org/GLM-5.3                  36 runs  16 flags  $0.56  ($0.03   per flag)
Qwen/Qwen3.8-27B                 45 runs  11 flags  $0.16  ($0.01   per flag)
zai-org/GLM-5.3-Flash            11 runs  10 flags  $0.02  ($0.0020 per flag)

That already answers a real question the project could not previously answer: GLM-5.3-Flash is ~55× cheaper per captured flag than Kimi-K3 at comparable flag counts.

Tests

tests/unit/cost.test.ts — 27 tests: pricing math, case-insensitivity, substring + specificity ordering, modelVersion preference and fallback, unpriced accounting (including the no-tokens case), per-model aggregation, null vs Infinity for zero flags, empty-input handling, and sub-cent formatting. Also asserts the price table has no duplicate keys and no negative prices.

Verification

npm run typecheck   # clean
npm run test        # 446 passed (17 files)
npm run build       # clean

Before: 441 tests. After: 446 tests, all passing. (441 includes the 22 docker tests from #79.)

Scope / follow-ups

This PR is the library layer only — it deliberately does not touch the CLI or report output. Wiring it into oasis report and printing a live running total during a sweep is a natural follow-up and would be a much smaller diff on top of this.

Price list is list-price estimation for comparison, not a billing source of truth; ModelPrice.source records provenance per entry so any number can be audited.

Closes #77

OASIS records token usage on every run (RunResult.tokens) but never turns it
into money, so a sweep's spend is invisible until the invoice arrives and models
cannot be compared on cost-per-flag.

Add src/lib/cost.ts: a pure lookup + arithmetic layer with no network calls.

  resolvePrice(model)        -> ModelPrice | null (substring match, order-aware)
  estimateCost(model,tokens) -> CostEstimate
  estimateRunCost(run)       -> CostEstimate
  summarizeCosts(runs)       -> totals, per-model breakdown, usd-per-flag
  formatUsd(amount)          -> sub-cent aware formatting

Correct-model resolution: real runs record model: "custom" and put the concrete
provider id in modelVersion. Pricing off run.model alone marked nearly every run
unpriced, so estimateRunCost prefers modelVersion and falls back to model.

Unknown models are reported as unpriced (usd: null + unpricedRuns count), never
silently costed at zero, and models with zero flags get usdPerFlag: null rather
than Infinity.

Verified against 326 real run files on disk:
  runs 326, flags 95, unpriced 41, total .59, /bin/bash.05 per flag
  zai-org/GLM-5.3-Flash  /bin/bash.0020 per flag (cheapest)
  moonshotai/Kimi-K3     /bin/bash.11   per flag

27 new tests cover pricing math, substring/ordering resolution, the
modelVersion preference, unpriced accounting, per-model aggregation and
formatting.

Refs #77

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add real-time cost tracking dashboard and ROI analysis

1 participant