Skip to content

M2: atomic reserve-before-work quota check - #25

Merged
man4ish merged 1 commit into
mainfrom
feature/v1-atomic-quota-reservation
Oct 3, 2026
Merged

man4ish merged 1 commit into
mainfrom
feature/v1-atomic-quota-reservation

Conversation

@man4ish

@man4ish man4ish commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator

Summary

Part of M2 (make quota enforcement real and race-free). The old quota check was check-then-later-decrement: quota_remaining() read the counter before forwarding to RAG, consume_quota() only decremented it after a successful response came back. Any number of concurrent requests could all observe "quota available" before any of them decremented, and all of them would succeed — overrunning a near-zero quota by however many were in flight at once.

  • reserve_quota() — atomic Redis DECR, compensated back (INCR) if it would go negative. Only as many concurrent reservations as there are units left can ever succeed; unmetered orgs (no quota key) are unaffected.
  • release_quota() — refunds a reservation when the forwarded request then fails upstream, since the reservation optimistically assumed success before knowing the outcome.
  • Both wired into the one shared _handle_billable_literature_call path /literature/answers and /literature/search already go through, so this closes the race for both endpoints at once.

Test plan

  • pytest — 355 passed, including new coverage for concurrent-reservation exhaustion (exactly one of two requests over a 1-unit quota succeeds, never both) and quota refund on upstream failure
  • ruff check . — clean

🤖 Generated with Claude Code

Part of M2. The old quota check was check-then-later-decrement:
quota_remaining() read the counter before forwarding to RAG,
consume_quota() only decremented it after a successful response. Any
number of concurrent requests could all observe "quota available"
before any of them decremented, and all of them would succeed --
overrunning a near-zero quota by however many were in flight at once.

Replaces that pair with reserve_quota() (atomic Redis DECR,
compensated back if it would go negative -- only as many concurrent
reservations as there are units left can ever succeed) called before
forwarding to RAG, and release_quota() (refund) called if the
forwarded request then fails upstream, since the reservation assumed
success before knowing the outcome. Both are wired into the one
shared _handle_billable_literature_call path /literature/answers and
/literature/search both already go through, so this closes the race
for both at once.

Tests: 355 passed, including new coverage for concurrent-reservation
exhaustion (exactly one of two requests over a 1-unit quota succeeds,
never both) and quota refund on upstream failure. ruff clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@man4ish
man4ish merged commit 045e2ab into main Oct 3, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant