Skip to content

feat: token-bucket rate limiting and concurrent-answer limits - #28

Merged
man4ish merged 1 commit into
mainfrom
feature/m12-token-bucket-concurrency-limits
Oct 4, 2026
Merged

man4ish merged 1 commit into
mainfrom
feature/m12-token-bucket-concurrency-limits

Conversation

@man4ish

@man4ish man4ish commented Oct 4, 2026

Copy link
Copy Markdown
Collaborator

Summary

Closes the last two bullets of design-audit gap #7:

  • Token bucket, not a fixed window. V1Store.hit_rate_limit now models a continuously-refilling bucket (capacity = the configured/plan-aware limit, refill rate = limit/60 tokens/sec) instead of a hard reset at each minute boundary -- the old window let a caller spend its full budget in the last second of one window and again in the first second of the next, for up to 2x the limit in under two seconds. Not atomic against a concurrent request for the same subject on a real Redis backend (plain GET/SET, no Lua script) -- a deliberate, documented, bounded tradeoff; nothing here is billing-critical the way reserve_quota's atomic DECR had to be.
  • Concurrent-answer limits. A new V1Store.acquire_concurrency_slot/release_concurrency_slot pair caps in-flight /v1/literature/answers calls per key and per organization (V1_MAX_CONCURRENT_ANSWERS, default 5, operator-tunable -- an overload-protection default, not a per-plan number). Enforced only on /answers (the expensive, LLM-invoking call the gap specifically names), not /search.

No omnibioai-billing change needed -- unlike the rate limit's plan-aware override, concurrency has no per-plan column; this is gateway-only overload protection.

Test plan

  • pytest tests/test_v1_literature.py -- 56 passed (49 existing + 7 new: token-bucket refill behavior, limit-of-zero edge case, concurrency-cap denial for both key and org, slot release on success/quota-denial, and confirming /search has no concurrency cap)
  • Full repo suite -- 370 passed

🤖 Generated with Claude Code

Replaces the fixed one-minute rate-limit window with a token bucket
(continuous refill, same capacity as the configured/plan-aware limit) --
closes the old window's double-burst-across-a-boundary gap. Also adds a
concurrency cap on in-flight /v1/literature/answers calls, enforced per
key and per organization the same way the rate limiter already is,
independent of (and alongside) request-frequency limiting. Both close
the two remaining bullets of design audit gap #7.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@man4ish
man4ish merged commit e630567 into main Oct 4, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant