Skip to content

M8: enforce /v1 rate limits per organization, not just per key - #27

Merged
man4ish merged 1 commit into
mainfrom
feature/v1-per-org-rate-limiting
Oct 4, 2026
Merged

man4ish merged 1 commit into
mainfrom
feature/v1-per-org-rate-limiting

Conversation

@man4ish

@man4ish man4ish commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator

Summary

Part of M8 (design audit gap #7's third bullet: "it is enforced per caller/key, not both per key and per organisation"). An organization holding several API keys previously got an independent rate-limit budget per key — N keys meant N times the effective limit, since each key's counter was tracked in total isolation.

_rate_limited now hits two counters per request: the existing per-subject (API key or user session) one, and a new one keyed by organization_id. Both share the same limit value (the plan-aware override from M7, or the global default); the request is rejected if either is exhausted. The response headers report whichever counter is actually binding (the smaller remaining count), so a key blocked purely by its organization's exhausted shared budget sees 0 remaining, not its own still-fresh per-key count.

Test plan

  • pytest — 363 passed
  • ruff check . — clean
  • New coverage: two different keys in the same organization share one budget, the organization counter doesn't leak across organizations, headers reflect the binding counter correctly

🤖 Generated with Claude Code

Part of M8 (design audit gap #7's third bullet: "it is enforced per
caller/key, not both per key and per organisation"). An organization
holding several API keys previously got an independent rate-limit
budget per key -- N keys meant N times the effective limit, since
each key's counter was tracked in total isolation.

_rate_limited now hits two counters per request: the existing
per-subject (API key or user session) one, and a new one keyed by
organization_id. Both share the same limit value (the plan-aware
override from the companion M7 work, or the global default); the
request is rejected if either is exhausted. The response headers
report whichever counter is actually binding (the smaller remaining
count), so a key blocked purely by its organization's exhausted
shared budget sees 0 remaining, not its own still-fresh per-key count.

Tests: 363 passed. ruff clean. New coverage: two different keys in the
same organization share one budget (a key with zero requests of its
own is still blocked once a sibling key exhausts the shared budget),
the organization counter doesn't leak across organizations, and the
reported headers reflect the binding counter correctly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@man4ish
man4ish merged commit d3dc3e9 into main Oct 4, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant