Track 20,710+ AI benchmark, eval, dataset, and data-quality records from 39 public sources, with linked evidence and daily updates.
-
Updated
Oct 11, 2026 - Python
Track 20,710+ AI benchmark, eval, dataset, and data-quality records from 39 public sources, with linked evidence and daily updates.
MyPhoneBench: Do Phone-Use Agents Respect Your Privacy?
20 runnable LLM agent design patterns in Python, benchmarked, traced, offline-compatible. ReAct, multi-agent, constitutional AI, and more.
The same AI agent pipeline built in Mastra and LangChain. Runs in parallel, measures everything.
splyntra
SnorkelAI Terminus2 task suite with deterministic evaluators, Dockerized environments, and reproducible agent-benchmarking primitives.
Deterministic multi-lane evaluation domain for SnorkelAI Terminus3. Cockpit, lane orchestration, agent benchmarking, reasoning-drift detection.
Correctness-gated measurement of mutation, edit payload, and semantic revisit across agent trajectories.
A Claude Agent SDK security benchmark project
To associate your repository with the agent-benchmarking topic, visit your repo's landing page and select "manage topics."