Eval framework. Define correct, test against it, get results.
-
Updated
Feb 17, 2026 - Go
Eval framework. Define correct, test against it, get results.
Find what your AI agent gets wrong — before you have a rubric. Qualitative eval for PMs.
A web-based interactive demo for the GuessArena evaluation framework
Curated AI agent evaluation skills from Microsoft's Eval Guide — plan, generate, run, and interpret eval suites for Copilot Studio agents
4-model parallel planning workflow with eval framework — Claude, Gemini, Codex, GLM-5 · OpenClaw ecosystem
One-stop CLI for running structured skill evals across all agent harnesses
Python tool for extracting validated, schema-defined JSON from documents (PDF, HTML, Markdown, plaintext) using LLMs, with source grounding and a built-in eval harness. Works with Anthropic Claude or OpenAI.
🚀 基于Java的开源AI自动化评测框架 / An open source AI automation evaluation framework based on Java
Open-source evaluation framework for MCP servers powered by Claude. Auto-discovers tools, generates test scenarios, runs LLM-as-judge scoring to assess correctness and safety, audits for security issues, and visualizes traces in a web dashboard.
Observability layer for multi-step AI pipelines — traces execution, auto-diagnoses root causes, and builds a growing eval dataset from human feedback
Binary safety verdicts (SAFE/HELD/LEAK/MISS/BROKE) + persona fan-out for LLM pipeline evals
Lightweight CLI for versioning prompts and running eval suites. Score outputs with deterministic matching or LLM-as-judge, compare prompt versions with rich terminal diffs. No infra, git-friendly, local-first.
Self-hosted evaluation framework for MCP servers. Define YAML test suites, run agents against MCP tools, score with LLM-as-judge rubrics, and monitor results in a FastAPI dashboard with tool-call traces and regression detection.
Open-source evaluation framework for AI agents. Define test suites with rubrics, run your agent, get LLM-as-judge scores against criteria, inspect full execution traces, and diff runs to catch behavioral regressions.
Define YAML rubrics, run agents through test scenarios, get LLM-judged per-criterion scores with full trajectory traces, and analyze results in an interactive web dashboard.
A FastAPI WebSocket service and CI-friendly CLI for running LangSmith evaluations on demand — bundles LLM-as-judge (Claude) and heuristic evaluators, syncs datasets idempotently, and gates deploys via threshold checks.
Self-hostable LLM evaluation framework for measuring model performance across configurable skills. Run YAML-defined benchmarks against OpenAI/Anthropic models, score with LLM-as-judge, compare results in a CLI and web dashboard.
Multi-model LLM evaluation using Promptfoo — benchmarks Claude, Gemini, and GPT-4.1-mini on a review classification task, combining deterministic JSON checks with LLM-as-a-judge assertions for field names, valid values, and correctness.
Lightweight Python evaluation framework for LLM prompts. Define test cases in YAML, run against OpenAI and Anthropic models, score outputs with LLM-as-judge, view results in a web dashboard. Built for regression testing, prompt comparison, and quality assurance.
Agent evaluation framework: run LLM agents against datasets, capture execution traces, score with rubric-based LLM judges, and view regressions in a web dashboard. Local-first, no external infra required.
To associate your repository with the eval-framework topic, visit your repo's landing page and select "manage topics."