Lighthouse for agents - score an agent run, then tell it how to get faster and cheaper.
-
Updated
Jun 20, 2026 - TypeScript
Lighthouse for agents - score an agent run, then tell it how to get faster and cheaper.
Quality and trigger-evaluation gate for reusable AI agent skills.
Pre-Execution Gate for AI Code. A deterministic, gradient-immune structural guard against reward hacking and hardcoding in RL training loops.
AI Agent Root Cause Analysis Protocol — A systematic five-question methodology adapted from Toyota's 5 Whys for diagnosing failures in AI agent systems. Battle-tested over 60+ production incidents.
Benchmarking the ability of large language models to detect semantic conflicts across domains, documents, and evolving knowledge bases.
LLM Agent quality metrics — structured recording and quality threshold testing for Function Calling agents
Professional cross-agent answer quality gate for improving AI responses: intent match, evidence, assumptions, verification, brevity, and usefulness.
SAGE — Skills for Agent-Grade Evaluation. Two Claude Code skills: assess data before you build, measure quality after. From 3 real projects. 智能体评估方法论:前置数据评估+后置质量评估。
Deterministic spec-driven development CLI
Field-tested QA validation gates for AI agent systems. Tiered gates, protocol gates, severity classification, and automated checks. Born from production failures.
Quality gate plugin for coding agents with automated tests, PASS/WARN/FAIL policies, and agent feedback.
Add a description, image, and links to the agent-quality topic page so that developers can more easily learn about it.
To associate your repository with the agent-quality topic, visit your repo's landing page and select "manage topics."