Preregistered Jacobian-lens reliability research on Gemma 4: 25k prompts, frozen probes, public traces, a prospective transfer miss, and an answer-identity confound.
-
Updated
Sep 11, 2026 - Python
Preregistered Jacobian-lens reliability research on Gemma 4: 25k prompts, frozen probes, public traces, a prospective transfer miss, and an answer-identity confound.
Catch AI code hallucinations—fabricated packages, invented APIs, and contradicted behavior—deterministically, without trusting another model.
Turn Chaos Into Structure. A Type-Safe AI Agent that extracts valid JSON from unstructured data using PydanticAI, FastHTML, and Gemini 2.5.
Interactive Phoenix LiveView demonstrations of the Crucible Framework - showcasing ensemble voting, request hedging, statistical analysis, and more with mock LLMs
Code for the ICML 2026 Main Track Paper "Cognitive Fatigue in Autoregressive Transformers: Formalization and Measurement"
Four tiny local LLMs decide a pixel-art character's mood, pose and dialogue instead of a hand-written state machine. 72% of the arbiter's vetoes cited rules that don't exist, and every test still passed. A small study in valid outputs with invented justifications. Runs fully offline.
Three small LLMs, one CPU, no cloud. Rigorous benchmarking, structured output validation, and head-to-head quality scoring via Ollama + FastAPI.
A Python to Prolog pipeline that bounds Elasticsearch LLM errors by grounding ES diagnostics analysis in verified metrics.
Collection of LLM failure modes used on failmodes.com
Extracts a driver profile from a call transcript behind a validating schema gate, then screens a load board before ranking by effective rate per mile.
TypeScript eval harness for measuring whether Grok answers stay grounded in source evidence
Reference implementation of CAAF — three-pillar agent framework with monotonic convergence.
Companion project for the TechnologyDig Academy tutorial on building reliable generative programs with Mellea.
Official implementation of Cascading Hallucination in Agentic RAG: The CHARM Framework for Detection and Mitigation.
Map where your bolted-on AI feature breaks before customers do. A free Claude Code tool: fragility map, reliability score across six dimensions, ranked gaps, and a 30-day plan. Built by a threat-intel practitioner.
Profile README. AI engineer working on agentic LLM systems, deterministic guardrails, and how models fail in long interactions.
CrucibleFramework: A scientific platform for LLM reliability research on the BEAM
When the agent breaks, which layer dropped the ball? Operator-facing failure-attribution taxonomy for AI agent estates: ten-way dictionary, MAST + AgentRx crosswalks, postmortem template. Part of the Spine catalog.
Public artifact bundle for the preprint 'Lightweight Evaluation and Operational Scorecards for Tool-Using AI Agents'
Make small local LLMs reliable at multi-step agent workflows. Constrained decoding + verify-before-commit: 96.5% end-to-end vs 2.0% unguarded on a 50-step run. Zero dependencies, runs on-prem.
To associate your repository with the llm-reliability topic, visit your repo's landing page and select "manage topics."