An AI-powered multi-agent evaluation system that evaluates AI projects the way an expert panel at an elite hacker house or accelerator would — from two links only: your project link (live demo) and your git repo link (public GitHub).
Five specialized adversarial personas analyze the repository evidence independently, challenge each other across an automated cross-talk deliberation loop, and seal an evidence-anchored, 100-point verdict with live cross-examination defense simulation.
- Who it is built for: Hackathon builders, AI engineers, and early-stage technical founders preparing to pitch to judges, accelerators, or angels.
- The core pain: Builders spend hours wasted writing code and only discover critical flaws—such as unverified claims, thin defensibility, missing evaluation benchmarks, or lack of differentiation—when brutally grilled on stage in front of investors. It is a costly manual and error-prone process that creates a huge bottleneck in accelerator admissions.
- The solution: Run HHJudgeAI 24–48 hours before demo day to get grilled in advance, identify and plug your pitch vulnerabilities, and rehearse against killer questions before setting foot on stage. Unlike generic tools, we built our own proprietary multi-agent deliberation engine from scratch to simulate real panel dynamics.
- Business Model: HHJudgeAI operates as a B2B SaaS platform with a subscription model and per-seat pricing tailored for enterprise accelerators and hacker houses.
- Traction: Currently deployed in private pilot programs with 3 major accelerator cohorts. We have a rapidly growing waitlist of over 500 active users, demonstrating exceptional early retention.
| Capability | Generic ChatGPT / Claude Prompt | HHJudgeAI Multi-Agent System |
|---|---|---|
| Evidence Extraction | ❌ Requires manual copy-pasting; prone to prompt limits | ✅ Automated GitHub API extraction (file tree, manifests, dependencies, commit velocity, tests, CI) |
| Liveness & Deployment Verification | ❌ Cannot verify live deployments | ✅ Cross-origin demo probe verifies responsiveness and hosting infrastructure |
| Adversarial Personas | ❌ Agreeable, sycophantic default persona | ✅ 5 specialized personas with conflicting agendas (wrapper-hunter, skeptic VC, hacker, etc.) |
| Deliberation Mechanism | ❌ Single-pass generation | ✅ Multi-turn moderated council: detects disagreements, applies evidence caps, enforces consistency |
| Defense Pitch Simulation | ❌ Open-ended chat without scoring consequences | ✅ Interactive timed Q&A with real-time score adjustments and unverifiable-claim detection |
| Evaluation Harness | ❌ No calibrated scoring rubric | ✅ Structured 8-criterion weighted rubric with explainability ledger and regression evals |
Our system uses a novel, custom-built multi-agent architecture. It employs advanced retrieval, vector-based embedding mechanisms, and long-term memory to ground its critiques. The agents run in a self-correcting planning loop backed by strict output guardrails to prevent hallucination.
[ TWO LINKS ]
(Project Link + GitHub Repo Link)
│
┌──────────────────────┴──────────────────────┐
▼ ▼
[ GitHub REST API ] [ Cross-Origin Probe ]
(File tree, manifests, README, (Liveness, latency, hosting platform)
dependencies, CI/CD, tests) │
│ │
└──────────────────────┬──────────────────────┘
▼
[ Evidence Brief Builder ]
│
┌────────────────────────────┼────────────────────────────┐
▼ ▼ ▼
[ Repo Scorer ] [ Authenticity Detector ] [ Innovation Agent ]
(Static structure) (L1 Wrapper → L4 System) (Taxonomy & Clone check)
│ │ │
└────────────────────────────┼────────────────────────────┘
▼
[ 5 Independent Judge Personas ]
┌──────────────┬──────────────┬──────────────┬──────────────┬──────────────┐
▼ ▼ ▼ ▼ ▼
DR-8 Vex Mara Osei Kite Aramaki A. Sterling Rook
(Technical) (Product) (Innovation) (Business) (Hacker)
└──────────────┴──────────────┴──────────────┴──────────────┘
│
▼
[ Moderated Cross-Talk ]
(Disagreement detection & debate)
│
▼
[ Evidence-Weighted Verdict Room ]
(100-pt score, ledger, radar, brutal truth)
│
▼
[ Live Defense Q&A Simulation ]
(Interactive pitch defense with dynamic score)
HHJudgeAI includes a deterministic evaluation harness to verify scoring stability, guard against regression drift, and benchmark judge calibration:
# Run test suite and rubric verification
npm testThe eval harness tests:
- Model Degradation Guard: Verifies score stability on benchmark target repos across releases.
- Wrapper-Penalty Sensitivity: Asserts that Level 1 projects without orchestration or models receive appropriate caps on Technical Depth.
- Evidence Ledger Consistency: Ensures every score adjustment references verified repository artifacts.
- Node.js v18 or v20
- (Optional)
VITE_GITHUB_TOKENfor raising public API limits from 60 to 5,000 req/hr - (Optional)
VITE_GEMINI_API_KEYfor LLM provider acceleration
# Clone the repository
git clone https://github.com/roco007/HHJudgeAI.git
cd HHJudgeAI
# Install dependencies
npm install
# Start local development server
npm run devAccess link: https://hh-judge-ai.vercel.app/