diff --git a/AGENTS.md b/AGENTS.md index 1d329eca..76ada0f3 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -43,18 +43,45 @@ readonly QUALITY_MIN_CHARS=500 Detailed reference material in `agents-docs/`: -- [Development](agents-docs/DEVELOPMENT.md) -- [Configuration](agents-docs/CONFIG.md) -- [Overview](agents-docs/OVERVIEW.md) -- [Semantic Health](agents-docs/SEMANTIC_HEALTH.md) +| Document | Path | Description | +|---|---|---| +| Assets | `agents-docs/ASSETS.md` | Visual assets and screenshots | +| Configuration | `agents-docs/CONFIG.md` | Configuration guide | +| Dependabot Auto-Merge SOP | `agents-docs/DEPENDABOT_AUTO_MERGE_SOP.md` | Dependabot auto-merge verification runbook | +| Deployment | `agents-docs/DEPLOYMENT.md` | Deployment guide | +| Development | `agents-docs/DEVELOPMENT.md` | Development guide | +| Known Issues | `agents-docs/ISSUES.md` | Known issues and audit findings | +| Overview | `agents-docs/OVERVIEW.md` | Project overview | +| Reference Index | `agents-docs/README.md` | Index for this directory | +| Releases | `agents-docs/RELEASES.md` | Release process | +| Semantic Health | `agents-docs/SEMANTIC_HEALTH.md` | Current semantic health summary | +| Semantic Health — June 2026 | `agents-docs/SEMANTIC_HEALTH_2026_06.md` | June 2026 archive | +| Semantic Health — July 2026 | `agents-docs/SEMANTIC_HEALTH_JULY_2026.md` | July 2026 archive | +| Semantic Health — August 2026 | `agents-docs/SEMANTIC_HEALTH_AUG_2026.md` | August 2026 archive | +| Semantic Health — August 2026 Summary | `agents-docs/SEMANTIC_HEALTH_AUG_2026_SUMMARY.md` | August 2026 analysis and issue summary | +| Semantic Health Issue | `agents-docs/SEMANTIC_HEALTH_ISSUE.md` | Cache hit latency and telemetry investigation | +| Semantic Health Summary | `agents-docs/semantic_health_summary.md` | August 2026 summary (legacy filename) | ## Skills -- `do-web-doc-resolver`: `.agents/skills/do-web-doc-resolver/` -- `do-wdr-cli`: `.agents/skills/do-wdr-cli/` -- `anti-ai-slop`: `.agents/skills/anti-ai-slop/` -- `readme-best-practices`: `.agents/skills/readme-best-practices/` -- `skill-creator`: `.agents/skills/skill-creator/` +| Skill | Path | Description | +|---|---|---| +| `agent-browser` | `.agents/skills/agent-browser/` | Browser automation for navigating, filling forms, and screenshots | +| `anti-ai-slop` | `.agents/skills/anti-ai-slop/` | Audit and fix UI, UX, and copy that reads as generic "AI slop" | +| `codacy` | `.agents/skills/codacy/` | Query Codacy analysis, triage issues, suppress false positives | +| `do-github-pr-sentinel` | `.agents/skills/do-github-pr-sentinel/` | Monitor a PR until merged, green, or blocked; diagnose and retry CI | +| `do-wdr-assets` | `.agents/skills/do-wdr-assets/` | Capture screenshots and visual assets for documentation | +| `do-wdr-cli` | `.agents/skills/do-wdr-cli/` | Use the compiled `do-wdr` CLI to resolve URLs and queries | +| `do-wdr-issue-impl` | `.agents/skills/do-wdr-issue-impl/` | Implement a single GitHub issue through to a merged PR | +| `do-wdr-issue-swarm` | `.agents/skills/do-wdr-issue-swarm/` | Batch-implement GitHub issues with parallel specialist agents | +| `do-wdr-release` | `.agents/skills/do-wdr-release/` | Manage releases, versioning, changelogs, and tags | +| `do-wdr-ui-component` | `.agents/skills/do-wdr-ui-component/` | Build CSS-only components for the `cli/ui` design system | +| `do-wdr-visual-resolver` | `.agents/skills/do-wdr-visual-resolver/` | Resolve scanned PDFs and JS-heavy SPAs via CLIP embeddings | +| `do-web-doc-resolver` | `.agents/skills/do-web-doc-resolver/` | Python resolver: full cascade, quality scoring, circuit breakers | +| `privacy-first` | `.agents/skills/privacy-first/` | Keep email addresses and personal data out of the codebase | +| `readme-best-practices` | `.agents/skills/readme-best-practices/` | Create and audit README files against 2026 best practices | +| `skill-creator` | `.agents/skills/skill-creator/` | Create, edit, and benchmark agent skills | +| `vercel-cli` | `.agents/skills/vercel-cli/` | Deploy and manage projects on Vercel from the command line | ## Coding Workflow @@ -66,7 +93,7 @@ Detailed reference material in `agents-docs/`: ### PR Checklist - Quality gate command passes: `./scripts/quality_gate.sh` -- Linting clean (`ruff`, `black`, `cargo fmt`, `cargo clippy`, `npm run lint`) +- Lint clean (`ruff`, `black`, `cargo fmt`, `cargo clippy`, `npm run lint`) - No new secrets added - `AGENTS.md` updated if repository structure or skills change diff --git a/README.md b/README.md index dcd076c8..f79c74d2 100644 --- a/README.md +++ b/README.md @@ -18,15 +18,39 @@ --- +## Table of Contents + +- [Overview](#overview) +- [What the Cascade Is](#what-the-cascade-is) + - [Query Cascade Order](#query-cascade-order) + - [URL Cascade Order](#url-cascade-order) +- [How to Install](#how-to-install) + - [Python](#python) + - [Rust CLI (`do-wdr`)](#rust-cli-do-wdr) + - [Web UI](#web-ui) +- [How to Run](#how-to-run) + - [Python CLI](#python-cli) + - [Python Module](#python-module) + - [Rust CLI (`do-wdr`)](#rust-cli-do-wdr-1) + - [Web UI](#web-ui-1) +- [Environment Variables Required](#environment-variables-required) +- [How to Run Tests](#how-to-run-tests) + - [Python Test Suite](#python-test-suite) + - [Rust Test Suite](#rust-test-suite) + - [Web UI Playwright Tests](#web-ui-playwright-tests) + - [Quality Gate Script](#quality-gate-script) + +--- + ## Overview -`do-web-doc-resolver` fetches web pages and executes search queries, stripping boilerplate and formatting the output into token-dense Markdown for LLM prompt context. It executes an execution cascade across free and paid providers, falling back automatically if a provider fails or returns low-density content. +`do-web-doc-resolver` fetches web pages and executes search queries, stripping HTML boilerplate and returning token-dense Markdown for LLM context windows. It evaluates results against content density and quality thresholds, executing a tiered fallback cascade across free and paid providers. --- ## What the Cascade Is -The resolution engine queries providers in tiered priority order, returning upon the first result that satisfies quality thresholds. +The resolver queries providers in priority tiers and halts upon receiving a response that satisfies quality scoring thresholds. ### Query Cascade Order @@ -36,7 +60,7 @@ The resolution engine queries providers in tiered priority order, returning upon ### URL Cascade Order -1. **Semantic Cache**: Pre-cached document lookup. +1. **Semantic Cache**: Local vector lookup. 2. **Free Static Tier**: `llms.txt` discovery. 3. **Free Direct & Lite Tier**: Direct HTTP fetch, Jina Reader, Firecrawl. 4. **Browser Tier**: Mistral Browser. @@ -57,7 +81,7 @@ pip install -r requirements.txt ### Rust CLI (`do-wdr`) -Requires Rust 1.80+. +Requires Rust 1.80 or higher. ```bash cd cli @@ -66,7 +90,7 @@ cargo build --release ### Web UI -Requires Node.js 18+. +Requires Node.js 18 or higher. ```bash cd web @@ -79,6 +103,8 @@ npm install --legacy-peer-deps ### Python CLI +Pass a target URL or query as the positional argument: + ```bash python -m scripts.cli "https://docs.python.org/3/" python -m scripts.cli "python asyncio taskgroup example" @@ -112,11 +138,11 @@ npm run dev ## Environment Variables Required -No API keys are required for zero-config operation using free providers. Optional API keys enable additional paid provider tiers: +Zero-configuration mode runs using free providers without API keys. Optional environment variables enable paid provider tiers: | Environment Variable | Provider | Required | Description | |---|---|---|---| -| `EXA_API_KEY` | Exa SDK | No | Enables Exa search and extraction | +| `EXA_API_KEY` | Exa SDK | No | Enables Exa search and content extraction | | `TAVILY_API_KEY` | Tavily | No | Enables Tavily web search | | `SERPER_API_KEY` | Serper | No | Enables Google Search via Serper | | `FIRECRAWL_API_KEY` | Firecrawl | No | Enables Firecrawl scraping |