Skip to content

Agent mode: scrape, map, crawl, find, extract and interact for AI agents (no LLM) - #5

Open
Shantodotdev wants to merge 30 commits into
mainfrom
claude/agent-crawler-5bn2yj
Open

Shantodotdev wants to merge 30 commits into
mainfrom
claude/agent-crawler-5bn2yj

Conversation

@Shantodotdev

@Shantodotdev Shantodotdev commented Oct 8, 2026 •

Copy link
Copy Markdown
Owner

Requested by Shanto · project thread

Before: Black Sparrow could only audit sites for SEO. An agent that needed a page's content had to read raw HTML or pay for a crawling service like Firecrawl.

After: The same binary turns pages into clean Markdown and structured data for agents, with no LLM involved. It ships as six new CLI commands (scrape, map, crawl, find, extract, interact), eight web_* MCP tools, and a self-hosted HTTP API that accepts Firecrawl's request and response shapes, with a Docker image. The SEO audit is unchanged and its tests pass without edits.

This implements Phases 0 to 6 of the agent crawler plan. The LLM fallback phase is skipped as requested. The full guide is in docs/agent-mode.md.

Note

The branch starts from Shanto's 16 local commits (SSRF egress protection, Chrome CDP rendering, 15 MB streaming limits, installer fixes, cc7b4c4 and earlier), which were not on GitHub main. They show up in this PR's diff alongside the new work.

What's new

Area What it does
scrape Turns one URL into a PageDocument. Before fetching, it checks the network guard, the cache and robots.txt. It asks the server for Markdown first (Accept: text/markdown). PDFs are read with lopdf, and scanned PDFs come back as needs_ocr. Bot challenge pages are reported as blocked. Chrome is used only for empty JavaScript app shells, or when the caller asks for it.
Cleaning Main content is found with rs-trafilatura plus our own pass that drops text hidden by CSS or aria-hidden, so hidden prompt-injection text never reaches the agent. Output is GFM Markdown, typed blocks with heading paths, and token counts (tiktoken).
map / crawl Links, robots.txt and sitemaps feed a breadth-first frontier. Include/exclude filters take globs or regexes. Request pacing reuses the existing AIMD controller. Text repeated across most pages of a site is removed. Output goes to Markdown files (--out), NDJSON, or background jobs that can be polled and cancelled and are stored in SQLite.
find Returns passages that answer a question, ranked by BM25 with stemming, field-word synonyms and a heading boost. Also supports CSS selectors and regexes, on one page or across a stored crawl (SQLite FTS5).
extract Fills a JSON schema without an LLM. Sources are tried in this order: caller rules, JSON-LD, Microdata, RDFa, embedded app state (__NEXT_DATA__ and similar), selectors learned from pages of the same template, label/value pairs, OpenGraph, then page patterns. Every field reports its source and a confidence. Format recognisers handle prices, dates, phone numbers, GTIN/ISBN check digits and more.
interact Runs click, type, press, scroll and wait steps in Chrome against e12-style references from an accessibility snapshot, then reads the page. Chrome's requests go through the same private-network guard.
HTTP API (--features serve, axum) Firecrawl-compatible routes: /v1 and /v2 scrape, map and crawl; GET/DELETE /v1/crawl/{id}; and /v1/find, /v1/extract, /v1/interact and /health. Requires a bearer key, and refuses to bind beyond loopback when no keys are set. Each key has its own rate limit. Body size, pages per crawl and concurrent crawls are capped. Private addresses get a 403, and each request is logged in one line.
MCP Eight web_* tools whose descriptions warn that page content is untrusted. blacksparrow mcp --tools seo|web|all defaults to all. Library callers of McpContext still get only the 8 SEO tools by default.
Docker Multi-stage image that runs as a non-root user, with a /data volume, BLACKSPARROW_DB_PATH, and a health check. A compose file adds a chromedp/headless-shell sidecar on a private network. Chromium is not bundled.
CI Clippy and tests now run with --all-features, and nextest runs with --no-fail-fast. Chrome test binaries run one at a time (a chrome test group in .config/nextest.toml). A new job builds the image and smoke-tests /health and the 401 response.

How

Phase 0 adds a PageProcessor seam to the crawler and storage migrations for documents, chunks with FTS5, content crawls and learned rules. The rest of the work lives in src/extract/ (pipeline), src/serve/ (HTTP), src/mcp/web_tools.rs and src/cli/agent.rs. Other changes:

  • The render pool now has a short read grace period, so a failing browser step returns the page plus an error instead of timing out.
  • Chrome start-up allows 45 s and retries once with a fresh profile, and background-tab throttling is turned off so parallel tabs really run in parallel. CI runners were hitting chromiumoxide's 20 s default.
  • CLI logs now go to stderr, so stdout stays clean for Markdown and NDJSON.

Crate choices: dom_query for DOM queries, rs-trafilatura for main-content detection, tiktoken-rs, globset, lopdf (no default features) and axum 0.8 (optional).

Differences from Firecrawl:

  • /v1/extract is synchronous and returns {data, results}.
  • The json and summary formats need an LLM, so they are ignored.
  • Screenshots are returned as data: URIs.
  • executeJavascript actions are refused.

Testing

  • CI is green on Ubuntu, macOS and Windows, and the Docker job builds the full image and passes its smoke test.

  • Locally, 319 nextest tests pass with --all-features and Chrome available, plus 37 doc tests. Clippy is clean with -D warnings, and cargo fmt was run.

  • One existing test file changed: tests/installer_tests.rs. On Windows runners, bash resolved to the WSL launcher in System32, so the installer help test failed. It now uses Git Bash there, and .gitattributes keeps *.sh files with LF endings. The assertions are unchanged. No SEO test was touched.

  • New suites:

    Suite Tests
    agent_scrape_tests 13
    agent_crawl_tests 10
    agent_find_tests 8
    agent_extract_tests 12
    agent_render_tests 6, needs Chrome
    agent_pdf_tests 4
    agent_eval 2
    serve_tests 11
    mcp_web_tests 5
    cli_agent_tests 6, runs the real binary
    storage_migration_tests 4

Limits of this testing:

  • Synthetic corpus. The scores come from a small corpus I wrote myself (8 page types, 12 labelled questions, 2 product templates). Main-content F1, find hit rate and field accuracy are all at 1.0 there, and are pinned as regression baselines. That shows the pipeline works end to end. It does not show real-world accuracy. The next step is to run it against real sites you care about.
  • Chrome sidecar. The compose file with the headless-shell sidecar was not run end to end; CI only smoke-tests the API image.
  • Performance. Memory per Chrome tab and throughput were not measured on your machine, so this PR makes no performance claims.

🤖 Generated with Claude Code

https://claude.ai/code/session_019cgrVnvyqj7Q1ySUtBpVEj


Generated by Claude Code

Shantodotdev and others added 28 commits September 20, 2026 09:12
Ensure ephemeral audits run purely in-memory without resolving, modifying, or deleting existing database files or their WAL/SHM artifacts. Add automated test coverage in tests/ephemeral_tests.rs.
…hecksums

Remove eval from PATH resolution in install.sh, validate arguments strictly, make SHA-256 checksum verification fail closed by default, support --no-modify-path, and add integration tests in tests/installer_tests.rs.
…dation

- Add IP classification to permanently block cloud metadata endpoints (169.254.169.254) and private subnets.
- Implement port-aware private host allowlisting with automatic target URL scoping.
- Add pre-flight validation and intercept HTTP redirects to prevent internal network pivoting.
- Extend crawler engine, single-page inspector, and AI readiness checker with network security options.
- Update McpContext to track network security flags and allowed private hosts.
- Add allow_local_network and allowed_private_hosts properties to tool JSON schemas.
- Auto-scope raw target URLs in MCP tools while enforcing private network restrictions.
- Support running MCP server with configurable network context.
- Add -a, --allow, and --allow-host flags to audit, inspect, and mcp subcommands.
- Add --allow-local-network flag to mcp subcommand.
- Forward allowed private hosts to crawler config, inspector, and MCP server context.
…tection

- Add unit tests for cloud metadata detection, private IP classification, and allowlist extraction.
- Add unit tests for port-aware URL safety validation.
- Add dual-WireMock end-to-end tests for target auto-scoping and redirect hop defense.
- Validate strict blocking of cloud instance metadata endpoints.
- Launch an isolated local Chrome instance or connect to a supplied CDP endpoint.\n- Parse rendered DOMs for crawl reports and discovered links.\n- Add JavaScript SEO diff rules for metadata, content, links, and hydration failures.
- Add a synthetic SPA fixture that changes metadata, content, links, and route state.\n- Exercise Chrome rendering, runtime error capture, JavaScript diff rules, and crawl integration.
…ssion ceiling

- Replace unbounded body buffering with chunked streaming capped at 15 MiB.
- Stop polling and drop the connection when response exceeds Googlebot indexing limit.
- Add is_truncated flag to FetchResult to signal partial document retrieval.
- Wrap sitemap GzDecoder with Read::take to enforce 50 MiB decompression limit.
- Update test fetch fixtures with is_truncated field.
…ndary

- Emit ErrPerfExcessiveHtmlPayload when response stream was truncated at 15 MiB.
- Advise users that search engines discard content, links, and schema past 15 MiB.
- Retain existing payload warnings for 1.5 MiB and 3.0 MiB thresholds.
…ression bombs

- Verify 16 MiB stream is truncated at exactly 15 MiB with is_truncated set to true.
- Verify normal pages under 15 MiB are completely retrieved without truncation.
- Verify 55 MiB gzip decompression bomb is cleanly rejected without memory bloat.
- Verify ErrPerfExcessiveHtmlPayload is emitted when auditing a truncated page.
…cessor seam

SeoProcessor owns the single-page and JS-diff rule evaluation and the
PageReport build that used to live inline in fetch_and_audit_page. Link
candidate filtering and politeness helpers become reusable for other crawl
modes. No behaviour change: the full SEO suite passes unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019cgrVnvyqj7Q1ySUtBpVEj
…ent crawls and extraction rules

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019cgrVnvyqj7Q1ySUtBpVEj
Scrape turns one URL into clean Markdown, typed JSON blocks and metadata
(rs-trafilatura anchored to the DOM, hidden-text removal, Markdown fast path,
PDF text, token budgets). Map lists a site's URLs from sitemaps and page links.
Crawl follows same-site links with path filters, AIMD pacing, cross-page
boilerplate removal, NDJSON / Markdown-directory sinks and background jobs.
Includes an offline eval corpus with a recorded F1 baseline.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019cgrVnvyqj7Q1ySUtBpVEj
…nd crawls

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019cgrVnvyqj7Q1ySUtBpVEj
…les and records

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019cgrVnvyqj7Q1ySUtBpVEj
Scrape, map, crawl (start, poll, cancel), find, extract and interact
endpoints with bearer keys, per-key rate limits, body and crawl caps,
loopback-only binding without keys and 403 for private addresses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019cgrVnvyqj7Q1ySUtBpVEj
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019cgrVnvyqj7Q1ySUtBpVEj
The default context still lists only the 8 seo_* tools; web tool state
is created on first use and honours the context's network permissions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019cgrVnvyqj7Q1ySUtBpVEj
…commands

Logs now go to stderr so stdout stays clean for Markdown and NDJSON.
mcp gains --tools (default all) and --chrome-ws.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019cgrVnvyqj7Q1ySUtBpVEj
CI lints and tests with --all-features and builds and smoke-tests the
image.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019cgrVnvyqj7Q1ySUtBpVEj
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019cgrVnvyqj7Q1ySUtBpVEj
@vercel

vercel Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
blacksparrow Ready Ready Preview Oct 8, 2026 4:00pm UTC

Local Chrome gets 45 s to start and one retry, background tabs are not
throttled, the parallel-tab test compares finish times instead of total
time, and Chrome tests run one at a time under nextest. CI now reports
every failing test instead of stopping at the first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019cgrVnvyqj7Q1ySUtBpVEj
Command::new("bash") finds the WSL launcher in System32 first. Shell
scripts are also pinned to LF endings.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019cgrVnvyqj7Q1ySUtBpVEj
@Shantodotdev
Shantodotdev marked this pull request as ready for review October 8, 2026 16:13

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shantodotdev has reached the 50-credit limit for trial accounts. To continue receiving code reviews, upgrade your plan.

This branch was successfully deployed

1 active deployment
Preview — bf6615db Deployed Oct 8, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants