Repository navigation
Apply web_fetch's max_bytes to the extracted text, not the markup - #54
Conversation
`max_bytes` cut the raw HTML before the extractor ran, so lowering it could destroy the page it was meant to bound. Measured: a 432,864-byte client-rendered document fetched with max_bytes 50000 was cut mid-DOM, and the readability pass recovered only its <title> — `extracted=48B_of_432864B`, reported as `status=200`. The same URL with no max_bytes yields `extracted=37243B_of_433638B`. Identical tool, identical extractor, 776x the content, one parameter. A caller setting a sensible cost bound therefore lost the document silently, and the knob it would then reach for — raising the cap — was the right knob turned too timidly, which is the worst case for learning anything. Convert first, bound second. Bounding the output is what the schema promises, and it costs no memory: `resp.text()` already materialises the whole body, so max_bytes never bounded a download. Only the extractor's input needs a ceiling, and that is for CPU — EXTRACTOR_INPUT_CEILING, 8 MiB, ~19x the largest real page seen. The ordering is lifted into `render_body` so it is directly testable, including a counterfactual that the old ordering loses the prose from the same document. Contract changes: the `max_bytes` schema description now says which side of the extractor it bounds, and the header reports `output_capped_at=` instead of `download_capped_at=` — nothing here ever capped a download, and the old name sent readers looking in the wrong place. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Tiny Sweeper reviewTiny Sweeper completed its review; deterministic results follow. State: Ready for maintainer review Review snapshot
Completeness: Complete What changedIn the fetch path, the pre-extraction byte cut of the downloaded body was removed and replaced by a post-conversion bound: the body is converted to markdown when HTML is detected, and only the rendered text is truncated to `max_bytes` (`crates/tinytools-std/src/network/web_fetch.rs#impl Tool for WebFetchTool {`). A new `render_body` function (cited as part of `crates/tinytools-std/src/network/web_fetch.rs#impl WebFetchTool {`) implements the convert-then-bound ordering, returns a `RenderedBody` recording the pre-cap extracted length, whether the output was capped, and whether the markup was cut at `EXTRACTOR_INPUT_CEILING` (8 MiB) before the extractor saw it. The fetch header now reports `output_capped_at` instead of the misnamed `download_capped_at`, and a new helper appends `markup_truncated_at=<ceiling>B` when markup was cut. Both the inline schema in the tool and the fixture schema reword `max_bytes` to state that it caps the returned text and never the markup the extractor reads. Features
Tests
FindingsPreviously reported and still active
Resolved this pass
Before merge
How this fits togetherflowchart LR
n0["WebFetchTool<br/>changed"]:::changed
n1["fetch"]:::impacted
n2["test_security"]:::impacted
n3["NetGate"]:::impacted
n4["Tool"]:::impacted
n5["..._the_schema_into_extra_optional_arguments"]:::impacted
n6["HtmlExtractor"]:::impacted
n0 -->|uses| n3
n0 -->|implements| n4
n0 -->|uses| n6
n1 -->|uses| n0
n5 -->|calls| n1
n5 -->|tests| n1
n5 -->|calls| n2
n5 -->|tests| n2
classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Agent review detailscritique
security
tests
commits
description
e2e
Evidence and run details
|
|
Warning Review limit reached
This review includes 3 billable files and costs up to $0.75. Or wait 56 minutes for your next included review. View limit detailsLimit details: You’ve used the included review currently available. Review configuration: ⚙️ Run configuration
📒 Files selected for processing (3)
Comment |
There was a problem hiding this comment.
tinysweeper found nothing blocking. Approving.
$0.0045 · 102,609 in / 5,813 out · 13,081 cached (13%) · gpt-5.6-luna, glm-5.3-flash
critique: $0.0025 · 44,856 in / 2,642 out · 6,265 cached (14%) · gpt-5.6-luna
security: $0.0018 · 35,951 in / 745 out · 3,744 cached (10%) · gpt-5.6-luna
tests: $0.0001 · 7,757 in / 796 out · 1,536 cached (20%) · glm-5.3-flash
description: $0.0001 · 7,853 in / 157 out · 1,408 cached (18%) · glm-5.3-flash
Add a test that drives an HTML body past the extractor's input ceiling and asserts the fetch header discloses the truncation while the extracted text still comes through. This locks in the reporting behaviour for oversized markup inputs. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Move the markup truncation header logic into a dedicated helper so the header can be built without a full fetch, and rewrite the corresponding test to exercise the helper directly instead of going through the async fetch path. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Remove an assertion checking that rendered content contains "visible" from the markup truncation test, since the test already verifies the truncation flag and header output. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The ceiling assertions in the web fetch tests now run inside a const block so they are checked at compile time rather than only when the test executes. The custom failure message was dropped since const assertions cannot format one. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Problem
max_bytescut the raw HTML before the extractor ran:So lowering it could destroy the page it was meant to bound. Measured on a 432,864-byte client-rendered document:
{url, max_bytes: 50000}extracted=48B_of_432864B— the<title>, nothing else{url}extracted=37243B_of_433638B— the real documentSame URL, same tool, same extractor. 776× the content, one parameter. Reported as
status=200both times.The cut lands mid-tag and mid-DOM, and a readability pass handed that wreckage recovers only the heading. A caller setting a sensible cost bound therefore lost the document silently — and the knob it would then reach for, raising the cap, was the right knob turned too timidly, which is the worst case for learning anything from a failure.
Change
Convert first, bound second.
Bounding the output is what the schema already promised, and it costs no memory:
resp.text()above materialises the whole body regardless, somax_bytesnever bounded a download. Only the extractor's input needs a ceiling, and that is for CPU —EXTRACTOR_INPUT_CEILING, 8 MiB, ~19× the largest real page observed and ~8× the defaultmax_response_size.The ordering is lifted into
render_bodyso it can be tested directly rather than through the network.Testing
cargo test -p tinytools-std --lib network::— 98 passed, 0 failed.The discriminating test places the prose past the cap offset, so the two orderings give different answers:
and a counterfactual proves it cannot pass vacuously — the old ordering, performed by hand on the same document, loses the prose:
Plus:
extractedis pre-cap so the header ratio describes the extraction rather than the truncation;raw: trueis still bounded by the same cap; the ceiling dwarfs any real page.cargo fmt -p tinytools-std -- --checkclean.Contract changes — please review these specifically
max_bytessemantics. Same parameter, different side of the extractor. A caller that passed a small value to bound cost now gets bounded output rather than a destroyed page. The schema description is updated to say so, and the pinned fixture infixtures/web_fetch.jsonwith it.download_capped_at=→output_capped_at=, plus a newmarkup_truncated_at=for the ceiling. Nothing here ever capped a download; the old name is why this went unnoticed.🤖 Generated with Claude Code