Skip to content

Apply web_fetch's max_bytes to the extracted text, not the markup - #54

Merged
senamakel merged 5 commits into
tinyhumansai:mainfrom
sanil-23:fix/web-fetch-cap-bounds-output
Oct 8, 2026
Merged

senamakel merged 5 commits into
tinyhumansai:mainfrom
sanil-23:fix/web-fetch-cap-bounds-output

Conversation

@sanil-23

@sanil-23 sanil-23 commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Problem

max_bytes cut the raw HTML before the extractor ran:

let (body, byte_capped) = if downloaded > max_bytes { /* cut */ };   // first
let content = if converted { self.html.to_markdown(&body) };          // then extract

So lowering it could destroy the page it was meant to bound. Measured on a 432,864-byte client-rendered document:

call result
{url, max_bytes: 50000} extracted=48B_of_432864B — the <title>, nothing else
{url} extracted=37243B_of_433638B — the real document

Same URL, same tool, same extractor. 776× the content, one parameter. Reported as status=200 both times.

The cut lands mid-tag and mid-DOM, and a readability pass handed that wreckage recovers only the heading. A caller setting a sensible cost bound therefore lost the document silently — and the knob it would then reach for, raising the cap, was the right knob turned too timidly, which is the worst case for learning anything from a failure.

Change

Convert first, bound second.

Bounding the output is what the schema already promised, and it costs no memory: resp.text() above materialises the whole body regardless, so max_bytes never bounded a download. Only the extractor's input needs a ceiling, and that is for CPU — EXTRACTOR_INPUT_CEILING, 8 MiB, ~19× the largest real page observed and ~8× the default max_response_size.

The ordering is lifted into render_body so it can be tested directly rather than through the network.

Testing

cargo test -p tinytools-std --lib network:: — 98 passed, 0 failed.

The discriminating test places the prose past the cap offset, so the two orderings give different answers:

let body = page_with_prose_after(4_000);
let rendered = render_body(&TestHtml, body, true, 1_000);
assert!(rendered.content.contains("the prose that matters"));

and a counterfactual proves it cannot pass vacuously — the old ordering, performed by hand on the same document, loses the prose:

assert!(!TestHtml.to_markdown(&body[..1_000]).contains("the prose that matters"));

Plus: extracted is pre-cap so the header ratio describes the extraction rather than the truncation; raw: true is still bounded by the same cap; the ceiling dwarfs any real page.

cargo fmt -p tinytools-std -- --check clean.

Contract changes — please review these specifically

  1. max_bytes semantics. Same parameter, different side of the extractor. A caller that passed a small value to bound cost now gets bounded output rather than a destroyed page. The schema description is updated to say so, and the pinned fixture in fixtures/web_fetch.json with it.
  2. Header field renamed — download_capped_at= → output_capped_at=, plus a new markup_truncated_at= for the ceiling. Nothing here ever capped a download; the old name is why this went unnoticed.

🤖 Generated with Claude Code

`max_bytes` cut the raw HTML before the extractor ran, so lowering it could
destroy the page it was meant to bound.

Measured: a 432,864-byte client-rendered document fetched with max_bytes 50000
was cut mid-DOM, and the readability pass recovered only its <title> —
`extracted=48B_of_432864B`, reported as `status=200`. The same URL with no
max_bytes yields `extracted=37243B_of_433638B`. Identical tool, identical
extractor, 776x the content, one parameter.

A caller setting a sensible cost bound therefore lost the document silently,
and the knob it would then reach for — raising the cap — was the right knob
turned too timidly, which is the worst case for learning anything.

Convert first, bound second. Bounding the output is what the schema promises,
and it costs no memory: `resp.text()` already materialises the whole body, so
max_bytes never bounded a download. Only the extractor's input needs a
ceiling, and that is for CPU — EXTRACTOR_INPUT_CEILING, 8 MiB, ~19x the
largest real page seen.

The ordering is lifted into `render_body` so it is directly testable, including
a counterfactual that the old ordering loses the prose from the same document.

Contract changes: the `max_bytes` schema description now says which side of the
extractor it bounds, and the header reports `output_capped_at=` instead of
`download_capped_at=` — nothing here ever capped a download, and the old name
sent readers looking in the wrong place.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@tinysweeper

tinysweeper Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Tiny Sweeper review

Tiny Sweeper completed its review; deterministic results follow.

State: Ready for maintainer review
Priority: none
Reviewed head: ac43d01ef709
Updated: 1791458641 (Unix time)

Review snapshot

Change surface Files Review signal Count
Production 1 Active findings 1
Tests 2 Noted findings 0
Documentation 0 Resolved findings 4
Configuration 0 Pending checks/questions 0

Completeness: Complete
Test assessment: Test coverage is assessed from changed tests and lane evidence; execution is not claimed without trusted check data.

What changed

In the fetch path, the pre-extraction byte cut of the downloaded body was removed and replaced by a post-conversion bound: the body is converted to markdown when HTML is detected, and only the rendered text is truncated to `max_bytes` (`crates/tinytools-std/src/network/web_fetch.rs#impl Tool for WebFetchTool {`). A new `render_body` function (cited as part of `crates/tinytools-std/src/network/web_fetch.rs#impl WebFetchTool {`) implements the convert-then-bound ordering, returns a `RenderedBody` recording the pre-cap extracted length, whether the output was capped, and whether the markup was cut at `EXTRACTOR_INPUT_CEILING` (8 MiB) before the extractor saw it. The fetch header now reports `output_capped_at` instead of the misnamed `download_capped_at`, and a new helper appends `markup_truncated_at=<ceiling>B` when markup was cut. Both the inline schema in the tool and the fixture schema reword `max_bytes` to state that it caps the returned text and never the markup the extractor reads.

Features

  • Modified — max_bytes now bounds returned text instead of downloaded markup: Callers setting a low `max_bytes` no longer silently lose page content: the body is converted to markdown first, and the cap is applied to the extracted text (or to the raw body with `raw:true`). The header reports `output_capped_at=<max_bytes>B` when the rendered text is cut, replacing the old `download_capped_at` token, and `extracted=<pre-cap len>B_of_<downloaded>B` still describes extraction rather than truncation. (crates/tinytools-std/src/network/web_fetch.rs#impl Tool for WebFetchTool {, crates/tinytools-std/src/network/web_fetch.rs#impl WebFetchTool {)
  • Added — Extractor input ceiling with header disclosure: A separate 8 MiB `EXTRACTOR_INPUT_CEILING` cuts the markup before extraction only when a document exceeds it, preventing unbounded CPU in `to_markdown`; when that happens, `markup_truncated_at=8388608B` is appended to the fetch header. This is independent of the caller's `max_bytes` and, per the added rationale, is far above the largest page observed (a ~434 KB document) and the default response size. (crates/tinytools-std/src/network/web_fetch.rs#impl WebFetchTool {)
  • Modified — Schema wording for max_bytes clarified in tool and fixture: The `max_bytes` description in the tool's inline schema and the JSON fixture now says the cap bounds the returned text (the extracted markdown, or the raw body with `raw:true`) and never the markup the extractor reads, so lowering it cannot discard content the page actually had. (crates/tinytools-std/src/network/web_fetch.rs#impl Tool for WebFetchTool {)

Tests

  • unit — The cap applies to the extracted text, not the markup: a document whose prose sits past the 1,000-byte cap offset still contains that prose in the rendered output.: Strong — lane evidence describes it as the regression test for the convert-then-bound fix, and notes it fails if the markup is cut first. (crates/tinytools-std/src/network/web_fetch_tests.rs#async fn a_redirect_is_still_reported_as_a_successful_result() {)
  • unit — Verifies the reported `extracted` length is the pre-cap extraction size while the returned content is capped and flagged, so the header ratio describes extraction rather than truncation.: Reasonable — cited in the diff as part of the new test block. (crates/tinytools-std/src/network/web_fetch_tests.rs#async fn a_redirect_is_still_reported_as_a_successful_result() {)
  • unit — An output within the cap is returned whole with no output-cap or markup-truncation flags set.: Reasonable — cited in the diff as part of the new test block. (crates/tinytools-std/src/network/web_fetch_tests.rs#async fn a_redirect_is_still_reported_as_a_successful_result() {)
  • unit — A body exceeding `EXTRACTOR_INPUT_CEILING` sets `markup_truncated` and the emitted fetch header discloses `markup_truncated_at=<ceiling>B`.: Strong — lane evidence states this resolves the earlier finding about the uncovered markup_truncated branch. (crates/tinytools-std/src/network/web_fetch_tests.rs#async fn a_redirect_is_still_reported_as_a_successful_result() {)
  • unit — With `raw: true` (conversion off), the cap applies to the body itself and the returned content is truncated accordingly.: Reasonable — cited in the diff as part of the new test block. (crates/tinytools-std/src/network/web_fetch_tests.rs#async fn a_redirect_is_still_reported_as_a_successful_result() {)
  • unit — Asserts at compile time that the extractor input ceiling is at least 8 MiB and far above the ~434 KB page size cited in the incident narrative.: Reasonable — cited in the diff as part of the new test block. (crates/tinytools-std/src/network/web_fetch_tests.rs#async fn a_redirect_is_still_reported_as_a_successful_result() {)
  • unit — Checks the tool's schema for `max_bytes` now contains wording saying the cap bounds the OUTPUT and never the markup.: Reasonable — cited in the diff as part of the new test block. (crates/tinytools-std/src/network/web_fetch_tests.rs#async fn a_redirect_is_still_reported_as_a_successful_result() {)

Findings

Previously reported and still active

  • Cover the markup\_truncated branch and its header flag

Resolved this pass

  • Cover the markup_truncated branch and its header flag
  • Cover the markup_truncated branch and its header flag
  • Cover the markup_truncated branch and its header flag
  • Cover the markup_truncated branch and its header flag

Before merge

  • Address carried finding Cover the markup\_truncated branch and its header flag.

How this fits together

flowchart LR
  n0["WebFetchTool<br/>changed"]:::changed
  n1["fetch"]:::impacted
  n2["test_security"]:::impacted
  n3["NetGate"]:::impacted
  n4["Tool"]:::impacted
  n5["..._the_schema_into_extra_optional_arguments"]:::impacted
  n6["HtmlExtractor"]:::impacted
  n0 -->|uses| n3
  n0 -->|implements| n4
  n0 -->|uses| n6
  n1 -->|uses| n0
  n5 -->|calls| n1
  n5 -->|tests| n1
  n5 -->|calls| n2
  n5 -->|tests| n2
  classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
  classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
  classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
  classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading
Agent review details

critique

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: The change adds coverage for output truncation, extractor-input truncation, header reporting, raw responses, and schema semantics. The earlier markup-truncation coverage finding is resolved, and this test change looks safe to merge. (1 earlier finding(s) still open) _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._

security

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: The added regression coverage verifies output caps, extractor input ceilings, truncation reporting, raw responses, and schema documentation. The change looks safe to merge. (1 earlier finding(s) still open) _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._

tests

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Positive: The revision adds direct tests for render_body and the markup-truncation header flag, resolving the earlier finding about the uncovered markup_truncated branch; the tests assert real discriminating behaviour, so the change looks sound and safe to merge.
  • Lane summary: The revision adds direct tests for render_body and the markup-truncation header flag, which resolves the earlier finding about the uncovered markup_truncated branch. The tests assert real discriminating behaviour (prose survival under the new convert-then-bound ordering, a counterfactual showing the old ordering fails, cap flagging, extractor-input ceiling reporting), so the change looks sound and safe to merge. (1 earlier finding(s) still open) _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._

commits

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: Nothing sensitive found in what this pull request commits.

description

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Positive: The revision fixes the earlier finding: `markup_truncated_input_is_reported_in_the_fetch_header` now drives the extractor input across the ceiling and asserts the header flag, alongside the counterfactual, pre-cap `extracted`, raw-path and schema tests; the test suite matches the described contract change and looks sound to merge.
  • Lane summary: The revision fixes the earlier finding: `markup_truncated_input_is_reported_in_the_fetch_header` now drives the extractor input across the ceiling and asserts the header flag, alongside the counterfactual, pre-cap `extracted`, raw-path and schema tests. The test suite matches the described contract change, and the change looks sound to merge. (1 earlier finding(s) still open) _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._

e2e

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: No end-to-end harness in this repository: no e2e test files and no e2e workflow.
Evidence and run details
  • Models: gpt-5.6-luna, glm-5.3-flash
  • Spend: $0.007382
  • Tokens: 94640 input · 5457 output · 7640 cached · 0 embedding
Head State Pass summary
dfcdf53f6474 ready for maintainer review 1 active finding(s), 0 resolved finding(s) (at 1791396632)
59a266eae4fb ready for maintainer review 0 active finding(s), 7 resolved finding(s) (at 1791458527)
ac43d01ef709 ready for maintainer review 0 active finding(s), 4 resolved finding(s) (at 1791458641)

tinysweeper 0.1.0

@coderabbitai

coderabbitai Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

  • Run on-demand review

This review includes 3 billable files and costs up to $0.75.

Or wait 56 minutes for your next included review.

Check out review usage here.

View limit details

Limit details: You’ve used the included review currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 9108024a-8f85-46cb-917c-da8fcdf97404
📥 Commits

Reviewing files that changed from the base of the PR and between 68163b1 and ac43d01.

📒 Files selected for processing (3)
  • crates/tinytools-std/src/network/fixtures/web_fetch.json
  • crates/tinytools-std/src/network/web_fetch.rs
  • crates/tinytools-std/src/network/web_fetch_tests.rs
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking. Approving.

             $0.0045 · 102,609 in / 5,813 out · 13,081 cached (13%) · gpt-5.6-luna, glm-5.3-flash
critique:    $0.0025 · 44,856 in  / 2,642 out · 6,265 cached (14%)  · gpt-5.6-luna
security:    $0.0018 · 35,951 in  / 745 out   · 3,744 cached (10%)  · gpt-5.6-luna
tests:       $0.0001 · 7,757 in   / 796 out   · 1,536 cached (20%)  · glm-5.3-flash
description: $0.0001 · 7,853 in   / 157 out   · 1,408 cached (18%)  · glm-5.3-flash

Comment thread crates/tinytools-std/src/network/web_fetch_tests.rs
@senamakel senamakel self-assigned this Oct 8, 2026
senamakel and others added 4 commits October 8, 2026 14:15
Add a test that drives an HTML body past the extractor's input ceiling and asserts the fetch header discloses the truncation while the extracted text still comes through. This locks in the reporting behaviour for oversized markup inputs.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Move the markup truncation header logic into a dedicated helper so the
header can be built without a full fetch, and rewrite the corresponding
test to exercise the helper directly instead of going through the async
fetch path.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Remove an assertion checking that rendered content contains "visible" from the markup truncation test, since the test already verifies the truncation flag and header output.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The ceiling assertions in the web fetch tests now run inside a const block so
they are checked at compile time rather than only when the test executes. The
custom failure message was dropped since const assertions cannot format one.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
@senamakel
senamakel merged commit 02343e5 into tinyhumansai:main Oct 8, 2026
12 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants