Skip to content

fix(english/wetriedtls): strip site header chrome, keep container content, harden chapter parsing - #2608

Closed
RibatTRW wants to merge 3 commits into
lnreader:masterfrom
RibatTRW:fm/lnreader-wetriedtls-improve
Closed

RibatTRW wants to merge 3 commits into
lnreader:masterfrom
RibatTRW:fm/lnreader-wetriedtls-improve

Conversation

@RibatTRW

@RibatTRW RibatTRW commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Improves the We Tried TLS plugin (plugins/english/wetriedtls.ts), version 1.0.4 → 1.1.0.

This PR was written by an AI agent (Claude); it has not been human-reviewed.

What changes

Chapter cleanup

  • Site header chrome is stripped from chapter bodies, using the chapter page's own data (series title, chapter name, chapter title) to recognise title lines. There is no title cache filled by parseNovel.
  • Header formats the old parser missed, found by sampling all 94 catalog novels:
    • WeTried Translations / WE TRIED TLS banners (no space, or wrapped in <span><strong>)
    • Translator/Editor:, Editors:, [Translator – X], Editor & TLC:, Discord: and Ko-Fi: lines
    • ◈ Series / Chapter N / ────── headers, including ones whose text differs slightly from the catalog title (The Wizard in the Poncho, Vol. 1 Ch. 1: ...)
    • horizontal-rule dividers
  • Story content is kept:
    • Lists, tables, blockquotes and divs that carry text are preserved. splitBlocks walks tags by depth, so <ul><li><p>one</p></li><li>two</li></ul> stays one block.
    • div / section / article wrappers are unwrapped so a header and the story inside one wrapper are judged line by line. Only short plain lines are trimmed as chrome.
    • A bold lead-in counts as header only when a divider follows it and it reads like a title (Chapter N, Vol., Prologue, ◈ ..., or a known title), so a bold scene or POV opening is kept.

Sanitizer

  • Entity-decoded script-URL check (&#106;avascript:, &colon;, whitespace smuggling, vbscript:), including xlink:href and formaction.
  • Event-handler and dangerous-element stripping.

Chapter list and covers

  • parseChapterList throws on a blank or malformed response, a missing meta.last_page, or a chapter entry without slug/name, instead of returning a truncated list.
  • Covers are served directly instead of through images.weserv.nl.

Relationship to #2604

#2604 is still open and edits the same file. I used it and the review findings on #2602 / #2588 as reference. The sanitizer, chapter-list validation and direct-cover changes follow its approach, but this PR is an independent rewrite of those parts, not a rebase. The two will conflict on wetriedtls.ts, and whichever lands first, the other needs a rebase. I have not commented on or touched #2604.

Follow-up review fixes (third commit)

A second AI review (Claude, not a human) found and fixed these, each reproduced first:

  • Sanitizer deleted prose. The event-handler regex also ran over text, so If one = 1 became If and two = 2 and online=true vanished. Attributes are now read by walking each tag the way an HTML parser tokenizes it, so text is never touched.
  • Sanitizer bypasses. An attribute directly after a quoted value (<img src="x"onerror=...>, <a title="t"href="javascript:...">) and SVG <animate values="javascript:..."> / <set to=...> were missed and are now caught. Comments, CDATA and raw-text elements (title, xmp, ...) are dropped so they cannot hide a following tag. The LNReader app also runs sanitize-html on chapter HTML, so this is defense in depth. A code comment claiming the reader renders it unsanitized was wrong and is corrected.
  • Header left on the newest Inept Mage chapters. The Inept Mage's Infinite Regression<br>Chapter 86–88: ... plus ────── stayed at the top because the page data has no chapter title. A short line above the divider that starts with a known title now counts as header. That also strips Addicted to Defeat (1) from Villain ch. 1 again.
  • Credit-line regex could eat an opening line such as Editor-in-chief Kim walked in. A dash after a role word now needs a following space. Every real credit format in the sample still matches.

Dropped, not changed: chapter-list validation on a malformed entry (no real response in the sample triggers it), direct covers (all 94 thumbnails are absolute https URLs), and the one-off The Little Prince of the Ossuary ch. 1 header (stripping it would need a rule broad enough to risk story lines).

Testing

What I ran:

  • npm run check:plugin -- plugins/english/wetriedtls.ts: all four checks PASS (12 novels, search, 885 chapters, chapter body).
  • parseNovel plus first / middle / last free chapter for all 94 catalog novels (282 chapters), plus Inept Mage ch. 60/75/85–87, before and after each change. No errors and no empty chapters. The only output changes were the intended header removals and wrapper-tag removals (12 chapters changed by a few dozen characters). The header residue listed above is gone, apart from a few nonstandard headers left in place (e.g. The Little Prince of the Ossuary).
  • Both first-review findings were reproduced on the earlier commit with synthetic pages (banner and story inside one <div> returned "empty"; a bold opening line followed by ────── was dropped), then confirmed fixed.
  • Spot checks: popularNovels (status filter, page past the end), searchNovels (hit and empty), parseNovel including paywalled chapters, and a direct cover URL (HTTP 200).
  • Ad-hoc adversarial inputs, not committed: encoded/SVG/onerror script vectors, unclosed <p>, lists/tables, image-only and bold-only chapters, blank / non-JSON / partial chapter-list responses.
  • npm run build:compile, and eslint / prettier --check on the changed file: clean apart from the pre-existing _isNovel unused-arg warning. npm run lint / format:check report unrelated existing failures in other files.

Not tested: on-device in the LNReader app, and paywalled chapter bodies (they only show the premium notice).

No regression test files are added: maintainers asked on #2575 for none until a test framework exists, so the cases above were run ad hoc only.

@greptile-apps

greptile-apps Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 3/5

[Medium risk] Refines content parsing for a web scraper plugin.

The PR should not merge until chapter cleanup preserves content-bearing containers and bold scene openings.

Findings

  1. P1 Container content gets discarded ▶
  2. P1 Bold scene openings disappear ▶
  3. P2 Chapter cleanup lacks regression tests ▶

Summary

The PR revises We Tried TLS chapter cleanup, retains nested content blocks, validates chapter-list responses more strictly, sanitizes returned HTML, and serves covers directly.

  • Header detection is stateless and uses titles from the chapter page.
  • Edge cleanup can discard story content inside a container or before an in-story divider.
  • The new parsing behavior lacks committed regression fixtures.

Reviews (1) · Last reviewed commit: "fix(english/wetriedtls): strip site head..."

Comment thread plugins/english/wetriedtls.ts Outdated
Comment on lines +541 to +545
const isEdgeJunk = (p: string) =>
isPromoParagraph(p) ||
isCreditLine(p) ||
isDivider(p) ||
isTitleRepeat(p, titles);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Container content gets discarded

If a chapter wraps its header and story in one <div>, splitBlocks keeps them together. A “We Tried Translations” banner anywhere in that block makes isPromoParagraph mark the whole block as site chrome, so the leading story paragraphs disappear. A leading translator credit can have the same effect. Check header nodes separately from the content they wrap.

Comment on lines +555 to +563
for (let i = start; i < Math.min(end, start + 8) && divider < 0; i++) {
if (isDivider(blocks[i])) divider = i;
}
let header = divider >= 0;
for (let i = start; header && i < divider; i++) {
header = isEdgeJunk(blocks[i]) || isBoldOnly(blocks[i]);
}
if (!header) break;
start = divider + 1;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Bold scene openings disappear

If a chapter starts with a genuine bold scene or POV line followed by an in-story divider within eight blocks, this loop treats both as a site header. It skips past the divider without checking that the bold line is a title, silently removing the chapter’s opening text.

Comment thread plugins/english/wetriedtls.ts Outdated
Comment on lines +538 to +545
const blocks = splitBlocks(body);
if (blocks.length === 0) return { status: 'empty' };
const titles = pageTitles(flight);
const isEdgeJunk = (p: string) =>
isPromoParagraph(p) ||
isCreditLine(p) ||
isDivider(p) ||
isTitleRepeat(p, titles);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Chapter cleanup lacks regression tests

The new splitting, header filtering, and sanitizing logic has no committed chapter HTML fixtures or regression tests. Add cases for a header and story inside one container and for a bold scene opening followed by a divider; without them, future parser changes can silently remove chapter text.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

…headers

- sanitize attributes by tokenizing tags as a browser does, so prose such
  as "one = 1" is no longer deleted and handlers or script URLs that
  directly follow a quoted value are no longer missed; also drop comments,
  raw-text elements and script URLs in SVG animation attributes
- treat a short header line that opens with a known title as header, so
  "Series<br>Chapter 88: ..." and "Title (1)" above a divider are stripped
  when the page data lacks the chapter title
- only accept a dash after a credit role word when it is spaced, so
  "Editor-in-chief ..." is not trimmed as a credit line
@RibatTRW RibatTRW closed this Sep 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant