fix(english/wetriedtls): strip site header chrome, keep container content, harden chapter parsing - #2608
fix(english/wetriedtls): strip site header chrome, keep container content, harden chapter parsing#2608RibatTRW wants to merge 3 commits into
Conversation
…tent, fail loudly on bad chapter lists
|
| const isEdgeJunk = (p: string) => | ||
| isPromoParagraph(p) || | ||
| isCreditLine(p) || | ||
| isDivider(p) || | ||
| isTitleRepeat(p, titles); |
There was a problem hiding this comment.
Container content gets discarded
If a chapter wraps its header and story in one <div>, splitBlocks keeps them together. A “We Tried Translations” banner anywhere in that block makes isPromoParagraph mark the whole block as site chrome, so the leading story paragraphs disappear. A leading translator credit can have the same effect. Check header nodes separately from the content they wrap.
| for (let i = start; i < Math.min(end, start + 8) && divider < 0; i++) { | ||
| if (isDivider(blocks[i])) divider = i; | ||
| } | ||
| let header = divider >= 0; | ||
| for (let i = start; header && i < divider; i++) { | ||
| header = isEdgeJunk(blocks[i]) || isBoldOnly(blocks[i]); | ||
| } | ||
| if (!header) break; | ||
| start = divider + 1; |
There was a problem hiding this comment.
If a chapter starts with a genuine bold scene or POV line followed by an in-story divider within eight blocks, this loop treats both as a site header. It skips past the divider without checking that the bold line is a title, silently removing the chapter’s opening text.
| const blocks = splitBlocks(body); | ||
| if (blocks.length === 0) return { status: 'empty' }; | ||
| const titles = pageTitles(flight); | ||
| const isEdgeJunk = (p: string) => | ||
| isPromoParagraph(p) || | ||
| isCreditLine(p) || | ||
| isDivider(p) || | ||
| isTitleRepeat(p, titles); |
There was a problem hiding this comment.
Chapter cleanup lacks regression tests
The new splitting, header filtering, and sanitizing logic has no committed chapter HTML fixtures or regression tests. Add cases for a header and story inside one container and for a bold scene opening followed by a divider; without them, future parser changes can silently remove chapter text.
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
…headers - sanitize attributes by tokenizing tags as a browser does, so prose such as "one = 1" is no longer deleted and handlers or script URLs that directly follow a quoted value are no longer missed; also drop comments, raw-text elements and script URLs in SVG animation attributes - treat a short header line that opens with a known title as header, so "Series<br>Chapter 88: ..." and "Title (1)" above a divider are stripped when the page data lacks the chapter title - only accept a dash after a credit role word when it is spaced, so "Editor-in-chief ..." is not trimmed as a credit line
Improves the We Tried TLS plugin (
plugins/english/wetriedtls.ts), version 1.0.4 → 1.1.0.This PR was written by an AI agent (Claude); it has not been human-reviewed.
What changes
Chapter cleanup
parseNovel.WeTried Translations/WE TRIED TLSbanners (no space, or wrapped in<span><strong>)Translator/Editor:,Editors:,[Translator – X],Editor & TLC:,Discord:andKo-Fi:lines◈ Series/Chapter N/──────headers, including ones whose text differs slightly from the catalog title (The Wizard in the Poncho,Vol. 1 Ch. 1: ...)splitBlockswalks tags by depth, so<ul><li><p>one</p></li><li>two</li></ul>stays one block.div/section/articlewrappers are unwrapped so a header and the story inside one wrapper are judged line by line. Only short plain lines are trimmed as chrome.Chapter N,Vol.,Prologue,◈ ..., or a known title), so a bold scene or POV opening is kept.Sanitizer
javascript:,:, whitespace smuggling,vbscript:), includingxlink:hrefandformaction.Chapter list and covers
parseChapterListthrows on a blank or malformed response, a missingmeta.last_page, or a chapter entry without slug/name, instead of returning a truncated list.Relationship to #2604
#2604 is still open and edits the same file. I used it and the review findings on #2602 / #2588 as reference. The sanitizer, chapter-list validation and direct-cover changes follow its approach, but this PR is an independent rewrite of those parts, not a rebase. The two will conflict on
wetriedtls.ts, and whichever lands first, the other needs a rebase. I have not commented on or touched #2604.Follow-up review fixes (third commit)
A second AI review (Claude, not a human) found and fixed these, each reproduced first:
If one = 1becameIf and two = 2andonline=truevanished. Attributes are now read by walking each tag the way an HTML parser tokenizes it, so text is never touched.<img src="x"onerror=...>,<a title="t"href="javascript:...">) and SVG<animate values="javascript:...">/<set to=...>were missed and are now caught. Comments, CDATA and raw-text elements (title,xmp, ...) are dropped so they cannot hide a following tag. The LNReader app also runssanitize-htmlon chapter HTML, so this is defense in depth. A code comment claiming the reader renders it unsanitized was wrong and is corrected.The Inept Mage's Infinite Regression<br>Chapter 86–88: ...plus──────stayed at the top because the page data has no chapter title. A short line above the divider that starts with a known title now counts as header. That also stripsAddicted to Defeat (1)from Villain ch. 1 again.Editor-in-chief Kim walked in.A dash after a role word now needs a following space. Every real credit format in the sample still matches.Dropped, not changed: chapter-list validation on a malformed entry (no real response in the sample triggers it), direct covers (all 94 thumbnails are absolute https URLs), and the one-off
The Little Prince of the Ossuarych. 1 header (stripping it would need a rule broad enough to risk story lines).Testing
What I ran:
npm run check:plugin -- plugins/english/wetriedtls.ts: all four checks PASS (12 novels, search, 885 chapters, chapter body).parseNovelplus first / middle / last free chapter for all 94 catalog novels (282 chapters), plus Inept Mage ch. 60/75/85–87, before and after each change. No errors and no empty chapters. The only output changes were the intended header removals and wrapper-tag removals (12 chapters changed by a few dozen characters). The header residue listed above is gone, apart from a few nonstandard headers left in place (e.g.The Little Prince of the Ossuary).<div>returned "empty"; a bold opening line followed by──────was dropped), then confirmed fixed.popularNovels(status filter, page past the end),searchNovels(hit and empty),parseNovelincluding paywalled chapters, and a direct cover URL (HTTP 200).onerrorscript vectors, unclosed<p>, lists/tables, image-only and bold-only chapters, blank / non-JSON / partial chapter-list responses.npm run build:compile, andeslint/prettier --checkon the changed file: clean apart from the pre-existing_isNovelunused-arg warning.npm run lint/format:checkreport unrelated existing failures in other files.Not tested: on-device in the LNReader app, and paywalled chapter bodies (they only show the premium notice).
No regression test files are added: maintainers asked on #2575 for none until a test framework exists, so the cases above were run ad hoc only.