Skip to content

fix(lightnovelwp): keep images and paragraph HTML in chapters - #2611

Open
RibatTRW wants to merge 2 commits into
lnreader:masterfrom
RibatTRW:fm/lnreader-centralnovel-2610
Open

RibatTRW wants to merge 2 commits into
lnreader:masterfrom
RibatTRW:fm/lnreader-centralnovel-2610

Conversation

@RibatTRW

@RibatTRW RibatTRW commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Closes #2610

What was wrong

Central Novel is a lightnovelwp multisrc source. Its "Ilustrações" chapters contain only images.

  • Trigger: fix(lightnovelwp): parse chapter text when bottomnav is absent #2593 (4673688) replaced the bottomnav-anchored regex in parseChapter with an htmlparser2 walk. The walk collected only the text of <p> elements and joined it with \n.
  • Symptom: a chapter made of <p><img …></p> has no text, so it came back as ''. Ordinary chapters on every lightnovelwp source also lost their inline images and <em>/<strong> formatting, and paragraphs came back as plain text.
  • Second, older gap: some Central Novel illustration pages (for example Grimgar Vol. 1) put the images in <div class="separator"> inside an <article>, with no <p> at all. Those pages were empty before fix(lightnovelwp): parse chapter text when bottomnav is absent #2593 as well, because the old regex also kept only <p> elements.

These are the results from running the parser on the real pages, fetched through a browser because Central Novel is behind Cloudflare:

Page Before #2593 (regex) master (#2593) This PR
Grimgar Vol. 17 – Ilustrações 13 <img> 0 chars 13 <img>
Grimgar Vol. 1 – Ilustrações 0 chars 0 chars 22 <img>
Grimgar Vol. 17 – Capítulo 2 (text) 174 <p>, 1 <img> 19,100 chars of plain text, no images 174 <p>, 1 <img> (same HTML as before #2593)

What changed

plugins/multisrc/lightnovelwp/template.ts, parseChapter only:

  • Select the first div.epcontent with cheerio, which the template already imports. The block ends where its element closes, so bottomnav is still not needed and a <script> inside the block does not cut it short. A page without epcontent still returns ''.
  • Remove script, style and noscript from the block.
  • Return each non-empty <p> as HTML, plus any <img> that sits outside a paragraph, in document order. Other elements inside the block (such as ad wrappers) are still dropped, as they were before.
  • Image src: if it is missing or a data: placeholder, use data-lazy-src or data-src. Any relative src (for example images/p.jpg, ../p.jpg, /p.jpg or //cdn/p.jpg) is resolved against the chapter page URL with new URL(src, site + chapterPath).
  • The app renders the returned HTML as is, so every element in the block loses its on* attributes. Any href or src using a javascript:, vbscript: or non-image data: scheme is removed. The check ignores case, whitespace and control characters.
  • The template version base changes from 1.1.10 to 1.1.11, so every lightnovelwp source gets the update.

Testing

  • Fixture comparison: I ran the generated CentralNovel[lightnovelwp].ts against the three saved Central Novel pages in the table, with fetch stubbed to serve the saved HTML, and compared the output with the pre-fix(lightnovelwp): parse chapter text when bottomnav is absent #2593 regex on the same HTML.
  • Synthetic edge cases: a <script> inside epcontent that contains </p><div class=bottomnav>, an ad <div>, an &nbsp; paragraph, lazy and relative image sources, and a comments <p> outside the block. There was also a page with no epcontent block, which returns ''.
  • Review follow-up (2baad3b):
    • The three Central Novel pages and the earlier synthetic cases give byte-identical output before and after the commit.
    • New synthetic case: images/page-1.jpg, ../up/p2.jpg and a lazy lazy/p3.jpg resolve against the chapter URL, and //cdn… becomes https:.
    • In the same case, onerror, ONLOAD and onclick are removed, and so are the " JaVa&#x09;Script:", vbscript: and data:text/html hrefs and a javascript: img src. An https: link and a data:image/png img are kept.
    • The live check was not re-run for this commit.
  • Live check: npm run check:plugin on the generated lightnovelwp plugins. parseChapter passes on BlumeVerse, DobyNovels, HyacinthinBloom, KnoxT, LazyGirlTranslations, TranslationWeaver and KodeksLibrary.
    • Central Novel itself and several other sources were INCONCLUSIVE: Cloudflare 403, DNS failures or connection resets.
    • Some sources FAIL before reaching parseChapter, with 404/403/503 responses or no novels/chapters. That is unrelated to this change.
  • Pre-existing, not addressed here: TC&Sega's parseChapter returns 0 chars both on master and with this PR. Its content block is <article class="epcontent"> rather than a div. I left it out of this PR to keep the change focused.
  • Lint, format, compile: npx prettier --check and npx eslint pass on the changed file, and npm run build:compile succeeds. npm run lint still reports 3 errors, all in other files that this PR does not touch (rewayatfans.ts, wetriedtls.ts, wntl.ts).
  • There is no plugin unit-test harness in the repo, so no regression test is committed.
  • Not tested in the LNReader app or the playground.

This PR was written by an AI agent (Claude Opus 5.5). A human has not reviewed it yet.

lnreader#2593 replaced the bottomnav-anchored regex with a walk that collected
only the text of <p> elements. Illustration chapters hold images and no
prose, so they came back empty, and ordinary chapters lost inline images
and formatting.

Select the first div.epcontent with cheerio and return its paragraphs as
HTML plus any images outside a paragraph, resolving lazy-load and
relative sources to absolute URLs. Scripts, styles and noscript are
dropped. A missing epcontent block still returns '' and bottomnav is not
needed. Bump the template version so every source picks up the fix.

Closes lnreader#2610
@greptile-apps

greptile-apps Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 4/5

[Medium risk] Changes how chapter content is extracted from HTML.

This PR is not safe to merge until chapter HTML is cleaned before it reaches the viewer.

Findings

  1. P1 Security Chapter HTML can run code ▶
  2. P2 Chapter-relative images fail to load ▶

Summary

The LightNovelWP chapter parser now keeps paragraph HTML and images, including in image-only chapters. It also fixes lazy and relative image URLs and raises the shared template version.

  • Selects chapter content from the first div.epcontent block.
  • Keeps formatted paragraphs and images in document order.

Reviews (1) · Last reviewed commit: "fix(lightnovelwp): keep images and parag..."

Comment on lines +473 to +475
if (!node.parents('p').length) blocks.push($.html(node));
} else if (node.text().trim() || node.find('img').length) {
blocks.push($.html(node));

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 security Chapter HTML can run code

A chapter page can put an onerror attribute on an image or a script URL in a link. The new parser keeps those values, and the in-repo chapter viewer inserts the returned HTML without cleaning it. Strip executable attributes and unsafe URLs before returning the HTML.

How this was verified: Source-controlled image attributes pass through $.html(node) into the viewer’s unsanitized HTML insertion.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 2baad3b. Before returning, every element in the kept block has its on* attributes removed. Any href/src using a javascript:, vbscript: or non-image data: scheme is also removed; the check ignores case, whitespace and control characters. I tested it with a synthetic page (onerror/ONLOAD/onclick, a tab-obfuscated JaVaScript: href, vbscript:, a data:text/html href, a javascript: img src). All were stripped, and https: links and data:image/* images were kept.

Comment on lines +461 to +463
if (src.startsWith('//')) src = 'https:' + src;
else if (src.startsWith('/')) src = this.site.replace(/\/$/, '') + src;
img.attr('src', src);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Chapter-relative images fail to load

An image URL such as images/page-1.jpg is relative to the chapter page, but this code makes only URLs starting with / absolute. The viewer receives the unchanged path, so it looks for the image relative to the viewer instead. Resolve these paths against the chapter URL too.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 2baad3b. Any relative image src is now resolved with new URL(src, site + chapterPath), so images/page-1.jpg and ../x.jpg resolve against the chapter page, and //cdn/... still becomes https:. Output for the three saved Central Novel pages is byte-identical.

Drop on* attributes and javascript:, vbscript: and non-image data: URLs
from href/src in the returned chapter HTML, which the app renders as is.
Resolve every relative image path against the chapter page URL instead
of only root-relative ones.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[centralnovel] Empty chapter: Hai to Gensou no Grimgar — Ilustrações

1 participant