Conversation
lnreader#2593 replaced the bottomnav-anchored regex with a walk that collected only the text of <p> elements. Illustration chapters hold images and no prose, so they came back empty, and ordinary chapters lost inline images and formatting. Select the first div.epcontent with cheerio and return its paragraphs as HTML plus any images outside a paragraph, resolving lazy-load and relative sources to absolute URLs. Scripts, styles and noscript are dropped. A missing epcontent block still returns '' and bottomnav is not needed. Bump the template version so every source picks up the fix. Closes lnreader#2610
|
| if (!node.parents('p').length) blocks.push($.html(node)); | ||
| } else if (node.text().trim() || node.find('img').length) { | ||
| blocks.push($.html(node)); |
There was a problem hiding this comment.
A chapter page can put an onerror attribute on an image or a script URL in a link. The new parser keeps those values, and the in-repo chapter viewer inserts the returned HTML without cleaning it. Strip executable attributes and unsafe URLs before returning the HTML.
How this was verified: Source-controlled image attributes pass through $.html(node) into the viewer’s unsanitized HTML insertion.
There was a problem hiding this comment.
Fixed in 2baad3b. Before returning, every element in the kept block has its on* attributes removed. Any href/src using a javascript:, vbscript: or non-image data: scheme is also removed; the check ignores case, whitespace and control characters. I tested it with a synthetic page (onerror/ONLOAD/onclick, a tab-obfuscated JaVaScript: href, vbscript:, a data:text/html href, a javascript: img src). All were stripped, and https: links and data:image/* images were kept.
| if (src.startsWith('//')) src = 'https:' + src; | ||
| else if (src.startsWith('/')) src = this.site.replace(/\/$/, '') + src; | ||
| img.attr('src', src); |
There was a problem hiding this comment.
Chapter-relative images fail to load
An image URL such as images/page-1.jpg is relative to the chapter page, but this code makes only URLs starting with / absolute. The viewer receives the unchanged path, so it looks for the image relative to the viewer instead. Resolve these paths against the chapter URL too.
There was a problem hiding this comment.
Fixed in 2baad3b. Any relative image src is now resolved with new URL(src, site + chapterPath), so images/page-1.jpg and ../x.jpg resolve against the chapter page, and //cdn/... still becomes https:. Output for the three saved Central Novel pages is byte-identical.
Drop on* attributes and javascript:, vbscript: and non-image data: URLs from href/src in the returned chapter HTML, which the app renders as is. Resolve every relative image path against the chapter page URL instead of only root-relative ones.
Closes #2610
What was wrong
Central Novel is a
lightnovelwpmultisrc source. Its "Ilustrações" chapters contain only images.bottomnav-anchored regex inparseChapterwith an htmlparser2 walk. The walk collected only the text of<p>elements and joined it with\n.<p><img …></p>has no text, so it came back as''. Ordinary chapters on every lightnovelwp source also lost their inline images and<em>/<strong>formatting, and paragraphs came back as plain text.<div class="separator">inside an<article>, with no<p>at all. Those pages were empty before fix(lightnovelwp): parse chapter text when bottomnav is absent #2593 as well, because the old regex also kept only<p>elements.These are the results from running the parser on the real pages, fetched through a browser because Central Novel is behind Cloudflare:
master(#2593)<img><img><img><p>, 1<img><p>, 1<img>(same HTML as before #2593)What changed
plugins/multisrc/lightnovelwp/template.ts,parseChapteronly:div.epcontentwith cheerio, which the template already imports. The block ends where its element closes, sobottomnavis still not needed and a<script>inside the block does not cut it short. A page withoutepcontentstill returns''.script,styleandnoscriptfrom the block.<p>as HTML, plus any<img>that sits outside a paragraph, in document order. Other elements inside the block (such as ad wrappers) are still dropped, as they were before.src: if it is missing or adata:placeholder, usedata-lazy-srcordata-src. Any relativesrc(for exampleimages/p.jpg,../p.jpg,/p.jpgor//cdn/p.jpg) is resolved against the chapter page URL withnew URL(src, site + chapterPath).on*attributes. Anyhreforsrcusing ajavascript:,vbscript:or non-imagedata:scheme is removed. The check ignores case, whitespace and control characters.1.1.10to1.1.11, so every lightnovelwp source gets the update.Testing
CentralNovel[lightnovelwp].tsagainst the three saved Central Novel pages in the table, withfetchstubbed to serve the saved HTML, and compared the output with the pre-fix(lightnovelwp): parse chapter text when bottomnav is absent #2593 regex on the same HTML.<script>insideepcontentthat contains</p><div class=bottomnav>, an ad<div>, an paragraph, lazy and relative image sources, and a comments<p>outside the block. There was also a page with noepcontentblock, which returns''.images/page-1.jpg,../up/p2.jpgand a lazylazy/p3.jpgresolve against the chapter URL, and//cdn…becomeshttps:.onerror,ONLOADandonclickare removed, and so are the" JaVa	Script:",vbscript:anddata:text/htmlhrefs and ajavascript:img src. Anhttps:link and adata:image/pngimg are kept.npm run check:pluginon the generated lightnovelwp plugins.parseChapterpasses on BlumeVerse, DobyNovels, HyacinthinBloom, KnoxT, LazyGirlTranslations, TranslationWeaver and KodeksLibrary.parseChapter, with 404/403/503 responses or no novels/chapters. That is unrelated to this change.parseChapterreturns 0 chars both onmasterand with this PR. Its content block is<article class="epcontent">rather than adiv. I left it out of this PR to keep the change focused.npx prettier --checkandnpx eslintpass on the changed file, andnpm run build:compilesucceeds.npm run lintstill reports 3 errors, all in other files that this PR does not touch (rewayatfans.ts,wetriedtls.ts,wntl.ts).This PR was written by an AI agent (Claude Opus 5.5). A human has not reviewed it yet.