Conversation
… site markup
The chapter pages now render the body inside
.tldariinggrissendiribrojangancopy .entry-content as
p.ds-markdown-paragraph elements, so the old
div:contains('Daftar Isi') + anchor resolves to nothing and
parseChapter returned an empty string ("tidak ada konten yang
bisa dibaca"). Extract from the new container, strip the
site promo paragraph, fall back to the legacy selector.
Closes lnreader#2571
Co-Authored-By: Muse Spark <noreply@muse-spark.ai>
|
| const contentSelectors = [ | ||
| '.tldariinggrissendiribrojangancopy .entry-content', | ||
| '.entry-content', | ||
| ]; | ||
| for (const selector of contentSelectors) { | ||
| const content = loadedCheerio(selector).first(); | ||
| if (!content.length) continue; | ||
| content.find("p:contains('Baca novel lain di sakuranovel')").remove(); | ||
| const chapterText = (content.html() || '').trim(); | ||
| if (chapterText) return chapterText; | ||
| } | ||
|
|
||
| return chapterText; | ||
| let paragraphs = ''; | ||
| loadedCheerio('p.ds-markdown-paragraph').each((i, el) => { | ||
| paragraphs += loadedCheerio(el).toString(); | ||
| }); | ||
| if (paragraphs.trim()) return paragraphs; | ||
|
|
||
| return loadedCheerio("div:contains('Daftar Isi') +").html() || ''; |
There was a problem hiding this comment.
New extraction lacks regression tests These container, paragraph, and legacy paths have no retained chapter-markup fixtures. The live chapter check was blocked, so adding fixtures for current and older markup would help catch empty chapters or broken fallbacks without relying on site access.
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
…ar Isi fallback On the retained legacy path, strip the first inner div (anti-scrape / navigation markup) before returning the sibling HTML, mirroring the v1.0.1 cleanup. Guarded against multi-class values. Addresses Greptile review on lnreader#2573. Co-Authored-By: Muse Spark <noreply@muse-spark.ai>
…allback The p.ds-markdown-paragraph fallback concatenated every match, including the site promo line. Skip paragraphs containing 'Baca novel lain di sakuranovel', mirroring the entry-content branches above. Co-Authored-By: Muse Spark <noreply@muse-spark.ai>
…oval Apply Clean Coder review: hoist content selectors and promo text to commented constants, scope the legacy class removal to the chapter container, and use map/get/join for paragraph assembly. Co-Authored-By: firstmate-crewmate <crewmate@firstmate.local>
… ad scripts On real archived chapter pages (Nov 2024 - Apr 2025), the body sits in a rotating-class wrapper after an empty "please stop scrape my site" decoy, with no .entry-content or ds-markdown paragraphs, so every tier returned whitespace. Take the first non-empty sibling between the top and bottom .entry-pagination instead of the Daftar Isi sibling, and only drop the first inner div when it is an empty ad wrapper (on that layout it is the text itself). Also remove script/style/ins/noscript up front: on the 2023 layout the .entry-content tier returned inline adsbygoogle scripts that v1.0.1 had stripped.
Closes #2571
Closes #2619
What was wrong
parseChapteranchored the chapter body ondiv:contains('Daftar Isi') +. The site now puts an empty anti-scrape decoy div (please stop scrape my site) right after the chapter navigation, so that selector returns whitespace and chapters open as "tidak ada konten yang bisa dibaca".What changed
.tldariinggrissendiribrojangancopy .entry-content, then.entry-contentp.ds-markdown-paragraphparagraphs.entry-paginationnav. This skips the decoy and doesn't depend on the wrapper class, which changes over time (readerss,asdasd, ...). The v1.0.1 inner-div cleanup now only runs when that div is an empty ad wrapper.script, style, ins, noscriptare removed before extraction.1.0.1->1.0.2.Review fixes (
e9f65cb)A follow-up review ran the v1.0.1 code, the previous PR head (
9db150f) and this head against real chapter pages archived by the Wayback Machine, plus small hand-written inputs:div.asdasdafter the decoy and usually has no.entry-contentor ds-markdown paragraphs. v1.0.1 and9db150fboth returned whitespace on 5 of the 6 archived chapters with this layout. The new tier 3 returns the full text for all 6..entry-contenttier returned 7 inlineadsbygoogle<script>blocks that v1.0.1 used to strip. Scripts are now removed up front..entry-contentor with a different class now extracts, where9db150freturned empty.What was run
npm run check:plugin -- plugins/indonesian/sakuranovel.ts: INCONCLUSIVE,siteReachabilityHTTP 403 (Cloudflare). sakuranovel.id served a Cloudflare challenge to curl, headless Chromium and a reader proxy, so nothing was checked against the live site. The current live markup was not seen during this review.eslintandprettier --checkon the file, andnpm run build:compile.No test files were added. The repository has no unit-test suite, and maintainers asked on #2575 not to add any until a test framework exists.
AI-authored (the original commits and this review fix), with no human review or testing beyond what is listed above. Please weight review accordingly.