Skip to content

fix(indonesian/sakuranovel): update chapter body selector for current site markup - #2573

Open
RibatTRW wants to merge 5 commits into
lnreader:masterfrom
RibatTRW:fm/lnreader-sakuranovel-2571-empty-chapter
Open

RibatTRW wants to merge 5 commits into
lnreader:masterfrom
RibatTRW:fm/lnreader-sakuranovel-2571-empty-chapter

Conversation

@RibatTRW

@RibatTRW RibatTRW commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Closes #2571
Closes #2619

What was wrong

parseChapter anchored the chapter body on div:contains('Daftar Isi') +. The site now puts an empty anti-scrape decoy div (please stop scrape my site) right after the chapter navigation, so that selector returns whitespace and chapters open as "tidak ada konten yang bisa dibaca".

What changed

  • Chapter body lookup, in order:
    1. .tldariinggrissendiribrojangancopy .entry-content, then .entry-content
    2. p.ds-markdown-paragraph paragraphs
    3. the first non-empty sibling between the top and bottom .entry-pagination nav. This skips the decoy and doesn't depend on the wrapper class, which changes over time (readerss, asdasd, ...). The v1.0.1 inner-div cleanup now only runs when that div is an empty ad wrapper.
  • script, style, ins, noscript are removed before extraction.
  • The "Baca novel lain di sakuranovel" promo line is removed on every path.
  • Version 1.0.1 -> 1.0.2.

Review fixes (e9f65cb)

A follow-up review ran the v1.0.1 code, the previous PR head (9db150f) and this head against real chapter pages archived by the Wayback Machine, plus small hand-written inputs:

  • Fixed: empty chapters on the archived Nov 2024 - Apr 2025 layout. The body is in div.asdasd after the decoy and usually has no .entry-content or ds-markdown paragraphs. v1.0.1 and 9db150f both returned whitespace on 5 of the 6 archived chapters with this layout. The new tier 3 returns the full text for all 6.
  • Fixed: ad scripts leaking into chapters. On an archived 2023 chapter, the .entry-content tier returned 7 inline adsbygoogle <script> blocks that v1.0.1 used to strip. Scripts are now removed up front.
  • Hand-written inputs: an empty body, a page with no navigation, and an images-only body all behave correctly. A wrapper without .entry-content or with a different class now extracts, where 9db150f returned empty.

What was run

  • npm run check:plugin -- plugins/indonesian/sakuranovel.ts: INCONCLUSIVE, siteReachability HTTP 403 (Cloudflare). sakuranovel.id served a Cloudflare challenge to curl, headless Chromium and a reader proxy, so nothing was checked against the live site. The current live markup was not seen during this review.
  • Offline before/after run with the repo's cheerio on the 7 archived chapter pages and 6 hand-written inputs described above.
  • eslint and prettier --check on the file, and npm run build:compile.

No test files were added. The repository has no unit-test suite, and maintainers asked on #2575 not to add any until a test framework exists.

AI-authored (the original commits and this review fix), with no human review or testing beyond what is listed above. Please weight review accordingly.

… site markup

The chapter pages now render the body inside
.tldariinggrissendiribrojangancopy .entry-content as
p.ds-markdown-paragraph elements, so the old
div:contains('Daftar Isi') + anchor resolves to nothing and
parseChapter returned an empty string ("tidak ada konten yang
bisa dibaca"). Extract from the new container, strip the
site promo paragraph, fall back to the legacy selector.

Closes lnreader#2571

Co-Authored-By: Muse Spark <noreply@muse-spark.ai>
@RibatTRW
RibatTRW marked this pull request as draft September 25, 2026 12:07
@greptile-apps

greptile-apps Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[Medium risk] Updates chapter content extraction for a novel scraper plugin.

The PR appears safe to merge, though retained regression coverage for the extraction paths would be useful.

Findings

  1. P2 New extraction lacks regression tests ▶

Summary

The plugin now tries current chapter-content containers and markdown paragraphs before the legacy selector, removes the known promotional paragraph, and updates its version to 1.0.2.

Reviews (5) · Last reviewed commit: "fix(indonesian/sakuranovel): name fallba..."

Comment thread plugins/indonesian/sakuranovel.ts Outdated
Comment thread plugins/indonesian/sakuranovel.ts Outdated
Comment on lines +140 to +158
const contentSelectors = [
'.tldariinggrissendiribrojangancopy .entry-content',
'.entry-content',
];
for (const selector of contentSelectors) {
const content = loadedCheerio(selector).first();
if (!content.length) continue;
content.find("p:contains('Baca novel lain di sakuranovel')").remove();
const chapterText = (content.html() || '').trim();
if (chapterText) return chapterText;
}

return chapterText;
let paragraphs = '';
loadedCheerio('p.ds-markdown-paragraph').each((i, el) => {
paragraphs += loadedCheerio(el).toString();
});
if (paragraphs.trim()) return paragraphs;

return loadedCheerio("div:contains('Daftar Isi') +").html() || '';

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 New extraction lacks regression tests These container, paragraph, and legacy paths have no retained chapter-markup fixtures. The live chapter check was blocked, so adding fixtures for current and older markup would help catch empty chapters or broken fallbacks without relying on site access.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

RibatTRW and others added 2 commits September 25, 2026 20:23
…ar Isi fallback

On the retained legacy path, strip the first inner div (anti-scrape /
navigation markup) before returning the sibling HTML, mirroring the
v1.0.1 cleanup. Guarded against multi-class values.

Addresses Greptile review on lnreader#2573.

Co-Authored-By: Muse Spark <noreply@muse-spark.ai>
…allback

The p.ds-markdown-paragraph fallback concatenated every match,
including the site promo line. Skip paragraphs containing
'Baca novel lain di sakuranovel', mirroring the entry-content
branches above.

Co-Authored-By: Muse Spark <noreply@muse-spark.ai>
@RibatTRW
RibatTRW marked this pull request as ready for review September 25, 2026 12:46
…oval

Apply Clean Coder review: hoist content selectors and promo text to
commented constants, scope the legacy class removal to the chapter
container, and use map/get/join for paragraph assembly.

Co-Authored-By: firstmate-crewmate <crewmate@firstmate.local>
@RibatTRW
RibatTRW marked this pull request as draft September 25, 2026 13:26
@RibatTRW
RibatTRW marked this pull request as ready for review September 25, 2026 13:35
@RibatTRW
RibatTRW marked this pull request as draft September 25, 2026 13:36
@RibatTRW
RibatTRW marked this pull request as ready for review September 25, 2026 13:40
@RibatTRW
RibatTRW marked this pull request as draft September 25, 2026 13:41
@RibatTRW
RibatTRW marked this pull request as ready for review September 25, 2026 13:43
… ad scripts

On real archived chapter pages (Nov 2024 - Apr 2025), the body sits in a
rotating-class wrapper after an empty "please stop scrape my site" decoy,
with no .entry-content or ds-markdown paragraphs, so every tier returned
whitespace. Take the first non-empty sibling between the top and bottom
.entry-pagination instead of the Daftar Isi sibling, and only drop the
first inner div when it is an empty ad wrapper (on that layout it is the
text itself).

Also remove script/style/ins/noscript up front: on the 2023 layout the
.entry-content tier returned inline adsbygoogle scripts that v1.0.1 had
stripped.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant