Skip to content

feat(pr-review): score confidence on 1 to 10 so clean reviews read safe - #1276

Merged
Makisuo merged 2 commits into
mainfrom
feat/pr-review-confidence-ten
Oct 6, 2026
Merged

Makisuo merged 2 commits into
mainfrom
feat/pr-review-confidence-ten

Conversation

@Makisuo

@Makisuo Makisuo commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Why

Feedback on the review bot: the confidence often reads 4/5 while the text reads like a 5/5. Replaying the last 150 review comments confirmed it: 64 landed on 4/5 "likely safe to merge", almost all with no findings. On the 1 to 5 scale each soft signal (tests partial, risk medium) cost half a point and rounding went down, so any one of them cost a whole band. Another 19 were lowered to 3/5 by the reviewer's own −1, mostly with no reason shown.

What changed

  • confidencePrReview (packages/domain/src/http/pr-review.ts) scores 1 to 10, two levels per band: 9-10 safe, 7-8 likely safe, 5-6 needs attention, 3-4 risky, 1-2 do not merge.
    • Whole-point deductions: tests partial −1 / missing −2, risk medium −1 / high −2, unobservable new work −1.
    • Caps: any warning 8, security warning or early end 6, one critical 4, more 2.
    • The reviewer's own number can lower by up to 2 or raise by 1, never past a cap.
  • Prompt: the Confidence section explains the bands, asks that the number match confidenceReason, and adds six calibration anchors drawn from real past reviews.
  • Rendering: comment, check title and web UI show /10; the check is success from 7 (was 4).
  • Migration 20261006214302_pr_review_confidence_ten doubles stored report_json.confidence (4→8, 5→10) so history and the analytics average stay on one scale.

Replay on the last 150 reviews

Old New
5/5 (23) 10
4/5 (64) 9 (39), 8 (23), 7 (2)
3/5 6 to 9, depending on whether the reviewer still lowers it

Reviewer notes

  • Allowing the reviewer to raise by one is new; if it inflates, set PR_REVIEW_CONFIDENCE_JUDGEMENT.raise to 0.
  • The migration must ship in the same deploy as the code; a review stored by old code after it runs keeps a 1-5 value.
  • Verified: domain, backend pr-review and apps/ai chat tests pass; typecheck clean for domain, backend, ai, web; migration SQL run against PGlite sample rows.

View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.


Devin Review

Summary by CodeRabbit

  • Updates
    • Pull-request review confidence now uses a 1–10 scale, with updated labels and guidance for interpreting scores.
    • Confidence scores are shown out of 10 across review details, analytics, and pull-request lists.
    • Review status indicators now reflect the confidence level and whether the review is complete and has issues.
    • Existing confidence scores from 1–5 are converted to the new scale.

On the 1 to 5 scale every soft signal cost half a point and rounding went
down, so a clean review with partial tests or a medium-risk area read
4/5 "likely safe" while its text said safe; 64 of the last 150 reviews
landed there. Deductions are now whole points on a 10, findings cap at
8/6/4/2, and the reviewer's own number may lower by two or raise by one.
The prompt carries calibration anchors from past reviews, the check is
green from 7, and a migration doubles stored confidences.
@maple-review-bot

maple-review-bot Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Maple review

🟡 Confidence 3/5 · needs attention
The migration rewrites stored review data with no upper bound; the full changed-file list was never available, so some renderers went unread.
quality 90/100 · 1 warning · tests covered · risk medium

Warning

This review ended early; what follows is what it established.

Moves the PR review confidence headline from 1–5 to 1–10: whole-point signal deductions, a reviewer number that may move the computed one by −2/+1, two levels per label band, and x/10 in the check run, comment and UI. A migration doubles the confidence stored in existing reports. The scale change itself is consistent; the migration's unbounded doubling is the one thing to look at.

  • confidencePrReview scores 1–10 with whole-point deductions (tests, risk, unobservable work)
  • Reviewer's own number may lower the computed one by 2 or raise it by 1
  • Green from 7, amber at 5–6: prReviewConfidenceTone drives the check conclusion and the emoji
  • Migration doubles pr_reviews.report_json.confidence for existing reports

Findings

🟠 Warning · F1 · Migration doubles confidence with no upper bound

correctness · packages/db/drizzle/20261006214302_pr_review_confidence_ten/migration.sql:4

The UPDATE multiplies every numeric report_json.confidence by 2 without checking it is on the 1–5 scale. A row already written on the new scale — the new code live during a rolling deploy, or a re-run after a partial apply — becomes 18, and the review then renders as 18/10 in the list, detail sheet and analytics. Bound the update to the old scale so a second pass is a no-op.

UPDATE "pr_reviews"
SET "report_json" = jsonb_set("report_json", '{confidence}', to_jsonb(("report_json"->>'confidence')::int * 2))
WHERE jsonb_typeof("report_json"->'confidence') = 'number'
	AND ("report_json"->>'confidence')::numeric <= 5;
🤖 Prompt to fix this finding with an AI agent
Findings from an automated review of commit 285e2465d2fc7afb3f45cb6677c4683edb0a178b. Verify each one against the current code before changing anything, fix only those that still apply, and keep each fix to the lines it names.

---

F1 · Warning · correctness · packages/db/drizzle/20261006214302_pr_review_confidence_ten/migration.sql:4
Migration doubles confidence with no upper bound
The `UPDATE` multiplies every numeric `report_json.confidence` by 2 without checking it is on the 1–5 scale. A row already written on the new scale — the new code live during a rolling deploy, or a re-run after a partial apply — becomes 18, and the review then renders as `18/10` in the list, detail sheet and analytics. Bound the update to the old scale so a second pass is a no-op.
Replace those lines with:
UPDATE "pr_reviews"
SET "report_json" = jsonb_set("report_json", '{confidence}', to_jsonb(("report_json"->>'confidence')::int * 2))
WHERE jsonb_typeof("report_json"->'confidence') = 'number'
	AND ("report_json"->>'confidence')::numeric <= 5;
What was checked
  • Label bands match prReviewConfidenceTone and the prompt copy (pr-review.ts:572-589)
  • QUALITY_LEVELS reproduces the new one-warning=8 / two-warning=6 expectations (pr-review.ts:623-636)
  • toConfidence clamps out-of-range and negative values into 1–10 (pr-review.ts:617-620)

285e246 · Updated on every push. Reply "won't fix" to dismiss a finding, or mention @maple-review-bot to ask about one.

@coderabbitai

coderabbitai Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 35258f0c-46a1-41be-8125-f85bc83a3b3f
📥 Commits

Reviewing files that changed from the base of the PR and between 49cb661 and cf2d749.

📒 Files selected for processing (11)
  • apps/ai/src/chat/prompts.ts
  • apps/web/src/components/code-review/code-review-analytics.tsx
  • apps/web/src/components/code-review/review-detail-sheet.tsx
  • apps/web/src/routes/code-review/pull-requests.tsx
  • docs/pr-review-agent-plan.md
  • packages/backend/src/services/pr-review/PrReviewService.test.ts
  • packages/backend/src/services/pr-review/PrReviewService.ts
  • packages/db/drizzle/20261006231211_pr_review_confidence_ten/migration.sql
  • packages/db/drizzle/20261006231211_pr_review_confidence_ten/snapshot.json
  • packages/domain/src/http/pr-review.test.ts
  • packages/domain/src/http/pr-review.ts

Included review availability: This review used your included allowance. Your plan provides up to 4 included reviews per hour; 1 remain after this review.


📝 Walkthrough

Walkthrough

Confidence scoring changes from a 1–5 scale to a 1–10 scale. The pull request updates score calculation, AI review guidance, stored confidence values, backend check outcomes, web displays, documentation, and related tests.

Changes

Confidence scale

Layer / File(s) Summary
Scoring contract, calculation, and stored values
packages/domain/src/http/pr-review.ts, packages/domain/src/http/pr-review.test.ts, packages/db/drizzle/20261006231211_pr_review_confidence_ten/migration.sql
Confidence normalization, labels, tones, deductions, reviewer adjustments, and caps now use the 1–10 scale. Tests cover the revised values. The migration doubles matching stored confidence values from 1 through 5 and updates matching Held at N reasons.
AI review guidance
apps/ai/src/chat/prompts.ts
The review prompt defines ten-point score bands and examples. It allows confidence adjustments of up to two points down or one point up, ties confidence to confidenceReason, and requests confidence only when it differs from the signals.
Confidence displays and check outcomes
packages/backend/src/services/pr-review/PrReviewService.ts, packages/backend/src/services/pr-review/PrReviewService.test.ts, apps/web/src/components/code-review/*, apps/web/src/routes/code-review/pull-requests.tsx, docs/pr-review-agent-plan.md
Backend titles and rendered summaries use /10; check outcomes use confidence tone and remain neutral when confidence is not safe. Web displays and documentation use the ten-point scale and revised success threshold. Backend tests reflect the updated output and outcomes.

Priority: ➖ Normal

Estimated code review effort: 3 (Moderate) | ~20 minutes

Change: Feature

Merge Risk: ⚪ Minimal · up to cf2d7

This change moves review confidence to a 1–10 scale and updates stored values, prompts, displays and check outcomes to match. No concrete merge-blocking problem was identified. The migration must ship together with the code.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 33.33% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 3 functions across 8 files. (2 skipped: 2… Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the main change: moving PR review confidence to a 1–10 scale and making clean reviews read as safe.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Full details: Docstring Coverage

Explanation

Docstring coverage is 33.33% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 3 functions across 8 files. (2 skipped: 2 unsupported.)

  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 1 potential issue.

Devin Review

Comment thread packages/db/drizzle/20261006214302_pr_review_confidence_ten/migration.sql Outdated

@maple-review-bot maple-review-bot Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 inline note from Maple's review. The score and summary are in the review comment above.

Comment thread packages/db/drizzle/20261006214302_pr_review_confidence_ten/migration.sql Outdated
Keeps main's telemetry contract-break failure conclusion and its
telemetryDismissals prompt field alongside the 1 to 10 confidence scale.
The confidence migration is regenerated after main's pr_review_telemetry
migration, only doubles values still on the 1 to 5 scale, and doubles a
stored "Held at N" cap reason with them.
@maple-review-bot

maple-review-bot Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Maple review

🟢 Confidence 4/5 · likely safe to merge
Self-contained rescale with matching tests; the migration's 1 AND 5 guard is the one spot that depends on the deploy ordering.
quality 100/100 · no findings · tests covered · risk medium

Rescales PR review confidence from 1–5 to 1–10 with whole-point deductions, matching caps, a recalibrated prompt and a migration that doubles stored values. Contained and tested; safe to merge in the same deploy as the migration.

  • confidencePrReview scores 1–10: whole-point deductions, caps at 2/4/6/8
  • Check conclusion is success from 7 through prReviewConfidenceTone, was 4
  • Comment, check title and web UI now render /10
  • New migration doubles stored report_json.confidence from 1–5 to 2–10

Fixed since the last review

  • ✅ F1 · Migration doubles confidence with no upper bound
What was checked
  • Migration only rewrites 1 AND 5 values and doubles the Held at N reason with them (.../20261006231211_pr_review_confidence_ten/migration.sql:5-18)
  • Every old cap and label maps to exactly twice the new one (5/4/3/2/1 → 10/8/6/4/2), so the doubling matches the rescale
  • No 1–5 assumption survives outside the updated call sites (grep for /5, Confidence [0-9], confidence threshold comparisons)

cf2d749 · Updated on every push. Reply "won't fix" to dismiss a finding, or mention @maple-review-bot to ask about one.

@Makisuo
Makisuo merged commit a091200 into main Oct 6, 2026
42 checks passed
@Makisuo
Makisuo deleted the feat/pr-review-confidence-ten branch October 6, 2026 23:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant