Repository navigation
feat(pr-review): score confidence on 1 to 10 so clean reviews read safe - #1276
Conversation
On the 1 to 5 scale every soft signal cost half a point and rounding went down, so a clean review with partial tests or a medium-risk area read 4/5 "likely safe" while its text said safe; 64 of the last 150 reviews landed there. Deductions are now whole points on a 10, findings cap at 8/6/4/2, and the reviewer's own number may lower by two or raise by one. The prompt carries calibration anchors from past reviews, the check is green from 7, and a migration doubles stored confidences.
Maple review🟡 Confidence 3/5 · needs attention Warning This review ended early; what follows is what it established. Moves the PR review confidence headline from 1–5 to 1–10: whole-point signal deductions, a reviewer number that may move the computed one by −2/+1, two levels per label band, and
Findings🟠 Warning · F1 · Migration doubles confidence with no upper boundcorrectness · The 🤖 Prompt to fix this finding with an AI agentWhat was checked
|
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (11)
Included review availability: This review used your included allowance. Your plan provides up to 4 included reviews per hour; 1 remain after this review. 📝 WalkthroughWalkthroughConfidence scoring changes from a 1–5 scale to a 1–10 scale. The pull request updates score calculation, AI review guidance, stored confidence values, backend check outcomes, web displays, documentation, and related tests. ChangesConfidence scale
Priority: ➖ Normal Estimated code review effort: 3 (Moderate) | ~20 minutes Change: Feature Merge Risk: ⚪ Minimal · up to This change moves review confidence to a 1–10 scale and updates stored values, prompts, displays and check outcomes to match. No concrete merge-blocking problem was identified. The migration must ship together with the code. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Docstring CoverageExplanation Docstring coverage is 33.33% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 3 functions across 8 files. (2 skipped: 2 unsupported.)
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Keeps main's telemetry contract-break failure conclusion and its telemetryDismissals prompt field alongside the 1 to 10 confidence scale. The confidence migration is regenerated after main's pr_review_telemetry migration, only doubles values still on the 1 to 5 scale, and doubles a stored "Held at N" cap reason with them.
Maple review🟢 Confidence 4/5 · likely safe to merge Rescales PR review confidence from 1–5 to 1–10 with whole-point deductions, matching caps, a recalibrated prompt and a migration that doubles stored values. Contained and tested; safe to merge in the same deploy as the migration.
Fixed since the last review
What was checked
|
Why
Feedback on the review bot: the confidence often reads 4/5 while the text reads like a 5/5. Replaying the last 150 review comments confirmed it: 64 landed on 4/5 "likely safe to merge", almost all with no findings. On the 1 to 5 scale each soft signal (
tests partial,risk medium) cost half a point and rounding went down, so any one of them cost a whole band. Another 19 were lowered to 3/5 by the reviewer's own −1, mostly with no reason shown.What changed
confidencePrReview(packages/domain/src/http/pr-review.ts) scores 1 to 10, two levels per band: 9-10 safe, 7-8 likely safe, 5-6 needs attention, 3-4 risky, 1-2 do not merge.confidenceReason, and adds six calibration anchors drawn from real past reviews./10; the check issuccessfrom 7 (was 4).20261006214302_pr_review_confidence_tendoubles storedreport_json.confidence(4→8, 5→10) so history and the analytics average stay on one scale.Replay on the last 150 reviews
Reviewer notes
PR_REVIEW_CONFIDENCE_JUDGEMENT.raiseto 0.apps/aichat tests pass; typecheck clean for domain, backend, ai, web; migration SQL run against PGlite sample rows.Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.Summary by CodeRabbit