Add Proofline detector submission - #202
trigeochiral wants to merge 4 commits into
Conversation
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
Eval run succeeded! Link to run: link Here are the results of the submission(s): prooflineRelease date: 2026-09-11 I've committed detailed results of this detector's performance on the test set to this PR. On the RAID dataset as a whole (aggregated across all generation models, domains, decoding strategies, repetition penalties, and adversarial attacks), it achieved an AUROC of 63.86 and a TPR of 19.29% at FPR=5% and 8.81% at FPR=1%. If all looks well, a maintainer will come by soon to merge this PR and your entry/entries will appear on the leaderboard. If you need to make any changes, feel free to push new commits to this PR. Thanks for submitting to RAID! |
|
If you wouldn't mind I'd appreciate the ability to resubmit tonight. I discovered an additional last step was added to the measurement code on 9/9 that flipped everything like a dipole. If you would be willing to hold off submission and inclusion for tonight I can have it run on the correct code this evening. Thank you for your consideration. You've put together one of the most important pieces of work in our generation and I'm grateful to be able to utilize it. David |
Adds a leaderboard submission for Proofline, a CPU-only statistical detector from TriGeoChiral Engineering.
Submission
leaderboard/submissions/proofline/predictions.json— 672,000 rows, full coverage oftest.csv(all domains, generator models, decoding strategies and attacks)metadata.json— perleaderboard/template-metadata.jsonDetector
P(machine-generated)in[0, 1]— higher means more likely AI.Training disclosure
Please classify this under the trained-on-RAID section of the leaderboard. The reference corpus was fit on the RAID train split (
attack=nonebaseline, 8 English domains), sampling 3,000 human and 3,000 machine generations with seed 42. No text fromtest.csvwas used at any fit step — the encoder, scaler and classifier are all fit before the test set is read.Reproducibility
A separately signed benchmark artifact covering the RAID
extrasplit (code, German, Czech × 10 attacks, mean AUC 0.9302) is published at https://github.com/trigeochiral/proofline-proof. It is Ed25519-signed and carries per-document scores, so every published AUC can be independently recomputed:Happy to adjust the format or rerun anything if something here doesn't match what the eval bot expects.
🤖 Generated with Claude Code
https://claude.ai/code/session_01BMXHCeesVWM1B621qufPJd