This repository accompanies our paper: Can We Trust LLM Security Research? A Reproducibility Study of Large Language Model Security Papers in Tier-1 Security Conferences
We conduct a systematic reproducibility study of LLM security research published at ACM CCS, USENIX Security, IEEE S&P, and NDSS between 2023 and 2025.
Our study includes 106 papers, among which 84 provide executable artifacts that were independently evaluated under a standardized reproduction protocol.
-
conference_papers.csv
The full list of the 106 papers included in our study. -
reproduction_results.csv
An anonymized artifact-level dataset covering the 84 evaluated artifacts. It includes reproduction outcomes, failure classifications, and checklist scores. -
checklist_demo.md
A description and example of the checklist used in our artifact evaluation.
Among the 84 evaluated artifacts, 33 were successfully reproduced under our evaluation protocol, corresponding to an observed reproduction rate of 39.29%.