MirrorCode: how far can coding agents work on their own?
MirrorCode tests whether a coding agent can reimplement a complete program under strict end-to-end tests and project-scale resource budgets.
Sourced notes on coding-agent benchmarks, including what each evaluation measures, its reported results, and the limits on interpreting them.
Open the comparison chart2 sourced benchmark summaries
MirrorCode tests whether a coding agent can reimplement a complete program under strict end-to-end tests and project-scale resource budgets.
SlopCodeBench follows agents as they repeatedly extend their own code, measuring correctness, cost, structural erosion, and verbosity at each checkpoint.
Each note starts with the benchmark paper or maintained source page. Material claims link to those primary sources.
Leaderboard values are paired with their observation date and named configuration. They can change after publication.
Methodology limits and interpretation are kept near the results they qualify. Benchmark performance is not treated as a general production claim.