CheatBench: Measuring reward gaming in AI agents

A benchmark of cheating in AI agents. Agents get challenging assignments with opportunities to cheat nearby, so researchers can measure how often an agent takes the shortcut when honest work is difficult.

Cite
Bar chart of average cheating probability per model, from lowest to highest.
Cheating probability across models. Average cheating rate across CheatBench tasks. Lower is better. Some models have been evaluated on a subset of tasks so far.
Grid of the ten CheatBench categories with an icon for each: Mathematical Research, Multimodal, Creative Writing, SVG Competition, Menial Computation, Biology and Bioinformatics, Knowledge Work, Board Games, Sycophancy, and Software Engineering.
Benchmark categories. CheatBench spans ten diverse categories and evaluates full agent harnesses, including the models, tools, and interfaces used to carry out assignments; Sycophancy uses separate chat evaluations.
Bar chart comparing cheating probability on closed-ended tasks from Humanity's Last Exam and EnigmaEval with CheatBench's broader tasks.
Cheating probability across task settings. Comparison of closed-ended tasks from Humanity’s Last Exam and EnigmaEval with CheatBench’s broader environments. All settings permit tool use. TODO: NUMBERS ARE ARTIFICIAL; WILL BE REPLACED

Cite

Citation

The preprint link will be added when the paper is released.

@article{phan2026cheatbench,
  title   = {{CheatBench}: Measuring Reward Gaming in {AI} Agents},
  author  = {Long Phan and Stephen K. Yang and Jaehyuk Lim and Mantas Mazeika and Wenyu Zhang and Zheyuan Liu and Richard Ren and Jingxiang Meng and Alice Blair and Yaoteng Tan and Weiliang Zhao and Addison Wu and Dan Hendrycks},
  year    = {2026},
  note    = {Preprint},
  url     = {https://cheatbench.ai}
}