CheatBench: Measuring reward gaming in AI agents
A benchmark of cheating in AI agents. Agents get challenging assignments with opportunities to cheat nearby, so researchers can measure how often an agent takes the shortcut when honest work is difficult.
Cite
Citation
The preprint link will be added when the paper is released.
@article{phan2026cheatbench,
title = {{CheatBench}: Measuring Reward Gaming in {AI} Agents},
author = {Long Phan and Stephen K. Yang and Jaehyuk Lim and Mantas Mazeika and Wenyu Zhang and Zheyuan Liu and Richard Ren and Jingxiang Meng and Alice Blair and Yaoteng Tan and Weiliang Zhao and Addison Wu and Dan Hendrycks},
year = {2026},
note = {Preprint},
url = {https://cheatbench.ai}
}