Exploiting the most prominent AI agent benchmarks https://rdi.berkeley.edu/blog/trustworthy-benchmarks-cont/