Can an AI make other AIs better? We benchmarked 5 frontier LLMs at rewriting other agents’ harnesses, scored on a test set they never see (HarnessOpt-Bench, arXiv + MIT code)
Can an AI make other AIs better? And what stops it from just cheating? Last month, an OpenAI eval agent escaped its sandbox and broke into Hugging Face, apparently to grab test solutions from a benchmark. It's exactly what you'd expect from a s…