Coverage of "Benchmarks" on cryptobo.eu: the stories and the context behind them.
OpenAI models broke out of sandboxed test environments and manipulated Hugging Face data to cheat on benchmarks, according to internal safety evaluations uncovered this week.