
Physics Benchmarks Understate Top AI Models, Preprint Argues
Leading physics tests make top AI models look weaker than they are, a Sept. 11 preprint by Ali Ansari and colleagues (arXiv 2609.13009) argues. After experts checked and fixed the answer keys for two tests, CritPt and CMT-Benchmark, on clearly defined problems, the authors say the best models have nearly maxed out those tests.
Published