Physics Benchmarks Understate Top AI Models, Preprint Argues

Physics Benchmarks Understate Top AI Models, Preprint Argues

Leading physics tests make top AI models look weaker than they are, a Sept. 11 preprint by Ali Ansari and colleagues (arXiv 2609.13009) argues. After experts checked and fixed the answer keys for two tests, CritPt and CMT-Benchmark, on clearly defined problems, the authors say the best models have nearly maxed out those tests.

Published

Read at another depth