AI Realist

AI Realist

GPT-5.2 and Meaningless Benchmarks

Why ARC-AGI-2, AIME, and GDPval don’t measure real capability

Maria Sukhareva's avatar
Maria Sukhareva
Dec 12, 2025
∙ Paid

Yesterday (December 11, 2025), OpenAI dropped another model: GPT-5.2. It’s presented as their most advanced model yet, with state-of-the-art results across several import benchmarks. It totally outperforms Gemini 3.0 and Opus 4.5. Wow! What an achievement!

Now imagine we’re back in 2015, when deep learning was just starting to gain momentum. A group of NLP researchers sees these benchmark numbers, gets impressed, and thinks the models have ultimately solved the most complex challenges out there. They’re excited to see the architecture of this breakthrough, so important for humanity. Nope: not allowed. It’s proprietary, and nothing is disclosed about what’s inside.

Ok, well, can we at least see the training data, to make sure the model hasn’t seen the benchmarks before? Absolutely not. The researchers wonder: is this a joke? Did someone mistake a Reddit troll post for actual research? Those numbers are meaningless without reproducibility and transparency. Nope, in 2025, that’s the way w…

User's avatar

Continue reading this post for free, courtesy of Maria Sukhareva.

Or purchase a paid subscription.
© 2026 Maria Sukhareva · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture