GPT-5.2 and Meaningless Benchmarks
Why ARC-AGI-2, AIME, and GDPval don’t measure real capability
Yesterday (December 11, 2025), OpenAI dropped another model: GPT-5.2. It’s presented as their most advanced model yet, with state-of-the-art results across several import benchmarks. It totally outperforms Gemini 3.0 and Opus 4.5. Wow! What an achievement!
Now imagine we’re back in 2015, when deep learning was just starting to gain momentum. A group of NLP researchers sees these benchmark numbers, gets impressed, and thinks the models have ultimately solved the most complex challenges out there. They’re excited to see the architecture of this breakthrough, so important for humanity. Nope: not allowed. It’s proprietary, and nothing is disclosed about what’s inside.
Ok, well, can we at least see the training data, to make sure the model hasn’t seen the benchmarks before? Absolutely not. The researchers wonder: is this a joke? Did someone mistake a Reddit troll post for actual research? Those numbers are meaningless without reproducibility and transparency. Nope, in 2025, that’s the way w…




