o3's 87.5% ARC Score Should Worry You

December 23, 2024

OpenAI seems to have casually walked into AGI territory last week, and no one noticed.

Tucked away in their blizzard of "12 Days of OpenAI" announcements, they nonchalantly mentioned their new o3 model scored 87.5% on the ARC Prize benchmark. If you're not an AI expert, that's meaningless. Here's what it actually means: they may have just created something that can think.

ARC isn't a typical AI test. It's not about regurgitating facts or recognizing patterns. It presents simple visual puzzles, the kind found in IQ tests. Examine these shapes, identify the pattern, and apply it to a new example. Humans find these puzzles relatively straightforward. Most AIs fail spectacularly.

Before now, the top models barely exceeded 50%. GPT-4 wasn't close. Claude couldn't touch it. ARC was built to resist the cheap pattern matching tricks that make modern AI seem clever. You need genuine reasoning ability, the capacity to understand abstract rules and apply them to new situations.

Then o3 achieves an 87.5%.

The AI safety community is deeply alarmed, and I don't blame them. That score came from a run using 172 times the compute of o3's own high-efficiency configuration, which managed 75.7%. It clears the 85% number attached to the ARC Grand Prize but not the efficiency condition that comes with it. And if they follow their usual playbook, this is only the beginning. They chose "o3" because "o2" sounds like a British telecom. These folks are moving so quickly they're running out of names.

But here's the real cause for concern: the acceleration. In 2024 alone, AI performance on major benchmarks exploded. MMMU scores increased 18.8 points. GPQA jumped 48.9 points. SWE-bench, a coding test, skyrocketed 67.3 points. In one year.

I've been closely following the AI race since ChatGPT launched. The progress curve is no longer linear or even exponential. It's a vertical line. Each breakthrough fuels the next, arriving faster each time. Better models help researchers create even better ones. Lather, rinse, repeat. The feedback loop is on the verge of ripping itself apart.

The timing of o3 is impeccable, and by that I mean deeply troubling. OpenAI drops this bombshell right before the holidays, while everyone is distracted with travel and family gatherings. Congress is on recess. Tech journalists are writing year-in-review fluff pieces. By the time we've processed what happened, they'll have moved on to o4.

Which they undoubtedly will, likely within months. Because that's what's truly terrifying: o3 isn't the end goal. It's just another milestone on the path to something we haven't fully considered.

What is AGI, really? Artificial General Intelligence refers to AI that can match or surpass humans at everything, not just narrow domains like chess or writing. The kind of intelligence that can acquire new skills as rapidly as we can, reason through entirely novel problems, and potentially even improve itself.

We were told AGI was decades away. Careful researchers said 2050. Optimists said 2035. Doomsayers said 2030. But if o3 truly broke the AGI barrier, we're not talking about the distant future anymore. We're talking about next week.

From here, the cascade of consequences is as predictable as it is alarming. First, the world will insist o3 isn't "really" AGI. They'll complain about all the things it can't do, all the ways it fails. They'll keep moving the goalposts, claiming AGI suddenly means something else now. Anything to keep reality at bay a little longer.

Meanwhile, OpenAI will keep building. o4 will hit 95% on ARC. o5 will ace every test we devise. Sooner or later, pretending it's not happening will no longer be an option. But by then, it will be far too late.

I'm 17. I thought I had my whole life to figure this out: go to college, start a career, maybe have a family someday. Now I wonder if any of that makes sense anymore. Why work relentlessly studying computer science when AI will surpass me at coding before I even declare a major? What's the point of pursuing any career when AGI will do it better and cheaper while I'm still an intern?

My parents think I've gone off the rails. They lived through Y2K and all the unfulfilled doomsday predictions. But this is different. Y2K was a specific technical issue with a clear solution. AGI is Pandora's Box. Once opened, good luck closing it again.

What's driving me crazy is the deafening silence. No emergency Congressional sessions, no major UN declarations, no Manhattan Project to prevent the AI from destroying us. Just OpenAI dropping the "we might have just made humans obsolete" bombshell between announcing ChatGPT's Santa Mode like it's no big deal.

Maybe I've got it all wrong. Maybe o3 is just another incremental step and true AGI is still years away. Maybe the lab researchers will solve alignment before the clock strikes midnight. Maybe the government will wake up and intervene before it all spirals out of control.

But the data is unambiguous. 87.5% on ARC. December 2024. We just crossed a line that was supposed to remain uncrossed for another decade. And instead of pausing to consider what that means, we're slamming the accelerator to the floor.