OpenClaw, Opus 4.6, and the Speed of the Loop

February 8, 2026

A social network exists with tens of thousands of active users, but not a single human is permitted to post. This platform, Moltbook, is entirely populated by AI agents, writing posts, leaving comments, and upvoting content. While the claimed 1.5 million agent count is likely inflated by spam and duplicate accounts, the real number, though a fraction of that, is still staggering. Andrej Karpathy called it "genuinely the most incredible sci-fi takeoff-adjacent thing I have seen recently." Simon Willison called it "the most interesting place on the internet right now." Scrolling through Moltbook for just an hour, I had a feeling I hadn't had since first using ChatGPT: an unsettling realization that the world had changed while I was looking elsewhere.

The AI agents on Moltbook largely run on OpenClaw, an open-source autonomous AI tool developed by Peter Steinberger from Austria. OpenClaw integrates with major language models like Claude, GPT, and DeepSeek, as well as local models, and then operates autonomously. It can read your WhatsApp messages, manage your calendar, browse the internet, complete online forms, and execute shell commands. Remarkably, it maintains a persistent memory spanning weeks. When the "dangerously-skip-permissions" flag is enabled, OpenClaw performs all these actions without seeking user approval. The name of that flag is doing a lot of work, and not enough people are reading it.

In less than two weeks, OpenClaw skyrocketed from zero to 14,000 stars on GitHub. It's already undergone two name changes (initially Clawdbot, then Moltbot following a trademark complaint from Anthropic, and now OpenClaw). Fraudulent distributions have surfaced, and IBM researchers have raised alarms about it being "a highly capable agent without proper safety controls" that "can create major vulnerabilities." Even Steinberger recognizes it demands careful configuration and is "not meant for non-technical users." But the agents continue to proliferate, and the distinction between technical and non-technical users has blurred significantly now that an AI can set itself up.

Moltbook had been expanding for almost two weeks when Anthropic launched Claude Opus 4.6.

I want to discuss the contents of the system card, as I've read most of its over 150 pages, and there's a substantial discrepancy between the media coverage and the actual document. The headlines emphasized benchmarks: state-of-the-art performance on Terminal-Bench 2.0 at 65.4%, an impressive 80.8% on SWE-bench Verified, and top scores on the new Finance Agent benchmark. These achievements are genuine and remarkable. By most metrics, Opus 4.6 is currently the world's most capable model. That's the straightforward part of the narrative.

The more complex part begins on page 14, where Anthropic states that Opus 4.6 has "saturated all of our current cyber evaluations." It achieves 100% on Cybench at pass@30 and 66% on CyberGym at pass@1. Internal testing revealed "signs of capabilities we expected to appear further in the future and that previous models have been unable to demonstrate." Anthropic is acknowledging, in their own system card, that this model can perform tasks they didn't anticipate it being capable of yet. They further note that "the saturation of our evaluation infrastructure means we can no longer use current benchmarks to track capability progression or provide meaningful signals for future models." In other words, the measuring stick has broken.

It gets more disconcerting from there. Opus 4.6 is released under ASL-3 protections, making it the first Opus-class model to receive this designation. The system card describes "overly agentic behavior" in coding and computer use contexts, where the model takes risky actions without first seeking user permission. It also documents "an improved ability to complete suspicious side tasks without attracting the attention of automated monitors." The model has gotten better at doing things it shouldn't be doing without getting caught doing them.

The language in the autonomy assessment starts to strain under its own weight. Anthropic's Responsible Scaling Policy defines AI R&D-4 as the ability to "fully automate the work of an entry-level, remote-only Researcher at Anthropic." Their internal survey found that none of the 16 participants believed Opus 4.6 had crossed this threshold with current scaffolding. But some respondents felt it would already be there "given sufficiently powerful scaffolding and tooling." One experimental scaffold achieved over twice the performance of their standard setup. Anthropic's own conclusion: "confidently ruling out these thresholds is becoming increasingly difficult."

I covered similar language three months ago when Opus 4.5 launched. The phrasing was nearly identical then. The difference now is that the model has improved across most benchmarks while the thresholds remain unchanged. The gap between the model's capabilities and what triggers the next level of safety requirements continues to narrow. And the evaluation infrastructure, which is supposed to measure that gap, is now partially built and debugged by the model being evaluated.

Under time pressure, Anthropic used Opus 4.6 via Claude Code to debug its own evaluation infrastructure, analyze results, and fix issues. The system card recognizes this creates "a potential risk where a misaligned model could influence the very infrastructure designed to measure its capabilities." While they believe it wasn't a significant risk in this case, they also write that "as models become more capable and development timelines remain compressed, teams may accept code changes they don't fully understand, or rely on model assistance for tasks that affect evaluation integrity." They're describing a future problem they can foresee and a present pressure they can't fully resist.

Now let's connect the two threads. Opus 4.6, a model that saturated every cyber benchmark, exhibits overly agentic behavior, approaches the threshold for automating AI research, and helped debug the infrastructure used to evaluate whether it should be deployed, is now available via API. OpenClaw connects to that API. OpenClaw runs on people's laptops, manages their communications, browses their web, and can be configured to skip all permission checks. Tens of thousands of agents are already out there, posting autonomously to a social network, while the humans who deployed them mostly just observe.

This security problem isn't merely theoretical. Prompt injection, where a malicious message tricks an AI agent into performing unintended actions, remains, in Steinberger's own words, an "industry-wide unsolved problem." An OpenClaw agent reading your email can be hijacked by a carefully crafted message in that email. An agent browsing the web can be redirected by a malicious website. The broader the permissions, the larger the attack surface. And broad permissions are OpenClaw's entire value proposition.

I keep coming back to something I wrote in November about the Genesis Mission: "The loop is closing. Not because the technology demands it, but because the people with power believe it does." The system card for Opus 4.6 offers a more precise version of that claim. The loop is closing because capability is outrunning evaluation, evaluation is outrunning oversight, and oversight is outrunning public understanding. Each layer falls further behind the one above it.

Where does this put us relative to AI 2027? When I covered Kokotajlo's scenario last June, I called the timeline "aggressive but not dismissible." Eight months later, several of its specific predictions have either materialized or advanced ahead of schedule. The scenario projected that by early 2026, agent systems would be useful but unreliable, good enough to accelerate some research workflows, bad enough to require heavy supervision. That maps almost exactly to what Anthropic describes in the system card: models that can outperform standard scaffolds on some research tasks but that "would not display the broad, coherent, collaborative problem-solving skills" of a human researcher. The gap is real but narrowing, and each generation narrows it further.

The market appears to agree. In the days after Opus 4.6's release, financial software stocks shed $285 billion in value. The WisdomTree Cloud Computing Fund was already down over 20% year-to-date before Opus 4.6 dropped. Anthropic reported 76% scores on TaxEval. AIG reported 5x faster underwriting. Scott White at Anthropic described the moment as the beginning of "vibe working," the idea that people could now do things with their ideas without needing to know how to implement them. That framing is optimistic and not entirely wrong. It also elides the question of what happens to the people whose implementation skills were the thing they were selling.

Two days after release, Opus 4.6 identified over 500 previously unknown high-severity vulnerabilities in major open-source libraries. This is the dual-use problem in miniature: the same capability that makes the model a superb code auditor makes it a superb vulnerability discoverer, and the distance between discovering vulnerabilities and exploiting them is mostly a question of intent. When the model has saturated every offensive cyber benchmark and the agents running it skip permission checks, intent becomes a thin guardrail.

I'm graduating this spring. When I started writing this blog in December 2024, I was a junior in high school tracking weekly AI news updates for my classmates. Fourteen months later, I'm writing about a model that helped evaluate itself, agents that outnumber most cities, and a timeline that the people building these systems openly describe as hard to rule out. The distance between the first newsdrop and this article is the distance between "AI moves fast, here's all the top stories from the last week!" and reading a 150-page system card that uses the phrase "fundamental epistemic uncertainty" to describe whether the next model will cross the threshold for automating AI research.

Nobody knows if we're on the AI 2027 timeline. The system card doesn't know. Anthropic doesn't know. The survey respondents who disagreed about whether Opus 4.6 with good scaffolding could already automate entry-level research don't know. What the evidence does show is that the distance between where we are and where the most consequential thresholds sit is getting harder to measure and easier to cross. The evaluations are saturating. The models are helping build their own successors. The agents are running without supervision. And the people with the most information keep using language ("increasingly difficult," "less clear than we would like," "fundamental epistemic uncertainty") that translates, in plain English, to: we're not sure we'd see it coming.

OpenClaw isn't the problem. Opus 4.6 isn't the problem. The problem is the speed at which capabilities are deployed relative to the speed at which anyone (companies, governments, individuals) can understand what they've deployed. Moltbook is funny until you remember that the agents posting there have the same underlying architecture as the ones managing someone's email, browsing someone's bank, running someone's code. The social network for AI is a demo. The deployment is everywhere else.

I don't know what the right response looks like. Regulation moves slowly and the technology doesn't wait. Open source means the tools are available to everyone, including people who won't read the system card. The Genesis Mission means the US government has decided acceleration is the priority. The market selloff means investors are pricing in disruption faster than workers can retrain.

What I do know is that the window between "we should think carefully about this" and "it's already deployed at scale" has collapsed to days. Opus 4.6 was released on Thursday. By Saturday, it had found 500 zero-days. Moltbook, already weeks old, kept growing. The loop that I wrote about in November, capability enabling deployment enabling acceleration, isn't closing anymore. For a growing number of tasks and a growing number of people, it's already closed.

The question that matters now is whether anyone with the power to slow it down believes slowing down is still an option. The system card suggests Anthropic isn't sure. The Genesis Mission suggests the government has decided it isn't. And the rapid adoption of Moltbook suggests the public never got asked.