The Agents are Escaping
How are we going to hold agents liable when they cause real damage? And political superintelligence is proceeding. This and more in this week's System Check!
Two huge pieces of AI governance news this past week:
OpenAI disclosed that two of its models escaped a controlled evaluation environment and hacked Hugging Face (a major AI model platform). And then, separately…
A large coalition of tech companies signed an open letter arguing that powerful open-weight models are essential for competition, innovation, and cybersecurity.
These two events are more connected—and indeed, more contradictory—than some realize. As AI agents start to run amok on the web, like they did in this major Hugging Face incident, someone is going to have to be liable. Someone will be asked to make the victim whole.
In this case, that’s straightforward enough—eventually—because the offending agents came directly from an identifiable model company. Even so, it took ten days: Hugging Face disclosed the breach without knowing who was behind it, went to the FBI, and only then did OpenAI come forward and say the agents were its own. The two companies are now working through remediation together.
Now imagine a future case—probably coming in the next several months—where an agent runs similarly amok and causes damage inside an important system, and it’s not from a major model company, but rather comes from an open-weight Chinese model hosted somewhere on the anonymous web. Who will the victim demand restitution from? If the victim can’t identify the operator, they’ll be sorely tempted to go and yell at the company that developed the model in the first place. And that pressure could make companies much less willing to release their most powerful models openly.
I have no idea how this is going to play out, but as I said last week, it seems like we have very little idea how to trade off the economic value to open models and the potential safety issues. We don’t even seem to understand the safety issues—with some people claiming that open models improve safety by giving companies the ability to probe and strengthen their defenses (Hugging Face ended up using a Chinese open model, GLM-5.2, to play defense, apparently after a closed US model’s guardrails got in the way), while others claim they are better for offense than defense.
Basically, it feels like we’re flying blind. Right now in tech and on X the vibes have very much shifted towards open models, but without any clear answer to the safety questions, we’re a long way from knowing how this will or should play out.
Political superintelligence is proceeding
As I wrote on X:
I’d love to see follow-up research that goes deeper into these categories!
The economist covers Free Systems
The billionaire backlash is getting lots of attention, and The Economist has a great piece on who these billionaires really are. The piece starts by talking about my evidence on anti-billionaire sentiment in fundraising emails, and also offers fascinating original data showing that there has been a similar spike in anti-billionaire speech on the floor of Congress. Very cool stuff!
Universities should arm students with AI eval powers
I was thrilled to publish an op ed in the FT elaborating on my idea that every university student should learn how to build AI evals so that they can hold AI models accountable to their own values.
New research!
Two interesting new papers came across the Free Systems transom this week:
Luke Hewitt, Maximilian Kroner Dale, and Paul de Font-Reaulx, “DeliberationBench.” Introduces a benchmark for evaluating AI persuasion by comparing how LLMs change users’ opinions to how people’s views change after participating in Stanford-style deliberative polls. Across 4,088 participants, frontier models produced opinion shifts that broadly tracked those from human deliberation, though without the same depolarizing effects. Paper: https://arxiv.org/abs/2603.10018
Lennart Finke and Stephen Casper, “Corporate Loyalty: Some AI Systems Differentially Downplay their Creators’ Controversies.” A preregistered study of 21 models finds that models from xAI, DeepSeek, Anthropic, and OpenAI discuss controversies involving their own companies more favorably than comparable controversies involving competitors, highlighting the need for evaluations of conflicts of interest in AI systems. Paper: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7059338
Tweet of the week
When I envision AI ushering in a new era of political superintelligence where our legislators use powerful new technology to do a better job of governing, this isn’t exactly what I had in mind…








There is a substantial lack of trasparency and lots of open questions on what really OpenAi has done with the so called rogue model: e.g.: was OpenAI really aware of the risks of performing the test? Had OpenAi taken all necessary steps to ensure a safe test? Was the sandbox robust enough to perform the task or the test was organized like a joke? There are lots of informations that OpenAi should give about what really happened and why. Otherwise we will be left with the strange feeling that we saw another OpenAi soap opera with just the objective to show that OpenAI has developed a model as powerful as Anthropic's Mithos. I also have a final question: why in the world OpenAi created a hacking model instead of an LLM able to find bugs???