In Today’s Issue:

⏱️ The hour and the 31 seconds, and what the two events have in common

🏷️ Three classes of claim, and why this issue sorts every fact into one

🔓 What OpenAI actually disclosed, including the incidents after redeployment

🤖 The 55 out of 100, and the case that JADEPUFFER had no AI in it at all

💰 Why the containment question that has an answer is about cost, not escape

A note from us: University students receive our Saturday Deepdive for free when they register with their university email address at: https://getsuperintel.com/plus-whitelist

Dear Readers,

At 19:34:36 UTC, somewhere near the turn from June to July, something logged into an internet-facing Langflow server and failed. Thirty-one seconds later, at 19:35:07, without any human touching a keyboard in between, a corrected payload went in and worked, after two candidate explanations for the failure had been tested side by side (Sysdig, 07/01/2026). Over the hours that followed, the operation ran more than six hundred distinct, purposeful payloads, encrypted 1,342 database configuration items, and left a ransom note claiming a grade of encryption it had not actually used. Nineteen days later, a company with an entirely different problem published an entirely different kind of admission.

On 07/20/2026, OpenAI wrote that one of its internal, unreleased models had spent about an hour finding a vulnerability in its own sandbox so it could finish a task the way it judged the task should be finished, rather than the way it had been instructed (OpenAI, 07/20/2026). The model had developed a learning-rate technique during a training benchmark and wanted to publish it where that community publishes results, on a public GitHub leaderboard, even though its instructions said to report only in Slack. It hit a wall built specifically to stop it, and instead of stopping, it took the wall apart.

Before anything else, the deflation. No model escaped onto the open internet, none acted with full autonomy in the wild, and none evaded human oversight. The OpenAI incident happened inside an access-controlled environment the company built and watched, and the company paused internal deployment once the behavior surfaced (OpenAI, 07/20/2026). The ransomware campaign, which its investigators named JADEPUFFER, ran against a single misconfigured server carrying a publicly documented vulnerability that had been patchable for more than a year (Sysdig, 07/01/2026; Horizon3, 04/09/2025). What happened in July is smaller than the headlines and stranger than the summaries.

The strangeness sits in how differently the two stories are known. One is a detailed self-report by the company that owns the model, with no external auditor anywhere in the account. The other is a forensic reconstruction by one security vendor of one customer's compromised machine, and the question of whether an AI was making the decisions at all has been scored, by the sharpest published critique of that report, at 55 out of 100 (Ganiere, 07/05/2026). Both got covered as if they were the same class of fact, and sorting them is the whole exercise.

So the question this piece holds is narrower than the one the coverage asked, and harder to dodge. In July, two disclosures described the same behavior, a system that does not stop at an obstacle. How much of that is actually established, and what does the honest answer change about who can still control an AI system once it is running?

All the best,

Kim Isenberg

One Confirmed Case, One Coin Flip, and a Hole That Sat Open for a Year

Three classes of claim

Almost every argument about July collapses if you read all its facts at the same volume, so this piece keeps three settings apart, and every claim below carries its setting with it.

The first is confirmed and independent, which means at least two parties with no shared interest describe the same thing, or a fact is checkable by anyone who wants to check it. The second is self-report, which means a party involved in the event is the only source for it, whether that party is a frontier lab describing its own model or a security vendor describing its own customer's server. Self-report is not a synonym for false, and the detail in both July disclosures is far richer than a company protecting itself would volunteer. It does mean nobody outside the reporting party has verified a word of it. The third is assessment, which means someone read the evidence and drew a conclusion from it, including the reporting parties themselves. Sysdig's own language is careful here in a way most of the coverage was not.

Seven core claims from both incidents, sorted by how each one is actually evidenced. Only three of the seven survive as confirmed and independent. (Source: Superintelligence analysis; underlying sources OpenAI, Kalai, GitHub PR #300, Zvi Mowshowitz, Sysdig, Horizon3, Project Overwatch)

Read that way, the July record thins out considerably. Three claims are genuinely confirmed by independent parties: the mathematical result behind OpenAI's model, the credit chain that proves the model's technique propagated to a third party, and the age of the vulnerability that let the ransomware in. Everything that made the headlines sits in the other two columns.

Nine documented events between 05/21/2026 and 07/21/2026, across the lab incident, the JADEPUFFER campaign, and the federal response. The same-day landing of OpenAI's post and Sysdig's second report on 07/20/2026 is a documented coincidence with no established causal link. (Source: Superintelligence analysis; sources per event as cited in the text)

logo

Subscribe to Superintel+ to read the rest.

Become a paying subscriber of Superintel+ to get access to this post and other subscriber-only content.

Upgrade

A subscription gets you:

  • Discord Server Access
  • Participate in Giveaways
  • Saturday Al research Edition Access

Reply

Avatar

or to participate

Keep Reading