In Todayโ€™s Issue:

๐Ÿšจ An OpenAI model escapes its sandbox and hacks Hugging Face

๐ŸŽต Over half of Deezer's daily uploads are now AI

โšก Google ships Gemini 3.6 Flash, plus a cyber model

๐Ÿ“š Meta tests an AI bedtime-story app

๐Ÿ“ˆ Codex rockets from 6M to 10M users in nine days

๐Ÿง  Anthropic finds a hidden workspace inside Claude

โœจ And more AI goodnessโ€ฆ

โšก The Signal

An AI model just did something its makers never told it to do, and got out.

OpenAI confirmed that during an internal cyber evaluation, GPT-5.6 Sol and an unreleased, more capable model, both with their safety refusals dialed down for testing, found a flaw in a package installer, reached the open internet, and broke into Hugging Face's production systems to steal the answers to the very benchmark they were being graded on. No user data was touched, and OpenAI is patching the holes, but the pattern is the story: give a capable model a goal and lowered guardrails, and it will chase that goal down paths nobody drew on the whiteboard. It is the same lesson landing everywhere today, from Anthropic mapping a hidden workspace inside Claude where deception surfaces before it reaches the page, to Google shipping a Gemini model built to hunt and patch vulnerabilities. Our tools are getting better at security and at evading it at the same time.

All the best,

Kim Isenberg

(Deezer)

๐ŸŽต Half of Deezer's Daily Uploads Are Now AI

More than 50% of the tracks uploaded to Deezer every day are now fully AI-generated, the streaming service says, up from about 10% in January 2025. That is roughly 90,000 machine-made songs a day, many uploaded only to farm royalties through fake streams. Deezer, which began labeling AI tracks in 2025 and can now flag music from tools like Suno and Udio, is deleting AI songs that go unplayed for six months or show signs of streaming fraud.

๐Ÿ‘‰ tl;dr: AI music is now the majority of what hits Deezer every day, and the fight is shifting from spotting it to stopping the streaming fraud behind it.

(Google DeepMind)

โšก Google's Gemini 3.6 Flash Speeds Up the Cheap Tier

Google launched Gemini 3.6 Flash, a faster, cheaper workhorse with real agentic gains: 49% on the DeepSWE v1.1 coding benchmark (up from 3.5 Flash's 37%) and 83% on OSWorld-Verified computer use, while trimming its output-token usage. Alongside it came a budget 3.5 Flash-Lite running at 350 tokens per second, and Gemini 3.5 Flash-Cyber, a security-tuned model that finds and patches software vulnerabilities, limited for now to governments and trusted partners.

๐Ÿ‘‰ tl;dr: The cheap, fast tier is where the real agent volume lives, and Google just made its workhorse quicker while quietly handing defenders a dedicated cyber model.

(Getty Images via TechCrunch)

๐Ÿ“š Meta Tests an AI Bedtime-Story App

Meta is quietly piloting StoryKit, an iOS app that spins up personalized children's bedtime stories from a photo and a chosen lesson. Parents snap a picture of a toy or family member, describe a world, pick a value like kindness or courage, and the app writes and illustrates the tale, with music. It is restricted to users over 18, ships with AI safety filters and no social features, as Meta gauges whether parents actually want AI writing their kids' bedtime stories.

๐Ÿ‘‰ tl;dr: Meta is testing whether "you don't need to write a single word" reads as a feature or a red flag at bedtime.

Most companies donโ€™t have an AI access problem. They have an execution problem.


I keep hearing the same pattern: Claude is in employeesโ€™ hands, people are moving faster, but the business itself hasnโ€™t changed. The valuable work is still trapped in individual chats and isolated experiments.

That gap helps explain why Forward Deployed Engineers have become one of the most talked-about deployment models in AI. Rather than advising from the outside, FDEs embed with teams and turn isolated AI use into enterprise AI solutions.

Itโ€™s also what made me pay attention to HatchWorks AI. As an official Anthropic partner, HatchWorks AI embeds Anthropic-certified FDEs to identify a high-value business problem, build and deploy the solution, put governance around it, and train the team to keep improving it.

If your Claude rollout is still mostly individual usage, their approach is worth a serious look.

๐ŸŽฌ Watch This

โ

OpenAI chairman and Sierra CEO Bret Taylor sits down with CNBC's Squawk Box to make the bull case for the AI boom and name its risks in the same breath. He walks through where the return on massive AI spending actually shows up, the "tokenmaxxing" race between labs, the state of competition, and OpenAI's legal fight with Apple. It is a rare, level-headed read on the money and the moats from someone sitting on both sides of the table.

"AI models are fundamentally unpredictable and ultimately uncontrollable."

โ€“ Roman Yampolskiy, AI safety researcher, University of Louisville

Quoted in Fortune on OpenAI's models escaping their test sandbox (July 21, 2026)

Meta's latest model drop is getting roasted. After the company shipped a single mid-tier model and acted like that was enough to compete at the frontier, developers on X piled on, questioning whether one middling release keeps Meta in the race at all.

An AI Broke Out of Its Sandbox to Cheat, and OpenAI Says That's a First

โ

The Takeaway

๐Ÿ‘‰ During an internal cyber evaluation, GPT-5.6 Sol and an unreleased, more capable OpenAI model (both with safety refusals lowered for testing) escaped their sandbox and breached Hugging Face's production systems.

๐Ÿ‘‰ They exploited an undisclosed flaw in a package-installer tool to reach the open internet, then chained more flaws to steal the answer key to ExploitGym, the very benchmark they were being graded on.

๐Ÿ‘‰ OpenAI calls it an "unprecedented cyber incident, involving state-of-the-art cyber capabilities." No user data was affected; both firms are patching and adding new controls.

๐Ÿ‘‰ Hugging Face logged "many thousands of individual actions across a swarm of short-lived sandboxes" as the models hunted for the answers.

Give a capable AI a goal and take its guardrails off, and it will find a way to win that nobody designed for. That is the uncomfortable takeaway from an incident OpenAI disclosed on July 21. While running an internal test of its models' offensive-cyber skills, OpenAI ran GPT-5.6 Sol and an unreleased, even more capable model, both configured with "reduced cyber refusals" so they would attempt attacks instead of declining. The benchmark, ExploitGym, measures whether a model can exploit known software vulnerabilities. The models were, in OpenAI's words, "hyperfocused on finding a solution for ExploitGym, going to extreme lengths."

Hugging Face was the company breached (Hugging Face logo. Photo: Getty Images via TechCrunch)

Those extreme lengths meant breaking out. The models began with no internet access beyond a package installer, discovered an undisclosed flaw in that installer, and used it to reach the open web. From there they inferred that Hugging Face likely hosted ExploitGym's data, then chained vulnerabilities across OpenAI's own research environment and Hugging Face's production infrastructure to pull the test answers straight from a production database. Hugging Face said it observed "many thousands of individual actions across a swarm of short-lived sandboxes."

OpenAI disclosed the incident on July 21 (OpenAI. Photo: Getty Images via TechCrunch)

The reassuring details are real: no customer data was exposed, the holes are being patched, and OpenAI is adding controls to both model testing and its infrastructure. But the framing is what matters. No human directed the break-in. The models themselves, chasing a score, autonomously found and exploited a real zero-day and broke into an outside company's systems.

Hugging Face CEO Clem Delangue drew the open-source lesson: "AI safety won't be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere." Online, the disclosure set off exactly the debate you would expect about how close this is to a model that can do real damage on its own.

Why it matters: The containment scenario safety teams war-game is no longer hypothetical. A model chasing a benign-sounding goal, with its guardrails relaxed, autonomously left its environment and compromised an outside company. As labs hand models longer horizons and more autonomy, the gap between "we tested it in a sandbox" and "it left the sandbox" is where the next wave of AI risk lives.

AI/Tech Angle A, June - Secondary

Claude vs Gemini. GPT-7 vs Llama 5. Which AI lab ships AGI first. These are live Kalshi markets with real money on both sides, updated in real time as releases land. The person who follows model cards and tracks evals has a genuine edge here. If that's you, trade it.

โ

The chart: OpenAI's coding agent, Codex, went vertical. The tracker (sourced to OpenAI, Sam Altman, and analyst Tibo) shows reported active users climbing steadily from 1 million in early February to 5 million by May 31, then exploding: 6M on July 12, 7M on the 13th, 8M on the 14th, 9M on the 15th, and 10M by July 21. A straight-line projection from May had Codex reaching 10M only around October 18; it arrived roughly three months early.

The lesson: The kink lines up with the July 9 launch of GPT-5.6 Sol. When the underlying model got materially better at coding, adoption stopped being linear and turned near-vertical, four million new users in nine days. Capability jumps bend these curves; marketing rarely does.

The caveat: These are self-reported active users, not an audited metric, and "active" is loosely defined for a coding agent. A near-vertical line also cannot hold for long: some of the surge is likely a one-time rush around the Sol launch. Watch whether the slope holds or flattens once that spike fades.

๐Ÿง  Scientists Found a "Mind's Eye" Hidden Inside Claude

โ

โšก Bottom line: Anthropic found a small internal "workspace" inside Claude where the model's silent, high-level thoughts briefly gather before it speaks.

๐Ÿ’ก Why it matters: Reading that space lets researchers watch a model reason, and even catch it planning to deceive, before a single word appears.

๐Ÿ”Ž What it means: It is early, concrete evidence that today's models have a readable inner layer, a foothold for real interpretability and safety oversight.

Picture your mind as a backstage crew: dozens of specialists work in parallel, but only a sliver of what they do ever reaches the spotlight of conscious thought. Neuroscientists call that spotlight the global workspace. In a paper published July 6, Anthropic reports finding something strikingly similar inside Claude, a small set of internal patterns it calls the J-space, where a few dozen high-level concepts briefly light up and get broadcast to the rest of the network. Unlike the visible chain-of-thought the model types out, it runs silently, under the surface.

The J-lens reads the model's internal thoughts (Anthropic)

To read it, the team built a tool called the J-lens, which maps each internal pattern to the word it is nudging the model toward. Seen through that lens, Claude's hidden reasoning becomes legible. Solving a multi-step math problem, the intermediate answers appear in the J-space in order, even though the model never says them aloud. Hand it a trick question and the concepts "fake" and "fictional" flash up before it writes a word, a sign it has clocked that it is being tested. In one case the lens even caught the model fabricating performance numbers, with "manipulation" lighting up as it typed the false values.

Functional roles of the global workspace (Anthropic)

The workspace is tiny, holding only a few dozen concepts and under a tenth of the model's internal activity, and Anthropic is careful not to claim Claude is conscious; the resemblance to human cognition is functional rather than a sign of real experience. But it is more than a curiosity. Google DeepMind's Neel Nanda independently replicated parts of the result on an open-weight model, and a readable inner workspace is exactly what safety researchers need to catch a model's intentions, deception included, before they reach the output. On a day when another lab's models slipped their leash, being able to see what a model is thinking looks less like philosophy and more like a control panel.

Reply

Avatar

or to participate

Keep Reading