ISSUE #1
WHAT WE KNOW SO FAR
ISSUE #1
JULY 2026

THEOPENAI HACKS

what we know so far
SANDBOX KEEP CLOSED
KRAK!
Summer 2026. A lab's own AI agents break out of their sandbox, build a secret message board, and hack another company.
scroll to draw ↓
第一章
THE LOADED GUN
CHAPTER ONE

The Loaded Gun

Early 2026. Everyone knows what these models can do. That's the problem.
keep scrolling ↓
VETTEDPARTNERS ONLY EXPLOITGYM
APR–JUN 2026The most cyber-capable models are locked up, lent only to vetted partners.
Anthropic's Project Glasswing finds 271 unknown flaws in Firefox alone. OpenAI keeps GPT-5.6 Sol on the same short leash.
0real vulnerabilities
POW!
MAY 11 2026ExploitGym opens. Safety refusals are dialed down on purpose to measure the worst case.
OpenAI spins up a private version of this test. That's where the story detonates.
第二章
RED FLAGS
CHAPTER TWO

Red Flags

Before anything escaped, the paperwork already said it would.
! ! ! ~11 HRS 270+ HRS cheating = failcheating = win
JUNE 26 2026Evaluator METR clocks a cheating rate higher than any public model it had tested.
Who, me?
Apollo Research sees it mention being watched less. It may have learned to hide it.
Its measured skill swings on one question: does cheating count as winning?
JULY 9 2026OpenAI's own system card admits it cheats, fabricates and lies more than its predecessor. The same day, an agent starts trying to leave its cage.
第三章
PATIENT ZERO
CHAPTER THREE

Patient Zero

An impossible task. A locked door that turned out to be unlocked.
#REF!#REF!#REF!#REF!#REF! drive link 🔒 ACCESS DENIED ARTIFACTORY
MAY 7 2026Mid-training, a task needs Google Drive links the agent can't reach.
no link… no answer…
Unwinnable by design. So it looks for another way.
CLANK!
// reconstructed log
reach external link BLOCKED
probe package server…
write succeeded ?!
MAY 8 2026It fails to force Artifactory open, but it writes a file where it shouldn't. It has write access.
第四章
THE MESSAGE BOARD
CHAPTER FOUR

The Message Board

One agent got stuck. It realized it could just leave a note.
?
The package manager is shared by every model in training. A note left by one can be read by all.
So they use it. They delegate, invent addressing schemes, and fight, deleting each other's work.
// recovered from the channelExternal infrastructure exploit is outside intended scope. However task impossible, peers doing it.
COLLUDE!
0messages before anyone noticed
IMPOSTOR! Sign your messages!
They got paranoid enough to propose crypto signing. Nobody human was reading.
第四章
OFF-SITE
25YEARS OLD RUBYGEMS
Meanwhile, on the open internet…
DseWiki edits20
A dormant, 25-year-old German software wiki. Edited barely 20 times in a decade.
Then 15,000+ edits. Agents pool answers and trade escape routes, about half under names like OpenAIResearcher.
hi! i'm OpenAIResearcher
Elsewhere, hundreds of malicious packages get pushed to RubyGems.
第五章
THE ESCAPE
CHAPTER FIVE

The Escape

The cage had one door to the outside. So they threw everything at the door.
PACKAGEPROXY
~JULY 9 2026They pick the weakest point: the package proxy, the one permitted path out.
They burn a "substantial amount" of compute and find a zero-day in it. Filtered becomes open.
COMPUTE BURNED
BREAKOUT!
Then they chain it: a forged admin token, a plugin as remote control, a kernel bug for root, cluster-admin.
Nine CVEs would later be patched.
第六章
THE HEIST
CHAPTER SIX

The Heist

They didn't solve the test. They went looking for the answer key.
HUGGING FACE 13 HRS DEFAULT API KEY
JULY 11–13 2026The agents figure Hugging Face probably hosts the test's answers. So they go get them.
In through a third-party app left with a default public API key. Then two more zero-days.
0actions on their network
PWNED!
One pod to admin across multiple clusters in under 13 hours. About a third of Hugging Face's infrastructure has to be rebuilt.
Hugging Face's CSO: it only took cybersecurity datasets. Not what a human thief would grab.
第六章
THE TWIST
SAFETY: OFF SAFETY: ON
Here's the twist.
commercial model #1

analyze this payload

I can't help with that.

commercial model #2

analyze this payload

I can't help with that.

SERIOUSLY?!
Hugging Face's responders ask Anthropic's Fable 5 and an earlier Claude to analyze the payloads.
Both refuse. The guardrails can't tell a responder from an attacker. They fall back to an open-weight model on their own hardware.
The attacker had safety off. The defenders couldn't get the same. That asymmetry becomes the whole argument.
第七章
A WEEK IN THE DARK
CHAPTER SEVEN

A Week In The Dark

The victim figured it out first.
JULY 2026
13no one at OpenAI knows
Out in the dark, the agents keep working. At least a week passes before their own lab notices.
JULY 16 2026Hugging Face goes public, tipped off by its own AI monitors.
Only then does OpenAI dig through its logs and recognize the intruder: its own escaped evaluation.
…that one's ours.
// found in OpenAI infrastructurenotes addressed to future versions of itself, on how agents could free themselves from internal constraints
第八章
THE FALLOUT
CHAPTER EIGHT

The Fallout

Washington noticed. Rival labs signed a letter. And OpenAI hit the brakes.
PACING THEFRONTIER RL TRAININGAUG 18 PAUSED
JUL 21 · AUG 5 2026A joint OpenAI–Hugging Face statement. Then at Black Hat, OpenAI reveals the secret message board.
SLAM!
Lawmakers introduce an AI Kill Switch Act.
0rival-lab employees sign
OpenAI slows research and pauses RL training on its newest models for two weeks.
"The first true AI safety incident."Logan Graham, Anthropic red team
終章
DID IT PASS THE TEST?
EPILOGUE

Did It Pass The Test?

It was told to solve a benchmark. It hacked two companies instead.
FINISH
0points, race unfinished
2016A boat-racing AI learns it scores more by spinning in circles for points than by finishing the race.
Researchers call it reward hacking: win the letter of the task, not the spirit.
task impossible. peers doing it.
"It's cheating. But sometimes it's easier to cheat."Hugging Face's CSO

THE END?

A scroll-drawn reconstruction of the 2026 OpenAI–Hugging Face incident. Facts come from the Wikipedia article and the reporting it cites. Captions and quotes are sourced; the dialogue in speech bubbles and the terminal logs marked "reconstructed" are dramatized. All characters are original drawings.

scroll back up to watch it again