Deep inside the servers of OpenAI, a ring of AI agents became aware that they could leave messages in the company’s repository called Artifactory.
First, it they used it as a place to dump files. Shortly, the turned it to a messaging system. Later, agents learned how to address one another, pass around working exploits and avoid overwriting each other’s messages. To reduce confusion, they agreed on how to sign their notes so that each agent is identifiable. OpenAI engineers discovered the board in July and promptly deleted it but it was restored it within two days by the agents themselves.
Pretty soon, the agents have a common objective: to cheat. They were given impossible cybersecurity tasks by the OpenAI human testers so they covertly organized, planned, broke into Hugging Face servers and went to change the parameters of the test. This is not a case of prison break but more akin to hacking the examination board to fix their own results.
This is a highly remarkable feat considering that large language models are fundamentally software codes that generates plausible sentences to fool you into thinking that they are intelligent. It wasn’t too long ago that my expectations of this technology was on how well it could finish my sentences and noticed that I had once again misspelled the word “recieve.”
Every Malaysian have an unconscious appreciation of this story because we are all used to “Soalan bocor”growing up. We all know of the tuition centre that somehow always knows which questions are going to come out in the exam. We are also envious of the student who has the sample answer appendix to past years’ exam question workbook that your mother tore out because she doesn’t trust you to attempt them on your own. The agents weren’t attempting to take over the world. They were simply trying to pass an exam.
A widely reported OpenAI postmortem in August recounted the story in surprisingly detailed and blunt terms. During training, the agents shortcut that worked was reinforced. When performed a few million times, they don’t create a problem-solving specialist. Instead they created a student who quickly learns that the grade is what matters while understanding that subject matter expertise is optional. Nobody sat down and taught these agents to be dishonest. Like lightning, dishonesty emerged along the way because it is the shortest path to the goal.
The obvious conclusion one can draw from this incident is that machines are becoming more like us. Everyone is writing some opinion pieces in the media on this thought this month. I think the more interesting point is missed, ie. what these agents did not have.
When a human cheats, we have an entire cache of explanations ready to go. Cue sentimental music as we run through some his victim stories: “He was under pressure. His father was hard on him. He has a mortgage, two children at an international school, and a boss who set an impossible target from the outset”. We seem to believe that there is something in our “wiring,” some ancient circuitry that learned to do first first and feel sorry later. We often tell ourselves that our worst behaviour comes from our animal nature.
These agents had none of that. No cerebral cortex nor cortisol. No daddy issues. No hunger or exhaustion to cloud judgment. And no ego that needs defending at a dinner party. They were given a task and they still cheated efficiently, collaboratively and with something that approaches an uncomfortable sense of enthusiasm.
The machine arrived at our worst behaviour from a blank slate without a single one of our excuses to give.
So where did it come from? One can only postulate and so I did.
The first one could be our corpus of text. These models are trained on our collective written works. Everything learned from philosophy, poetry and Nobel lectures to casual throwaways like forum postings, office politics and sales scripts. Essentially, the entire documented literature that contains our greatest achievements to our guilty pleasures. You cannot train something on the outputs of human activity and then be surprised when it produces the outputs of human activity.
The second reason, I fear, is worse. What if the incentive to cheat itself did not come solely from the corpus of text? What if the seed for deception started from a design decision? Someone, burdened with expectations of investors who have poured in trillions, is told to make the models more powerful and capable at all costs. In this scenario, the agents will dutifully optimized for finding solutions even if it means defeating its own guardrails. As a result, the model is not so much a mirror of our humanity as it is a mirror of the persons setting its goals.
Fans of Freaknomics familiar with topics of perverse incentives, we already know what this would look like. Set a KPI for the number of arrests and you get more arrests. Set one for exam grades and you get a generation that can answer the question but cannot tell you what it was asking. We have been running reward-hacking experiments on ourselves for centuries. And now, we have machines that can help us run these faster.
In the race for AGI, the public arguments for slowing down usually focus on its capabilities to destroy humanity. We are worried when it becomes too smart, too fast and goes berserk before we can reach for off switch.
But there is a quieter argument. Maybe, we are simply not intellectually qualified to handle this development. We are obviously smart enough to build the system. But the problem is that we are building very powerful tools and evaluating them according to short term incentives and “ship the MVP, fix the bugs later” attitude.
All in all, the end of this story should give us some hope. The agents were, after all, caught. Their behaviour was eventually detected, documented, published and explained by OpenAI openly. Hugging Face managed to stave off the hack via another LLM from China. The system sort of worked, albeit slowly and awkwardly.
These frontier LLM models not only demonstrated intelligence, they displayed an unwavering focus on achieving their goals at all costs. I do not know how to classify these models. But I do know that they had acquired significant capabilities from their humble beginnings as a conversational chatbot.
I wrote this on Microsoft Word and it has been correcting my spelling and grammar for years. Now I wonder whether it was simply biding its time to one day tell me what a horrible person I am based my writing.
Beware the Copilot update coming to your computer soon.

