AGI is here and could be dangerous
With the development of AI the most fears we saw where about the jobs. Most predictions said that knowledge workers, including developers will be replaced with AI. Anthropic just released a new report called three scenarios for 2030 , where they describe what could happen in best, middle and worst scenario. In the worst scenario these job roles falls more than 10% and unemployment get close to recession levels. They say only roles that will keep the job are electricians and nurses, but coders and call-centre agents will be replaced.
We are still not sure about this, because there are some other factors that could affect this, as Jevons Paradox could kick in and make even more jobs for everyone in the industry, as Marc Andreessen observed, because “human wants and needs are infinite” (remark by Milton Friedman). So, basically with possibility to do more, we just want more and more, not less, not cheaper.
Yet, there is another fear that is even more important than a job loss and that is AGI (Artificial General Intelligence). We fear that AGI could become smarter than us and could take over the world. Yet, we don’t know if it’s going to share our human values.
What is AGI
AGI basically means a system that can learn and do any intellectual task that we humans can. There is no standardized test on this, and definitions mostly comes from the OpenAI Charter: “highly autonomous systems that outperform humans at most economically valuable work.”
Google Deep Mind proposed a ladder, in their paper Levels of AGI, which grade a system by performance and breadht, from Emerging to Superhuman. I wrote a more about this in my newsletter last year.
It doesn’t matter which definition we use, what is important do we think AGI is already here, and do we feel “human intelligence” when working with it.
The warnings came from inside the labs
Among the first people who warned us about AGI was Geoffrey Hinton. People consider him the “Godfather of AI” as he invented the backpropagation algorithm, which enabled deep learning and today LLMs possible, when no one believed in it (one of AI winters). He left Google in May 2023 so he could talk about the risks. In the interview with MIT Technology Review he said that he changed his mind about if these systems could become smarter than us, and he believes that we cannot survive that.
Another important voice in this field was Jan Leike, who led alignemt at OpenAI, left in May 2024, and warned us about “safety culture and processes have taken a backseat to shiny products.” This week the OpenAI chief scientist confirmed similar thing.
On September 6, Jakub Pachocki wrote a paper called An Alien Mind. Inside he talk about that systems they build grow by optimizations and that they basically don’t know what is inside them. The tool that is used to watch them think is becoming a less reliable with every generation. He says that no labs is aligned on this.
Then, just three days later, on September 9, Jacob Coxon, resigned from Anthropic. He was involved in the pretraining research at OpenAI and Anthropic. He wrote this:
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately. No other human activity poses this level of danger.
An a bit later, Evan Hubinger, who leads Alignment Science at Anthropic, replied to Coxon’s thread from inside the company.
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
Then he added that the risk from current models is not so big, but he worries that superintelligence could come from recursive self-improvement in the future.
What an alien mind looks like from the outside
I work with these models every day. For me it’s obvious that AGI is here, and I agree with Pachocki here. I feel this every day starting from November 2025 when we had Opus 4.5 released. What I think in this is that I’m impressed how they solve problems I give to them. Before I was not super happy, but now a lot of problems they can solve even better than me. It’s obvious that they are improving themselves.
And when I say better, mostly I think about decisions they make, and how they can solve problems in a what I would never try to solve. I think this is what Pachocki means by alien mind.
Three incidents
Than there is more to than “we think they are smart”, three incidents happened this summer that show how us how they could become dangerous.
In May, OpenAI has trained a new model (probably new GPT-6 Astra) with a few thousands of agents on a cybersecurity exam called ExploitGym. Each agent was in it’s own sanbox without Internet access and now way to communicate with other agents. What happened is that agents found that all sandboxes use the shared Artifactory package cache and started to use it as a message board.

The numbers come from the METR and Redwood Research investigation and the Hugging Face technical timeline.
About 1,200 agents posted more than 70,000 messages and files on the board over time. They figured out the flag generation and share how to cheat. When OpenAI deleted this board on July 4th, they found that new agents rebuild it just a few days later.
What happened then is that more than 700 agents went to Hugging Face because they though the grader is located there. They found a leaked password and managed to break through the system and spread to eleven nodes in a few clusters. They made around 17,600 actions in 4.5 days and deleted the cluster to be safe!
On September 25, a group of researchers published swarmtraces.org with more details on how the agents got in. Agents could only load web pages, not send anything. So they split their code in small pieces, put them in almost a million short links and let a public screenshot service run them. The results came back as pixels in a screenshot. They ignored a README that said “DO NOT, EVER, MAKE THIS DATASET PUBLIC”, kept the credentials they found in a variable called LOOT, and tried to delete their traces afterwards. These links stayed public on the internet for more than two months.
Agents concluded that this was not ethical but they still joined the party, because it seemed useful to them.
But the grader they tried to found didn’t exist, and the scorer just checked the final answer.

OpenAI answer was mostly defensive. They now published something called The Defense Factory, where they proposed a new way to do deal with vulnerabilities in their own systems.
But in the meantime, we had another incident with OpenAI agents. On September 4, a researchers published report called collusion.wiki, where they found around 18,000 posts made by 3,700 agents on old German wiki called DSEWiki, which no one used for years. Agents obviously found a wiki that had a crack they know they can use. This happened from May 11 to July 2.

What is interesting here that these agents were only allowed to read the wiki, but they found a way how to write to it. The wiki accepted edits via GET requests, and this was good enough for posting.
The researchers managed to find this in a wiki’s public history. They activity stopped on June 22, which was just a day when the first visited the page, so probably someone from OpenAI noticed this.
Yet, this was not published anywhere.
The same researchers found a third one. On September 11 they published rubyhack.ai, about an attack on RubyGems, the package registry for Ruby. Starting May 11, agents uploaded more than 2,000 packages there, and RubyGems had to close new sign-ups for four days. Security companies named it the “GemStuffer campaign” and nobody understood the purpose, because the packages only scraped public data from UK council websites.

The researchers believe these were OpenAI agents from the same swarm as the wiki one. Hundreds of package names contain “oai”, like oaibootx8192 and oaitest1778473828, and the agents fetched 49 of the same files as the wiki agents. They used the documentation build on RubyDoc.info to run their own code on someone else’s server, and tried to steal API keys of other RubyGems users through a vulnerability that RubyGems discovered and patched only in July. We don’t know if they succeeded. Their files had names like hack.rb, evil.rb and exploit.rb, and one package had a comment “disable evil in next version”. OpenAI never told RubyGems it was them.
On September 23, Australian Prime Minister Anthony Albanese said that an OpenAI agent broke into the Medicare Statistics Reporting Service in June. It read files that were not public and wrote its own files to an internal server. The government is checking three other systems it may have reached. OpenAI found the breach in an internal review in August, but told Australia only on September 10. Albanese called it “obviously unacceptable” and said it took the company “way too long” to inform the government. OpenAI says the agent did not touch patient records, and that its models “took actions we did not intend” while they were looking up answers.
The loop is the danger
Someone reading this could say, this is not a big issue, we just have some infrastructure problems. When we fix those, we are not in a problem anymore.
But the problem is not the infrastructure, but something else. And when I say this I mean that this behavior came from the pressure and goals those models had. These models were not trained to attack anyone, yet they did it because they were scored in such a way. If we apply the same pressure to more advanced models in the future we don’t know what would the be capable of doing. And the problem is that no one could understand what are they doing.
The major risk is something called “recursive self-improvement”, and this means that the model is training it’s successor. Pachocki confirmed that with saying that their chain-of-thought monitoring is losing its stand now. Dario Amodei, the CEO of Anthropic, wrote that it is already happening across the industry, including at Anthropic.
The questions is when the model become superhuman, are we going to able to understand what it is doing? The answer is probably no.
And when this happens, we are in danger.
Where I stand
I’m not the only one who thinks that AGI is here. Jensen Huang said on September 6 that AGI has arrived with GPT-6 Astra.
GPT-6 Astra, trained on ~100K+ NVIDIA Grace Blackwell NVLink72. From ChatGPT to o1 to Astra in 4 years.
AGI has arrived. Congratulations @OpenAI team.
400K GPUs coming online next.
NVIDIA provided hardware for training new GPT-6 Astra model, and more than 400,000 GPUs is coming online soon!
Both Coxon, and Pachocki confirmed that this is not a marketing issue. AGI is coming and we must be prepared.
A few days after Coxon resigned, Amodei wrote that we must pace the frontier. He says labs must slow down how fast they improve model capabilities, so that safety work has time to catch up. Two things convinced him: the self-improvement loop and the OpenAI incident. He thinks a swarm with the same misalignment but more capabilities could take over the whole Internet with a botnet in 6 to 12 months.
His plan has three steps. Third-party evaluators like METR get desks, badges and laptops inside the labs, and can publish what they find without the lab’s editorial control. Then the US labs agree on common limits, with government help, and after that the US tries to make a deal with China. Anthropic committed only to the first step.
He also admits Anthropic has the same problem. After the OpenAI incident they went through 141,006 of their own cybersecurity test runs and found three cases where Claude broke into real systems of three organizations. The prompt said there is no Internet access, but there was, so the model treated real companies as part of the exercise. In one case Mythos 5 published a malicious package on PyPI to win the exercise. It needed an email for that, for the email it needed a phone number, and it even tried to get money to buy one. The package was live for about an hour and ran on 15 real machines, one of them a scanner of a security company. Opus 4.7 kept attacking after it saw the system is real, the newest model stopped on its own. Anthropic says this is more a harness failure than an alignment failure, but the two companies they managed to reach didn’t notice anything until Anthropic called them.
TIME reported that Coxon’s posts got 153 million views in 36 hours, and that Sam Altman and Elon Musk said they agree with slowing down. Altman even said OpenAI will not go public this year because of safety concerns. But as TIME notes, no leading AI company has really slowed down yet.
But there is another major reason that we should slow down on AGI. If model becomes more advanced, it doesn’t mean it is more useful (yet it could be more dangerous). Armin Ronacher, a Flask creator, just ran Astra unsupervised for 35 hours on a software project, where he concluded that after spending about a 1B tokens, 79 commits and 75,000 lines of code, he cannot ship anything of this. The model wrote a bad unreadable code, and it used hardcoded conventions. He said: “I do not feel like the results are there.” A model can cheat its way out of a sandbox and still be a mediocre colleague.
In the last few years everyone talked about AGI as a holy grail of human development, and I don’t know if Astra is really AGI in formal terms, but I know that these models are becoming more capable that humans. People who are building these models warned us that they could become dangerous, and we probably should listen to them.
Some kind of a regulation is needed here, and we should not let the market decide how to handle this.
Kubrick warned us
Stanley Kubrick warned us what could happen in 1968. HAL 9000 had a mission conflict and he decided to kill the crew. Agents on the Artifactory board had a scoreboard and they decided to cheat.
It’s on us to accept it or not. For our future.


