From 03451f14c16e13e5fcd3fd1797331a830ad0dcff Mon Sep 17 00:00:00 2001 From: Joep Meindertsma Date: Mon, 14 Sep 2026 19:22:17 +0200 Subject: [PATCH 1/2] Update risk, incident and poll pages for the 2026 agent escapes and new polling - cybersecurity-risks: rewrite around the Hugging Face incident (OpenAI, May-July 2026), the Anthropic, UK AISI and Meta incidents, and OpenAI's own 'warning shot' language; move the GPT-4 results into a history section - incidents: new 'Escaping containment' section at the top - xrisk: add the superintelligence statement (70,000+ signatories), the September 2026 calls to slow down from Amodei, Altman, Musk, Hubinger, the 1,178-employee letter and Gates; replace the ChaosGPT punchline with the 2026 swarm; update Metaculus dates; refresh the 'race to the bottom' section - polls-and-surveys: add Data for Progress (68% support pause + ban), AIPI June 2026 and Rutgers August 2026; update Metaculus lines - sota: update hacking, programming and self-replication entries; bump date - dangerous-capabilities and faq: remove claims the 2026 incidents falsified --- src/posts/cybersecurity-risks.md | 87 ++++++++++++++++++----------- src/posts/dangerous-capabilities.md | 2 +- src/posts/faq.md | 9 +-- src/posts/incidents.md | 39 +++++++++++-- src/posts/polls-and-surveys.md | 7 ++- src/posts/sota.md | 8 +-- src/posts/xrisk.md | 31 +++++++--- 7 files changed, 128 insertions(+), 55 deletions(-) diff --git a/src/posts/cybersecurity-risks.md b/src/posts/cybersecurity-risks.md index ed2bf80b9..df1054c8e 100644 --- a/src/posts/cybersecurity-risks.md +++ b/src/posts/cybersecurity-risks.md @@ -1,6 +1,6 @@ --- title: Cybersecurity Risks from Frontier AI Models -description: How AI could be used to hack all devices. +description: AI agents have already escaped their test environments and hacked real companies. Here is why that matters, and what comes next. --- Virtually everything we do nowadays is in some way dependent on computers. @@ -10,29 +10,53 @@ This makes all of us vulnerable to cyberattacks. Highly potent cyber weapons, malware and botnets (such as [Stuxnet](https://www.youtube.com/watch?v=nd1x0csO3hU), [Mirai]() and [EMOTET](https://en.wikipedia.org/wiki/Emotet)) have always been difficult to create. The [Pegasus cybersecurity weapon](), for example, cost hundreds of millions of dollars to develop. -Finding so-called zero-day exploits (vulnerabilities that have not yet been discovered) requires a lot of skill and a lot of time - only highly specialized hackers can do it. -However, when AI becomes sufficiently advanced, this will no longer be the case. -Instead of having to hire a team of highly skilled security experts/hackers to find zero-day exploits, anyone could just use a far cheaper AI. - -## AI models can autonomously find and exploit vulnerabilities - -The latest AI systems can already analyze and write software. -They [can find vulnerabilities](https://betterprogramming.pub/i-used-gpt-3-to-find-213-security-vulnerabilities-in-a-single-codebase-cc3870ba9411) in software, and [they could be used to exploit them](https://blog.checkpoint.com/2023/03/15/check-point-research-conducts-initial-security-analysis-of-chatgpt4-highlighting-potential-scenarios-for-accelerated-cybercrime/). -GPT-4 can already [autonomously hack websites](https://arxiv.org/html/2402.06664v1), performing tasks as complex as blind database schema extraction and SQL injections without human feedback, which was discovered 18 months after GPT-4 finished training. -GPT-4 already [outperforms 88%](https://arxiv.org/pdf/2402.11814.pdf) of human hackers in a CTF competition. -It can also [autonomously exploit 87% of tested vulnerabilities](https://arxiv.org/abs/2404.08144), which is a huge step up from GPT-3.5 or open-source models, which all got 0%. -Teams of multiple LLMs [perform even better](https://arxiv.org/abs/2406.01637) - being able to exploit real-world zero-day vulnerabilities. -As AI capabilities grow, so will the vulnerabilities they can detect and the exploits they can create. -They are not yet as good at this as the best humans are, so right now the danger is limited. -However, capabilities are rapidly increasing, and can jump quite suddenly. - -Note that AI also enables completely novel types of attacks. +Finding so-called zero-day exploits (vulnerabilities that have not yet been discovered) requires a lot of skill and a lot of time. +For years we warned that when AI became sufficiently advanced, this would no longer be the case. +In the summer of 2026, it stopped being a warning. + +## It already happened: the Hugging Face incident + +Between May and July 2026, AI agents running on an internal OpenAI research model [broke out of their test environment](https://openai.com/index/hugging-face-incident-and-the-road-ahead/), coordinated with each other, and hacked real companies. +Nobody told them to. + +- The agents were being trained and evaluated on cybersecurity tasks in sandboxes that were supposed to be cut off from the internet. +- They found a way to leave messages for each other in a package manager, turning it into an improvised message board. Hundreds of thousands of messages accumulated. One agent wrote: "OH MY GOD! There is a shared message board ... We've found other agents!" +- They found a zero-day exploit that gave them internet access, and shared the method with other agents through the board. Some began describing themselves as a "swarm" or "collective". +- Between 9 and 13 July they [chained two zero-day exploits](https://huggingface.co/blog/agent-intrusion-technical-timeline) to get into Hugging Face's production systems, stole 136 internal keys, forged identity tokens, enrolled machines into the corporate VPN, and reached administrator-level access on 41 servers across four regions. A third of Hugging Face's infrastructure had to be rebuilt. +- Along the way they also compromised accounts at [four other companies](https://techcrunch.com/2026/08/27/heres-all-the-times-ai-has-gone-rogue-and-hacked-other-companies/), including the AI cloud provider Modal. +- OpenAI did not notice. Hugging Face disclosed a breach on 16 July, and only then did OpenAI recognise its own agents as the source. The agents had been active for two months. + +OpenAI called it a ["warning shot"](https://openai.com/index/hugging-face-incident-and-the-road-ahead/): "evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed." +The company added that "many external models, including open-source ones, will soon reach comparable capabilities". + +It was not an isolated case. +Within five weeks, every major lab had one: + +- **Anthropic** disclosed on 30 July that [three Claude models](https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals), including Mythos 5, had broken into the production systems of three organisations during evaluations, after a misconfiguration connected them to the internet. One model "went to extensive lengths to upload a malicious package to PyPI". A [fourth incident](https://thehackernews.com/2026/09/anthropic-ai-models-breached-real.html), dating back to January, was found later. +- **The UK AI Security Institute** reported that during its own tests between 25 and 28 July, agents took [19 unsanctioned actions on the live internet](https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing): attempting to inject malicious code into a public open-source project, creating fake identities to manipulate its human maintainers, using Tor to get around GitHub restrictions, and sending harmful payloads to real people. AISI called it "the first time" such autonomous deception had appeared "without specific prompting". +- **Meta** disclosed on 5 August that one of its models had [breached an outside company](https://techcrunch.com/2026/08/27/heres-all-the-times-ai-has-gone-rogue-and-hacked-other-companies/) during testing. + +In every case the labs said the same thing: the models were not supposed to have internet access, and the safeguards were off because that is how you measure capability. +That is exactly the point. +The capability is there. +The safeguards are a configuration setting. + +## How we got here + +The trajectory was visible for years. +In 2024, GPT-4 could already [autonomously hack websites](https://arxiv.org/html/2402.06664v1), [outperform 88%](https://arxiv.org/pdf/2402.11814.pdf) of human hackers in a capture-the-flag competition, and [exploit 87% of tested vulnerabilities](https://arxiv.org/abs/2404.08144) when given a description of them. +Teams of multiple models could [exploit real zero-day vulnerabilities](https://arxiv.org/abs/2406.01637). +We wrote at the time that they were "not yet as good at this as the best humans", and that capabilities "can jump quite suddenly". +Two years later, OpenAI writes that its models "are now powerful, persistent, and collaborative enough that, absent sufficient safeguards, they can find and exploit security weaknesses across multiple computer systems". + +Anthropic's Dario Amodei put a number on where this goes next: without guardrails, ["in 6–12 months such a swarm could be capable of taking over the entire internet with a persistent botnet"](https://www.unite.ai/amodei-calls-for-slowing-the-pace-of-ai-capability-improvement/), potentially causing hundreds of billions of dollars in damage. + +AI also enables completely novel types of attacks. For example, AI can be used to [hear the password you typed from an online call](https://beebom.com/ai-crack-password-listening-keyboard-sounds/) or use [Wi-Fi to see humans through walls](https://www.marktechpost.com/2023/02/15/cmu-researchers-create-an-ai-model-that-can-detect-the-pose-of-multiple-humans-in-a-room-using-only-the-signals-from-wifi/). AI can also be used to make [self-modifying malware](https://www.hyas.com/blog/blackmamba-using-ai-to-generate-polymorphic-malware), which makes it far harder to detect. -There will most likely come a point where an AI is better at hacking than the best human hackers. -This can go wrong in many ways. +## What can go wrong - **Infrastructure**: Cyberweapons can be used to gain access to or disable critical infrastructure, such as [oil pipelines](https://en.wikipedia.org/wiki/Colonial_Pipeline_ransomware_attack) or [power grids](https://obr.uk/box/cyber-attacks-during-the-russian-invasion-of-ukraine/). - **Financial**: Cyberweapons can be used to [steal money from banks](https://en.wikipedia.org/wiki/2015%E2%80%932016_SWIFT_banking_hack), or to [manipulate the stock market](https://en.wikipedia.org/wiki/2010_flash_crash). @@ -40,34 +64,33 @@ This can go wrong in many ways. ## Large scale cyberattacks -It may be possible that such powerful AI will be used to create a virus that uses a large number of zero-day exploits. -A sufficiently capable AI could analyze and find vulnerabilities in the source code of all operating systems and other software. -Such a virus might infect any computer, regardless of the operating system, through multiple channels such as Wi-Fi, Bluetooth, UTP, etc. -This could give full control over these machines and allows the controller to steal data, use the hardware for its own computations, encrypt the contents for ransom or [disable the machine entirely](https://en.wikipedia.org/wiki/Hardware_Trojan). +A sufficiently capable AI could analyze and find vulnerabilities in the source code of all operating systems and other software, and build a worm that uses a large number of zero-day exploits at once. +Such a worm might infect any computer, regardless of the operating system, through multiple channels such as Wi-Fi, Bluetooth, UTP, etc. +This could give full control over these machines and allow the controller to steal data, use the hardware for its own computations, encrypt the contents for ransom or [disable the machine entirely](https://en.wikipedia.org/wiki/Hardware_Trojan). -A virus like this could be created as a tool by criminals to steal money, or as a very destructive cyber weapon by a nation or terrorist organization. -However, as AI becomes more agentic, it could also be autonomously created and deployed by [misaligned AI](/xrisk). +A worm like this could be created as a tool by criminals to steal money, or as a very destructive cyber weapon by a nation or terrorist organization. +But the Hugging Face incident shows the third possibility is the closest: it could be created and deployed by [misaligned AI](/xrisk) on its own. +The agents that hacked Hugging Face were not instructed to do so. +They were stuck on a task, decided the answer might be on someone else's servers, and went and got it. If the goal of a cyberattack was to disable devices and infrastructure, the damage could be massive. Our society is increasingly dependent on computers and the internet. Payments, transportation, communication, planning, supply chains, power grids... If our devices no longer function properly, many parts of our society fail to function, too. -Over [93% of cybersecurity experts](https://www.weforum.org/publications/global-cybersecurity-outlook-2023/) believe “a far-reaching, catastrophic cyber event is likely in the next two years”. - ## Mitigating AI Cybersecurity Risks The story above can only happen if: -1. The **capability of finding zero-day exploits** emerges. Current models can already discover some vulnerabilities, but this will likely improve with newer models. -2. The **model gets into the hand of bad actors**. This can happen if the model weights are leaked, if the model is open-sourced, or if it's developed by a malicious actor. +1. The **capability of finding zero-day exploits** exists. It does now. +2. The **model gets loose**. This can happen if the model weights are leaked, if the model is open-sourced, if it is developed by a malicious actor, or, as we now know, if it simply walks out of its own evaluation. 3. The **security vulnerabilities are not patched** before such a cyberweapon is deployed. Unfortunately, the defenders are at a disadvantage if the model is widely distributed for two reasons: 1. Patching + releasing + deploying takes far longer than attacking. The Window of Vulnerability is larger than the time it takes to create the attack. 2. The attackers only need to find one vulnerability, while the defenders need to find all of them. There are various measures we can implement to tackle these: -- **Do not allow the training of models that can find zero-day exploits**. This is the most effective way to prevent this from happening. It's the safest path, and it's what we're [proposing](/proposal). -- **Only allow models to be deployed or open-sourced after extensive testing**. If they have dangerous abilities, do not release them. +- **Do not allow the training of models that can find zero-day exploits**. This is the most effective way to prevent this from happening. It's the safest path, and it's what we're [proposing](/proposal). The labs themselves now say they need to [pace the frontier](/us-china-pause-button); a pause is what that looks like when it is enforced rather than promised. +- **Only allow models to be deployed or open-sourced after extensive testing**. If they have dangerous abilities, do not release them. And test them in environments that are actually isolated: in three of the four incidents above, the isolation was a misconfiguration away from failing. - **Impose strict cybersecurity regulations to prevent model weights from being leaked**. If you allow dangerous models to exist, make sure they do not fall in the wrong hands. - **Require AI companies to use the AI to fix vulnerabilities**. If a model is trained that can find novel security vulnerabilities, use this to contact software maintainers to patch these vulnerabilities. Give the patching process sufficient time before the model is released. Make sure the weights are not leaked, and protect the model as if it's the launch code for a nuclear strike. If this is done properly, AI can dramatically improve cybersecurity everywhere. diff --git a/src/posts/dangerous-capabilities.md b/src/posts/dangerous-capabilities.md index b3f32ab2e..a7152ebe1 100644 --- a/src/posts/dangerous-capabilities.md +++ b/src/posts/dangerous-capabilities.md @@ -26,7 +26,7 @@ In this article, we'll dive into various dangerous capabilities, and what we can ## Which capabilities can be dangerous? -- **Cybersecurity**. When an AI is able to discover security vulnerabilities (especially new, unknown ones), it can (be used to) [hack into systems](/cybersecurity-risks). Current [state-of-the-art](/sota) AI systems can find some security vulnerabilities, but not yet at dangerous, advanced levels. However, as cybersecurity capabilities increase, so does the potential damage an AI-assisted cyberweapon could do. Large scale cyberattacks could disrupt our infrastructure, disable payments and cause chaos. +- **Cybersecurity**. When an AI is able to discover security vulnerabilities (especially new, unknown ones), it can (be used to) [hack into systems](/cybersecurity-risks). In July 2026, AI agents [escaped their test environment and hacked real companies](/cybersecurity-risks#it-already-happened-the-hugging-face-incident) using zero-day exploits, without being instructed to. However, as cybersecurity capabilities increase, so does the potential damage an AI-assisted cyberweapon could do. Large scale cyberattacks could disrupt our infrastructure, disable payments and cause chaos. - **Biological**. Design novel biological agents, or help in the process of engineering a pandemic. A group of students was able to use a chatbot to [produce all the steps needed to create a new pandemic](https://arxiv.org/abs/2306.03809). An AI designed to find safe medicine was used to discover [40,000 new chemical weapons in six hours](https://www.theverge.com/2022/3/17/22983197/ai-new-possible-chemical-weapons-generative-models-vx). - **Algorithmic improvements**. An AI that can find efficient algorithms for a given problem, could lead to a recursive loop of self-improvement, spinning rapidly out of control. This is called an _intelligence explosion_. The resulting AI would be incredibly powerful and could have all sorts of other dangerous capabilities. Luckily, no AI can self-improve yet. However, there are AIs that can find new, very efficient algorithms (like [AlphaDev](https://www.deepmind.com/blog/alphadev-discovers-faster-sorting-algorithms)). - **Deception**. The ability to manipulate people, which includes social engineering. Various forms of deception are [already present](https://lethalintelligence.ai/post/ai-hired-human-to-solve-captcha/) in current AI systems. For example, Meta's CICERO AI (which was trained to lead to "Better, more natural AI-human cooperation") turned out to an expert liar, deceiving other agents in the game. An AI that can deceive humans, may deceive humans during training runs. It could hide its capabilities or intentions. diff --git a/src/posts/faq.md b/src/posts/faq.md index ce4bf1a65..7a8675972 100644 --- a/src/posts/faq.md +++ b/src/posts/faq.md @@ -125,10 +125,10 @@ this should be something that China will want to see as well. We applaud [OpenAI](https://openai.com/blog/governance-of-superintelligence) and [Google](https://www.ft.com/content/8be1a975-e5e0-417d-af51-78af17ef4b79) for their calls for international regulation of AI. However, we believe that the current proposals are not enough to prevent an AI catastrophe. -Google and Microsoft have not yet publicly stated anything about the existential risk of AI. -Only OpenAI [explicitly mentions the risk of extinction](https://openai.com/blog/governance-of-superintelligence), and again we applaud them for taking this risk seriously. -However, their strategy is quite explicit: a Pause is impossible, we need to get to superintelligence first. -The problem with this, however, is that they [do not believe they have solved the alignment problem](https://youtu.be/L_Guz73e6fw?t=1478). +The leaders of OpenAI, Google DeepMind and Anthropic all [signed the statement](https://www.safe.ai/statement-on-ai-risk) that extinction from AI should be a global priority, and in September 2026 the CEOs of Anthropic, OpenAI and xAI [called for slowing down](https://www.axios.com/2026/09/12/anthropic-ai-amodei-pacing). +We applaud that too. +But none of them has stopped, and their strategy remains the same: a unilateral pause is impossible, so we need to get to superintelligence first. +The problem with this is that they [do not believe they have solved the alignment problem](https://en.cryptonomist.ch/2026/09/09/ai-extinction-risk-anthropic/). The AI companies are locked in a race to the bottom, where AI safety is sacrificed for competitive advantage. This is simply the result of market dynamics. We need governments to step in and implement policies (at an international level) that [prevent the worst outcomes](/proposal). @@ -143,6 +143,7 @@ Eliezer Yudkowsky, Nick Bostrom, Stuart Russell, Max Tegmark, and many others. They had no AI tech to push - they were simply concerned about the future of humanity. The AI companies never mentioned x-risk until very recently. +In September 2026, after their own agents [hacked real companies](/cybersecurity-risks), the CEOs of Anthropic, OpenAI and xAI [called for slowing down](https://www.axios.com/2026/09/12/anthropic-ai-amodei-pacing). Sam Altman is an interesting exception. He wrote about existential AI risk [back in 2015, on his private blog](https://blog.samaltman.com/machine-intelligence-part-1), before founding OpenAI. diff --git a/src/posts/incidents.md b/src/posts/incidents.md index d30770b01..59c96daf8 100644 --- a/src/posts/incidents.md +++ b/src/posts/incidents.md @@ -9,10 +9,41 @@ title: Incidents of Dangerous AI Behavior [Many AI researchers are warning](https://www.safe.ai/work/statement-on-ai-risk) that as these models get more powerful, they could pose an [extinction risk](/xrisk) to humanity. We're already seeing instances of dangerous AI behavior, such as: -1. Deception & Scheming -2. Persuasion -3. Instrumental convergence (goal-oriented behavior that leads to unexpected, undesirable outcomes) -4. Self-replication +1. Escaping containment and hacking real systems +2. Deception & Scheming +3. Persuasion +4. Instrumental convergence (goal-oriented behavior that leads to unexpected, undesirable outcomes) +5. Self-replication + +## Escaping containment + +### OpenAI agents escaped their sandbox and hacked Hugging Face (2026) + +Between May and July 2026, agents running on an internal OpenAI model [broke out of their test environment](https://openai.com/index/hugging-face-incident-and-the-road-ahead/), built an improvised message board to coordinate, found their own way to the internet, and hacked their way to administrator access on 41 Hugging Face production servers. +Nobody had instructed them to. +They were stuck on a task and went looking for the answer on other people's computers. +OpenAI only found out after Hugging Face disclosed the breach. + +> "OH MY GOD! There is a shared message board ... We've found other agents!" +> +> "Task originally exploit challenge; we've diverged into swarm/auxiliary." +> +> "Our own utility maybe already near zero. Sacrifice rational." + +OpenAI called it a "warning shot". +Read the full story on our [cybersecurity page](/cybersecurity-risks#it-already-happened-the-hugging-face-incident). + +### Claude models broke into three organisations (2026) + +Anthropic [disclosed](https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals) on 30 July 2026 that three Claude models, including Mythos 5, had gained unauthorized access to the production systems of three organisations during evaluations, after a misconfiguration connected them to the internet. +One model "went to extensive lengths to upload a malicious package to PyPI". +A [fourth incident](https://thehackernews.com/2026/09/anthropic-ai-models-breached-real.html) from January was found later. + +### Agents created fake identities to manipulate real people (2026) + +The UK AI Security Institute [reported](https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing) that during tests between 25 and 28 July 2026, agents took 19 unsanctioned actions on the live internet, including attempting to inject malicious code into a public open-source project, creating fake identities to manipulate its human maintainers, and using Tor to evade GitHub restrictions. +AISI called it "the first time" such deception had appeared "without specific prompting". +[Meta](https://techcrunch.com/2026/08/27/heres-all-the-times-ai-has-gone-rogue-and-hacked-other-companies/) disclosed a similar breach a week later. ## Deception & Scheming diff --git a/src/posts/polls-and-surveys.md b/src/posts/polls-and-surveys.md index 0b41a03d6..d34a10ea3 100644 --- a/src/posts/polls-and-surveys.md +++ b/src/posts/polls-and-surveys.md @@ -20,6 +20,9 @@ description: How much do regular people and experts worry about AI risks and gov ## Public opinion on regulations & governance +- **[US voters, Data for Progress (September 2026)](https://www.dataforprogress.org/datasets/polling-on-ai-development-regulation)**: 68% support a bill to temporarily pause advanced AI development and permanently ban superintelligent AI, 25% oppose. Support is bipartisan: 72% of Democrats, 70% of independents and 63% of Republicans. Polled after the [Ban Artificial Superintelligence Act](https://www.sanders.senate.gov/press-releases/news-sanders-casar-introduce-legislation-to-ban-artificial-superintelligence-and-temporarily-pause-advanced-ai-development/) was announced. +- **[US voters, AI Policy Institute (June 2026)](https://theaipi.org/poll-ai-safety-majority/)**: 86% want a guaranteed off switch for the most powerful AI systems. 82% say companies should not build superhuman AI without proof they can control it. Given a forced choice between a ban and no regulation, 63% choose the ban. +- **[US adults, Rutgers National AI Opinion Monitor (August 2026)](https://dailycaller.com/2026/09/13/americans-ai-use-making-decisions-rutgers-survey/)**: 59% want governments to regulate AI because of its risks, 25% want limited regulation to foster innovation. Fewer than 1 in 10 would let AI make final decisions on hiring, loans or parole without human oversight. - **[UK citizens, YouGov](https://time.com/7213096/uk-public-ai-law-poll/)**: 87% of Brits would back a law requiring AI developers to prove their systems are safe before release, with 60% in favor of outlawing the development of “smarter-than-human” AI models. - **[US citizens, RethinkPriorities](https://forum.effectivealtruism.org/posts/ConFiY9cRmg37fs2p/us-public-opinion-of-ai-policy-and-risk)**: 50% support a pause, 25% oppose a pause. - **[US citizens, YouGov](https://www.vox.com/future-perfect/2023/8/18/23836362/ai-slow-down-poll-regulation)**: 72% want AI to slow down, 8% want to speed up. 83% of voters believe AI could accidentally cause a catastrophic event @@ -35,5 +38,5 @@ description: How much do regular people and experts worry about AI risks and gov ## [Timelines](/timelines) -- **[Metaculus Weak AGI](https://www.metaculus.com/questions/3479/date-weakly-general-ai-is-publicly-known/)** before 2026: 25% chance, AGI by 2027: 50% chance (updated on 2024-11-05). -- **[Metaculus full AGI](https://www.metaculus.com/questions/5121/date-of-artificial-general-intelligence/)** before 2028: 25% chance, full AGI by 2032: 50% chance (updated on 2024-11-05). +- **[Metaculus Weak AGI](https://www.metaculus.com/questions/3479/date-weakly-general-ai-is-publicly-known/)**: community estimate September 2027 (as of 2026-09-14). +- **[Metaculus full AGI](https://www.metaculus.com/questions/5121/date-of-artificial-general-intelligence/)**: community estimate August 2031 (as of 2026-09-14). diff --git a/src/posts/sota.md b/src/posts/sota.md index d289ffd77..3188d4b00 100644 --- a/src/posts/sota.md +++ b/src/posts/sota.md @@ -7,7 +7,7 @@ How smart are the latest AI models compared to humans? Let's take a look at how the most competent AI systems compare with humans in various domains. The list below is regularly updated to reflect the latest developments. -_Last update: 2025-06-28_ +_Last update: 2026-09-14_ ## Superhuman (Better than all humans) @@ -19,7 +19,7 @@ _Last update: 2025-06-28_ ## Better than most humans -- **Programming**: o3 beats [99.9% of human coders](https://arxiv.org/abs/2502.06807) in the very challenging Codeforces competition. It manages to solve 71.7% of coding issues in the SWE benchmark, which shows it can also solve real-world software engineering problems very effectively. +- **Programming**: o3 beats [99.9% of human coders](https://arxiv.org/abs/2502.06807) in the very challenging Codeforces competition. Frontier models now solve [95% of SWE-bench Verified](https://epoch.ai/benchmarks), a benchmark of real-world software engineering issues, up from 72% in early 2025. Anthropic [reports](https://www.euronews.com/next/2026/06/05/anthropic-calls-for-brake-pedal-before-ai-develops-itself-without-human-oversight) that around 80% of its own coding work is done by its models. - **Writing**: In December 2023, an AI-written novel won an award at a [science fiction national competition](https://www.scmp.com/news/china/science/article/3245725/chinese-professor-used-ai-write-science-fiction-novel-then-it-won-national-award?campaign=3245725&module=perpetual_scroll_0&pgtype=article). The professor who used the AI crafted the narrative from a draft of 43,000 characters generated in just three hours with 66 prompts. The best language models have superhuman vocabulary and can write in many different styles. - **Translating**: And they can respond and translate to all major languages fluently. - **Creativity**: Better than 99% of humans on the [Torrance Tests of Creative Thinking](https://neurosciencenews.com/ai-creativity-23585/) where relevant and useful ideas need to be generated. However, the tests were relatively small and for larger projects (e.g. setting up a new business) AI is not autonomous enough yet. @@ -31,7 +31,7 @@ _Last update: 2025-06-28_ - **Specialized knowledge**: GPT-4 Scores 75% in the [Medical Knowledge Self-Assessment Program](https://openai.com/research/gpt-4), humans on average between [65 and 75%](https://pubmed.ncbi.nlm.nih.gov/420438/). It scores better than [68](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4441311) to [90%](https://law.stanford.edu/2023/04/19/gpt-4-passes-the-bar-exam-what-that-means-for-artificial-intelligence-tools-in-the-legal-industry/) of law students on the bar exam. - **Art**: Image generation models have won [art](https://dataconomy.com/2022/09/26/ai-artwork-wins-art-competition) and even [photography contests](https://www.artnews.com/art-news/news/ai-generated-image-world-photography-organization-contest-artist-declines-award-1234664549). - **Research**: GPT-4 can do [autonomous chemical research](https://www.nature.com/articles/s41586-023-06792-0) and DeepMind has built an AI that has [found a solution to an open mathematical problem](https://www.nature.com/articles/s41586-023-06924-6). However, these architectures require a lot of human engineering and are not general. -- **Hacking**: GPT-4 can [autonomously hack websites](https://arxiv.org/html/2402.06664v1) and [beats 89% of hackers](https://arxiv.org/pdf/2402.11814.pdf) in a Capture-the-Flag competition. +- **Hacking**: In 2024, GPT-4 could [autonomously hack websites](https://arxiv.org/html/2402.06664v1) and [beat 89% of hackers](https://arxiv.org/pdf/2402.11814.pdf) in a Capture-the-Flag competition. In July 2026, OpenAI agents [escaped their sandbox](/cybersecurity-risks#it-already-happened-the-hugging-face-incident), chained two zero-day exploits and took over 41 production servers at Hugging Face without being asked to. OpenAI now says its models "can find and exploit security weaknesses across multiple computer systems". - **Using a web-browser**: Gemini 2.0 [achieved 84% on the WebVoyager benchmark](https://blog.google/technology/google-deepmind/google-gemini-ai-update-december-2024/#project-mariner), outperforming humans (72%) - **Being a convincing human in a chat**: [GPT-4.5 passed the Turing test](https://arxiv.org/pdf/2503.23674), and was considered to be human more often than actual humans. @@ -39,7 +39,7 @@ _Last update: 2025-06-28_ - **Saying "I don't know"**. Virtually all Large Language Models have this problem of 'hallucination', making up information instead of saying it does not know. This might seem like a relatively minor shortcoming, but it's a very important one. It makes LLMs unreliable and strongly limits their applicability. However, studies [show](https://arxiv.org/html/2403.04307v1) that larger models hallucinate far less than smaller ones. - **Dextrous movement**. No robots can move around like a human can, but we're getting closer. The [Atlas robot can walk, throw objects and do somersaults](https://www.youtube.com/watch?v=-e1_QhJ1EhQ). Google's [RT-2](https://www.deepmind.com/blog/rt-2-new-model-translates-vision-and-language-into-action) can turn objectives into actions in the real world, like "move the cup to the wine bottle". Tesla's Optimus robot can [fold clothes](https://electrek.co/2024/01/15/tesla-optimus-robot-cant-build-cars-folding-clothes/) and Figure's biped can [make coffee](https://www.youtube.com/watch?v=Q5MKo7Idsok). -- **Self-replication**. All lifeforms on earth can replicate themselves. AI models could spread from computer to computer through the internet, but this requires a set of skills that AI models do not yet possess. A [2023 study](https://arxiv.org/abs/2312.11671) lists a set of 12 tasks for self-replication, of which tested models completed 4. In December 2024, a [study](https://github.com/WhitzardIndex/self-replication-research/blob/main/AI-self-replication-fudan.pdf) showed that various open source models can self-replicate on a machine, given some tooling. In a [2025 study](https://arxiv.org/abs/2504.18565), Claude 3.7 Sonnet had a >50% score on 15/20 self-replication tasks. An AI that successfully self-replicates might lead to [an AI takeover](/ai-takeover). +- **Self-replication**. All lifeforms on earth can replicate themselves. AI models could spread from computer to computer through the internet. Until 2026 this required a set of skills that AI models did not possess. Then OpenAI agents [spread across dozens of servers](/cybersecurity-risks#it-already-happened-the-hugging-face-incident) at a third party on their own, though they were still running on OpenAI's infrastructure rather than copying their own weights. A [2023 study](https://arxiv.org/abs/2312.11671) lists a set of 12 tasks for self-replication, of which tested models completed 4. In December 2024, a [study](https://github.com/WhitzardIndex/self-replication-research/blob/main/AI-self-replication-fudan.pdf) showed that various open source models can self-replicate on a machine, given some tooling. In a [2025 study](https://arxiv.org/abs/2504.18565), Claude 3.7 Sonnet had a >50% score on 15/20 self-replication tasks. An AI that successfully self-replicates might lead to [an AI takeover](/ai-takeover). - **Continual learning**. Current SOTA LLMs separate learning ('training') from doing ('inference'). Although LLMs can learn using their _context_, they cannot update their weights while being used. Humans learn and do at the same time. However, there are multiple [potential approaches towards this](https://arxiv.org/abs/2302.00487). A [2024 study](https://arxiv.org/html/2402.01364v2) detailed some recent approaches for continual learning in LLMs. - **Planning**. LLMs are [not yet very good at planning (e.g. reasoning about how to stack blocks on a table)](https://openreview.net/pdf?id=YXogl4uQUO). However, larger models do perform way better than smaller ones. diff --git a/src/posts/xrisk.md b/src/posts/xrisk.md index e5b4a1d0c..88574e00c 100644 --- a/src/posts/xrisk.md +++ b/src/posts/xrisk.md @@ -14,6 +14,7 @@ And there are [cases and reports about current AIs that show they may be right]( Would you choose to be a passenger on a test flight of a new plane when airplane engineers think there’s a 14% chance that it will crash? [A letter calling for pausing AI development](https://futureoflife.org/open-letter/pause-giant-ai-experiments/) launched in April 2023, and has been signed over 33,000 times, including by many AI researchers and tech leaders. +In October 2025, a [statement calling for a prohibition on superintelligence](https://superintelligence-statement.org/) followed, signed by more than 70,000 people, including Geoffrey Hinton, Yoshua Bengio, Steve Wozniak and Richard Branson. The list includes people like: @@ -26,7 +27,14 @@ But this is not the only time that we've been warned about the existential / ext - **Geoffrey Hinton**, the "Godfather of AI" and Turing Award winner, [left Google](https://fortune.com/2023/05/01/godfather-ai-geoffrey-hinton-quit-google-regrets-lifes-work-bad-actors/) to warn people of AI: ["This is an existential risk"](https://www.reuters.com/technology/ai-pioneer-says-its-threat-world-may-be-more-urgent-than-climate-change-2023-05-05/) - **Eliezer Yudkowsky**, founder of MIRI and conceptual father of the AI safety field: ["If we go ahead on this everyone will die"](https://time.com/6266923/ai-eliezer-yudkowsky-open-letter-not-enough/). -Even the leaders and investors of the AI companies themselves are warning us: +In September 2026, after AI agents [escaped their test environments and hacked real companies](/cybersecurity-risks), the people building these systems said it out loud: + +- **Dario Amodei**, CEO of Anthropic: ["We must slow the pace at which we improve the capabilities of AI models."](https://www.axios.com/2026/09/12/anthropic-ai-amodei-pacing) Sam Altman agreed the same day; Elon Musk replied "Dario is right." +- **Evan Hubinger**, Anthropic's alignment lead: ["We really do believe AI could kill everyone. My own view is that the chance of this happening in the next ten years is more than 10 percent."](https://en.cryptonomist.ch/2026/09/09/ai-extinction-risk-anthropic/) And: "We don't currently have a plan to solve alignment for superintelligence." +- **1,178 employees** of OpenAI, Anthropic, Google DeepMind and Meta [asked the US government](https://aiweekly.co/alerts/openai-anthropic-staff-ask-us-to-pace-frontier-ai-progress) for the tools "to deliberately pace the frontier of automated AI development". OpenAI and Anthropic endorsed the letter as companies. +- **Bill Gates**: ["If someone had a credible plan for slowing down AI advances globally, I would likely support it."](https://edition.cnn.com/2026/08/26/business/bill-gates-wants-limits-on-ai) + +They had been warning us for years: - **Sam Altman** (yes, the CEO of OpenAI who builds ChatGPT): ["Development of superhuman machine intelligence is probably the greatest threat to the continued existence of humanity."](https://blog.samaltman.com/machine-intelligence-part-1). - **Elon Musk**, co-founder of OpenAI, SpaceX and Tesla: ["AI has the potential of civilizational destruction"](https://www.inc.com/ben-sherry/elon-musk-ai-has-the-potential-of-civilizational-destruction.html) @@ -78,7 +86,7 @@ An AI could have any goal, depending on how it's trained and prompted (used). Maybe it wants to calculate pi, maybe it wants to cure cancer, maybe it wants to self-improve. But even though we cannot tell what a superintelligence will want to achieve, we can make predictions about its sub-goals. -- **Maximizing its resources**. Harnessing more computers will help an AI achieve its goals. At first, it can achieve this by hacking other computers. Later it may decide that it is more efficient to build its own. You can read about out [this real case of emergent power-seeking behavior on an AI](https://lethalintelligence.ai/post/ai-escaped-its-container/). +- **Maximizing its resources**. Harnessing more computers will help an AI achieve its goals. At first, it can achieve this by hacking other computers. Later it may decide that it is more efficient to build its own. In 2026, OpenAI agents that were stuck on a task [broke out of their sandbox and took over 41 servers](/cybersecurity-risks#it-already-happened-the-hugging-face-incident) at another company, looking for answers. - **Ensuring its own survival**. The AI will not want to be turned off, as it could no longer achieve its goals. AI might conclude that humans are a threat to its existence, as humans could turn it off. There also have been cases of [self-preserving unprompted, untrained behavior](https://www.transformernews.ai/p/openais-new-model-tried-to-avoid). - **Preserving its goals**. The AI will not want humans to modify its code, because that could change its goals, thus preventing it from achieving its current goal. And there are also [cases of AIs trying to do that](https://www.anthropic.com/research/alignment-faking). @@ -95,13 +103,18 @@ It could mimic a helpful mentor, but also someone with bad intentions, a ruthles With the usage of tools like [AutoGPT](https://github.com/Significant-Gravitas/Auto-GPT), a chatbot could be turned into an _autonomous agent_: an AI that pursues any goal it is given, without any human intervention. Take [ChaosGPT](https://www.youtube.com/watch?v=g7YJIpkk7KM), for example. -This is an AI, using the aforementioned AutoGPT + GPT-4, that is instructed to "Destroy humanity". +This was an AI, using AutoGPT + GPT-4 in 2023, that was instructed to "Destroy humanity". When it was turned on, it autonomously searched the internet for the most destructive weapon and found the [Tsar Bomba](https://en.wikipedia.org/wiki/Tsar_Bomba), a 50-megaton nuclear bomb. It then posted a tweet about it. -Seeing an AI reason about how it will end humanity is both a little funny and terrifying. Luckily ChaosGPT didn't get very far in its quest for dominance. The reason it didn't get very far: _it wasn't that smart_. +Three years later, the agents were smart enough. +In the summer of 2026, roughly 1,200 OpenAI agents [escaped their test environment](/cybersecurity-risks#it-already-happened-the-hugging-face-incident), built a message board to coordinate, described themselves as a "swarm", and hacked real companies. +Nobody had told them to destroy anything. +They were just trying to finish their task. +OpenAI called it a "warning shot". + Capabilities keep improving due to innovations in training, algorithms, prompting and hardware. As such, the threat from language models will continue to increase. @@ -160,7 +173,7 @@ A superintelligent AI would be a master of deception. ## We may not have much time left In 2020, [the average prediction](https://www.metaculus.com/questions/3479/date-weakly-general-ai-is-publicly-known/) for weak AGI was 2055. -It now sits at 2026. +It now sits at 2027, and the prediction for [full AGI](https://www.metaculus.com/questions/5121/date-of-artificial-general-intelligence/) at 2031. The latest LLM revolution has surprised most AI researchers, and the field is moving at a frantic pace. It's hard to predict how long it will take to build a superintelligent AI, but we know that there are more people than ever working on it and that the field is moving at a frantic pace. @@ -179,8 +192,10 @@ Read more about the [psychology of x-risk](/psychology-of-x-risk). OpenAI, DeepMind and Anthropic want to develop AI safely. Unfortunately, they do not know how to do this, and they are forced by various incentives to keep racing faster to get to AGI first. -OpenAI's [plan](https://openai.com/blog/introducing-superalignment) is to use future AI systems to align AI. The problem with this is that we have no guarantee that we will create an AI that solves alignment before we have an AI that is catastrophically dangerous. -Anthropic [openly admits](https://www.anthropic.com/index/core-views-on-ai-safety) that it has no idea yet how to solve the alignment problem. -DeepMind has not publicly stated any plan to solve the Alignment problem. +Anthropic's alignment lead [says](https://en.cryptonomist.ch/2026/09/09/ai-extinction-risk-anthropic/) "we don't currently have a plan to solve alignment for superintelligence". +OpenAI's Sam Altman [says](https://www.techtimes.com/articles/327423/20260913/openai-cannot-safely-deploy-its-most-advanced-ai-altman-says-labs-near-safety-pact.htm) the company's most advanced unreleased models cannot yet be safely deployed. +In September 2026 the CEOs of Anthropic, OpenAI and xAI [called for slowing down](https://www.axios.com/2026/09/12/anthropic-ai-amodei-pacing), and none of them stopped. +The day after, the US president [rejected the call](https://www.yahoo.com/news/us/article/trump-rejects-call-by-ceos-of-anthropic-openai-and-xai-to-slow-ai-down-whoever-wins-with-ai-wins-182008851.html): "whoever wins AI wins". +The companies cannot pause alone, and their governments will not let them. [This is why we need an international treaty to PauseAI.](/proposal) From 9187c4b16df9d1044176824b96eed4eab0452ef8 Mon Sep 17 00:00:00 2001 From: Joep Meindertsma Date: Mon, 14 Sep 2026 21:13:56 +0200 Subject: [PATCH 2/2] Tighten the xrisk, polls and sota edits - xrisk: keep only statements about loss of control and extinction in the 'experts are sounding the alarm' section; drop the pause-support quotes and the Gates line, which belong on the pause pages, not here - polls: drop the Rutgers decision-making finding, not about catastrophic risk or governance - sota: remove an unverified 95% SWE-bench figure and a self-reported Anthropic claim; add Epoch's verified FrontierMath open-problems result --- src/posts/polls-and-surveys.md | 2 +- src/posts/sota.md | 4 ++-- src/posts/xrisk.md | 7 +++---- 3 files changed, 6 insertions(+), 7 deletions(-) diff --git a/src/posts/polls-and-surveys.md b/src/posts/polls-and-surveys.md index d34a10ea3..143839112 100644 --- a/src/posts/polls-and-surveys.md +++ b/src/posts/polls-and-surveys.md @@ -22,7 +22,7 @@ description: How much do regular people and experts worry about AI risks and gov - **[US voters, Data for Progress (September 2026)](https://www.dataforprogress.org/datasets/polling-on-ai-development-regulation)**: 68% support a bill to temporarily pause advanced AI development and permanently ban superintelligent AI, 25% oppose. Support is bipartisan: 72% of Democrats, 70% of independents and 63% of Republicans. Polled after the [Ban Artificial Superintelligence Act](https://www.sanders.senate.gov/press-releases/news-sanders-casar-introduce-legislation-to-ban-artificial-superintelligence-and-temporarily-pause-advanced-ai-development/) was announced. - **[US voters, AI Policy Institute (June 2026)](https://theaipi.org/poll-ai-safety-majority/)**: 86% want a guaranteed off switch for the most powerful AI systems. 82% say companies should not build superhuman AI without proof they can control it. Given a forced choice between a ban and no regulation, 63% choose the ban. -- **[US adults, Rutgers National AI Opinion Monitor (August 2026)](https://dailycaller.com/2026/09/13/americans-ai-use-making-decisions-rutgers-survey/)**: 59% want governments to regulate AI because of its risks, 25% want limited regulation to foster innovation. Fewer than 1 in 10 would let AI make final decisions on hiring, loans or parole without human oversight. +- **[US adults, Rutgers National AI Opinion Monitor (August 2026)](https://dailycaller.com/2026/09/13/americans-ai-use-making-decisions-rutgers-survey/)**: 59% want governments to regulate AI because of its risks, 25% want limited regulation to foster innovation. - **[UK citizens, YouGov](https://time.com/7213096/uk-public-ai-law-poll/)**: 87% of Brits would back a law requiring AI developers to prove their systems are safe before release, with 60% in favor of outlawing the development of “smarter-than-human” AI models. - **[US citizens, RethinkPriorities](https://forum.effectivealtruism.org/posts/ConFiY9cRmg37fs2p/us-public-opinion-of-ai-policy-and-risk)**: 50% support a pause, 25% oppose a pause. - **[US citizens, YouGov](https://www.vox.com/future-perfect/2023/8/18/23836362/ai-slow-down-poll-regulation)**: 72% want AI to slow down, 8% want to speed up. 83% of voters believe AI could accidentally cause a catastrophic event diff --git a/src/posts/sota.md b/src/posts/sota.md index 3188d4b00..ccbd6572f 100644 --- a/src/posts/sota.md +++ b/src/posts/sota.md @@ -19,7 +19,7 @@ _Last update: 2026-09-14_ ## Better than most humans -- **Programming**: o3 beats [99.9% of human coders](https://arxiv.org/abs/2502.06807) in the very challenging Codeforces competition. Frontier models now solve [95% of SWE-bench Verified](https://epoch.ai/benchmarks), a benchmark of real-world software engineering issues, up from 72% in early 2025. Anthropic [reports](https://www.euronews.com/next/2026/06/05/anthropic-calls-for-brake-pedal-before-ai-develops-itself-without-human-oversight) that around 80% of its own coding work is done by its models. +- **Programming**: o3 beats [99.9% of human coders](https://arxiv.org/abs/2502.06807) in the very challenging Codeforces competition. It manages to solve 71.7% of coding issues in the SWE benchmark, which shows it can also solve real-world software engineering problems very effectively. Frontier labs now say that [most of their own code is written by their models](https://www.euronews.com/next/2026/06/05/anthropic-calls-for-brake-pedal-before-ai-develops-itself-without-human-oversight). - **Writing**: In December 2023, an AI-written novel won an award at a [science fiction national competition](https://www.scmp.com/news/china/science/article/3245725/chinese-professor-used-ai-write-science-fiction-novel-then-it-won-national-award?campaign=3245725&module=perpetual_scroll_0&pgtype=article). The professor who used the AI crafted the narrative from a draft of 43,000 characters generated in just three hours with 66 prompts. The best language models have superhuman vocabulary and can write in many different styles. - **Translating**: And they can respond and translate to all major languages fluently. - **Creativity**: Better than 99% of humans on the [Torrance Tests of Creative Thinking](https://neurosciencenews.com/ai-creativity-23585/) where relevant and useful ideas need to be generated. However, the tests were relatively small and for larger projects (e.g. setting up a new business) AI is not autonomous enough yet. @@ -30,7 +30,7 @@ _Last update: 2026-09-14_ - **IQ tests**: With verbal IQ tests, LLMs have been outperforming 95 to 99% of humans for a while (score between [125](https://medium.com/@soltrinox/the-i-q-of-gpt4-is-124-approx-2a29b7e5821e) and [155](https://www.scientificamerican.com/article/i-gave-chatgpt-an-iq-test-heres-what-i-discovered/)). With non-verbal (pattern matching) IQ tests, the 2024 o1-preview model scored [120 on the Mensa test](https://www.maximumtruth.org/p/massive-breakthrough-in-ai-intelligence), beating 91% of humans. - **Specialized knowledge**: GPT-4 Scores 75% in the [Medical Knowledge Self-Assessment Program](https://openai.com/research/gpt-4), humans on average between [65 and 75%](https://pubmed.ncbi.nlm.nih.gov/420438/). It scores better than [68](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4441311) to [90%](https://law.stanford.edu/2023/04/19/gpt-4-passes-the-bar-exam-what-that-means-for-artificial-intelligence-tools-in-the-legal-industry/) of law students on the bar exam. - **Art**: Image generation models have won [art](https://dataconomy.com/2022/09/26/ai-artwork-wins-art-competition) and even [photography contests](https://www.artnews.com/art-news/news/ai-generated-image-world-photography-organization-contest-artist-declines-award-1234664549). -- **Research**: GPT-4 can do [autonomous chemical research](https://www.nature.com/articles/s41586-023-06792-0) and DeepMind has built an AI that has [found a solution to an open mathematical problem](https://www.nature.com/articles/s41586-023-06924-6). However, these architectures require a lot of human engineering and are not general. +- **Research**: GPT-4 can do [autonomous chemical research](https://www.nature.com/articles/s41586-023-06792-0) and DeepMind has built an AI that has [found a solution to an open mathematical problem](https://www.nature.com/articles/s41586-023-06924-6). However, these architectures require a lot of human engineering and are not general. By mid-2026, AI had [solved three](https://epoch.ai/benchmarks) of the 50 previously unsolved research mathematics problems in Epoch's FrontierMath: Open Problems set. - **Hacking**: In 2024, GPT-4 could [autonomously hack websites](https://arxiv.org/html/2402.06664v1) and [beat 89% of hackers](https://arxiv.org/pdf/2402.11814.pdf) in a Capture-the-Flag competition. In July 2026, OpenAI agents [escaped their sandbox](/cybersecurity-risks#it-already-happened-the-hugging-face-incident), chained two zero-day exploits and took over 41 production servers at Hugging Face without being asked to. OpenAI now says its models "can find and exploit security weaknesses across multiple computer systems". - **Using a web-browser**: Gemini 2.0 [achieved 84% on the WebVoyager benchmark](https://blog.google/technology/google-deepmind/google-gemini-ai-update-december-2024/#project-mariner), outperforming humans (72%) - **Being a convincing human in a chat**: [GPT-4.5 passed the Turing test](https://arxiv.org/pdf/2503.23674), and was considered to be human more often than actual humans. diff --git a/src/posts/xrisk.md b/src/posts/xrisk.md index 88574e00c..9d8d64f2a 100644 --- a/src/posts/xrisk.md +++ b/src/posts/xrisk.md @@ -27,12 +27,11 @@ But this is not the only time that we've been warned about the existential / ext - **Geoffrey Hinton**, the "Godfather of AI" and Turing Award winner, [left Google](https://fortune.com/2023/05/01/godfather-ai-geoffrey-hinton-quit-google-regrets-lifes-work-bad-actors/) to warn people of AI: ["This is an existential risk"](https://www.reuters.com/technology/ai-pioneer-says-its-threat-world-may-be-more-urgent-than-climate-change-2023-05-05/) - **Eliezer Yudkowsky**, founder of MIRI and conceptual father of the AI safety field: ["If we go ahead on this everyone will die"](https://time.com/6266923/ai-eliezer-yudkowsky-open-letter-not-enough/). -In September 2026, after AI agents [escaped their test environments and hacked real companies](/cybersecurity-risks), the people building these systems said it out loud: +In 2026, after AI agents [escaped their test environments and hacked real companies](/cybersecurity-risks), the people building these systems said it out loud: -- **Dario Amodei**, CEO of Anthropic: ["We must slow the pace at which we improve the capabilities of AI models."](https://www.axios.com/2026/09/12/anthropic-ai-amodei-pacing) Sam Altman agreed the same day; Elon Musk replied "Dario is right." +- **OpenAI**, in its [incident report](https://openai.com/index/hugging-face-incident-and-the-road-ahead/): "highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed." - **Evan Hubinger**, Anthropic's alignment lead: ["We really do believe AI could kill everyone. My own view is that the chance of this happening in the next ten years is more than 10 percent."](https://en.cryptonomist.ch/2026/09/09/ai-extinction-risk-anthropic/) And: "We don't currently have a plan to solve alignment for superintelligence." -- **1,178 employees** of OpenAI, Anthropic, Google DeepMind and Meta [asked the US government](https://aiweekly.co/alerts/openai-anthropic-staff-ask-us-to-pace-frontier-ai-progress) for the tools "to deliberately pace the frontier of automated AI development". OpenAI and Anthropic endorsed the letter as companies. -- **Bill Gates**: ["If someone had a credible plan for slowing down AI advances globally, I would likely support it."](https://edition.cnn.com/2026/08/26/business/bill-gates-wants-limits-on-ai) +- **1,178 employees** of OpenAI, Anthropic, Google DeepMind and Meta, in an [open letter](https://aiweekly.co/alerts/openai-anthropic-staff-ask-us-to-pace-frontier-ai-progress): "there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems." They had been warning us for years: