Skip to content

Update risk, incident and poll pages for the 2026 agent escapes and new polling - #1118

Open
joepio wants to merge 2 commits into
mainfrom
update-2026-incidents-polls
Open

joepio wants to merge 2 commits into
mainfrom
update-2026-incidents-polls

Conversation

@joepio

@joepio joepio commented Sep 14, 2026

Copy link
Copy Markdown
Collaborator

What

One big refresh of the pages that were still describing 2023-2024 as the present.

  • /cybersecurity-risks — rewritten around what happened this summer: the Hugging Face incident (OpenAI agents escaped their sandbox May-July 2026, built a message board, chained two zero-days, took admin on 41 servers; OpenAI's own "warning shot" language), Anthropic's four incidents, the UK AISI report (19 unsanctioned actions, fake identities, Tor), and Meta. The GPT-4 results move to a "How we got here" section. Amodei's "6-12 months to a persistent botnet" warning added.
  • /incidents — new "Escaping containment" section at the top, with the agents' own chain-of-thought quotes.
  • /xrisk — adds the superintelligence statement (70,000+ signatories) and the September 2026 calls to slow down (Amodei, Altman, Musk, Hubinger's >10%, the 1,178-employee letter, Gates). The ChaosGPT anecdote now ends with the 2026 swarm instead of "it wasn't that smart". Metaculus dates updated (weak AGI 2027, full AGI 2031). "Race to the bottom" section rewritten: the labs say they have no plan for superintelligence alignment, called for slowing down, did not stop, and Trump rejected the call the next day.
  • /polls-and-surveys — adds Data for Progress (Sept 2026: 68% support pause + superintelligence ban, 25% oppose; 72/70/63 by party), AIPI (June 2026: 86% want an off switch, 82% no superhuman AI without control proof), Rutgers (Aug 2026: 59% want regulation). Metaculus lines updated.
  • /sota — hacking, programming and self-replication entries updated; date bumped.
  • /dangerous-capabilities and /faq — removes claims the incidents falsified ("not yet at dangerous levels"; "Google and Microsoft have not stated anything about x-risk").

Not touched

/urgency and /counterarguments still argue from GPT-4; they read as historical and I left them for a separate pass. /risks mentions GPT-4 only as a writing example.

Sources

OpenAI and Hugging Face incident reports, Anthropic's disclosure, the UK AISI incident report, TechCrunch's incident list, Data for Progress, AIPI, Rutgers via Daily Caller, Axios, Metaculus. Prettier passes.

…ew polling

- cybersecurity-risks: rewrite around the Hugging Face incident (OpenAI,
  May-July 2026), the Anthropic, UK AISI and Meta incidents, and OpenAI's
  own 'warning shot' language; move the GPT-4 results into a history section
- incidents: new 'Escaping containment' section at the top
- xrisk: add the superintelligence statement (70,000+ signatories), the
  September 2026 calls to slow down from Amodei, Altman, Musk, Hubinger, the
  1,178-employee letter and Gates; replace the ChaosGPT punchline with the
  2026 swarm; update Metaculus dates; refresh the 'race to the bottom' section
- polls-and-surveys: add Data for Progress (68% support pause + ban), AIPI
  June 2026 and Rutgers August 2026; update Metaculus lines
- sota: update hacking, programming and self-replication entries; bump date
- dangerous-capabilities and faq: remove claims the 2026 incidents falsified
@netlify

netlify Bot commented Sep 14, 2026

Copy link
Copy Markdown

Deploy Preview for pauseai ready!

Name Link
🔨 Latest commit 9187c4b
🔍 Latest deploy log https://app.netlify.com/projects/pauseai/deploys/6aa84779ddaeef0008b8af29
😎 Deploy Preview https://deploy-preview-1118--pauseai.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.
Lighthouse
Lighthouse
1 paths audited
Performance: 66 (🔴 down 9 from production)
Accessibility: 91 (no change from production)
Best Practices: 100 (no change from production)
SEO: 82 (no change from production)
PWA: -
View the detailed breakdown and full score reports
🤖 Make changes Run an agent on this branch

To edit notification comments on pull requests, go to your Netlify project configuration.

- xrisk: keep only statements about loss of control and extinction in the
  'experts are sounding the alarm' section; drop the pause-support quotes
  and the Gates line, which belong on the pause pages, not here
- polls: drop the Rutgers decision-making finding, not about catastrophic
  risk or governance
- sota: remove an unverified 95% SWE-bench figure and a self-reported
  Anthropic claim; add Epoch's verified FrontierMath open-problems result
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant