Reason iconReasonSep 22, 2026 ~7 min source read

How the AI 'safety' movement may be reducing safety by concentrating power and hiding risk

Arguments about catastrophic AI risk and a specific belief cluster inside the AI community are driving policies and product choices that make the technology less transparent, more consolidated, and harder for the public to monitor or defend against.

The 'AI Safety' Movement Is Making AI Less Safe

Share this story

Send the public story page.

Useful takeaways from this story.

A set of beliefs labeled TESCREAL — including effective altruism and fears of an imminent superintelligence — is shaping corporate and regulatory choices about AI.

Calls to slow or centralize AI research risk consolidating capabilities in a few frontier labs, reducing transparency and public ability to test or mitigate harms.

Some safety-driven design choices and technical measures can make systems harder to inspect, increasing the chance of unexpected behavior or misuse.

The useful part

Employees at Anthropic's competitor, OpenAI, have been voicing similar sentiments. Many AI frontier researchers and "AI doomers" alike are animated by an almost religious belief that "artificial superintelligence" will soon supplant humanity. Enforced through tyrannical means, a ban on uncontrolled research would deprive the public of both an understanding of the state of AI and the technology needed to defend against its harmful uses.

How it works

  • For all the speculation that China wants an AI-powered police state, that country's AI industry is much more focused on open-source research and practical industrial applications.
  • Spurred on by Amodei himself, the industry and politicians have decided to try to " pace the frontier " of AI research.
  • It has already negotiated an AI safety testing framework with the frontier labs that would keep information about the process and results secret.
  • If more knowledge about how to create intelligence can lead to a runaway superintelligence, then the only way to stop the process is to hide the knowledge.
  • Concentrating everything behind closed doors in just a few organizations doesn't work.

What to take from it

Because we defended ourselves with an open model," he told CBS, explaining that HuggingFace used a remixed American version of a Chinese open-source AI model for cybersecurity. During the recent war with Iran, military planners killed 123 children at a school and nearly attacked a Chinese ship falsely accused of carrying nuclear weapons parts, both in part because of AI tools. These disasters came not because of a "misaligned superintelligence," but because human beings were relying too much on a system that they believed was so and whose behavior they did not understand.

Example or evidence

  • Will " recursive self-improvement " (the AI building itself) lead to an uncontrollable "intelligence explosion"?
  • While the frontier AI labs chase government contracts that would put their software in control of critical systems, some researchers at the same labs are treating that software as more human, pushing it to...
  • The second "E" in TESCREAL stands for "effective altruism," a philosophical movement popular among computer nerds over the past decade.
  • Anthropic's Dario Amodei and the leadership of OpenAI were both heavily influenced by effective altruism.

Details worth keeping

After Anthropic employee Jacob Coxon resigned over fears that its product could destroy humanity itself earlier this month, the company's alignment science lead, Evan Hubinger, publicly agreed that his product has a 10 percent chance of causing human extinction in the next decade. (So was infamous cryptocurrency scammer Sam Bankman-Fried.) Many members of this movement have become convinced that the most important question of the 21st century is keeping AI "aligned" with human values. Perhaps the clearest effective altruist vision is laid out in AI 2027, a pair of future scenarios written by several big names in the "AI safety" community.

Related coverage

  • Fastcompany: Some investors, analysts, and industry watchers believe the big AI labs, especially Anthropic, are taking advantage of the current AI safety scare by inviting regulations that only the largest labs can...
  • Theverge: Remember when tech leaders would tell their employees to "move fast and break things"?

More context around this story.

Making AI safer makes it more dangerous
Boingboing iconBoingboingSep 18, 2026

Making AI safer makes it more dangerous

Watermarking was supposed to make AI easier to identify; instead, it has made it easier for AI to jump the rails and escape its little box. The system du jour intended to make AI safer and its use more accountable can apparently make it less safe. — Read the rest The post Making AI safer makes it more dangerous appeare

The AI Superintelligence Slowdown
Theverge iconThevergeSep 17, 2026

The AI Superintelligence Slowdown

Remember when tech leaders would tell their employees to “move fast and break things”? It seemed that would be the way of AI too. But after a summer where rogue AI agents became reality, and researchers warned that AI could kill us all, a number of leading US AI companies are publicly suggesting it’s time […]

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app