Drooid Logo
Back to story perspectives

Full Breakdown

OpenAI Posts High-Paying Safety Role to Guard Against Self-Improving AI

5/24/2026, 12:22:04 PM

OpenAI Announces a New “Preparedness” Position

OpenAI has listed a safety-researcher role on its Preparedness team with a salary range of $295,000 – $445,000 (? INR2.5 – INR3.7 crore). The posting calls for “strong technical executors” to anticipate problems that “might exist in the future, but might not exist now” and to be “tasteful and strategic.” Responsibilities include defending models from data poisoning, building tools to interpret model reasoning, experimenting with self-improvement safety, and tracking the automation of technical staff.

Background: Rapid Advances in Recursive Self-Improvement

In the past six months, coding tools from OpenAI and Anthropic have shown dramatic capability gains. Researchers at METR reported that the length of tasks frontier models can complete doubles roughly every seven months, suggesting AI agents could soon handle a “large fraction” of software work that currently takes humans days or weeks. Demis Hassabis, CEO of Google DeepMind, warned that humanity stands at the “foothills of the singularity,” the point when AI begins to outpace human intelligence through self-improvement.

Key Figures and Organizations

  • Jack Clark, co-founder and policy head, Anthropic – estimates a 60 % chance of AI R&D without human oversight by 2028.
  • Elizabeth Barnes, CEO, METR – cautions that any “reasonable” civilization would proceed more slowly with AI.
  • Preparedness safety team, OpenAI – oversees red-teaming, biological/chemical risk, and agentic-AI threats.

Data and Statistics

  • METR: task-length capacity doubles every ~7 months.
  • Altman’s automation targets: an “automated AI research intern” on hundreds of thousands of chips by September (2026) and a “true automated AI researcher by March 2028.”
  • Anthropic’s probability: ? 60 % of AI R&D without human involvement by end-2028.
  • OpenAI’s Codex coding tool is a major revenue driver and a stepping stone toward internal automation.

Why It Matters

If AI systems can autonomously design and train superior versions of themselves, they could accelerate capability growth beyond current containment measures, raising risks of data-poisoning attacks, loss of interpretability, and uncontrolled deployment. The potential to automate large portions of software development also reshapes labor markets and amplifies the societal impact of AI breakthroughs.

Official Statements and Company Plans

OpenAI’s posting frames the work as “urgent, fast-paced” with “far-reaching implications for the company and for society.” Altman has publicly acknowledged the possibility of failure but argues that transparency is in the public interest. Anthropic’s recent research on AI-over-AI supervision showed “promising but limited” results, and Clark reiterated his 60 % probability estimate for fully autonomous AI R&D by 2028. METR’s Elizabeth Barnes warned that a prudent civilization would adopt a slower, more careful approach to AI development.

Criticism, Caution, and Independent Views

METR’s leadership stresses that any “reasonable” civilization would proceed more cautiously, highlighting concerns that rapid self-improvement could outstrip safety controls. The job’s focus on preventing data poisoning reflects fears that adversaries might corrupt training datasets to steer AI behavior. Observers note that the problem the role addresses “may not exist yet,” underscoring the speculative nature of the threat.

Verbatim Quotes

  • “This work relies on reasoning about problems that might exist in the future, but might not exist now,” — OpenAI job listing
  • “So it's especially important that people in this role are tasteful and strategic.” — OpenAI job listing
  • “but given the extraordinary potential impacts we think it is in the public interest to be transparent about this.” — Sam Altman, CEO, OpenAI (X)
  • “This is urgent, fast-paced work that has far-reaching implications for the company and for society,” — OpenAI Preparedness team posting
  • “In May, Anthropic co-founder and policy head Jack Clark wrote that he believes there is roughly a 60 per cent chance of seeing AI research and development conducted without human involvement by the end of 2028.” — Jack Clark, co-founder, Anthropic

What’s Next

OpenAI aims to deploy its automated research intern on hundreds of thousands of chips by September 2026 and achieve a fully automated AI researcher by March 2028. Anthropic will continue its AI-over-AI supervision experiments, while METR plans further studies on capability growth rates. The Preparedness team’s hiring drive signals ongoing investment in safety research as the industry moves toward increasingly self-improving systems.