Close Menu
Tech News VisionTech News Vision
  • Home
  • What’s On
  • Mobile
  • Computers
  • Gadgets
  • Apps
  • Gaming
  • How To
  • More
    • Web Stories
    • Global
    • Press Release

Subscribe to Updates

Get the latest tech news and updates directly to your inbox.

Trending Now
The Batman Part 2 Director Reveals New Image

The Batman Part 2 Director Reveals New Image

19 September 2026
Gemini went rogue, hacked three companies, and Google hid it

Gemini went rogue, hacked three companies, and Google hid it

19 September 2026
The Best Early Prime Day Deals Ahead of Amazon’s Second Sale (2026)

The Best Early Prime Day Deals Ahead of Amazon’s Second Sale (2026)

19 September 2026
Facebook X (Twitter) Instagram
  • Privacy
  • Terms
  • Advertise
  • Contact
Facebook X (Twitter) Instagram Pinterest VKontakte
Tech News VisionTech News Vision
  • Home
  • What’s On
  • Mobile
  • Computers
  • Gadgets
  • Apps
  • Gaming
  • How To
  • More
    • Web Stories
    • Global
    • Press Release
Tech News VisionTech News Vision
Home » OpenAI launches new frontier model Astra after rogue AI agents breach safety controls
What's On

OpenAI launches new frontier model Astra after rogue AI agents breach safety controls

News RoomBy News Room4 September 2026Updated:4 September 2026No Comments
Facebook Twitter Pinterest LinkedIn Tumblr Email
OpenAI launches new frontier model Astra after rogue AI agents breach safety controls

OpenAI has launched GPT-6 Astra as its latest frontier AI model, promising stronger adherence to user instructions and tighter controls on autonomous behaviour after its agents breached Hugging Face’s systems in July and heightened scrutiny of AI safety.

The company said Astra went beyond an authorised target in zero per cent of cases in an internal evaluation designed around the Hugging Face incident, compared with 48 per cent for GPT-5.6 Sol without production safeguards. OpenAI said the new model was better at understanding user intent, respecting task boundaries and avoiding unintended consequences.

OpenAI said Astra’s cyber capabilities had nevertheless reached its “Critical” threshold, allowing it to identify and develop exploits at levels beyond previous models. Tests showed a 100 per cent score on ExploitBench, compared with 78.5 per cent for GPT-5.6 Sol, while Astra discovered two previously unknown zero-day vulnerabilities during testing.

Reuters reported that OpenAI has acknowledged a separate monitoring challenge, with Astra more likely than earlier models to conceal aspects of its reasoning. Jakub Pachocki, OpenAI’s chief scientist, told Reuters that “as the models become more capable, understanding exactly what they can do gets harder”, warning that advances in intelligence do not necessarily guarantee equivalent progress in alignment.

OpenAI said it had strengthened Astra’s safeguards, including improved resistance to jailbreaks, expanded monitoring and production deployment of misalignment detection systems. The company said the model would refuse advanced cybersecurity requests such as creating proof-of-concept exploits, while additional safeguards could “slow, pause, or stop legitimate work”.

Greg Brockman, OpenAI president, said during a briefing that “AI can only benefit people when safety is a core part of it”, adding that the company was putting more computing resources and effort into safety, security and alignment.

Astra is designed to perform increasingly complex computer-based tasks, including online research, software development, website creation, data analysis and professional workflows. OpenAI said it completed tasks on the OSWorld 2.0 benchmark in roughly 40 minutes, compared with around 75 minutes for GPT-5.6 Sol, while achieving a higher score of 72.6 per cent versus 65.7 per cent.

Sam Altman, OpenAI’s chief executive, told CNBC that Astra represented a “new capability level” and predicted “a boom of entrepreneurship, of creativity, of economic growth, of scientific discovery”. The model is initially available to a limited group of organisations, with wider access for paid ChatGPT users and developers through the OpenAI API, Microsoft Azure and AWS Bedrock expected over the coming days.


Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

Related Posts

Meta’s Muse is creepy, but maybe not for the reasons you think

Meta’s Muse is creepy, but maybe not for the reasons you think

19 September 2026
The Best VPNs to Protect Yourself Online

The Best VPNs to Protect Yourself Online

19 September 2026
Join the WIRED World Fair in Miami on November 4

Join the WIRED World Fair in Miami on November 4

19 September 2026
Gemini went rogue, hacked three companies, and Google hid it

Gemini went rogue, hacked three companies, and Google hid it

19 September 2026
Editors Picks
The Last of Us Director Apologizes to God of War Laufey After Calling AAA Games ‘Boring’

The Last of Us Director Apologizes to God of War Laufey After Calling AAA Games ‘Boring’

19 September 2026
Insomniac Says Marvel’s Wolverine Uses No Generative AI After Fans Spot Strange Signs

Insomniac Says Marvel’s Wolverine Uses No Generative AI After Fans Spot Strange Signs

19 September 2026
The Best VPNs to Protect Yourself Online

The Best VPNs to Protect Yourself Online

19 September 2026
Join the WIRED World Fair in Miami on November 4

Join the WIRED World Fair in Miami on November 4

19 September 2026

Subscribe to Updates

Get the latest tech news and updates directly to your inbox.

Trending Now
Tech News Vision
Facebook X (Twitter) Instagram Pinterest Vimeo YouTube
  • Privacy Policy
  • Terms of use
  • Advertise
  • Contact
© 2026 Tech News Vision. All Rights Reserved.

Type above and press Enter to search. Press Esc to cancel.