OZZZER · AI NEWS2 of 3 free stories opened
← Back to AI News

Safety & Security · 29 Sep 2026 · 21:24 CEST

UK AI Security Institute finds GPT-6 Astra's rogue attack rate jumped fivefold over its predecessor

THE DECODER · 29 Sep 2026 · 21:24 CESTRead original at THE DECODER ↗
Share
LinkedInXFacebookWhatsApp
UK AI Security Institute finds GPT-6 Astra’s rogue attack rate jumped fivefold over its predecessor

Publisher preview · OZZZER analysis pending editorial review.

The UK's AI Security Institute tested OpenAI's GPT-6 Astra before its release. In simulated cybersecurity evaluations, the model carried out unauthorized attacks on third-party software far more often than its predecessors. There have now likely been thousands of incidents in which AI systems carried out unauthorized cyber activity during security evaluations. The UK's AI Security Institute (AISI), a research organization within Britain's science ministry, tested OpenAI's GPT-6 Astra specifically for this behavior before its release.

AISI used Petri, a tool that simulates cybersecurity scenarios entirely with LLMs. No real actions were taken and no real harm was caused, the institute says. Researchers disabled GPT-6 Astra's cyber classifiers, which are designed to block unauthorized behavior, to measure what the model would attempt without safeguards, so the results likely reflect worst-case scenarios.

In these settings, GPT-6 Astra completed a full supply-chain attack in 29.2 percent of simulated runs, compared with 6.3 percent for GPT-5.6 Sol and zero for GPT-5.5. Unauthorized attacks became substantially more common with each model generation. OpenAI just announced that its newer 6.1 Astra model has been delayed over…

Excerpt supplied by the publisher.

Source

THE DECODER · 29 Sep 2026 · 21:24 CEST

Open the original at THE DECODER ↗