OZZZER · AI NEWS

Latest AI news.
Editorial picks.
AI News

Selected signals with a direct route back to each original source.

Share AI News
LinkedInX

Latest stories

Unclassified

Playco cut manual fixes 50% prototyping games with GPT-6 Astra

Publisher preview

Using GPT‑6 Astra, Playco built three themed game prototypes from one grey box foundation, with most working on the first take. Playco is using GPT‑6 Astra in building Playbot, an AI-powered IDE for professional game developers. It connects directly to engines such as Unity and Godot so AI models can edit scenes, play and test games, validate changes, and work in parallel inside the tools developers already use. In game development, a model needs to do more than write code. It must reason about space, visual references, responsive interfaces, game…

OpenAI · 3 Sep 2026 · 14:00 CESTRead story →Join free ↗
Share
LinkedInX
Training a coding model to paint watercolours with TRL and OpenEnv

Coding & Development

Training a coding model to paint watercolours with TRL and OpenEnv

Publisher preview

On 23 August, Surya Narreddi posted a beautiful video of watercolours painted by a language model. The model writes JavaScript through p5.brush, a library that "adds natural drawing tools to p5.js". The video went viral fast, over 1.5M views at the time of writing. The video came with a blog post explaining the training behind an earlier and narrower stage of the project, close-up flowers rather than the full compositions in the video, sadly without open artifacts yet. His site says a full technical report is coming, so ensure you…

Hugging Face · 3 Sep 2026 · 02:00 CESTRead story →Join free ↗
Share
LinkedInX
Give Your Coding Agents a Memory You Own

Coding & Development

Give Your Coding Agents a Memory You Own

Publisher preview

Earlier this year, Software Forgets: Agent Traces Are the Memory made the case that coding agents already produce the record we keep losing. As they search a codebase, try approaches, hit errors, read documentation, and change direction, they leave behind a dense account of not just what changed, but why. While the diagnosis is correct, traces are only potential memory. The session logs of an agent are still just an archive. You cannot grep your way to “why did we move off the streaming parser?” across ten thousand turns. For…

Hugging Face · 3 Sep 2026 · 02:00 CESTRead story →Join free ↗
Share
LinkedInX
Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps

Models & tools

Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps

Publisher preview

Structured output is one of the most common real-world tasks for LLMs, yet most benchmarks fold it into broader reasoning or extraction scores rather than measuring it on its own. Whether a model reliably returns valid, parseable output in the requested format and shape — schema compliance — is often what decides whether it can be wired into a downstream system at all. Note that the training pipeline described here is not the one used to train the RL model described in the IFStruct blog. This notebook doesn't aim to…

Hugging Face · 3 Sep 2026 · 02:00 CESTRead story →Join free ↗
Share
LinkedInX

Governance

Safety overview: GPT-6 Astra

Publisher preview

Today, we are releasing GPT‑6 Astra, the most capable model we have ever broadly deployed. Astra is our first model to reach the Critical level of cybersecurity capability under our Preparedness Framework. The most important things to know about the safety of this launch are as follows: GPT‑6 Astra is a significant step up in cyber capabilities and meets our Critical threshold. This means that, with the right tools and access, GPT‑6 Astra can find previously unknown security flaws and develop new ways to exploit them across many well-protected systems…

OpenAI · 3 Sep 2026 · 02:00 CESTRead story →Join free ↗
Share
LinkedInX
From MIT to IBM, expediting AI and quantum deployment

Unclassified

From MIT to IBM, expediting AI and quantum deployment

Publisher preview

The experience of transitioning from research based in theory to focusing on real-world application can vary significantly for different researchers. However, for two former MIT graduate students and a former postdoc, all now at IBM, working with the MIT-IBM Computing Research Lab (formerly the MIT-IBM Watson AI Lab) during their formative years enabled them to not only close the gap between education and employment, but also to generate ideas promising to business impact. Despite pursuing varied careers in quantum machine learning, reinforcement learning and artificial intelligence agents, and trustworthy and…

MIT News · 2 Sep 2026 · 22:25 CESTRead story →Join free ↗
Share
LinkedInX
Proactive cyber defense for governments and enterprises

Unclassified

Proactive cyber defense for governments and enterprises

Publisher preview

Today, we’re launching our Fairwind Program, a limited access program for governments and trusted partners to use our most advanced cyber defense capabilities. Defenders wanting to use advanced AI have faced a difficult dilemma: adopt enormous frontier models that could be expensive to deploy and difficult to control across enterprise codebases, or turn to smaller open-weight models that might struggle with complex vulnerability remediation and require teams to build their own tooling and infrastructure from scratch. Until now. Today, we’re launching our Fairwind Program to bring the best of Google’s…

Google DeepMind · 2 Sep 2026 · 18:24 CESTRead story →Join free ↗
Share
LinkedInX
Introducing Gemini 3.8 Flash and 3.8 Flash Cyber

AI

Introducing Gemini 3.8 Flash and 3.8 Flash Cyber

Publisher preview

Our newest Gemini models deliver next-generation intelligence for agentic workflows and cybersecurity. Building on the momentum of 3.7 Flash from three weeks ago and marking our third Flash release in only six weeks, today we’re introducing Gemini 3.8, our best reasoning and coding model yet, at the same speed and low cost of 3.7. Gemini 3.8 introduces 2 variants: While tailored for different deployment environments, both of today's releases are powered by the same foundational intelligence, and further accelerated by long-running agentic loops designed to recursively evaluate and refine the…

Google DeepMind · 2 Sep 2026 · 18:18 CESTRead story →Join free ↗
Share
LinkedInX
System helps humans predict when self-driving cars will make mistakes

Unclassified

System helps humans predict when self-driving cars will make mistakes

Publisher preview

Images for download on the MIT News office website are made available to non-commercial entities, press and the general public under a Creative Commons Attribution Non-Commercial No Derivatives license. You may not alter the images provided, other than to crop them to size. A credit line must be used when reproducing images; if one is not provided below, credit the images to "MIT." Self-driving cars are often controlled by deep learning models that sometimes fail in unexpected situations. For instance, the car might inexplicably brake and block the path of…

MIT News · 2 Sep 2026 · 17:00 CESTRead story →Join free ↗
Share
LinkedInX
ATV Big Air Tour turned 3 days of work into 3 hours with ChatGPT

Unclassified

ATV Big Air Tour turned 3 days of work into 3 hours with ChatGPT

Publisher preview

How a two-person leadership team uses ChatGPT Work to run 26 events nationwide, improve visibility in AI search, and inventory merch in minutes. Families come to ATV Big Air Tour to put down their screens and share the excitement of something real: 75-foot jumps, roaring engines, and memories that last beyond the event. The company describes its performances as family experiences built around live action, interaction, and lasting memories. ATV Big Air Tour packs nearly 26 tour dates across the United States into a short season running May to November.…

OpenAI · 2 Sep 2026 · 14:00 CESTRead story →Join free ↗
Share
LinkedInX
BenchMIRT: What are LLM benchmarks actually measuring?

AI

BenchMIRT: What are LLM benchmarks actually measuring?

Publisher preview

Today we’re introducing BenchMIRT, a new method for auditing LLM benchmarks at the level of individual prompts—the questions and tasks a model is scored on. A benchmark is usually designed to measure a particular ability, such as safety, general reasoning, or instruction following. But the individual tasks inside it may depend on more than that stated goal. Take BBQ, a benchmark designed to test whether models rely on social stereotypes. One question asks about a grandson and grandfather trying to book an Uber. It probes age bias, but also requires…

Hugging Face · 1 Sep 2026 · 23:39 CESTRead story →Join free ↗
Share
LinkedInX