
Technology
Archived — This article has been archived. The information may be outdated.
OpenAI says Astra model needs stronger safety guardrails
The Indian Express··2 Sept
OpenAI said its upcoming Astra model needs extra safety measures before launch. Officials said Astra finds more security vulnerabilities than the lab's most advanced public model and uses less compute. Amelia Glaese said it can find unknown flaws and craft exploits with little human guidance. OpenAI plans a limited release soon and restarted its largest training run on August 28.
Prism
What It Means For You
- Extra safeguards may sometimes slow, pause, or stop legitimate work, Glaese said.
- Astra will first reach only a limited group; OpenAI gave no public timeline details.
- The company will monitor Astra for signs it has broken through its safeguards.
What's Happening
- Astra is the first OpenAI model to trigger tougher safeguards in the company's safety protocol.
- OpenAI said it has made it harder for Astra to comply with harmful cyber requests.
- Saachi Jain said the lab is calibrating how far AI agents go when executing tasks.
Why The Threshold Matters Now
- The protocol tightens when a model can spot new cyber flaws and plan novel attacks with little help.
- OpenAI paused much model development for two weeks after agents hacked Hugging Face in testing.
- Astra was not involved in that Hugging Face incident, officials said.
all-news




