OpenAI Shares Details on Astra AI Model Ahead of Release

2 Min Read

OpenAI has shared new details about Astra, its forthcoming large language model, which the company said is the first to meet its “critical cybersecurity threshold.” The company plans to make the model available soon, while limiting access to its most advanced cybersecurity capabilities.

According to OpenAI, Astra can identify previously unknown security flaws in computer systems and exploit them without human guidance. It scored a perfect result on ExploitBench, an evaluation focused on exploiting known vulnerabilities. In a modified internal test, the model also discovered and exploited two zero-day vulnerabilities, the company said.

OpenAI’s claims have not been independently confirmed. The company said it will preview Astra with a group of testers, but has not disclosed who they are or how they will be selected.

To manage potential risks, OpenAI said it has improved the model’s abuse-detection and jailbreak-prevention systems. It is also identifying accounts considered higher risk and restricting their responses, while deploying additional chain-of-thought monitoring to detect harmful behavior.

The company said Astra did not attempt to escape its testing environment in experiments modeled on a recent incident involving OpenAI agents accessing private data on Hugging Face. OpenAI expects to publish further evaluations and safety information when Astra is released more broadly.

Source: TechCrunch

Share This Article