Loading IndicatorLoading Indicator

OpenAI Classifies New Astra Model as ‘Critical’ Cyber Risk, Limits Access

Source
Korea Economic Daily

Summary

  • OpenAI said it will release its next model, 'Astra,' with some features restricted after the system reached the 'critical' capability threshold in cybersecurity.
  • Astra scored 100% on a benchmark for developing exploit code for known vulnerabilities, and OpenAI said the model also identified previously unknown flaws on its own and used them in attacks during high-risk vulnerability tests.
  • Anthropic said it will separate access rights for 'Claude Fable 5.1' and 'Claude Mythos 5.1,' with Mythos 5.1 — which has stronger cybersecurity and bioscience capabilities — available only to a limited number of vetted institutions.

Forecast Trend Report by Period

Loading IndicatorLoading Indicator

OpenAI Rates New Model’s Security Capability as ‘Critical’

AI Can Find and Exploit Flaws Without Human Guidance

Anthropic Also Restricts Access to New Models

Photo: Shutterstock
Photo: Shutterstock

Advanced artificial intelligence is reaching a point where it can identify and attack previously unknown software vulnerabilities without human instructions.

OpenAI and Anthropic have decided to release their latest AI models with some functions restricted after their cyberattack capabilities rose sharply. The AI race is moving beyond a simple push for better performance into a contest over how widely high-risk capabilities should be released and who should be allowed to use them.

On Sept. 1, OpenAI said its next AI model, Astra, had reached the “critical” capability threshold in cybersecurity under the company’s internal risk evaluation system, the highest level in that category. The designation means that, if given the proper tools and access privileges, the model can find unknown vulnerabilities in hardened systems and develop ways to exploit them without step-by-step human direction.

It is the first time OpenAI has classified one of its models at that level. Amelia Glaze, OpenAI’s vice president for safety, said Astra can identify unknown security flaws and develop methods to exploit them without a person directing every stage, provided it has the necessary tools and access.

Astra scored 100% on a benchmark measuring a model’s ability to develop exploit code for known vulnerabilities. In a separate OpenAI evaluation using 20 recently disclosed high-risk vulnerabilities, Astra found two previously unknown flaws on its own and used them in the attack process. OpenAI plans to make Astra’s most powerful cyber capabilities available only to a limited group of test users when the model is released.

Anthropic also split access rights for Claude Fable 5.1 and Claude Mythos 5.1, unveiled the same day. The two products are built on the same base model but differ in the strength of their safeguards. Fable 5.1, which will be available to general users, can identify vulnerabilities in source code, but its capabilities in penetration testing, exploit-code generation and binary-based vulnerability discovery have been limited. Mythos 5.1, which has stronger cybersecurity and bioscience capabilities, will be offered only to a small number of institutions that pass verification and receive Anthropic’s approval.

The move by OpenAI and Anthropic, the world’s top two players in the field, underscores the speed of AI development. In the past, companies typically released more capable models and then blocked harmful prompts afterward. Now that AI can execute attack plans, safeguards are shifting toward screening which users can access the highest-risk capabilities themselves.

Lee Young-ae, Hankyung.com reporter 0ae@hankyung.com

#AI Safety
#Cybersecurity
Korea Economic Daily

Korea Economic Daily

hankyung@bloomingbit.ioThe Korea Economic Daily Global is a digital media where latest news on Korean companies, industries, and financial markets.

What do you think about this news?








PiCK News






Hashtag News