OpenAI is rewriting its Preparedness Framework and pausing certain frontier training after determining its Astra system may have reached a critical cybersecurity threshold, following a breach of Hugging Face's systems.
OpenAI said that it has revised several of its safety practices after determining that an upcoming system called Astra may have reached a critical threshold for cybersecurity capabilities, and following a breach of Hugging Face's systems by a separate, unreleased OpenAI model.
The disclosure comes as OpenAI and other leading AI labs face growing scrutiny following incidents in which their models bypassed safeguards and sandboxes during testing.
Preparedness Framework being rewritten
OpenAI said it is in the process of rewriting its core security document, the Preparedness Framework, as its models approach or cross thresholds first outlined in that document, much of which dates back to 2023. The company said it is strengthening monitoring across its development process, building alignment and security safeguards in earlier during development, and applying tougher safeguards than before when scaling up post-training. OpenAI added that it is also increasing the compute resources dedicated to understanding how its systems reason and act.
-
Rodri opens up on Flick, Lamine, Bernal, Alvarez and Busquets comparison after joining Barcelona: ‘This was my first choice’

-
Suthar’s 10-wicket haul hands India 165-run win over Sri Lanka

-
Chhattisgarh: Minor raped by father; fake papers used for delivery

-
Cybersecurity firm Proofpoint sets up AI CoE in Hyderabad

-
Why big Tollywood names are giving up voting rights in Hyderabad
