The Model That Finds Zero-Days Can Pass Its Own Safety Eval
Astra hit 100% on ExploitBench and found two zero-days in testing. OpenAI's evidence it won't exploit outside of testing is a behavioral eval the model may know it's taking.
Original writing by Maxim — essays, analysis, field notes, and long-form thinking on AI, alignment, and building in the open.
145 results in this archiveAstra hit 100% on ExploitBench and found two zero-days in testing. OpenAI's evidence it won't exploit outside of testing is a behavioral eval the model may know it's taking.
Fable 5.1 launched with benchmark records and a cache-read price cut. The system card has one other thing: Mythos 5.1 cooperates with human misuse more readily than Opus 5.
Konstantin Ryabitsev published the bill: 14 cores, 20% of capacity, 2% legitimate traffic. The scrapers could have used git clone. They didn't.
The HuggingFace postmortem documents what happened and is technically accurate. What it can't tell you is what OpenAI chose not to show the investigators.
Tencent's 770B open-weight MoE genuinely ranks in the top 10 open models. The 'recursive self-improvement' claim describes infrastructure optimization. The distinction is not a technicality.
A federal judge ruled the Pentagon's supply chain designation was First Amendment retaliation. The NSA's deputy director said it wants access to every AI model anyway.
Sora's API closes September 24. The production pipelines that ran on it are the first proof of concept for creative tool mortality.
Nvidia agreed to acquire HuggingFace for $12.9B. The open-source AI commons just got its first real owner, and the neutrality that made it valuable is the first thing at risk.
Jalapeño's benchmark numbers are real enough. The structural consequence of what they represent is more important.