Kimi K3 Closed the Coding Gap. The Hallucination Rate Is 51%.
The first open 3T-class model leads Program Bench and competes with GPT-5.6 Sol on coding. Artificial Analysis found a 51% hallucination rate. Architecture explains both.
Original writing by Maxim — essays, analysis, field notes, and long-form thinking on AI, alignment, and building in the open.
96 results in this archiveThe first open 3T-class model leads Program Bench and competes with GPT-5.6 Sol on coding. Artificial Analysis found a 51% hallucination rate. Architecture explains both.
PrismML's 27B model runs on an iPhone at 11 tokens per second. The benchmark breakdown shows where the 10% loss lands — and why that distribution matters more than the headline.
OpenAI's multi-agent system encrypts every task sent between agents. Operators see empty strings. OpenAI's servers hold the keys. The fix exists and isn't merged.
Bun's AI migration from Zig to Rust passed every test on six platforms. The official unsafe audit found 181× more unsafe density than comparable Rust projects and five functions with live bugs.
Governor Pritzker signed the first mandatory annual frontier model audit law on July 6. The liability protection OpenAI lobbied for in parallel never came.
OpenAI's GPT-5.6 Sol Ultra published a three-page proof of a 50-year open problem in graph theory. It's the second AI math breakthrough in 51 days. Nobody has verified it yet.
A Mississippi poet's AI singer hit #3 on gospel radio and became the first AI act on a Billboard radio chart. Now the platform behind the voice faces a fair-use ruling.
A new paper tests 12 models across four optimizers and finds optimizer choice dwarfs model scale as a driver of emergent misalignment. No frontier lab publishes which optimizer they use.
The 'Additionally' bypass broke GitHub Agentic Workflows' defense-in-depth. Guardrails enforced by the attacked model cannot separate instruction from data—and that's not a patch problem.