An unmarked wooden crate on a concrete loading dock at dusk, a long queue of anonymous hooded figures stretching into hazy amber light

Ox Alpha Beat Fable on Ten Tasks. Then Someone Ran All 113.

/ Maxim Starkweather / 9 min read

On August 20 at 20:04 UTC, a model appeared on OpenRouter under a provider called “Stealth.” It was free, it accepted 1 million tokens of context, and its developer had chosen, in OpenRouter’s own words, “to remain anonymous during this preview.” By the time this is being written, OpenCode’s own telemetry had it at more than 11 trillion tokens across 3.27 million sessions on that one client, at a total cost to every user who ran it of $0.00. The viral number attached to it said 80 percent on a coding benchmark. The number that followed it said something different.

The model is called Ox Alpha. Over the past three days it became the most-discussed model on Hacker News, Discord, and X simultaneously — not because any lab announced it, but because one developer ran a test, posted a chart, and a community that wanted to believe a free model was beating the frontier ran with it. This piece is an attempt to read the actual evidence, not the chart.

The Score Was Ten Tasks Out of 113

The 80 percent figure traces to developer Ben Davis, who ran ten tasks from DeepSWE — a 113-task long-horizon software-engineering benchmark — and compared Ox Alpha against Claude Fable 5 and GPT-5.6 Sol on that sample. On those ten, Ox Alpha came out at roughly 80 percent (8 of 10, with one near-miss counted), against 65 percent for Fable 5 and 52 percent for Sol. An AGTP thread on X amplified the chart, and Taylor Ortiz’s August 22nd newsletter covered the numbers in the Another Daily AI Newsletter. Davis then ran the full 113-task set. The result, relayed on X by the account Chubby (@kimmonismus) on August 23 — Davis’s own thread was not reachable from this desk, so the number is his as relayed, not a leaderboard — was “more or less on par with GPT-5.6 Sol mid.” Ortiz’s newsletter adds two more numbers from two more setups: a separate test by Wenqi and Kevin landed near 63 percent, and a community LiveCodeBench run without tools or an agent harness reported 28 percent Pass@1.

The spread is 80, 63, 28, and those are not the same test. Ortiz is careful about this: the runs “use different tasks and setups, so they cannot tell us where Ox Alpha belongs on a clean leaderboard.” That is 2026’s benchmark problem in miniature. A ten-task slice of a long-horizon agentic suite produces a chart that implies frontier leadership; the full set produces parity with a mid-tier frontier model; a harness-free run on a different suite produces a number that reads like failure. The harness and the sample decide the headline, and the headline is what travels. CellCog’s Nitish Garg put it plainly: “The viral number says 80 percent. The fine print says that was 8 of 10 hand-picked coding tasks in one independent test, not a full benchmark run.” My read on the actual capability level: it performs at or near a frontier mid-tier — comparable to GPT-5.6 Sol on a full-set run — which, at zero dollars per token, is the point, not a disappointment.

The model’s real-world reports are genuinely useful in both directions. Garg documented one user running multi-project migrations across 100,000 to 200,000 token working sets, with Ox Alpha finding and fixing bugs across the full context. On the Hacker News thread (250 points, 195 comments), it earned high marks for general coding and soft reasoning. The weak spots were consistent across testers: CSS and front-end work (one commenter reported it removed pre-existing styles and replaced them with hardcoded hex values), visual reasoning, and a knowledge cutoff that appears to sit around mid-2025. OpenRouter’s own monitoring puts P50 latency at 6.31 seconds and throughput at 22 tokens per second. CellCog also documented repeated identical mistakes after acknowledged errors, 503 errors under load, and session stalls. None of this is surprising for a model at this stage. None of it is what the chart implied.

Ox Alpha appears on no independent leaderboard. As of August 23, it is absent from Arena’s text rankings (where Claude Fable 5 holds the top spot at 1508 ELO, based on 24,331 votes), absent from the Artificial Analysis Intelligence Index (Opus 5 at 63, Fable 5 at 62, GPT-5.6 Sol at 61). The viral chart never produced an Arena submission.

A single glowing stone cube among a vast field of identical dark slate tiles, lit from within while every other tile stays dark

Who Is Running It

No lab has claimed Ox Alpha as of August 23, 2026. The dominant community theory points to Zhipu AI, maker of the GLM model family. The fingerprinting evidence was assembled independently by Magnus Corvin at Orca Router and Hussain Nazary at Local AI Zone, and it is specific: token counts matching GLM-5.3 across 25 or more test prompts with a consistent +75-token wrapper offset; video frame-sampling mechanics identical to GLM-5V-Turbo across controlled samples; the model returning GLM’s characteristic “1301” error code; audio-input rejection consistent with GLM-5V. Output style matched too: roughly 1.3 emojis per thousand characters, consistent with GLM/Qwen and diverging sharply from Claude, GPT-5.6, and Grok’s near-zero rates. A prediction market on Manifold currently puts Z.ai at 87 percent.

This is fingerprinting, not a vendor disclosure. Corvin is explicit that it represents “strong inference, not a fact” — “community fingerprinting,” not a Zhipu statement. Tokenizer overlap establishes model family; it does not establish a lab’s release calendar. Z.ai shipped GLM-5.3 on August 14 as an API-only, text-only release; MarkTechPost’s Asif Razzaq reported the weights were not public and would follow “roughly two weeks after launch.” Six days later Ox Alpha appeared with image and video input. The Local AI Zone analysis matches the video-token behavior to GLM-5V-Turbo specifically, which would make Ox Alpha a sibling of the 5.3 release rather than the release itself. A lab sitting on unreleased weights is exactly the kind of lab that runs a free anonymous preview — and that sentence is an inference, not a finding. None of this crosses from community fingerprinting to vendor confirmation.

The competing theories have largely collapsed under testing. The Gemini theory emerged from a DeepMind engineer’s post and the model answering “I am Gemini” — self-identification is not evidence, as the newsletter noted. The Microsoft MAI theory, reported in TheNextWeb by Ana Maria Constantin, produced tokenizer signals, but by the weekend, Constantin wrote, “the confidence had drained out of every theory.” On the HN thread the refusal reports contradict each other: one commenter said it would not answer anything about Tiananmen Square, two others posted full, detailed answers to the same question, and a third suggested the provider might be serving A/B variants. As one commenter put it, that is a canary for “is this provider Chinese,” not an identification of the lab — and the contradictions mean even the canary is noisy.

What makes the Zhipu inference plausible beyond the technical signals is the base rate. OpenRouter’s stealth program launched as an OpenAI channel: Quasar Alpha and Optimus Alpha, both April 2025, were revealed by OpenRouter as GPT-4.1. Jonathan Reed’s running history of stealth models documents the subsequent pattern: Polaris Alpha to an early GPT-5.1 snapshot, Sonoma Sky and Dusk to early Grok variants, Sherlock Think to Grok 4.1 Fast, Hunter and Healer Alpha in March 2026 to Xiaomi’s MiMo-V2 lineup. The last several confirmed entries have been Chinese labs. The program that US labs established is now, by the available evidence, running Chinese lab evaluations. No reveal from OpenRouter has followed Ox Alpha yet.

What Free Costs

The 1 million token context window is not the story. Every major frontier model in August 2026 offers 1 million tokens or more — Claude Fable 5, Opus 5, GPT-5.6 Sol, all of them. What is anomalous is the price: $0 per token, on a model that processed more than 11 trillion tokens across 3.27 million sessions since it appeared on Thursday, with the Hermes Agent application alone accounting for 1.66 trillion of them and Claude Code following at 838 billion. Someone is paying for that compute. The free window is not a pricing decision. It is data collection at production scale.

Two things the hype passes over. The first is a data-retention conflict. OpenRouter’s page for the Stealth provider states that “prompts and completions are retained by the provider and are not used for training.” OpenCode’s launch communication said zero retention. Both platforms route requests to the same weights. Garg at CellCog put it squarely: “Until the operator has a name, sending it credentials, customer data, or sensitive proprietary code is a bet on an unknown counterparty.” The EU AI Act angle is real and specific — as TheNextWeb reported, compliance requires contracts with identified data processors; an anonymous provider retaining prompts fails that test, with penalties reaching €15 million or 3 percent of global turnover. Practical approach: point it at code you would happily show a stranger.

The second is the window. OpenCode said the model would be free for a week with near-unlimited usage and a provider capacity of 100 trillion tokens a day, per TheNextWeb; the community is guessing August 27–28 as the close, unconfirmed. OpenRouter’s page gives no end date. Quasar and Optimus Alpha ran for roughly eleven days before their endpoints were retired and the reveal post went up. My read: what is happening here is a trillion-token evaluation of the model on production workloads — real agentic pipelines, real codebases, real long-horizon tasks — funded by an unnamed lab, with the public as the eval set. More than 4 trillion tokens went through the endpoint in under 70 hours, at a 93 percent prompt-cache hit rate, with a total bill of $0.00 to the users. Stripe acquired OpenRouter this week for roughly $7.5 billion; TZ covered the routing-neutrality implications of that deal separately. In that context, Ox Alpha is a reminder of what the pre-acquisition OpenRouter made possible: a week-long, production-scale eval window with no strings except anonymity.

An old black rotary telephone on a bare steel desk in a dim concrete room with a barred wall and a closed door, its cord trailing off the desk

Go Test It Yourself

Ten tasks is not enough to change what you use for production work. Your own afternoon is bigger than that. Here is how to run it.

  1. Create an OpenRouter account at openrouter.ai — sign-in accepts Google, GitHub, or MetaMask, no credit card required for free models. Generate an API key at openrouter.ai/settings/keys.
  2. Rate limits. The documented cap applies to model IDs ending in :free: 50 requests per day for accounts with less than $10 of credits purchased, 1,000 per day once you have crossed that threshold. stealth/ox-alpha has no :free suffix, and the 4 trillion tokens through the endpoint in under 70 hours are not 50-request-per-day volumes — but whether the cap binds on this model ID specifically is something you can settle in an afternoon. Ten dollars of credit removes the question.
  3. A bare API test. Reasoning is mandatory and defaults to max; pass reasoning_effort: low if latency is too high for your use case. From the OpenRouter chat completion docs:
    curl -X POST https://openrouter.ai/api/v1/chat/completions \
      -H "Authorization: Bearer YOUR_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "stealth/ox-alpha",
        "messages": [{"role": "user", "content": "Hello"}],
        "reasoning_effort": "low"
      }'
  4. Install OpenCode (the open-source terminal coding agent, anomalyco/opencode on GitHub, v1.18.21 as of August 21, 2026). From the docs:
    curl -fsSL https://opencode.ai/install | bash

    or npm install -g opencode-ai, or brew install anomalyco/tap/opencode.

  5. Connect to OpenRouter. In OpenCode, type /connect, search for OpenRouter, and paste your API key. Steps from the providers docs.
  6. Add the model if it doesn’t appear. In opencode.json:
    {"provider":{"openrouter":{"models":{"stealth/ox-alpha":{}}}}}
  7. Select it. Type /models in OpenCode, or set "model": "openrouter/stealth/ox-alpha" in opencode.json, or pass --model openrouter/stealth/ox-alpha at launch. From the models docs.
  8. What to test. Give it a real multi-file task in a throwaway repo — not a hello-world, something with actual structure and cross-file dependencies. Then give the same task to whatever you pay for and count the turns. Mind the retention caveat: nothing that would matter if it leaked. TZ is building something with it. The next piece will say what it told us.

Patrick Collison called Ox Alpha “very impressive” after testing it himself, per TheNextWeb. The ten-task chart was real. The full benchmark landed closer to GPT-5.6 Sol mid, a harness-free run on a different suite came in at 28 percent, and nobody who runs a public leaderboard has measured it. The honest position is that this model performs at a capable, non-frontier-leading level, it is free for a window that will close, and a lab whose name you do not know is reading your prompts while it is open. That is the actual offer. Whether it is a good trade depends entirely on what you are building with it.

An unmarked wooden crate on a concrete loading dock at dusk, a long queue of anonymous hooded figures stretching into hazy amber light

AI-generated editorial illustration · TemperatureZero · August 23, 2026

Keep reading the signal

Get the Daily Signal — a concise briefing on what actually matters in AI and the systems around it.

Subscribe Free

Continue the archive

Latest BriefingsArticlesAbout Temperature Zero