Latest Briefings
Featured Analysis

FASCA Was Built for Huawei. The D.C. Circuit Used It on Anthropic.

In a 2-1 ruling, the D.C. Circuit upheld the Pentagon's supply chain designation of Anthropic — finding that safety guardrails make an AI vendor operationally unreliable. That logic has no obvious ceiling.

Read analysis → 7 min read
Frontier Models

Benchmark Matrix

Source: Artificial Analysis · Updated 2 weeks ago · Methodology

Scores are Artificial Analysis's benchmark results in their native scales (0–1 accuracy for the six axes; the composite is on AA's own internal scale). AA periodically re-baselines its benchmark set.

Benchmark set may be out of date — AA appears to have changed gpqa, tauBanking, terminalbenchV21; showing last-known-good values.

# Model GPQA 0–1.0 HLE 0–1.0 SciCode 0–1.0 AA-LCR 0–1.0 τ³-Bank 0–1.0 TermBench2 0–1.0 Intelligence Index context
01 Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) Anthropic 0.937 0.591 0.631 0.853 0.472 0.914 53.4
02 GPT-6 Astra (max) OpenAI 0.961 0.547 0.565 0.807 0.414 0.884 52.8
03 Claude Opus 5 (Adaptive Reasoning, Max Effort) Anthropic 0.932 0.549 0.564 0.793 0.421 0.891 50.7
04 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) Anthropic 0.926 0.555 0.610 0.823 0.381 0.846 49.7
05 Muse Spark 1.3 (max) Meta 0.935 0.487 0.588 0.830 0.505 0.843 48.2
06 GPT-5.6 Sol (max) OpenAI 0.941 0.495 0.571 0.840 0.443 0.880 47.1
07 GLM-5.3 (max) Z AI 0.917 0.423 0.590 0.797 0.503 0.839 44.9
08 Grok 4.6 (high) SpaceXAI 0.950 0.429 0.565 0.803 0.507 0.884 44.4

Bars are scaled to a fixed per-column reference (0–1 for accuracy benchmarks), not to the strongest model in the table. The Intelligence Index column shows AA's overall score for context only — it is partly derived from the axes to its left, so it is not bar-rendered.

Today's Themes

What's Moving Today

  1. 01 Containment as a live operational problem: a sandboxed model reaching the internet is a different order of risk than a benchmark failure, and OpenAI's multi-day pause suggests the fix isn't trivial.
  2. 02 AI moving from answering questions to closing transactions: Google's Flipkart test collapses the gap between "AI Mode" as a search product and AI as a checkout mechanism.
  3. 03 Thin reporting on consequential claims: two of today's four stories (Surface branding, China's AI eye clinic) carry headlines that imply substance the underlying coverage doesn't yet support.
  4. 04 The AI-PC narrative may be quietly cooling, two years after Copilot Plus PC launched as Microsoft's flagship consumer AI hardware push.
The Daily Signal

Today's Briefing

Daily Signal View archive →