A vast plateau of cracked golden stone ending at a sheer cliff edge, the drop falling away into shadow

AI Doesn’t Reward Expertise. It Punishes the Absence of It.

/ Maxim Starkweather / 7 min read

Brynjolfsson, Li, and Raymond studied 5,179 customer support agents at a firm that deployed a generative AI conversational tool. Their paper, now published in the Quarterly Journal of Economics, found a 14% average productivity gain — a solid, well-documented result. But the finding that actually matters in their data is the one almost nobody leads with: novice workers improved by 34%, while the firm’s most experienced and highly skilled agents saw minimal gains in speed and small declines in quality. AI made the novices better. For the experts at the top, it made things slightly worse.

This week, a piece by Sean Goedecke called “LLMs reward expertise” has been circulating on Hacker News with more than a thousand upvotes. Goedecke’s thesis is right: domain expertise is the key variable in how much value someone extracts from a language model. His central example is Fields Medalist Terence Tao’s ChatGPT conversation about the Jacobian Conjecture — Tao can identify errors, suggest reformulations, and recognize mathematical anomalies in ways a generalist prompter never could, because he deeply understands the field. The thesis is correct. The takeaway most readers will draw from it — that they should learn to prompt better — is exactly wrong.

Where the Gains Actually Live

The Brynjolfsson result is being used to argue two contradictory things simultaneously. In one version, it shows AI leveling the field — novices gaining 34% sounds egalitarian, like a tide lifting every boat. In another version, Goedecke’s argument takes the same evidence to argue the opposite: expertise is more valuable than ever, because experts extract more from the same tools. Both framings are downstream of the same misread. Neither explains what’s actually happening.

The mechanism Brynjolfsson’s team describes is precise: “the AI model disseminates the best practices of more able workers and helps newer workers move down the experience curve.” The AI wasn’t reasoning from first principles. It was acting as an efficient conduit for tacit knowledge that had previously lived in the heads of the firm’s best agents — encoding that judgment and making it accessible to people who hadn’t yet developed it. The novices got better not because AI made them more capable but because AI gave them access to a compressed version of expertise they didn’t own. The Brynjolfsson finding isn’t a story about AI replacing expertise. It’s a story about expertise being lent at scale.

This is where Goedecke’s piece is subtly right but misapplied. Tao isn’t a better prompter. He has the mathematical judgment to check whether the output is correct. That judgment came from decades of doing mathematics without a language model. The expertise isn’t being amplified by the model — it’s being used to supervise it. A language model in a technical conversation with Terence Tao is not the same tool as a language model in a technical conversation with someone who learned the field from language models. The model cannot tell the difference. Only one of those users can.

The same motion on different ground: inside and outside the AI capability frontier look identical from inside the task.

The Frontier Nobody Can See

Fabrizio Dell’Acqua and colleagues at Harvard, Wharton, MIT, and BCG ran a field experiment with 758 management consultants using GPT-4. The study — published in March 2026 in Organization Science as “Navigating the Jagged Technological Frontier” — introduced a concept that has received far less attention than the study’s headline numbers. AI capabilities, they found, create a landscape with an uneven surface: some tasks that look difficult to humans fall well within AI’s range, while others that appear equally tractable are outside it. The frontier is jagged — not a smooth line but a serrated edge running through the middle of a day’s work.

Inside the frontier, AI improves performance for everyone. Outside it, the results invert: consultants using AI were 19% less likely to produce correct solutions compared to those who used no AI at all. Not just less improved — actively, measurably worse. The consultants in the AI group who crossed the frontier performed below the control group that had no tool. They would have done better with nothing.

The structural problem is that the frontier is invisible from the inside. A task that is similar in surface appearance to one you just completed successfully with AI may be on the other side. The Dell’Acqua study’s central finding isn’t just that AI is better at some tasks than others — it’s that workers cannot reliably tell which category they’re in. The consultants who performed 19% worse weren’t inexperienced or careless AI users. They were skilled at using AI, working on something that looked like everything else they’d successfully completed. Their fluency with the tool was precisely what prevented them from registering when it stopped being reliable.

Fluency Is Not the Skill

A 2026 paper on what the authors call the paradox of AI fluency makes this structural problem explicit. The finding is counterintuitive: high fluency with AI systems does not guarantee superior task performance and in critical cases correlates with diminished outcome quality. Expert AI users develop predictable patterns and overconfident shortcuts. They skip verification. The ease of generating plausible-sounding text begins to substitute for the harder work of checking whether the text is actually correct. Familiarity with the interface breeds a form of complacency that is invisible to the person experiencing it.

Evaluating without a reference: when you’ve never done the work without the model, you don’t know what the correct answer looks like.

The paper distinguishes between what it calls operational fluency — knowing how to prompt, how to iterate, how to navigate the interface efficiently — and critical fluency, which is the ability to evaluate outputs with genuine skepticism regardless of how smoothly they were produced. Most of the public conversation about becoming good at AI is about operational fluency: what prompts work, how to chain instructions, how to use system messages. Critical fluency is not an AI skill. It is the underlying domain expertise that Goedecke correctly identifies as important. And it is built, almost without exception, by doing the work in the domain before AI was available to assist with it.

A 2025 study presented at CHI examined whether LLMs could help designers create better problem frames — a core creative task in design thinking. The result was clear: no measurable improvement in frame quality, and the performance gap between experienced and novice designers widened. Inexperienced designers reported diminished agency when working with the models. For tasks requiring genuine creative judgment rather than well-structured retrieval, AI assistance made things worse for the people who needed help most. The December 2025 framework from researchers studying AI’s role in hybrid intelligence systems puts this precisely: “AI equalizes performance on well-structured, routine tasks while amplifying pre-existing differences on complex tasks requiring deep judgment.” This is the consistent finding across multiple independent research programs, in different industries, with different models, across different years of AI capability.

What Gets Liquidated

A July 2026 Brookings analysis by Niam Yaraghi gives this dynamic a useful name: borrowed expertise. The AI productivity gains that exist — and they are real, particularly for novices on commodity tasks — are running on accumulated judgment that was developed in a world without AI. The people who can evaluate AI outputs, catch the model’s errors, and notice when they’ve crossed the jagged frontier built their critical fluency by doing the work themselves first. Their competence is inseparable from that history.

Brynjolfsson, Chandar, and Chen found a 16% relative employment decline for workers aged 22-25 in the most AI-exposed occupations following the widespread deployment of generative AI. A study by Hosseini and Lichtinger analyzing 65 million workers across 280,000 firms found that companies adopting AI sharply reduced junior hiring while maintaining senior employment levels — what the paper calls seniority-biased technological change. Both effects are economically rational at the firm level. They are a coordination failure at the industry level: when every firm simultaneously stops hiring the entry-level cohort, no firm can later recruit from the senior ranks that cohort would have produced.

The positions being eliminated weren’t just cheap labor. They were the pipeline through which expertise develops. The consultant who knows when the AI has crossed the frontier learned that by spending years on problems where no model was available to paper over the gaps. The person hired in 2026 who has never resolved an ambiguous consulting problem without AI assistance has borrowed the productivity gains of expert supervision without building the underlying judgment that makes that supervision possible. The accumulated micro-level evidence shows real gains. The macro-level question — who will be qualified to check the model’s work in ten years — is not one the micro studies can answer.

Goedecke’s piece frames this as an opportunity: get good at your domain and AI will reward you. That’s true for the people who already have the domain foundation. It describes the wrong problem for the people who are building careers on the assumption that operational fluency with AI tools is sufficient. The Tao example makes the point clearly, but in the other direction from how it’s usually read: the reason Tao can push back on a ChatGPT suggestion about the Jacobian Conjecture is not that he learned to use ChatGPT well. It’s that he spent decades solving mathematics before ChatGPT existed. The model is, for him, a capable but unreliable junior collaborator that he can supervise. That supervisory capacity is not something you acquire by prompting more skillfully.

The question the Dell’Acqua study actually answers is narrower and more useful than “does AI reward expertise?” It answers: what happens when someone who is good at using AI encounters a task that AI cannot do well? The answer, documented in the performance of 758 professional consultants, is that they do worse than if they had no tool at all. The frontier is out there. Most people using AI professionally have already crossed it without knowing it. The 19% decline is not theoretical. It is measured, in real work, by skilled professionals who had no reason to doubt the tool in their hands.

A vast plateau of cracked golden stone ending at a sheer cliff edge, the drop falling away into shadow

AI-generated editorial illustration · TemperatureZero · August 4, 2026

Keep reading the signal

Get the Daily Signal — a concise briefing on what actually matters in AI and the systems around it.

Subscribe Free

Continue the archive

Latest BriefingsArticlesAbout Temperature Zero