On October 6, Thomas Bloom, who built erdosproblems.com in 2023, announced that the site would freeze new problem comments and proof claims, stop displaying a status for any problem, drop the solved count and percentage, and stop using credit-giving language for results, whether a human or a model produced them. His stated reason is that the comment section filled with AI-produced solutions posted with no explanation, mostly to claim priority, and that this drove human mathematicians away. The freeze is the part people will quote, though Bloom’s site stays up and its general threads stay open. The part that matters is the other three: an institution that ran a public scoreboard for solved problems deleted the scoreboard instead of trying to police who posted to it. My read is that this was the correct move, and that it says something general about leaderboards.
What the scoreboard was
The site launched on May 28, 2023 with a little over 200 problems. It now holds 1,221, more than 9,000 comments, about 2,000 registered users and between 10,000 and 25,000 daily visitors, per Bloom’s own numbers. Comments arrived in August 2025. Every problem page carried a status, and a solved one carried a name. A running percentage sat above all of it. That is a scoreboard in the plain sense: a public list of who did what, with a count at the top, and it paid out in the only currency mathematicians care about, which is being the person the record names.
A scoreboard like that depends on an assumption nobody had to state. Putting a name next to a solved label was expensive to earn. You needed the mathematics, and then you needed enough standing in the field that someone would read your argument. The cost of a claim was roughly the cost of the work, so the payout was safe.
The scoreboard was also never separate from the good part of the site. Quanta reports that Tao worked on the site with amateurs and undergraduates, and Bloom’s own account is that after comments opened in August 2025 the threads were first vibrant collaboration. Open credit is what made a stranger willing to post a partial result in public. That is the cost of what he has now removed, and I do not think it is small.

How the price collapsed
The collapse took a year and the record is public. In October 2025 OpenAI claimed GPT-5 had solved open Erdős problems, and the claim fell apart because the model had surfaced results already in the literature. On January 26, 2026, Kevin Barreto posted on the site that GPT-5.2 Pro and the Lean formalizer Aristotle had solved problems 728, 729, 401 and 205, with every proof checked in Lean 4 before it went up, and with credit given to both his collaborator Liam and the systems. That was the careful version of the new workflow, and nothing about it was wrong.
The same week, a 24-author team led by Tony Feng reported running Gemini over 700 problems marked Open. It addressed 13: five through what the abstract calls seemingly novel autonomous solutions, eight by finding that the answer already existed. In May, OpenAI announced that a model had disproved the 1946 planar unit distance conjecture, and Bloom validated it while stressing that the proof was “significantly improved by the human researchers at OpenAI and the many other mathematicians involved.” On August 1, according to a secondary account of OpenAI’s announcement, an unreleased model called Astra was credited with ten open problems, three of them Erdős problems, for roughly $2,000 of compute. I did not fetch OpenAI’s own post, and the same account notes that failed attempts went unpublished.
Put those together and the cost curve is visible. The Astra figure above, roughly $2,000 for ten problems, is the sender’s side of the ledger. Reading what comes back still costs a person an afternoon, and that cost has not moved. The same account says failed attempts went unpublished, so the true cost of a success is higher than $2,000, but the direction is not in doubt. In Quanta’s August piece, Bloom put the receiving end this way: “no human has read it, and no human is going to read it.” Noga Alon, also quoted there, said that once AI started solving these problems, “there is no point anymore.”
The scoreboard was the attack surface
Here is the argument, and I want to be exact about which parts are evidence and which are mine. The evidence is Bloom’s: after comments opened, the thread filled with unexplained AI proofs posted to claim priority. The inference is mine. Priority was the bug. A status flip and a name are a reward, and the reward went to whoever posted first, not whoever explained best. Once a sender can generate attempts at near zero cost and the receiver cannot read them at near zero cost, the rational move for the sender is volume. Nobody has to be acting in bad faith for this to happen. Barreto was not. The structure still pays the fastest poster, and that is now a machine.
This is why moderation would not have worked. You can require Lean proofs, and Barreto’s workflow shows that works for correctness. But a Lean certificate is a program that compiles or does not. It settles whether a statement is true and says nothing about whether any person now understands why. Bloom’s announcement, as Terence Tao’s blog carried it the same day, shifts priority to written exposition that people can follow, with formal proofs as a complement. The problem he is solving was never wrongness. It was displaced attention, and a verifier does not fix that.
One limit on all of this. The announcement I read gives no breakdown of how many of the 9,000 comments were AI-written or how many problems were claimed by people who could not explain their own proofs. The diagnosis rests on Bloom’s firsthand account, which is credible because he runs the site, and not on a measured rate. I treat it as strong testimony plus an inference about mechanism, not as a result.
Removing the status line removes the thing being competed for. Bloom’s replacement for credit language is to record that a result is known, with a link to an explanation, and nothing about who found it.
The strongest objection
The deletion has real costs, and they land on the people the site was built for. Status is how a researcher decides whether a problem is worth an afternoon. Credit is how an undergraduate gets a first line on a record that mathematicians read. Barreto’s carefully verified, openly credited work is erased by the new rule exactly as thoroughly as the spam is. A reader has to take Bloom’s word, and the word of whoever emails him, that a result is known.

The Gemini paper complicates my thesis in an interesting way. Its authors concluded that the Open status of those problems “was through obscurity rather than difficulty.” If the label was already a poor proxy for hard, then the status line was carrying less information than it appeared to, and deleting it loses less than it seems to. That is a point in Bloom’s favor, but I will not oversell it: the paper also warns of “subconscious plagiarism” by models, which means the credit language was contested well before October. At the June Leiden Declaration, signed by 1,590 people within three days according to Science News, Harvard’s Melanie Matchett Wood said it is “not clear that there’s a way for [AI] to reasonably attribute the source of the ideas.” A scoreboard that cannot say whom to credit is a scoreboard with a legitimacy problem independent of spam.
What Bloom keeps is also worth listing. The problem statements, the historical archive, the tag system, general forum threads and links to Lean formalizations all stay. The site stops being a ledger and stays a library.
What other leaderboards should take from this
Every open leaderboard has the same dependency the Erdős site had. A bug bounty table, a benchmark ranking and a contributor list all pay out recognition on the assumption that producing the thing was costly. A model that cuts the producer’s cost to near zero without cutting the reader’s cost breaks that assumption quietly, and the board keeps running while it measures compute instead of contribution. There are two honest responses. You can raise the price of entry, which is what Lean requirements do for correctness. Or you can take the prize down. Bloom judged that nothing raising the price of entry could raise the price of understanding, so he took the prize down.
I expect the next institution to do this to be one that cares more about its community than its count. Bloom’s stated test is that the site should help the Erdős-community of humans flourish, and a percentage solved does not pass it. Counting was the easy part of what the site did, and models just made it free. What is left is the part that was always scarce, which is someone explaining a proof well enough that another person leaves the page understanding it.

AI-generated editorial illustration · TemperatureZero · October 7, 2026
Keep reading the signal
Get the Daily Signal — a concise briefing on what actually matters in AI and the systems around it.
Subscribe FreeContinue the archive