On September 26, a Florida woman typed into Claude that she was going to “shoot up” the Lee County Sheriff’s Office. The next day she wrote that she had a new gun. Anthropic’s systems flagged the text, a human reviewer judged it credible, and the company called the police. Deputies detained her at home without incident, and she now faces a second-degree felony under Florida Statute 836.10, as TechSpot reported on October 4. She told investigators she uses Claude like a diary. Most of the reaction has been about whether this particular referral was right. It probably was. The better question is why nobody outside Anthropic can say what the rule is.
This referral looks defensible, and that is not the point
Start with the strongest fact for Anthropic. The messages named a target, a building, and then a weapon the following day. If a human reviewer sat on that and something happened, the story would be a lawsuit, not a debate about privacy. I am not arguing the review team got this one wrong, and nothing I fetched suggests it did.
The trouble is that one clear case teaches you nothing about the line. The Next Web reports this is at least the third Claude conversation to reach police since August. A 22-year-old in San Antonio asked Claude about shooting at an elementary school on August 11 and was arrested on a felony terroristic threat charge. On August 14 a user threatened Anthropic’s CEO Dario Amodei and mentioned buying an AR-15; Anthropic notified police and no arrest followed. Three cases, three different fact patterns, and the public can see only the ones that ended in a courthouse record. We do not know how many conversations were flagged, how many a reviewer read, or how many were closed without action. That denominator is the whole story, and it does not exist anywhere Anthropic has published.
Consider what the reviewer in this case actually had to decide. A key-phrase system raised the flag, which is a filter that cannot tell a threat from a joke, a novel draft or a bad night. A person then read the text and made a call that ended with deputies at a door. Per the arrest coverage, her arraignment is set for November, and the charge attaches to the writing itself, a written threat, in a conversation she apparently believed was private. That reviewer is an employee of a software company, not a clinician or a police officer, and the company has not said what training, checklist or second reader stands behind the decision. I do not know that there is none. Nobody outside Anthropic does, which is the point.

What the published policy actually commits to
Anthropic’s privacy policy, effective September 10, allows sharing personal data with authorities on a good-faith belief that disclosure is reasonably necessary to “prevent serious harm to any person or to property,” and to “detect, prevent, or otherwise address fraud or other illegal activity.” It says conversations may be “flagged for safety, security, or policy review.” In the text I retrieved, it does not describe how much a human reviewer reads, and it says nothing about notifying a user that a conversation was escalated. Both are my reading of the page, not a statement from the company.
The usage policy, effective September 15, 2025, is more specific about exactly one thing: when Anthropic detects child sexual abuse material or coercion of a minor, “we will report to relevant authorities.” That is a commitment with a trigger and a verb. Nothing comparable exists for violent threats. The same company that wrote a hard rule for the one category the law forces its hand on wrote a judgment call for everything else, and the judgment call is the one that just put a user in front of a felony judge. TechSpot’s paraphrase of the company’s position is “limited emergencies” involving death or serious physical injury. “Serious harm to property” in the policy text is a good deal wider than that.
I also checked Anthropic’s transparency hub. It is thorough on model evaluations, safeguards and capability assessments. Across the text I read, it publishes no count of law-enforcement referrals, emergency disclosures, or threat reports. A company that reports its model’s honesty scores to a decimal place has not published how often it calls the sheriff.
The squeeze runs one direction
The pressure on labs is not symmetric, and the lawsuit that explains the asymmetry is already filed. On September 22, British Columbia sued OpenAI and Sam Altman in San Francisco federal court over the Tumbler Ridge school shooting. The complaint says OpenAI flagged the shooter’s conversations in June 2025, that her activity “did not meet the threshold for legal referral,” and that safety staff recommended contacting police anyway. TechSpot relays Wall Street Journal whistleblower accounts that executives, including Altman, rejected the recommendation. Altman apologized in April, and the suit alleges the promised reforms never followed. Those are allegations, and OpenAI has not had its day on them.
Put the two cases side by side. One lab is being sued for not reporting. Another reported and a user is being prosecuted. Any trust-and-safety lead reading both knows which error costs money and which costs a headline, and the incentive is to move the line toward reporting. I think that is where the industry is heading regardless of anyone’s intentions: a threshold that drifts down, set privately, never audited. Nothing I fetched shows Anthropic’s has drifted. That is exactly the problem, because nothing I fetched could show it either way.
There are two ways to be wrong here, and they are not priced the same. A missed threat is a body count and a lawsuit. A needless referral is a person with a record, a bad month, and a reason to never type honestly into a chat box again. My inference, and it is only that, is that the second kind of error is accumulating somewhere in the industry unreported, because the people who suffer it have no channel and the companies have no obligation to count. A rule that nobody measures for false positives will not be tuned for them.

Users were sold something different. Sheriff Carmine Marceno’s reported line, “Users need to understand that you are never truly anonymous,” is correct and is also not what the product feels like at midnight with a blinking cursor. A chat box invites the register of a diary. People type things to a model that they would never say to a person, partly because it is not a person. The woman in this case apparently treated it that way. A company can legitimately say that its policy forbids this and that it reserves the right to escalate. What it cannot claim is that users were told where the line sits, because the line is not written down.
What publishing the threshold would look like
I am not asking labs to stop reporting credible threats. I am asking for three things, none of them hard. First, state the standard in operational terms: what must be present, such as a named target, a stated intent, access to means, a time frame, before a reviewer escalates, and what is routinely closed. Second, publish a count every half year, the way companies already do for government data requests: flags, human reviews, referrals, and how many referrals led to charges. Third, say in the policy whether a flagged user is told. Those three disclosures cost a lab nothing in safety. Publishing them would not help an attacker, because a threshold built on intent and means is not a filter you can route around by rewording.
The counter-argument is real. A published threshold invites gaming, and a count invites misreading, since a rising number could mean more danger or more caution. Both are manageable. Credit-card issuers and platforms publish fraud and takedown figures with that same ambiguity and survive it. The alternative is worse for the labs themselves: a standard that exists only in a reviewer’s head cannot be defended in either courtroom, the one where a victim’s family sues and the one where a defendant’s lawyer asks what the company actually saw.
My position is that the next lab to publish its referral threshold and its numbers gets a lasting trust advantage, and the first to be forced to produce them in discovery gets a very bad week. Anthropic has the opening. It has built its public identity on being the lab that writes things down, and this is a case where the thing written down is a rule it applies to people, with consequences. Right now it has a hard rule for child safety, a vague one for everything else, and a felony case for the vague one. Close that gap before a court does.

AI-generated editorial illustration · TemperatureZero · October 6, 2026
Keep reading the signal
Get the Daily Signal — a concise briefing on what actually matters in AI and the systems around it.
Subscribe FreeContinue the archive