The full story
An AI Detector Flagged a Berkeley Professor's Op-Ed as 33% AI
She had already admitted it — and that is the part no one has a rule for
The column, and the irony sitting on top of it
The piece that started this was not, on its face, about artificial intelligence at all. Stankova's argument in the San Francisco Standard was about arithmetic: that a teaching professor at one of the country's most selective public universities now spends office hours explaining fractions, and that the students in front of her are, by her measure, years behind where a calculus class assumes they are.
She reached for numbers to make the case. Before 2020, she wrote, a fraction of one per cent of her incoming students tested below basic algebra; by the 2022-23 diagnostics, more than a third showed what she called a severe deficit. She tied the timing to the University of California's decision to stop considering standardized tests, arguing that the system "discarded the one standardized baseline it once had." Those figures are hers, drawn from Berkeley's own diagnostic exams and framed by her own initiative to bring testing back; they are her claims, and we are reporting them as claims rather than certifying them.
What made the story spread was not the arithmetic. It was the sentence, added days later, that the column warning about students who cannot do their own work had itself been edited with a machine. It is a genuinely funny juxtaposition, and it is also the reason the story is worth slowing down on, because almost everyone drew the wrong lesson from it.
A 33% score, and what that number is actually measuring
Start with the tool, because the whole affair turns on a misunderstanding of what it said. The Daily Californian reported that Pangram flagged the op-ed as "33% AI-generated or AI-assisted." That phrasing does a lot of quiet damage, because it welds together two things the detector itself keeps apart.
Pangram, a company founded in 2023 by two former Stanford researchers, does not return a single yes-or-no verdict. You feed it text and it sorts the result into one of four bands — Human-Written, Lightly AI-Assisted, Moderately AI-Assisted, or Fully AI-Generated — and attaches a numerical assistance score. A "33%" reading lives near the assisted end of that scale. It is the tool's way of saying the text bears the marks of a person working with a machine, not its way of saying a machine produced a third of the words. The distinction is not pedantic. It is the difference between the scandal people thought they were reading and the thing that actually happened.
And it matters that the detector is a credible one. AI detectors have earned a bad reputation, most of it deserved: as a class they misclassify real human writing, and the cost of a wrong flag lands on a person who often has no way to appeal it. One widely cited finding is that 2023-era detectors falsely flagged 61.3% of essays written by non-native English speakers as machine-made. That is a real and ongoing problem — we will come back to who pays for it — but it is not what happened here. On the specific benchmark that matters, Pangram is an outlier in the right direction: evaluations summarised by researchers at the University of Chicago Booth School put its false-positive rate at essentially zero on medium and long passages. A column of this length is exactly the case where the tool is most trustworthy.
There is one more caveat worth stating plainly, because it cuts against tidy conclusions in both directions. A single detector's number is not reproducible across the field: run the same text through a different tool and you will often get a different score, because each is trained on different data and tuned to a different threshold. That is an argument for treating any one figure as a signal rather than a measurement — and it is also why the good detectors publish their false-positive rates and the marketing-driven ones do not. Pangram's number is more trustworthy than most. It is still one instrument's reading, not a physical fact about the document.
So the two reflexes the internet reached for both fail. It was not a false positive; the detector was, on the evidence, right. And it did not unmask a fraud; the reading it produced — assisted, not generated — matches the account the author gave in her own words.
"About 80 hours are my own"
That account is worth reading closely, because it is unusually specific. Stankova did not issue the standard non-denial. She said the column was "the result of several hundred person-hours of intensive human work and deliberation, of which about 80 hours are my own," that a faculty team and several journalists had worked over multiple drafts across three weeks, and that AI had been used "to help edit the piece" and to "locate numerous documents and articles" behind it. "All analysis," she told the Guardian, "is the result of the team members." She also called the question of how AI was used "orthogonal to and a distraction from" the initiative the piece was arguing for.
You can find that last line evasive — a lot of people did — and still notice that it is not the posture of someone who was hiding something. The detector did not extract a confession. The confession was already available; the detector just gave a headline a hook to hang it on. When the disputed fact is one that the accused volunteered, the story is not "caught." The story is "so what."
The rule everyone assumed exists
Which brings us to the actual, unresolved thing here, the one a percentage can never adjudicate: is using AI to edit an opinion piece against the rules?
The honest answer is that there is no settled rule, and the loudest reactions assumed one into being. The San Francisco Standard's own policy, as it described it, is that "while AI may assist, our expectation is that humans are behind every article we publish and take responsibility for every word." Read that against what Stankova said she did — human authorship, machine editing, a named person taking responsibility — and it is hard to locate the breach. The policy appears to permit precisely the workflow that was supposed to be the scandal.
That gap between assumed rules and written ones is becoming the defining problem of this whole subject. In some settings the line is now drawn in law: the European Union's AI Act has begun forcing a plain disclosure in specific contexts — the obligation, at its simplest, to tell someone "you are talking to AI". In others, platforms are drawing it themselves: Spotify's approach to synthetic music is not a ban but a label, flagging the AI act while leaving the human listener to decide. Both are attempts to answer the same question this op-ed raised — when does a machine's involvement have to be declared? — and neither of them reaches an opinion column in a newspaper. There, the norm is still being made up in real time, one embarrassment at a time.
It is fair to think a professor arguing that students should do their own work owes her readers a note when she has had a machine polish her own. That is a reasonable expectation about disclosure. It is simply not, yet, a rule she can be said to have broken, and the two should not be blurred together — which is what a detector's tidy-looking number invites you to do.
Part of why the vacuum persists is that opinion writing has always been a collaborative form. Op-eds are routinely shaped by editors, ghost-drafted for public figures, and fact-checked by desks the byline never names; the reader is trusted to know that a signed column is a claim of responsibility, not a claim of solitary composition. AI slots into that older ambiguity without a settled place. The question is not really new — it is who did the work behind a byline — but the tool is new enough that no one has decided how much of it a machine may quietly do.
Who actually pays when a detector is wrong
The reason to get this case right is not Stankova, who has tenure, a platform and eighty documented hours of her own. It is everyone the same tools are pointed at who has none of those things.
Every day, detectors like Pangram are run not on op-eds by named professors but on essays by students who cannot email a reporter a three-paragraph rebuttal. For them the assisted-versus-generated distinction is not a debating point; it is the difference between a normal grade and an integrity hearing. And when a detector is wrong — rarer with the good ones, routine with the bad — the burden of proof inverts onto the accused, who is asked to demonstrate a negative about how their own sentences came to exist. It is the same shape as an automated system confidently misidentifying an ordinary person and leaving them to prove it wrong, the pattern a facial-recognition tool produced when a supermarket's cameras flagged the wrong shopper. The technology is different; the asymmetry is identical.
That is the exposure worth taking from this story, and it survives whatever you conclude about the professor. A number that looks like a verdict is not one. "33%" is a probability wearing the costume of a fact, and in the Berkeley case the reader could go and check it against an admission freely given. In the far more common case — a student, a flagged paper, a detector's score and nobody's confession to compare it against — there is no such check, and the same tidy number carries the same false air of having settled something it did not. The lesson is not that the detectors do not work. It is that even when they work, they answer a smaller question than the one people ask them.
Sources and verification
- The Guardian: Stankova's admission that she used AI to "help edit the piece" and to locate documents; her "orthogonal … distraction" remark; and the San Francisco Standard's statement that "while AI may assist … humans are behind every article … and take responsibility for every word."
- San Francisco Standard: the op-ed itself (15 August 2026), its "five to eight years" and severe-deficit claims, and the absence of any AI note in the published piece.
- The Daily Californian: Francis Luo's 18 August report that Pangram flagged the article as "33% AI-generated or AI-assisted," and Stankova's "about 80 hours are my own" response.
- Cryptobriefing: Pangram's four output bands (Human-Written / Lightly / Moderately AI-Assisted / Fully AI-Generated) and assistance score, its 2023 founding, and its 99.98% detection / 1-in-10,000 false-positive claims.
- GradPilot: comparative false-positive rates (Pangram near zero on medium and long passages, per University of Chicago Booth), and the class-level finding that 2023-era detectors falsely flagged 61.3% of non-native English essays.
- University of California press room: the Board of Regents' 21 May 2020 decision phasing out the SAT/ACT requirement.
- University of California / news.uci.edu: the 2 September 2020 court ruling that UC could no longer consider SAT/ACT scores in admissions.
- Stankova's UC Berkeley faculty page: confirms she is a teaching professor of mathematics at UC Berkeley and founder of the Berkeley Math Circle.