Stat audit / AI

Three "AI projects fail" stats.
Not one is a study.

87%, 80%, 85%. You have seen these in every AI keynote and boardroom deck. I traced all three to their primary sources. The finding was not fabrication. It was degradation, and a citation loop where each number borrows its authority from the next.

Run with Kit 003: the primary-source-only fact-check brief, the four-verdict rubric, and the STRIP checklist. Every claim was judged against the original source, never against a creator or journalist repeating it.
87%
"never reach production"
Distorted
80%
credited to RAND
Distorted
85%
credited to Gartner
Distorted

The receipts

"87% of data science projects never make it into production."
Credited to → "VentureBeat," as if it were research
"Why do only 13% of data science projects, or just one out of every 10, actually make it into production?"
Deborah Leff, IBM, spoken on a panel at VentureBeat Transform 2019. 100 minus 13 is where the 87 comes from.
What's missing
No sample, no methodology, no paper. And it recedes further: on that stage, Leff attributed the 13% to CIO Dive magazine, so even the person everyone quotes was repeating a secondary source. A spoken figure, one citation removed from anything measurable, now appears thousands of times as research. The units drifted too, from "data science projects" to "AI projects" to "ML models."
STRIP
I, Inflated inference. A verbal aside was upgraded into a hard statistic the speaker never claimed to have measured.
Distorted, bordering Unverifiable
"More than 80% of AI projects fail."
Credited to → RAND Corporation (2024)
"By some estimates, more than 80 percent of AI projects fail."
RAND, The Root Causes of Failure for AI Projects (Ryseff, De Bruhl, Newberry, 2024). Read the first four words.
What's missing
RAND is citing outside estimates, not reporting its own. RAND's actual work was 65 qualitative interviews about why projects fail. It never produced the 80% number. The headlines hand a rigorous institution a figure that institution explicitly attributed to "some estimates."
STRIP
P, Pinned to the wrong thing. The number is real. The attribution is not.
Distorted
"85% of AI projects fail."
Credited to → Gartner (2018)
"Through 2022, 85 percent of AI projects will deliver erroneous outcomes due to bias in data, algorithms or the teams responsible for managing them."
Gartner, 2018 forecast, research VP Jim Hare. "Erroneous outcomes," not "fail."
What's missing
Two edits turned a forecast into a fact. "Deliver erroneous outcomes" became "fail," and "through 2022," a bounded prediction about a future window, was dropped so it reads as a measurement of what already happened.
STRIP
I plus T. The verb was inflated, and the condition was trimmed.
Distorted

The worst thing

The finding

The problem is not one bent number. It is that all three cite each other. Each borrows authority from the next, so the pile looks well-sourced while none of it traces to a single clean measurement. Fabrication is rare. Degradation is the standard, and in a closed loop it compounds.

This matters most in financial services, where one slide reading "85% of AI projects fail" gets used to kill a budget or justify one. The fix is not blanket skepticism about obvious nonsense. It is one habit for the numbers that sound clean: ask whose units, whose conditions, whose subgroup, and whether a primary source exists at all.

The method

How each stat was graded

The rubric gives every claim exactly one verdict, judged only against its primary source. All three numbers here landed on the same one.

True
Matches the primary source within rounding.
Distorted
All three landed here. A real source exists, but the number or framing is materially wrong.
Unverifiable
No primary source findable after a genuine search.
False
The primary source flatly contradicts it.

STRIP, the five ways a true number goes bad

Two minutes, no agent needed. Of 15 claims audited from a top-1% creator, only one invented a number. Five took a real number and bent it, and every bend was one of these.

  • S, Swapped units. Tokens become words, medians become means. Is this the unit the study measured?
  • T, Trimmed conditions. The lab constraint disappears. "In exams" becomes "at work." Under what conditions was it measured?
  • R, Rounded to the ceiling. The best-case subgroup's number becomes everyone's. Average, or ceiling?
  • I, Inflated inference. The study's noun survives, its verb gets upgraded. Did the study test what the sentence now claims?
  • P, Pinned to the wrong thing. One group's number gets handed to another. Whose number is this, exactly?

Verify it yourself

Do not take this page's word for it either. That is the whole point. Every source is one click away, and the STRIP checklist takes about two minutes per number.