
The difference between an AI pilot and a proof-of-concept
A proof of concept is TRL-3. A pilot is TRL-7, running in an operational environment. The US government codified the difference decades ago. Most AI vendors ignore it.
Fewer than one in ten construction and transportation businesses reports using AI. Manufacturing sits at roughly 12%. The national rate is 20%, and professional services runs at about 40%. Those are the Federal Reserve Bank of Minneapolis’s readings of Census data from April 2026.
The reflex is to call these industries slow. Talk to the people running them and you find something different: most have already tried something. They ran a thing, it demoed well, and then it stopped.
That is a structural problem, and the structure has a name. The US government codified the difference between what most companies ran and what they should have run, decades before anyone put a language model in an ERP.
The federal engineering standard nobody applied to AI
The Department of Energy publishes Technology Readiness Levels, an ordinal scale for how far a technology has actually progressed. It comes out of NASA and defence acquisition, and any manufacturer or contractor who has filled in a federal grant application has met it.
Look at where the two terms sit.
| Level | Department of Energy’s own definition |
|---|---|
| TRL-3 | “Analytical and experimental critical function and/or characteristic proof of concept.” Validation achieved at this level. |
| TRL-4 | Component validation in a laboratory environment. Alpha prototype. |
| TRL-5 | Component validation in a relevant environment. Beta prototype. |
| TRL-6 | Prototype demonstration in a relevant environment, partially integrated with existing systems. |
| TRL-7 | System prototype demonstration in an operational environment. “Integrated pilot (system).” |
| TRL-8 | Actual system completed and qualified through test and demonstration. |
A proof of concept is level three, validated analytically and experimentally. A pilot is level seven, and the standard’s own words for it are “integrated pilot,” running in an operational environment.
Four levels separate them. Every one of those levels is a step toward the messy reality of a live business: real data, real systems, real people, real exceptions.
So when a vendor demonstrates a model against a clean sample extract and calls the result a pilot, the honest name for what was delivered is TRL-3. The scale your own procurement department uses says so.
What a pilot has
Four things, and all four exist before it starts.
A baseline. How long the process takes today, how often it errors, what it costs, measured on the current process before anything is configured. Once configuration begins, the baseline becomes unrecoverable, and the result becomes unreadable.
Live production data and live systems. Your actual ERP, your actual Procore instance, your actual VMS, your actual CMMS, with the mess in them. The exceptions are the point. A process that works on the clean 80% of records has proved nothing about the 20% that consume the labour.
A metric set fixed in advance. Named, agreed, and written down before anyone builds. Which numbers move, by how much, for this to be worth continuing.
A scheduled decision, in the calendar before the work starts. A date on which a named person decides what happens next, against the metrics set at the beginning.
The National Archives’ own guidance for pilot projects puts two of these in writing: “Establish the success criteria for the pilot, with input from all stakeholders,” and “An initial baseline analysis will help you to understand the concerns of participants.” Criteria first. Baseline first. Then build.
Here is the whole difference in one table.
| Proof of concept | Pilot | |
|---|---|---|
| Readiness level | TRL-3 | TRL-7, integrated, operational |
| Data | A curated extract | Your live production systems |
| Baseline | Captured afterwards, if at all | Captured before configuration begins |
| Success criteria | Agreed when the results arrive | Fixed before the work starts |
| Exceptions | Out of scope | The point |
| Ends with | A demonstration | A dated decision by a named person |
| Question it answers | Can this work? | Does this work here, and is it worth scaling? |
Why it always breaks at the same moment
The RAND Corporation interviewed 65 people with at least five years of AI and machine learning experience, 50 of them practitioners across more than 50 organisations, and published the root causes in 2024. RAND has no product in this market, which makes it the most useful source available.
The leading cause, cited by 84% of interviewees, is this:
“Industry stakeholders often misunderstand, or miscommunicate, what problem needs to be solved using AI. Too often, trained AI models are deployed that have been optimized for the wrong metrics or do not fit into the overall business workflow and context.”
— RAND Corporation, The Root Causes of Failure for Artificial Intelligence Projects
Optimised for the wrong metrics. Which is what happens when the metrics get chosen after the build, by looking at what the build turned out to be good at.
RAND then names the moment it surfaces: “These kinds of errors often become obvious only after the data science team delivers a completed AI model and attempts to integrate it into day-to-day business operations.”
That is the meeting every executive reading this has sat in. The demo was impressive. Integration started. And a set of questions arrived that nobody had asked, about the exception cases, the permissions, the handoffs and the six systems the workflow actually touches.
RAND’s finding on what happens next is the sharpest line in the report for anyone running a portfolio of experiments: “Organizations that quickly move from prototype to prototype often find that they are completely blind to failures that arise after the AI model has been completed and deployed.”
Prototype to prototype. Demo to demo. Every one successful on its own terms, and none of them producing evidence about whether the business changed.
Anushree Verma, a Senior Director Analyst at Gartner, said the same thing about agentic AI in June 2025: “Most agentic AI projects right now are early stage experiments or proof of concepts that are mostly driven by hype and are often misapplied.”
The number your board has heard, and why it is wrong
Somebody has probably told you that 95% of enterprise AI pilots fail. It comes from an MIT NANDA working paper published in July 2025, and it says something different from what it is quoted as saying.
The report’s actual sentence is: “Despite $30-40 billion in enterprise investment into GenAI, this report uncovers a surprising result in that 95% of organizations are getting zero return.” Organisations getting zero return. Not pilots failing.
The report’s own funnel is more interesting than its headline. For general-purpose tools like ChatGPT, 80% of companies investigated, 50% piloted, and 40% deployed, which the report describes as a pilot-to-implementation rate of about 83%. For custom, task-specific tools, 60% investigated, 20% piloted, and 5% implemented.
So the 95% is measured against a denominator in which most companies never ran a pilot at all.
“Saying that 95% of them were failing is like saying 95% of Tinder users have failing marriages, when 80% of the people you’re talking about have never even gone on a date in the first place.”
— Rob Wiblin, 80,000 Hours
The paper was never peer-reviewed, and its authors were working on agentic AI frameworks positioned as the answer to the problem it diagnoses, with no conflict-of-interest disclosure.
None of which means the underlying picture is rosy. It is not. S&P Global Market Intelligence surveyed more than 1,000 organisations across North America and Europe and found the share abandoning most of their AI initiatives had risen to 42%, up from 17% the year before, with the average organisation scrapping 46% of its AI proofs-of-concept before they reached production.
Forty-six percent of proofs of concept scrapped is a real, measured number. It is also entirely consistent with the structural argument here: a TRL-3 demonstration was never going to survive contact with an operational environment, because it was never asked to.
The gate, and why executives avoid it
The go/no-go gate is not a new idea in industry. Robert Cooper built Stage gate: a decision point in a project where it is formally evaluated and either allowed to proceed, sent back, or stopped. in the mid-1980s out of research into hundreds of new product launches, and its vocabulary has been standard in manufacturing product development ever since. In Cooper’s own description, gates are “quality-control checkpoints” preceding each stage, where the output is a decision and a resource commitment, judged against criteria that are set out in advance.
The reason gates get skipped in AI work is political rather than technical. A binary go/no-go turns into a referendum on whether the person who sponsored the project was wrong, and nobody schedules their own trial.
There is a better structure, and it comes from federal education research on pilot design. A pilot ends in one of three outcomes:
- Adapt. Identify the changes needed and make them.
- Adopt. Determine the steps to scale it up.
- Abandon. Decide whether the challenges are too costly to overcome.
Three doors, decided against criteria fixed before the work started. That is a routing decision rather than a verdict, which is precisely why it gets held rather than quietly postponed. Most well-run pilots end in adapt, and adapt is a success.
What separates the ones that convert
Two independent findings, both from sources your board already reads.
McKinsey surveyed 1,719 participants across 97 nations in May and June 2026 and found that only 37% attribute any EBIT impact at all to AI. The high performers, the small group attributing 5% or more of EBIT to it, differ in three ways: nearly three-quarters report fundamentally redesigning workflows, they are twice as likely to say senior leaders demonstrate commitment, and they are twice as likely to report that their organisations have defined processes to measure the impact of those initiatives.
Measurement discipline, correlating with profit, in McKinsey’s own data.
Gartner reached the same place from a different angle in April 2026, finding that organisations reporting successful AI initiatives invest up to four times more, as a share of revenue, in foundations: data quality, governance, AI-ready people and change management.
And the intent-versus-impact gap is now measured at national scale. The Census Bureau’s April 2026 working paper on AI diffusion, drawn from nationally representative data collected between November 2025 and January 2026, found:
An NBER working paper synthesising firm panels across the US, UK, Germany and Australia found more than 90% reporting no impact of AI on their own employment and 89% reporting no impact on labour productivity over the previous three years.
Adoption is happening. Impact is not. The gap between those two facts is the gate that never got held.
What this looks like when it is run properly
A pilot with all four properties is not a longer proof of concept. It is a different shape, and it is short.
Assessment, one to two weeks. Map the workflow as it actually runs today, find where it breaks, and set the specific metrics the pilot will be measured against. This is the baseline everything after it gets compared to.
Workflow pilot, four to six weeks. One named workflow, live against real data and real systems, worked by a small pod embedded in the team that owns it.
Evaluation and gate, one to two weeks. Adapt, adopt or abandon, decided on results measured against the metrics set in the assessment.
Then, and only if the gate says so, a build that runs in stages with manual steps retired as each piece goes live.
That is the model behind it, and it starts at $20,000, priced to match the pilot’s scope rather than the size of the ambition. For a company whose first question is adoption rather than one workflow, the equivalent is a bounded first step on one department, from $15,000, which produces the same thing in a different shape: a measured before-and-after and a real decision at the end.
Set that against what the industry currently spends to learn nothing. S&P Global found the average organisation scrapping 46% of its proofs of concept before production, and Gartner puts full generative AI deployment costs in the millions for transformation-scale programmes. A gated pilot starting at $20,000 is a different thing from a cheaper way to get the same answer. It is the only structure that produces an answer at all.
The reason to start small is modesty about evidence rather than about ambition. A small thing with a baseline produces proof. A large thing without one produces a slide.
Six questions before you approve the next one
- What is the baseline, who measured it, and when?
- Which specific numbers have to move, and by how much, for this to continue?
- Is it running against production data, including the exceptions?
- What date is the decision, and who makes it?
- What happens on each of the three outcomes, and has anyone written that down?
- If this is adopted, which manual step gets retired?
A proposal that answers all six is a pilot. A proposal that answers none of them will demonstrate beautifully in eight weeks and change nothing.
RAND’s own conclusion is a fair test of whether something deserves funding at all: “If an AI project is not worthy of such a long-term commitment, it most likely is not worth committing to at all.”
This is the same discipline problem visible across every industry that runs the physical world, where the tools land in the function that demos best and the money stays where nobody is counting. Once a pilot has converted, the next question is whether the rollout behind it actually stuck, which is what a rollout scorecard actually measures.
If the honest answer is that the metrics arrived with the results, the next one is worth structuring differently, and that decision costs nothing to make.
Get StartedWhat is the difference between a pilot and a proof of concept?
The Department of Energy's Technology Readiness Levels place a proof of concept at TRL-3, validated analytically and experimentally, and an integrated pilot at TRL-7, running in an operational environment. Four levels separate them. Practically, a pilot has a baseline captured before configuration, runs on live production data and systems, has success metrics fixed in advance, and ends at a scheduled decision by a named person. A proof of concept demonstrates that something can work.
Why do AI pilots fail to reach production?
The RAND Corporation's 2024 study of 65 AI practitioners found the leading cause, cited by 84% of interviewees, is that organisations misunderstand or miscommunicate the problem, and deploy models optimised for the wrong metrics or that do not fit the business workflow. RAND notes these errors typically surface only when a completed model is integrated into day-to-day operations, which is after the demonstration has already been judged a success.
Is it true that 95% of AI pilots fail?
No. That figure comes from an MIT NANDA working paper whose actual sentence is that 95% of organisations are getting zero return, measured across a population in which most companies never ran a pilot at all. The same report puts the pilot-to-implementation rate for general-purpose tools at about 83%. The better-measured figure is from S&P Global Market Intelligence, which found organisations scrapping an average of 46% of their AI proofs-of-concept before production, with 42% abandoning most of their AI initiatives, up from 17% the year before.
What should a go/no-go gate actually decide?
Three outcomes rather than two. Federal guidance on pilot design frames the decision as adapt, adopt or abandon: identify the changes needed and make them, determine the steps to scale up, or decide the challenges are too costly to overcome. A three-way decision is a routing choice rather than a verdict on whoever sponsored the work, which is why it actually gets held.
How long should an AI pilot take?
One to two weeks to establish the baseline and the metrics, four to six weeks to run one named workflow against live data and systems, and one to two weeks to evaluate and decide. Anything longer without a gate is a programme, and anything shorter without a baseline is a demonstration.
SOURCES (14)
- Federal Reserve Bank of Minneapolis, "AI adoption in business grows steadily but unevenly (Erick Garcia Luna)", 29 May 2026
- US Department of Energy, "Technology Readiness Levels (EERE R 540.112-02)", n/a
- US National Archives and Records Administration, "Guidance for Proof of Concept Pilot", n/a
- RAND Corporation, "The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed (RR-A2680-1)", 13 August 2024
- Gartner, "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027 (Anushree Verma quotation)", 25 June 2025
- MIT NANDA, "The GenAI Divide: State of AI in Business 2025", July 2025
- 80,000 Hours, "Podcast commentary on the MIT figure (Rob Wiblin)", n/a
- CIO Dive, reporting S&P Global Market Intelligence, "AI project failure data", June 2025
- Cooper and Edgett, "Stage-Gate and the Critical Success Factors for New Product Development", n/a
- Regional Educational Laboratory Appalachia at SRI International, for IES / US Dept. of Education, "Learning Before Going to Scale: An Introduction to Conducting Pilot Studies", May 2021
- McKinsey, "The State of AI", fielded 4 May-8 June 2026
- Gartner, "Organizations With Successful AI Initiatives Invest Up to Four Times More in Data and Analytics Foundations (Rita Sallam)", 16 April 2026
- US Census Bureau, "The Microstructure of AI Diffusion (CES-WP-26-25)", April 2026
- National Bureau of Economic Research, "Firm Data on AI (working paper w34836)", February 2026, revised March 2026


