Why the Numbers Disagree
Automation returns are reported at 100–200% in the first year and failure rates at 50 to 80% in the same period, by sources that are all describing something real.
Five mechanisms produce that, and only the last involves anyone shading anything. A product-side example of how workforce software approaches this topic is available in this reference.
One: survivorship
Vendor figures come from reference customers. A reference customer is one who implemented successfully, stayed a customer, and agreed to be named — three filters, each removing failures. For broader background and an independent point of comparison, see Atlassian Automation.
That is not a fabricated sample. It is an accurate description of a population selected by success, presented as a description of the population generally.
The same applies to case studies, testimonials and conference talks. Nobody publishes the implementation they turned off, so the visible record of this field is composed almost entirely of the successful minority.
Two: different subjects
RPA, agentic AI, systems integration and workflow tooling are distinct things with distinct track records, and figures move between them freely in retelling.
A 2020 RPA failure rate quoted about a 2026 AI agent deployment is comparing different technologies at different maturity, and it happens constantly in both directions.
Three: different definitions
Failure can mean abandoned, over budget, late, delivered but unused, or delivered with no measurable benefit. The published range from 30% to 80% is mostly this, and the strictest definitions produce roughly double the loosest.
Success is worse. Almost nobody defines it, and where it is defined it usually means the project completed rather than the benefit arrived.
Four: different populations
Nearly all serious data is enterprise. RAND's 80.3%, Gartner's 28%, the 350-executive survey at companies over $1bn revenue.
Quoted at a company of forty people, those figures describe a different world: different governance, different process complexity, different tolerance for an eighteen-month programme. There is no published base rate for small implementations at all.
Five: incentive
Last, smallest, and the only one people usually mention.
A vendor chooses favourable assumptions. An analyst firm has commercial relationships. A consultancy names causes that point at its own service line. None of this requires anyone to lie, and treating it as the main explanation obscures the four larger mechanisms above.
What to do in a field like this
Weight by who could be wrong in your favour. Deloitte attributing 37% of failures to change management is a consultancy naming a cause that sells its own services — and it is also consistent with independent work, which makes it stronger rather than weaker.
Look for agreement across opposed interests. Vendors and researchers disagree about the failure rate and agree about the causes. That agreement is the most reliable thing in the entire field.
Combine rather than choose. Vendors report the return achieved by successful projects; researchers report how many projects succeed. Both true, and together they say: a minority succeed, and that minority does well.
And measure your own. Your exception rate, your volume, your variance are knowable in two weeks with a tally sheet, and no published figure substitutes for them.
The honest position
Nobody can tell you your probability of success, and anyone who states one confidently is quoting a population you are not in.
What can be said is narrower and more useful. A large minority of projects do not deliver what was promised. The causes are consistently organisational. The successful ones share two properties, neither of which is a purchase. And the figure that most predicts your own outcome is one you can measure yourself before spending anything.
The short version
- Five mechanisms: survivorship, different subjects, different definitions, different populations, and incentive — in that order of size
- Vendor figures come from reference customers, a population filtered three times by success
- The 30–80% failure range is mostly definitional; strict definitions produce roughly double the loose ones
- Nearly all serious data is enterprise, and there is no published base rate for small implementations
- Incentive is the smallest mechanism and the only one people mention
- Look for agreement across opposed interests — vendors and researchers agree on causes, and that is the field's most reliable signal