AI · 17 August 2026

Your AI Trial Worked. So Why Is Everyone Still Doing It The Old Way?

Almost every business owner I speak to has run some kind of AI trial by now. Very few of them can tell me what changed in the business because of it.

The trial itself usually went well. Someone on the team, normally the person who was already curious, took a real task and did it a new way. It worked. Everyone in the room agreed it was impressive. Then a big job landed, two people took leave, the quarter closed, and six months later that task is being done exactly the way it was done in 2024.

Nobody decided to abandon it. That is the part I find interesting.

The Failure Rate Is Real, And It Is Not About The Technology

The numbers floating around this year are brutal. MIT’s work on generative AI in organisations found roughly 95 per cent of pilots produced no measurable return on the profit and loss. Various 2026 surveys put the share of AI pilots that reach genuine production somewhere between 12 and 20 per cent. BCG has been making the same point for a while with its 10-20-70 split: about 10 per cent of the value sits in the algorithms, 20 per cent in data and technology, and 70 per cent in people, process and the way the organisation actually behaves.

Those are enterprise studies, and I would be careful about reading them straight across to a business with fifteen staff. But the shape of the problem travels perfectly well, and in a smaller business it is often sharper, because there is no transformation office to keep something alive out of institutional habit. If the owner stops paying attention, it stops.

Meanwhile Australian adoption keeps climbing. Somewhere north of 40 per cent of Australian SMEs now report using AI in some form, and the ones who use it are growing faster than the ones who do not. Both things are true at once: usage is up, and most of it has not changed how the work gets done.

A Pilot Proves A Capability. A Business Runs On Defaults.

This is the gap, and once you see it you cannot unsee it.

A pilot answers one question: can this be done? It is a demonstration, run by a motivated person, on a good day, with everyone watching. Of course it works. That is what a demonstration is for.

A business does not run on capabilities. It runs on defaults. What happens automatically when nobody is thinking about it. Which screen someone opens at 8am, which template gets copied, which spreadsheet the quote comes out of, who gets the email when a job is booked. Those defaults were set years ago, most of them by accident, and they are the actual operating system of your business.

A successful pilot changes what is possible. It changes nothing about what is default. So the moment attention moves elsewhere, the default reasserts itself, because the old path is still sitting there, still open, still free.

Four Ways It Dies Quietly

When I look back at trials that went nowhere, the same handful of causes come up.

  • Nobody owned it after the demo. The person who ran the trial had a full-time job already. The new way of working was extra, and extra loses.
  • The old path stayed open. If the previous process still works, and it is what everyone knows, that is what gets used the first week things get hectic. Which is every week.
  • The pilot ran on tidy data. Someone hand-picked twenty good examples. Then it met the real thing: three versions of the same customer, notes living in an inbox, exceptions nobody ever wrote down. The model did not degrade, it just met your business. That is a data readiness problem wearing an AI costume.
  • Nobody defined what success meant beforehand. So afterwards the verdict was “yeah, pretty good”, which is not a result anyone can act on. Deciding what you are measuring after you have seen the output is how you end up arguing about vibes.

Notice that none of those are technical. You could swap in a better model and hit exactly the same wall.

The Question I Ask Instead Of “Did It Work?”

Owners want to tell me whether the trial worked. It is the wrong question and it always gets a yes.

What I want to know is: if we turned this off tomorrow morning, who would notice, and how long would it take them?

If the honest answer is that nobody would notice for a month, you did not adopt anything. You ran a very good demonstration. There is no shame in that, plenty of useful learning comes out of a demonstration, but it should not be counted as progress on the balance sheet or in your own head.

If the answer is that three people would be stuck by lunchtime, then it has become part of how the business runs. That is adoption. It is a much less exciting moment than the demo, and it is the only one that pays.

What Making It Stick Actually Involves

None of this is clever. It is mostly the unglamorous work that nobody puts in a case study.

  • Pick one recurring task with a real cost. Hours, errors, or a delay that annoys customers. One. Five shallow experiments give you five stories and no change.
  • Name the owner before you start. A person, not a team. Someone whose actual job now includes this working.
  • Set the decision date and the rule up front. “We decide on the 12th, and we adopt if it saves the estimating team four hours a week without more rework.” Written down before you see any output.
  • Close the old path deliberately. This is the step almost everyone skips, and it is the one that decides the outcome. Retire the old template. Change the checklist. Make the new way the way, not the option.
  • Keep a human reviewing the output. Permanently, not as training wheels. AI outputs are probability, not truth, and the expert in your business is the one who directs it, checks it and refines it.
  • Write down how it works. One page. If the knowledge only exists in the head of the person who set it up, you have swapped one dependency for another.

A fortnight of that beats another six months of interesting experiments.

This Is An Execution Problem Wearing A Technology Costume

For me this is the most under-appreciated thing about AI in business right now. The bottleneck moved. Two years ago the hard part was whether the technology could do the job. Today, for a huge range of ordinary business tasks, it can. The hard part is a business changing its own habits, and that has been the hard part since long before any of this existed.

Which is why the businesses pulling ahead are rarely the ones with the best tools. They are the ones who can decide something and then make it real, and build capability around the tools rather than accumulating licences. It is the same muscle that determines whether a strategy day turns into anything, or whether a delegation decision survives contact with a busy fortnight.

Most business problems are thinking problems, and the downstream symptoms keep repeating until something upstream changes. A dead AI pilot is a symptom. The thing upstream is usually that the business has no reliable way of turning a good idea into a new default. Fix that and you do not just rescue one AI project, you fix the reason the last three initiatives fizzled too.

Where That Leaves The Trial You Ran In March

It is probably still sitting in someone’s browser tab. The person who ran it still thinks it was great. Nobody has said a word about it since April.

That trial already gave you its answer. The technology can do the job. What it could not tell you is whether your business can absorb a change, and that question was never going to be answered by a model. It was always going to be answered by whether somebody closed the old path.

Frequently Asked Questions

Why do most AI pilots fail to go anywhere?

Because a pilot proves a capability and a business runs on defaults. In almost every case the technology did what it was asked to do, and then nobody changed whose job it was, which system it lived in, or what happens by default on a Monday morning. The old way of working stayed available and free, so people went back to it the first week things got busy. Research consistently puts the cause in people and process rather than in the model itself.

How long should an AI pilot run in a small business?

Two to four weeks on one real, recurring task is usually enough to learn what you need. Anything longer tends to be avoidance rather than evaluation. Set the decision date before you start and decide in advance what result would make you adopt it, because a trial with no end date and no decision rule turns into a permanent experiment that nobody has to own.

What is the difference between an AI pilot and actually adopting AI?

A pilot is one person doing one task a new way while the old way still exists. Adoption is when the new way is the only way that task gets done, it has a named owner, it survives that person being on leave, and someone would notice within a day if it stopped. If you cannot describe what breaks when the tool is switched off, you have not adopted anything.

Why did our AI trial work on sample data but not on the real thing?

Pilots are usually run on a tidy, hand-picked set of examples, and the real business is far messier than that. Inconsistent file names, three versions of the same customer, notes that live in someone’s inbox, exceptions that were never written down. The model did not get worse when it hit production, it just met your actual data for the first time. Fixing the inputs is usually the cheaper half of the job.

Should a small business start with one AI use case or several?

One, and finish it. Businesses that run five shallow experiments at once end up with five interesting stories and no change to how the work gets done. Take a single recurring task that has a real cost in hours or errors, make the new way the default, keep a human reviewing the output, and only then look at the next one. Momentum on one embedded change beats activity across five.

Josh Horneman is a business coach and AI guide based in Perth, Western Australia. He works with business owners and leaders across Australia and globally through one-on-one coaching, the HOWLL platform, and structured consulting engagements.

Learn more

Turn The Trial Into How You Work

If you have run an AI experiment that impressed everyone and changed nothing, the fix is rarely a different tool. Let’s look at the one task worth embedding properly, and what has to change around it.