Most automation plans start with the tool. They should start with a list of numbers: what happens in this business by hand every week, how often, and how long it takes. That is less exciting than a demo, and it is the only thing you can check afterwards.
Repetitive work is hard to see because the people doing it have grown used to it. Nobody walks in to report that they spend six hours a week retyping. What you do hear is "I just do that myself", "I run that export every Monday" and "I know that by heart". Those are not complaints, they are locations.
In almost every SME, manual work leaves the same five traces. Once you know where to look, you have a first list within half a day.
Something already exists digitally and is copied by hand into a second system.
ask: which screen is being read from here?The package shows almost, but not quite, what someone needs, so the answer is kept beside it.
ask: what is missing from the report?PDFs, receipts and attachments that someone opens, renames and files in the right folder.
ask: how many per week?Status questions from customers or colleagues that someone answers manually each time.
ask: where does the answer really live?Somebody periodically scans a list for errors without it existing anywhere as a task.
ask: what breaks when they are ill?It is tempting to turn this list straight into a decision. After all, everyone has a sense of where most of the time goes. The problem is that this sense deviates measurably from reality, and not by a little.
The sharpest example comes from a METR study published in July 2025. Sixteen experienced developers completed 246 tasks, half with AI assistance and half without. Afterwards they estimated they had been roughly twenty percent faster with AI. Measured, they took nineteen percent longer. That is not a story about developers or about AI, it is a story about how badly people judge their own use of time, even while they are in the middle of it.
That same blind spot explains part of the disappointment you see everywhere now. Gartner expects more than forty percent of all agentic AI projects to be cancelled before the end of 2027, and earlier MIT research produced the figure that roughly 95 percent of corporate generative AI pilots never deliver measurable returns. That is rarely down to the technology. It is down to there never having been a baseline, so nobody can say afterwards whether anything improved.
The mood in the market has shifted accordingly. Where vendors spent early 2026 launching agent builders, the conversation these months is about orchestration, hard limits and verifiability: a few well defined applications with proof beat ten that nobody checks. That is exactly the right move, and it starts with counting.
You do not need a time tracking system and you do not need a consultancy. You need one week in which everyone who works by hand briefly notes what they do, and one person who adds it up afterwards.
For one week, every action that feels like manual work: name, how often, how long, in which systems. Two lines each time.
All lines per task together, extrapolated to a year. Now you have numbers instead of impressions.
The top ten against four questions: frequency, predictability, available data, consequence of an error.
One task, not three. The rest goes on the list and waits for what you learn from the first.
The outcome is surprising in almost exactly the same way every time. The task that irritates people most is rarely at the top, because irritation scales with how unpleasant something is, not with how often it happens. At the top you usually find something dull: a three minute action that occurs seventy times a week and that nobody has ever added up.
You do not pick from the list on size alone. A task is only a good first candidate when it scores well on all four of these points:
These four questions do something else too: they separate genuine automation work from data work. If a task fails on the third question, that is no reason to drop it, but it is a sign that something else has to happen first.
Often complex, rarely frequent, and the data sits half in people's heads.
Unremarkable, high frequency, and fully reconstructable from systems.
Starting with the dull task is not modesty, it is risk management. The first automation mainly has to produce something you can build on: a working example, a measured saving, and an organisation that has seen it hold up. Only after that does it make sense to take on the complicated task everyone pointed at from the start.
Lay the results side by side and the same thing stands out almost every time: the manual work sits on the seams between systems. Retyping happens because two packages do not talk to each other. A spreadsheet appears because the data someone needs is spread across three places. A status question is answered by hand because the answer exists nowhere in full.
Your measurement week is therefore not only an automation list, but also a map of where your data falls apart. That is exactly the pattern we describe in five signs your business has outgrown its CRM, and the reason an open data layer in-house delivers more than a series of point to point connections. How that foundation works is set out on the page about the AI Native Data Layer.
Once you have chosen your task and want to know how to actually automate it, from fixed rules to an agent within limits, automating business processes with AI is the follow-up to this article. If the data you need is still locked inside a package that will not cooperate, a move comes first: the data migration plan describes how to do that in a controlled way.
By the traces it leaves: data retyped from one system into another, spreadsheets kept alongside a software package, folders of files sorted by hand, recurring questions answered from scratch every time, and checks somebody runs each week on the side. Manual work rarely announces itself as a problem, because the people doing it have grown used to it.
Because estimates of where time goes deviate systematically from measurements. A METR study from July 2025 had experienced developers complete tasks with and without AI assistance: afterwards they believed they had been about twenty percent faster, while in reality they took nineteen percent longer. Without a baseline you cannot tell whether your automation delivers anything or merely feels like it does.
The one that occurs often, runs predictably, relies on data that is already digitally available, and where a mistake is recoverable. Not the task that irritates people most, and not the one with the highest theoretical saving. A task that happens three times a year and where an error costs you a customer is a poor first candidate, however annoying it is.
Then that is the outcome of the measurement, and a useful one. Work that depends on knowledge recorded nowhere cannot be automated until that knowledge is recorded somewhere. That is usually not an automation project but a data question: first make sure the information the work runs on exists somewhere complete and usable.
A few minutes a day for the participants, and half a day for whoever adds it up. That is deliberately low: a measurement week that feels like a project never gets finished. What you get back is a list of real numbers that serves for years as the baseline against which you test whether automation delivers.
Everything above comes down to one reversal. Instead of deciding what you are going to automate and then looking for evidence, measure for a week and let the outcome decide where you begin. That costs a week of elapsed time and half a day of arithmetic, and it is the difference between a project that can justify itself and a pilot that quietly disappears within a year.
And if the measurement week shows that most of your manual work stems from systems that do not talk to each other, you do not have an automation question but a data question. That is not a bad outcome. That is the right order.
In a free advisory call we walk through your processes, look at which manual work can be removed fastest and what that requires in your systems. On site or online.
Book a free advisory call