The organisations that get value from agentic AI in the first quarter are not the ones with the best model or the biggest budget. They are the ones that chose a first use case the operation could feel, the finance team could sign, and the agent could actually do. This guide is the method we use to find that use case with a customer, written down.
Start with a list of jobs, not a list of ideas
Ask each function head one question: where does one person reconcile several systems by hand every week? The answers are the candidate list, and they arrive in the right vocabulary: the claim approver matching invoices to check sheets, the parts manager deciding reorder quantities from a spreadsheet, the service coordinator working out who is due a visit, the credit controller chasing from an ageing report. Ideas that start "what if AI could…" go on a different list, for later.
Write each candidate in the same four-box form — the person, their problems, the flow you want, the impact you would measure. We have described why that format works; here it is the unit of comparison.
Four questions that sort the list
How much of the flow is arithmetic? Ageing a receivable, netting a stock pipeline, checking a claim against policy, matching a contract to a schedule: these are rules over your own records. They can run alone, tonight, deterministically, and every output is checkable. The larger this share, the sooner the use case pays back and the less it depends on the model being clever. Rank candidates by it.
Where is the human click? Every flow has a step where money moves or the business is committed. Find it. If the flow ends there — a validated claim in a queue, a drafted order awaiting approval, a call list ready for the morning — the use case is governable and the finance team will sign it. If the proposal needs the agent to take that step alone to deliver its value, move it down the list. It may be right later; it is not the first one.
Is the data already in one place? A use case over records the platform already holds — claims, stock, contracts, invoices — can be configured. One that depends on a source you do not yet integrate, telematics or an external portal, needs the connector first. Not a reason to drop it, but a reason to sequence it second.
Will someone feel it on Monday? The queue that is full when the approver logs in. The reorder that arrives with its arithmetic shown. The call list that is ready. The first use case has to change one person's Monday visibly, because that person becomes the programme's advocate. Back-office savings that appear in a quarterly report do not recruit anyone.
The autonomy ladder, and what to expect at each rung
Every agentic flow sits somewhere on a short ladder, and knowing which rung you are buying sets expectations correctly.
- Reads. The agent answers questions over your records, as the asking user, with citations. Zero risk, immediate value, and the foundation for everything above it. If your people cannot yet ask the business a question and get a grounded answer, start here for a fortnight.
- Computes. Deterministic engines and nudges run on schedules: the overdue receivable, the scheme gap, the purchase line the pipeline already covers, the vehicle falling due. They explain themselves and wait in a queue. This rung runs alone, safely, because the arithmetic is inspectable.
- Drafts. The agent writes the reminder, assembles the claim, proposes the parts list, generates the suggested order. Nothing persists until a person says yes. This is where most first use cases land, and where the model earns its place.
- Acts with approval. One click applies the draft — sends the reminder, places the order, approves the clean claims in bulk. The click is the control; the audit trail is the record.
- Acts alone. Reserved for the deterministic rung. Engines that compute from rules run unattended; agents that generate do not commit the business without a human on the trigger. Vendors who promise this rung for judgement work are promising something your auditor will not accept.
Three starting points that reliably pay back
Claims. The deterministic share is enormous — eligibility, policy, consistency are rules — and the human click is natural: approve the queue. The channel feels faster reimbursement within the first cycle.
Replenishment. Suggested order quantities with the arithmetic shown, and a nudge that catches the line the pipeline already covers. The parts manager feels it on the first order, and the cash impact is visible in a month.
Service leads. Contracts and schedules already say who is due. Generating the reminder leads overnight and ranking the call list is pure computation; drafting the outreach is the model's small, safe contribution. Service revenue moves in a quarter.
All three share a property worth noticing: the agent's job is to fill a queue a person already owns. Nobody's role changes; their Monday does.
What to measure, and when
Decide the measure before the pilot, and keep numbers off the proposal until there is a baseline. Claims: cycle time from submission to decision, and the share approved without manual review. Replenishment: stockouts and stock days. Service leads: leads generated, contacted and booked. Measure the four weeks before and the four weeks after. Put the number on the slide then, and not before — a claimed percentage without a baseline costs more credibility than it buys.
What to leave for the second wave
Anything whose data is not yet integrated. Anything whose value depends on the agent acting alone on judgement. Anything that starts with a channel — "an AI on the phone" — rather than a job; the plumbing under such things is worth building first, and we have written about what that plumbing is. And anything nobody in the room could describe in the person's own words. Those are not bad use cases. They are second ones, and by then the first has taught the organisation how to run an agent.
The use cases xMatix runs today are written up in the same four-box form at agentic AI use cases; hold your candidate list against them and see which are already configuration.
Related: Agentic AI use cases · Rolling out AI nudges your team will trust · What we let agents do alone
