Most AI automation projects fail the ROI test — not because the technology doesn't work, but because no one defined what "working" actually means before they started. Six months in, the agent is running, the team is vaguely positive about it, and the CFO asks for numbers. Nobody has them.
We've deployed agents for dozens of SMEs. Every engagement now starts with the same question: what is the measurable value this agent must produce in 90 days to justify its existence? If you can't answer that before you build, you're not ready to build.
Here is the framework we use — the metrics, the formulas, and the benchmarks that separate a good deployment from an expensive experiment.
The 5 ROI Metrics That Actually Matter
Ignore vanity metrics like "number of tasks automated" or "percentage of workflow covered." These are means, not ends. The metrics that matter are the ones that show up in your P&L.
Hours per week that the agent handles tasks previously done by humans. Multiply by fully-loaded hourly cost to get weekly savings. This is the most direct and defensible ROI signal.
Agents don't get tired, distracted, or skip steps. Measure your pre-automation error rate on the target process, compare to post-automation. For document processing and data entry, typical reduction is 70–90%.
Time between trigger (lead arrives, invoice received, support request submitted) and first action. Agents compress this from hours to seconds. For sales workflows, faster first-touch typically increases conversion 10–30%.
How much more volume can your team handle with the same headcount? An agent that processes 200 invoices a day instead of 40 is a 5× throughput multiplier — without hiring. This is especially valuable during growth phases or seasonal spikes.
The most underestimated metric. When your senior account manager stops manually updating the CRM, what do they do instead? If the answer is higher-value activities — upsell calls, client relationships, product thinking — the ROI compounds beyond the direct time saved.
The ROI Calculation Formula
Once you have the five metrics, the aggregate ROI calculation is straightforward:
+ Response Time Value + Throughput Gain) × 12
Annual Cost = Agent Infrastructure + Management Fee + API Costs
ROI % = ((Annual Savings − Annual Cost) ÷ Annual Cost) × 100
A well-deployed agent at an SME typically runs €400–900/month in total cost. Annual savings of €3,000–8,000/month are common for the right process. That's an ROI of 300–900% in the first year.
Benchmarks by Process Type
Not all automation has the same payoff profile. Here are the benchmarks we see across our client base:
| Process | Typical Time Saved | Error Reduction | Payback Period | Rating |
|---|---|---|---|---|
| Invoice processing | 3–6 hrs/week | 80–95% | 2–4 months | Excellent |
| Lead qualification & routing | 5–10 hrs/week | 60–80% | 1–3 months | Excellent |
| Customer support triage | 8–15 hrs/week | 40–70% | 2–5 months | Excellent |
| Report generation | 2–4 hrs/week | 70–90% | 3–6 months | Good |
| Contract review (assist) | 2–5 hrs/week | 30–50% | 4–8 months | Good |
| Complex creative work | 0–1 hrs/week | <20% | >12 months | Avoid |
| Novel decision-making | 0 hrs/week | N/A | Never | Not viable |
The Measurement Framework: Before, During, After
ROI measurement is only possible if you establish baselines before deployment. This is where most companies fail — they build the agent, it goes live, and then they try to reconstruct what things were like before. You can't do it accurately.
Phase 1 — Before (Week −2 to −1)
Track manually for two weeks the exact process you're automating. Log: time spent, volume processed, errors found, time-to-first-action. These become your baseline. Export everything to a spreadsheet you can reference after go-live.
Phase 2 — During (Weeks 1–4 post-launch)
Run the agent alongside the manual process for the first two weeks if possible — shadow mode. Compare outputs. This surfaces hallucinations, edge cases, and schema mismatches before you fully hand over. Track the same metrics as your baseline.
Phase 3 — After (Month 2 onwards)
Monthly reporting on the five metrics. Look for: drift (is accuracy declining?), coverage gaps (what is the agent passing back to humans?), and opportunity creep (is the team actually using the freed time productively?). The best deployments improve over 6 months; bad ones decay.
What Good Looks Like at 90 Days
A deployment we consider successful at the 90-day mark hits these thresholds:
- Error rate ≤ 5% (vs. human baseline, which is typically 8–15% on repetitive tasks)
- Coverage rate ≥ 85% (the agent handles 85%+ of cases without human escalation)
- ROI already positive — the agent has paid for itself in direct labour savings
- Team sentiment neutral to positive — no active resistance from the people whose work changed
- At least one expansion use case identified — the team sees where to go next
If you're at 90 days and none of these are true, you have a problem worth diagnosing — not doubling down on.
The One Number That Matters Most
If you have to reduce everything to a single metric: fully-loaded cost per transaction. Before automation: how much does it cost your business to process one invoice, qualify one lead, handle one support ticket? After automation: what does that same transaction cost?
If that number has dropped by 60% or more within 90 days, you've built something defensible. If it hasn't moved, you've automated the wrong thing — or implemented it wrong.
The companies getting durable ROI from AI agents aren't doing anything exotic. They're picking high-volume, well-defined processes, measuring rigorously, and iterating when something isn't working. The math always catches up to whether you did it right.