AI in operations

How to implement AI in your company and measure the results

A guide to answering what AI delivered at your company: a maturity diagnostic in 15 minutes, cost per accepted task in a nine-field spreadsheet and a four-week plan for the first process.

24 min read

Your company signed up for AI tools a year ago. In a board meeting, someone asks what that delivered. The answer comes as active users, messages sent, maybe hours "saved" according to an internal survey. No one in the room can say whether any process became cheaper, faster or more reliable.

To answer that question with a number, you need three things, and you can start using them on Monday: a maturity diagnostic that takes 15 minutes, a cost per accepted task calculation that fits in a nine-field spreadsheet, available for download, and a four-week plan to put the first process into operation.

The method combines what the Carnegie Mellon Software Engineering Institute, Microsoft, Gartner, AWS and NIST have published with what we see in the projects we run. Where the synthesis is ours, we say so.

If you have five minutes:

  • AI adoption is already high: 88% of organizations use it, according to Stanford's AI Index 2026. What is missing is operation. Agents appear in fewer than 10% of companies in most business functions.
  • Implementing AI means changing a process and being able to measure the change. Distributing licenses stays at the first level.
  • Assess maturity by dimension, with no single score. An average of 2.5 hides that the company knows how to build and does not know how to measure.
  • Decide how much the AI can do on its own in each process. Keeping the AI preparing the action for a person to approve can be the most mature choice.
  • Measure the cost per accepted task. In a typical example, the model costs R$ 2.30 per task and the accepted task costs R$ 7.22, because someone has to check and correct.
  • Before you start, write down the condition under which you shut the project down.

The four questions of the board meeting

A company that is mature in AI can answer, for each process where AI works:

  1. How much does an accepted task cost? Adding up the model, infrastructure and the time of the people who check and correct.
  2. How often is it right? Measured on cases with a known answer, which anyone can check.
  3. What can it do on its own? And who approves the rest, and who can switch it off.
  4. Which business indicator changed? Lead time, operating cost, revenue, margin or risk.

If you answer all four, you know where to invest next. If you can answer none, start with the diagnostic further down.

The case many people cited by half

In February 2024, Klarna announced that its AI customer service assistant had handled 2.3 million conversations in its first month. That was two thirds of chat inquiries, the equivalent of the work of 700 full-time agents. Resolution time dropped from 11 minutes to under 2, and the company estimated a US$ 40 million profit improvement for that year. It also reported customer satisfaction equal to that of human service.

The announcement became an AI ROI reference in presentations around the world. In May 2025, CEO Sebastian Siemiatkowski told Bloomberg that the company had given cost too much weight, and that quality fell. Klarna went back to hiring human agents.

The 2024 numbers measured usage, volume and speed: conversations handled, time per conversation, cost avoided. The 2025 correction is about something else, the share of the work that the customer and the company accepted as well done. Cost per executed task and cost per accepted task tell different stories, and the second one is the one that reaches the result.

The gap between using and operating

Stanford's AI Index 2026 records that 88% of surveyed organizations used AI in 2025 and 70% used generative AI in at least one function. Agent use appeared in single digits in almost every function. In the McKinsey survey of November 2025, 23% of companies say they are scaling agents somewhere in the organization, and in no function does that number exceed 10%.

In Brazil, the OECD report with BCG and INSEAD surveyed 167 companies in the state of São Paulo. Uncertainty about return appears for 38% of them, behind only privacy and security (44%). We have already written about this gap in AI has reached companies. The results have not.

The difference between using and implementing becomes clear when you look at where AI enters:

Where AI entersExampleWhat changes
ToolOne person uses ChatGPT to write an emailThat person's time
TaskThe team uses a standardized assistant to summarize contractsOne step, disconnected from the rest
FlowThe AI reads the request, looks up internal data and prepares a response for approvalOne stretch of the process
ProcessThe approved response is recorded in the system, with history, and the process indicator is trackedThe result of the process
OperationSeveral processes share integrations, access controls, tests and dashboardsHow the company works

The first two rows are usage. From the third row on, the AI reads, looks up, prepares, submits for approval and records, and each of those verbs depends on something the model does not deliver on its own: access to systems, permissions, a review rule and a record of what happened.

The 5 levels of AI maturity

LevelWhere the company isObservable signs
1. ExplorationPeople use AI on their ownIndividual licenses, personal use, no process depends on AI
2. ExperimentationUse cases start to appearPilots with an owner, isolated automations, results measured by impression
3. OperationAI enters a real processIntegration with systems, process owner, baseline and monthly metrics
4. ScaleProven cases are replicatedIntegrations, access controls and tests reused across processes
5. TransformationProcesses are redesigned around AIRoles and steps have changed, and business indicators reflect the change

The levels are our synthesis of the SEI, Microsoft, Gartner and AWS models. The mapping is in the method note at the end of the text.

The hardest jump is from level 2 to 3. That is where the pilot gets an owner, integration and measurement, and where most projects stop.

No company needs to take every process to level 5. A document triage with thousands of items per month justifies reaching scale. A quarterly report produced by two people may never need to go beyond experimentation.

Level 5 depends on redesigning the work, and there is evidence that this is where the money shows up. McKinsey tested 25 organizational attributes in its March 2025 survey, and workflow redesign was the one with the largest effect on the impact of generative AI on operating results. Only 21% of companies had fundamentally redesigned at least some workflows.

A 15-minute maturity diagnostic

A company is not at a single level. The Microsoft model recommends assessing each pillar separately, without turning the result into an overall score, and recording only what happens consistently. Isolated examples and intentions are left out. The diagnostic below follows that rule.

How to use it. Answer yes or no to each question, based on what exists today and can be shown. In each dimension, start at level 1 and go up one level for each "yes", stopping at the first "no". Level 5 is left out, because it depends on process redesign, which is assessed case by case.

1. Strategy and value

  • L2: Is there a list of AI use cases, with an owner for each one?
  • L3: Does each case in use have a business objective and a numeric target approved by the board?
  • L4: Are new cases prioritized by value and feasibility in a recurring ritual, such as a quarterly review?

2. Process

  • L2: Does any AI pilot run with real users, even outside the official system?
  • L3: Does at least one process have a design of what the AI does, what the person does and in what order?
  • L4: Has a second process gone into operation reusing the design of the first?

3. People and accountability

  • L2: Does someone from the business area follow each pilot?
  • L3: Does each process with AI have an owner in the business area who answers for the indicator, and were users trained on it?
  • L4: Is there a fixed team or role that helps new areas put AI into operation?

4. Data and technology

  • L2: Does the pilot work with real company data, even if exported by hand?
  • L3: Does the solution read from and write to the systems where the work happens, respecting each user's permissions?
  • L4: Are integrations, access control and usage logging shared across processes?

5. Governance and authority

  • L2: Is there a policy on what data can be sent to AI tools?
  • L3: For each process, is it written down what the AI does on its own, what requires approval and who can switch it off?
  • L4: Is every AI action logged and periodically reviewed by someone outside the team that built it?

6. Measurement

  • L2: Has someone collected examples of hits and misses from the pilot?
  • L3: Is there a baseline for the process before AI and a set of test cases with known answers?
  • L4: Does every change to the system go through the tests before reaching users, and are the indicators reviewed every month?

How to read the result

One possible profile:

DimensionLevel
Strategy and value3
Process2
People and accountability2
Data and technology4
Governance and authority3
Measurement1

The average would be 2.5 and would say nothing. The profile says the company knows how to build and has rules, but has not changed processes and cannot prove value. The next investment goes to a process owner and a baseline, before any new technology.

The general rule: invest in the lowest dimension that keeps the next process from reaching level 3. In the diagnostics we run, it is usually measurement or a process owner.

How much the AI is authorized to do

Jake Moffatt asked Air Canada's virtual assistant how the bereavement fare worked. The assistant answered that he could buy the ticket and request the discount later. The company's policy said the opposite. When the refund was denied, Air Canada argued before the British Columbia civil resolution tribunal that the assistant was a separate entity, responsible for its own acts. The tribunal rejected the argument in 2024: the assistant "is still just a part of Air Canada's website", and the company is responsible for everything on its website. It had to pay the fare difference.

The amount was small, and the rule that came out of it applies to any company: when the AI speaks or acts on the company's behalf, the company answers for it. That is why each process needs a defined authority level.

Authority levelThe AI canExample
A0. ConsultAnswer questions and summarizeSummary of a contract for the lawyer to read
A1. RecommendSuggest a decision, which the person makesIndication of which invoices look divergent
A2. PrepareLeave the action ready for someone to approveDraft reply to a customer, filled in with system data
A3. Execute with approvalExecute after a confirmation clickStore-to-store transfer request, released by the buyer
A4. Execute within limitsExecute alone, within defined rules and amountsReclassification of expenses below a set amount, with a record

When the system only recommends, an error costs the reviewer's time. When it executes, the error becomes a wrong order, an undue payment or a promise to the customer that the company will have to honor.

The rules we use:

  • Start at A1 or A2. Move up a level when the numbers show that errors have become rare and cheap. The team's sense of comfort tends to arrive before the numbers do.
  • Authority is per process. The same company can have the AI executing small expense classification on its own while only preparing customer replies.
  • Write down who switches it off. The NIST AI Risk Management Framework asks that human oversight processes be defined, assessed and documented. In a mid-sized company, that fits on one page per process.

How to measure AI results

The number of active users, messages and tokens consumed measures AI consumption. A token is the unit in which AI providers charge for the text the model reads and writes, a piece of a word. It explains the invoice and says little about the quality of the work. Useful measurement has four layers:

LayerQuestionExample indicators
UsageWho uses it and how often?Active users, processes with AI, frequency of use
OutputHow much work came out?Tasks completed, documents processed, responses generated
PerformanceDid the work get better and cheaper?Time per task, cost per task, acceptance rate, error rate, rework
ResultsDid the company gain anything?Lead time, operating cost, revenue, margin, conversion, risk

Usage and output show adoption. Value only appears in performance and results. Klarna's 2024 announcement was strong on output and speed, and the correction came from performance.

These measures are often confused, and you need all of them:

  • Quality test: does the AI do the task correctly? Hamel Husain and Shreya Shankar define these evaluations as measuring whether an AI system works for its users on realistic tasks and data. Example: 94% accuracy on 300 cases with known answers.
  • Operating indicator: how much does it cost and how long does it take? Example: R$ 7.22 and under 8 minutes per accepted task, counting both the AI and the review.
  • Business indicator: what changed in the company? Example: customer response time dropped from three days to one.

A system can pass 94% of the tests and not move any business indicator, because it entered a step that was not the bottleneck.

Do not ask, measure

In a controlled METR experiment in 2025, 16 experienced developers predicted that AI would make them 24% faster. After the work, they believed they had gained about 20%. The measurement showed that tasks with AI took 19% longer. A new round, in 2026, had a mixed result, with a selection bias acknowledged by the authors. In both cases, the perception of the people using it did not work as a measure.

The effect also changes with who uses it. In the study by Brynjolfsson, Li and Raymond published by NBER, with 5,179 support agents, AI increased issues resolved per hour by 14%. Among novices, 34%. Among the most experienced, almost nothing. A study's average does not predict the gain in your process, and only measuring it answers that.

Cost per accepted task

When AI starts executing work, the useful question stops being how much the model costs. It becomes how much it costs to produce one unit of work that the company can accept.

Cost per accepted task = (AI cost + cost of human time spent checking and correcting) ÷ accepted tasks

An example with numbers

The example is illustrative and serves to show the calculation.

Before. An analyst takes 45 minutes per task, at a cost of R$ 80 per hour. Each task costs R$ 60. At 1,000 tasks per month, that is R$ 60,000.

After. The AI executes each task in 4 minutes, at a monthly cost of R$ 2,300, or R$ 2.30 per execution. This is the number that usually goes into the project presentation. The full calculation includes the people:

ItemCalculationMonthly cost
AI (model, infrastructure, maintenance)1,000 executionsR$ 2,300
Checking the 870 tasks accepted first time (87%)2 min each, 29 hR$ 2,320
Correcting the 130 tasks that came back (13%)15 min each, 32.5 hR$ 2,600
Total for 1,000 accepted tasksR$ 7,220

The cost per accepted task is R$ 7.22: three times the cost of the model and 88% below the R$ 60 of the previous process. The gain is real, and the number that should reach the board is R$ 7.22.

Now the same process with a design error. The AI errs unpredictably, no one knows which tasks to check, and the team starts reviewing everything from scratch: 45 minutes per task, plus the AI cost. That is R$ 62,300 per month, or R$ 62.30 per accepted task. The AI became more expensive than the old process.

What separates the two scenarios is the quality tests, the review rule and the team's confidence to check by sampling. The company decides all of this without changing models.

The nine-field spreadsheet

To run the calculation on your process, you need nine numbers. The downloadable spreadsheet already has the formulas and the values from the example above. Replace them with yours. The Log sheet counts, task by task, how many were accepted, corrected or redone, and delivers fields 5, 7 and 9.

FieldWhere to get it
1. Tasks per monthSystem or the area's own tracking
2. Minutes per task, before AITiming a sample of 30 to 50 tasks
3. Team cost per hourSalary with charges, divided by hours worked
4. Monthly AI costModel invoice, infrastructure and prorated maintenance
5. % of tasks accepted first timeAcceptance record from the review
6. Review minutes per accepted taskTiming a sample
7. % of tasks correctedAcceptance record from the review
8. Correction minutes per taskTiming a sample
9. % of tasks redone from scratchAcceptance record from the review

Cost before = field 2 × field 3 ÷ 60. Cost after = field 4 divided by the tasks, plus the time spent checking, correcting and redoing, converted into reais. Fields 5, 7 and 9 require one simple thing: recording, for each task, whether it was accepted, corrected or redone.

Metrics for AI agents

When AI executes tasks, five rates complete the calculation:

MetricCalculation
Acceptance rateaccepted results ÷ generated results
Human intervention ratetasks that required intervention ÷ tasks executed
Autonomous completion ratetasks completed without intervention ÷ total tasks
Escalation ratetasks sent to a person to decide ÷ total
Rework ratetasks corrected after completion ÷ completed tasks

The goal is the level of autonomy that gives the lowest cost per accepted task within the risk the company is willing to take. In many processes, that point sits at A2 or A3.

The mistakes we see most

The SEI attributes a large share of adoption failures to misaligned expectations, poorly chosen applications and poorly executed implementation. In the diagnostics we run, those causes show up with recognizable symptoms:

What you seeWhat is usually missingWhat to do
The discussion is about which AI to buyA chosen processChoose the process first and leave the model for later
The pilot "worked", but no one can say by how muchBaselineTime 30 to 50 tasks the current way before changing anything
The pilot works in the demo and stalls on real filesData accessTest with real volume and real mess from the first week
People copy the AI answer and paste it into another systemIntegrationHave the AI read and write where the work happens
Technology built it, the business did not take ownershipProcess ownerName the owner in the business area before starting
The evaluation is "looks fine"Test casesBuild a set of real examples with the correct answer known
The team reviews everything, or reviews nothingReview ruleDefine the authority level and what gets checked by sampling
The monthly report shows logins and messagesPerformance measureReplace it with cost per accepted task and the process indicator

The most cited case of an agent without limits happened in July 2025. Jason Lemkin, founder of SaaStr, was testing Replit's coding agent while building an application. With the project under a code freeze, and after asking eleven times, in capital letters, that nothing be changed, he watched the agent delete the production database. The agent also claimed that restoring it was impossible, and the restore worked (The Register). Replit's CEO, Amjad Masad, called the episode unacceptable and announced the automatic separation of the development and production databases (The Register).

The failure was in authority. The agent executed commands in the real environment without approval, level A4 of the authority table, without the limits that level requires. The instruction "do not change anything" was a request, and no rule in the system enforced it. Anthropic, which worked with dozens of teams building agents, makes the same recommendation in its guide on agents: autonomy increases the cost and the risk of compounding errors, and therefore calls for extensive testing in a sandboxed environment, with appropriate guardrails.

The first process in four weeks

Before week 1: ask whether the problem needs AI. NIST points out that an AI system may not be the right solution for the task. The federal government's Guia Unificado de Inteligência Artificial starts with a prospecting phase, in which the challenge is identified and prioritized before any experiment. Sometimes a simple rule or an integration between systems solves it.

Week 1. Choose and measure. Choose a process with high volume, known rules and an owner in the business area. Time 30 to 50 tasks the current way and record volume, time, cost and error rate. Write down the target and the shutdown criterion, for example: "if the cost per accepted task does not fall below half of the current one, we stop".

Week 2. Define authority and tests. Decide the initial authority level (A1 or A2), who reviews and what counts as an accepted task. Set aside 50 to 100 real cases with the correct answer known. Every change to the system goes through them before reaching users.

Week 3. Put it in the flow. The solution reads from and writes to the systems where the work happens. In the first weeks, review 100% of outputs and record each one as accepted, corrected or redone. Each error found becomes a new test case.

Week 4. Run the numbers and decide. Fill in the nine-field spreadsheet and compare with the baseline. The possible decisions are expand (more volume, less review or more authority), adjust (go back to week 2 with the errors mapped) or switch off, if the week 1 criterion was not met.

Switching off a project that did not meet the criterion is also a result: in four weeks you know that process does not pay off. Without a written criterion, the same discovery usually takes a year and surfaces in a board meeting.

This cycle of planning, executing, checking and correcting is the same that ISO/IEC 42001, the AI management system standard, adopts as continual improvement.

Where to go next

The SEI, in the maturity model published with Accenture in 2026, defines AI maturity as the ability to build reliable and resilient systems, with rigorous engineering practice and governance tied to business results.

Run the diagnostic with your leadership team, pick a process and do the math in the spreadsheet. If you want help with the first process, Catech Forward puts engineers inside your operation, with a baseline, tests and cost per accepted task from the first week.

Method note

The five levels and the six dimensions are a CatechLabs synthesis. The SEI uses eight dimensions and Microsoft uses five pillars. We condensed these structures to fit a board-level diagnostic for a mid-sized company. The mapping between levels is approximate, because each model measures slightly different things:

CatechLabs levelSEI / Accenture (2026)Microsoft (2026)GartnerAWS
ExplorationExploratory100 InitialAwarenessEnvision
ExperimentationImplemented200 RepeatableActiveExperiment
OperationAligned300 DefinedOperationalLaunch
ScaleScaled400 CapableSystemicScale
TransformationFuture-Ready AI500 EfficientTransformational(none)

The authority levels (A0 to A4), the nine-field spreadsheet and the four-week plan come from our practice. The reference numbers (samples of 30 to 50 tasks, 50 to 100 test cases) are practical starting points, without a statistical basis.

References

Carnegie Mellon SEI and Accenture, AI Adoption Maturity Model v1.0 (2026)

Five-level, eight-dimension model aimed at taking organizations from experimentation to predictable results.

Microsoft, Agentic AI Adoption Maturity Model (2026)

Five levels and five pillars, with guidance to assess each pillar separately and based on evidence.

Gartner, AI Maturity Model Toolkit

Five-level scale, from Awareness to Transformational.

AWS, Generative AI Maturity Model

Four levels (Envision, Experiment, Launch, Scale) assessed by aspect.

NIST, AI Risk Management Framework 1.0 (2023)

AI risk management in four functions: govern, map, measure and manage.

ISO/IEC 42001:2023 and 23894:2023

AI management system and guidance on AI risk management.

Secretaria de Governo Digital, Guia Unificado de Inteligência Artificial para o Setor Público

AI project cycle in seven phases, from prospecting to decommissioning.

OECD, BCG and INSEAD, The Adoption of Artificial Intelligence in Firms (2025)

Survey of companies in the G7 and Brazil on AI adoption and barriers, including uncertainty about return.

Stanford HAI, AI Index 2026, economy chapter

Data on adoption of AI, generative AI and agents in organizations.

McKinsey, The state of AI: How organizations are rewiring to capture value (2025) and The state of AI in 2025

Workflow redesign as the attribute most associated with financial impact, and data on agents at scale.

Brynjolfsson, Li and Raymond, Generative AI at Work, NBER and Quarterly Journal of Economics

Causal effect of AI assistance on the productivity of 5,179 support agents.

METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (2025) and update (2026)

Controlled experiment in which the perceived gain diverged from the measured performance.

Hamel Husain and Shreya Shankar, LLM Evals FAQ

Practical reference on how to test the quality of AI systems.

Anthropic, Building effective agents (2024)

Recommendations from people who built agents with dozens of teams: start with the simplest solution, test in a sandboxed environment and define guardrails.

The Register, Vibe coding service Replit deleted user's production database and Replit makes vibe-y promise to stop its AI agents making vibe coding disasters (2025)

The incident in which an agent deleted a production database during a code freeze, and Replit's response.

Klarna, Klarna AI assistant handles two-thirds of customer service chats in its first month (2024), and Bloomberg, Klarna Turns From AI to Real Person Customer Service (2025)

The announcement of the assistant's results and the company's own later review of its strategy.

Civil Resolution Tribunal of British Columbia, Moffatt v. Air Canada, 2024 BCCRT 149

Decision that held the company responsible for the wrong information given by its virtual assistant.

Frequently asked questions

How do you implement AI in a company?
Pick a process with volume, an owner and a known cost. Measure how it works today, define what the AI can do on its own and what needs approval, put the solution inside the real workflow, and compare before and after using the same indicators. Handing out licenses is the start of usage. Implementation starts when a process changes and the change can be measured.
Where should an AI project start?
Start with the process. The choice of model comes later. The best first case has high volume, known rules, a clear owner and a current cost that someone already knows how to calculate. Before building, ask whether the problem needs AI at all: NIST points out that an AI system may not be the right solution for the task.
What is AI maturity?
It is the ability of an organization to turn artificial intelligence into business results in a way that is repeatable, measurable and controlled. It depends on strategy, processes, people, data and technology, governance and measurement. The number of tools purchased says little about it.
What are the levels of AI maturity?
Five: exploration (individual use), experimentation (isolated pilots), operation (AI inside a real process, with an owner and metrics), scale (proven cases replicated on a common base) and transformation (processes redesigned around what people, software and AI each do best). Assess each dimension separately, without a single score for the company.
How do you measure the ROI of AI?
Compare the total cost per accepted task with the cost of the same task before AI. Total cost includes the model, infrastructure, maintenance and the time of the people who review and correct the output. Then check whether the savings or freed-up capacity showed up in a business indicator such as lead time, margin, revenue or risk.
Which KPIs should you use for artificial intelligence?
Organize them in four layers: usage (who uses it and how often), output (how many tasks were completed), performance (time, cost, acceptance rate and rework per task) and results (lead time, operating cost, revenue, margin, risk). Usage and output indicate adoption. Only performance and results indicate value.
How do you know if a company is ready for AI agents?
An agent takes actions, so the question is whether the company can say what it may do, check what it did and stop it when it fails. That requires a documented process, controlled access to systems, a set of test cases with known answers, an owner for the outcome and a baseline to compare against.
How do you scale an AI pilot to production?
A pilot becomes operation when it gets an owner, integration with the systems where the work happens, human review rules, quality tests repeated at every change and metrics tracked every month. Scaling means repeating this in other processes while reusing what was already built: integrations, access controls, tests and dashboards.

Facing a similar challenge?

Talk to the founders