How to implement AI in your company and measure the results
A guide to answering what AI delivered at your company: a maturity diagnostic in 15 minutes, cost per accepted task in a nine-field spreadsheet and a four-week plan for the first process.
Your company signed up for AI tools a year ago. In a board meeting, someone asks what that delivered. The answer comes as active users, messages sent, maybe hours "saved" according to an internal survey. No one in the room can say whether any process became cheaper, faster or more reliable.
To answer that question with a number, you need three things, and you can start using them on Monday: a maturity diagnostic that takes 15 minutes, a cost per accepted task calculation that fits in a nine-field spreadsheet, available for download, and a four-week plan to put the first process into operation.
The method combines what the Carnegie Mellon Software Engineering Institute, Microsoft, Gartner, AWS and NIST have published with what we see in the projects we run. Where the synthesis is ours, we say so.
If you have five minutes:
- AI adoption is already high: 88% of organizations use it, according to Stanford's AI Index 2026. What is missing is operation. Agents appear in fewer than 10% of companies in most business functions.
- Implementing AI means changing a process and being able to measure the change. Distributing licenses stays at the first level.
- Assess maturity by dimension, with no single score. An average of 2.5 hides that the company knows how to build and does not know how to measure.
- Decide how much the AI can do on its own in each process. Keeping the AI preparing the action for a person to approve can be the most mature choice.
- Measure the cost per accepted task. In a typical example, the model costs R$ 2.30 per task and the accepted task costs R$ 7.22, because someone has to check and correct.
- Before you start, write down the condition under which you shut the project down.
The four questions of the board meeting
A company that is mature in AI can answer, for each process where AI works:
- How much does an accepted task cost? Adding up the model, infrastructure and the time of the people who check and correct.
- How often is it right? Measured on cases with a known answer, which anyone can check.
- What can it do on its own? And who approves the rest, and who can switch it off.
- Which business indicator changed? Lead time, operating cost, revenue, margin or risk.
If you answer all four, you know where to invest next. If you can answer none, start with the diagnostic further down.
The case many people cited by half
In February 2024, Klarna announced that its AI customer service assistant had handled 2.3 million conversations in its first month. That was two thirds of chat inquiries, the equivalent of the work of 700 full-time agents. Resolution time dropped from 11 minutes to under 2, and the company estimated a US$ 40 million profit improvement for that year. It also reported customer satisfaction equal to that of human service.
The announcement became an AI ROI reference in presentations around the world. In May 2025, CEO Sebastian Siemiatkowski told Bloomberg that the company had given cost too much weight, and that quality fell. Klarna went back to hiring human agents.
The 2024 numbers measured usage, volume and speed: conversations handled, time per conversation, cost avoided. The 2025 correction is about something else, the share of the work that the customer and the company accepted as well done. Cost per executed task and cost per accepted task tell different stories, and the second one is the one that reaches the result.
The gap between using and operating
Stanford's AI Index 2026 records that 88% of surveyed organizations used AI in 2025 and 70% used generative AI in at least one function. Agent use appeared in single digits in almost every function. In the McKinsey survey of November 2025, 23% of companies say they are scaling agents somewhere in the organization, and in no function does that number exceed 10%.
In Brazil, the OECD report with BCG and INSEAD surveyed 167 companies in the state of São Paulo. Uncertainty about return appears for 38% of them, behind only privacy and security (44%). We have already written about this gap in AI has reached companies. The results have not.
The difference between using and implementing becomes clear when you look at where AI enters:
| Where AI enters | Example | What changes |
|---|---|---|
| Tool | One person uses ChatGPT to write an email | That person's time |
| Task | The team uses a standardized assistant to summarize contracts | One step, disconnected from the rest |
| Flow | The AI reads the request, looks up internal data and prepares a response for approval | One stretch of the process |
| Process | The approved response is recorded in the system, with history, and the process indicator is tracked | The result of the process |
| Operation | Several processes share integrations, access controls, tests and dashboards | How the company works |
The first two rows are usage. From the third row on, the AI reads, looks up, prepares, submits for approval and records, and each of those verbs depends on something the model does not deliver on its own: access to systems, permissions, a review rule and a record of what happened.
The 5 levels of AI maturity
| Level | Where the company is | Observable signs |
|---|---|---|
| 1. Exploration | People use AI on their own | Individual licenses, personal use, no process depends on AI |
| 2. Experimentation | Use cases start to appear | Pilots with an owner, isolated automations, results measured by impression |
| 3. Operation | AI enters a real process | Integration with systems, process owner, baseline and monthly metrics |
| 4. Scale | Proven cases are replicated | Integrations, access controls and tests reused across processes |
| 5. Transformation | Processes are redesigned around AI | Roles and steps have changed, and business indicators reflect the change |
The levels are our synthesis of the SEI, Microsoft, Gartner and AWS models. The mapping is in the method note at the end of the text.
The hardest jump is from level 2 to 3. That is where the pilot gets an owner, integration and measurement, and where most projects stop.
No company needs to take every process to level 5. A document triage with thousands of items per month justifies reaching scale. A quarterly report produced by two people may never need to go beyond experimentation.
Level 5 depends on redesigning the work, and there is evidence that this is where the money shows up. McKinsey tested 25 organizational attributes in its March 2025 survey, and workflow redesign was the one with the largest effect on the impact of generative AI on operating results. Only 21% of companies had fundamentally redesigned at least some workflows.
A 15-minute maturity diagnostic
A company is not at a single level. The Microsoft model recommends assessing each pillar separately, without turning the result into an overall score, and recording only what happens consistently. Isolated examples and intentions are left out. The diagnostic below follows that rule.
How to use it. Answer yes or no to each question, based on what exists today and can be shown. In each dimension, start at level 1 and go up one level for each "yes", stopping at the first "no". Level 5 is left out, because it depends on process redesign, which is assessed case by case.
1. Strategy and value
- L2: Is there a list of AI use cases, with an owner for each one?
- L3: Does each case in use have a business objective and a numeric target approved by the board?
- L4: Are new cases prioritized by value and feasibility in a recurring ritual, such as a quarterly review?
2. Process
- L2: Does any AI pilot run with real users, even outside the official system?
- L3: Does at least one process have a design of what the AI does, what the person does and in what order?
- L4: Has a second process gone into operation reusing the design of the first?
3. People and accountability
- L2: Does someone from the business area follow each pilot?
- L3: Does each process with AI have an owner in the business area who answers for the indicator, and were users trained on it?
- L4: Is there a fixed team or role that helps new areas put AI into operation?
4. Data and technology
- L2: Does the pilot work with real company data, even if exported by hand?
- L3: Does the solution read from and write to the systems where the work happens, respecting each user's permissions?
- L4: Are integrations, access control and usage logging shared across processes?
5. Governance and authority
- L2: Is there a policy on what data can be sent to AI tools?
- L3: For each process, is it written down what the AI does on its own, what requires approval and who can switch it off?
- L4: Is every AI action logged and periodically reviewed by someone outside the team that built it?
6. Measurement
- L2: Has someone collected examples of hits and misses from the pilot?
- L3: Is there a baseline for the process before AI and a set of test cases with known answers?
- L4: Does every change to the system go through the tests before reaching users, and are the indicators reviewed every month?
How to read the result
One possible profile:
| Dimension | Level |
|---|---|
| Strategy and value | 3 |
| Process | 2 |
| People and accountability | 2 |
| Data and technology | 4 |
| Governance and authority | 3 |
| Measurement | 1 |
The average would be 2.5 and would say nothing. The profile says the company knows how to build and has rules, but has not changed processes and cannot prove value. The next investment goes to a process owner and a baseline, before any new technology.
The general rule: invest in the lowest dimension that keeps the next process from reaching level 3. In the diagnostics we run, it is usually measurement or a process owner.
How much the AI is authorized to do
Jake Moffatt asked Air Canada's virtual assistant how the bereavement fare worked. The assistant answered that he could buy the ticket and request the discount later. The company's policy said the opposite. When the refund was denied, Air Canada argued before the British Columbia civil resolution tribunal that the assistant was a separate entity, responsible for its own acts. The tribunal rejected the argument in 2024: the assistant "is still just a part of Air Canada's website", and the company is responsible for everything on its website. It had to pay the fare difference.
The amount was small, and the rule that came out of it applies to any company: when the AI speaks or acts on the company's behalf, the company answers for it. That is why each process needs a defined authority level.
| Authority level | The AI can | Example |
|---|---|---|
| A0. Consult | Answer questions and summarize | Summary of a contract for the lawyer to read |
| A1. Recommend | Suggest a decision, which the person makes | Indication of which invoices look divergent |
| A2. Prepare | Leave the action ready for someone to approve | Draft reply to a customer, filled in with system data |
| A3. Execute with approval | Execute after a confirmation click | Store-to-store transfer request, released by the buyer |
| A4. Execute within limits | Execute alone, within defined rules and amounts | Reclassification of expenses below a set amount, with a record |
When the system only recommends, an error costs the reviewer's time. When it executes, the error becomes a wrong order, an undue payment or a promise to the customer that the company will have to honor.
The rules we use:
- Start at A1 or A2. Move up a level when the numbers show that errors have become rare and cheap. The team's sense of comfort tends to arrive before the numbers do.
- Authority is per process. The same company can have the AI executing small expense classification on its own while only preparing customer replies.
- Write down who switches it off. The NIST AI Risk Management Framework asks that human oversight processes be defined, assessed and documented. In a mid-sized company, that fits on one page per process.
How to measure AI results
The number of active users, messages and tokens consumed measures AI consumption. A token is the unit in which AI providers charge for the text the model reads and writes, a piece of a word. It explains the invoice and says little about the quality of the work. Useful measurement has four layers:
| Layer | Question | Example indicators |
|---|---|---|
| Usage | Who uses it and how often? | Active users, processes with AI, frequency of use |
| Output | How much work came out? | Tasks completed, documents processed, responses generated |
| Performance | Did the work get better and cheaper? | Time per task, cost per task, acceptance rate, error rate, rework |
| Results | Did the company gain anything? | Lead time, operating cost, revenue, margin, conversion, risk |
Usage and output show adoption. Value only appears in performance and results. Klarna's 2024 announcement was strong on output and speed, and the correction came from performance.
These measures are often confused, and you need all of them:
- Quality test: does the AI do the task correctly? Hamel Husain and Shreya Shankar define these evaluations as measuring whether an AI system works for its users on realistic tasks and data. Example: 94% accuracy on 300 cases with known answers.
- Operating indicator: how much does it cost and how long does it take? Example: R$ 7.22 and under 8 minutes per accepted task, counting both the AI and the review.
- Business indicator: what changed in the company? Example: customer response time dropped from three days to one.
A system can pass 94% of the tests and not move any business indicator, because it entered a step that was not the bottleneck.
Do not ask, measure
In a controlled METR experiment in 2025, 16 experienced developers predicted that AI would make them 24% faster. After the work, they believed they had gained about 20%. The measurement showed that tasks with AI took 19% longer. A new round, in 2026, had a mixed result, with a selection bias acknowledged by the authors. In both cases, the perception of the people using it did not work as a measure.
The effect also changes with who uses it. In the study by Brynjolfsson, Li and Raymond published by NBER, with 5,179 support agents, AI increased issues resolved per hour by 14%. Among novices, 34%. Among the most experienced, almost nothing. A study's average does not predict the gain in your process, and only measuring it answers that.
Cost per accepted task
When AI starts executing work, the useful question stops being how much the model costs. It becomes how much it costs to produce one unit of work that the company can accept.
Cost per accepted task = (AI cost + cost of human time spent checking and correcting) ÷ accepted tasks
An example with numbers
The example is illustrative and serves to show the calculation.
Before. An analyst takes 45 minutes per task, at a cost of R$ 80 per hour. Each task costs R$ 60. At 1,000 tasks per month, that is R$ 60,000.
After. The AI executes each task in 4 minutes, at a monthly cost of R$ 2,300, or R$ 2.30 per execution. This is the number that usually goes into the project presentation. The full calculation includes the people:
| Item | Calculation | Monthly cost |
|---|---|---|
| AI (model, infrastructure, maintenance) | 1,000 executions | R$ 2,300 |
| Checking the 870 tasks accepted first time (87%) | 2 min each, 29 h | R$ 2,320 |
| Correcting the 130 tasks that came back (13%) | 15 min each, 32.5 h | R$ 2,600 |
| Total for 1,000 accepted tasks | R$ 7,220 |
The cost per accepted task is R$ 7.22: three times the cost of the model and 88% below the R$ 60 of the previous process. The gain is real, and the number that should reach the board is R$ 7.22.
Now the same process with a design error. The AI errs unpredictably, no one knows which tasks to check, and the team starts reviewing everything from scratch: 45 minutes per task, plus the AI cost. That is R$ 62,300 per month, or R$ 62.30 per accepted task. The AI became more expensive than the old process.
What separates the two scenarios is the quality tests, the review rule and the team's confidence to check by sampling. The company decides all of this without changing models.
The nine-field spreadsheet
To run the calculation on your process, you need nine numbers. The downloadable spreadsheet already has the formulas and the values from the example above. Replace them with yours. The Log sheet counts, task by task, how many were accepted, corrected or redone, and delivers fields 5, 7 and 9.
| Field | Where to get it |
|---|---|
| 1. Tasks per month | System or the area's own tracking |
| 2. Minutes per task, before AI | Timing a sample of 30 to 50 tasks |
| 3. Team cost per hour | Salary with charges, divided by hours worked |
| 4. Monthly AI cost | Model invoice, infrastructure and prorated maintenance |
| 5. % of tasks accepted first time | Acceptance record from the review |
| 6. Review minutes per accepted task | Timing a sample |
| 7. % of tasks corrected | Acceptance record from the review |
| 8. Correction minutes per task | Timing a sample |
| 9. % of tasks redone from scratch | Acceptance record from the review |
Cost before = field 2 × field 3 ÷ 60. Cost after = field 4 divided by the tasks, plus the time spent checking, correcting and redoing, converted into reais. Fields 5, 7 and 9 require one simple thing: recording, for each task, whether it was accepted, corrected or redone.
Metrics for AI agents
When AI executes tasks, five rates complete the calculation:
| Metric | Calculation |
|---|---|
| Acceptance rate | accepted results ÷ generated results |
| Human intervention rate | tasks that required intervention ÷ tasks executed |
| Autonomous completion rate | tasks completed without intervention ÷ total tasks |
| Escalation rate | tasks sent to a person to decide ÷ total |
| Rework rate | tasks corrected after completion ÷ completed tasks |
The goal is the level of autonomy that gives the lowest cost per accepted task within the risk the company is willing to take. In many processes, that point sits at A2 or A3.
The mistakes we see most
The SEI attributes a large share of adoption failures to misaligned expectations, poorly chosen applications and poorly executed implementation. In the diagnostics we run, those causes show up with recognizable symptoms:
| What you see | What is usually missing | What to do |
|---|---|---|
| The discussion is about which AI to buy | A chosen process | Choose the process first and leave the model for later |
| The pilot "worked", but no one can say by how much | Baseline | Time 30 to 50 tasks the current way before changing anything |
| The pilot works in the demo and stalls on real files | Data access | Test with real volume and real mess from the first week |
| People copy the AI answer and paste it into another system | Integration | Have the AI read and write where the work happens |
| Technology built it, the business did not take ownership | Process owner | Name the owner in the business area before starting |
| The evaluation is "looks fine" | Test cases | Build a set of real examples with the correct answer known |
| The team reviews everything, or reviews nothing | Review rule | Define the authority level and what gets checked by sampling |
| The monthly report shows logins and messages | Performance measure | Replace it with cost per accepted task and the process indicator |
The most cited case of an agent without limits happened in July 2025. Jason Lemkin, founder of SaaStr, was testing Replit's coding agent while building an application. With the project under a code freeze, and after asking eleven times, in capital letters, that nothing be changed, he watched the agent delete the production database. The agent also claimed that restoring it was impossible, and the restore worked (The Register). Replit's CEO, Amjad Masad, called the episode unacceptable and announced the automatic separation of the development and production databases (The Register).
The failure was in authority. The agent executed commands in the real environment without approval, level A4 of the authority table, without the limits that level requires. The instruction "do not change anything" was a request, and no rule in the system enforced it. Anthropic, which worked with dozens of teams building agents, makes the same recommendation in its guide on agents: autonomy increases the cost and the risk of compounding errors, and therefore calls for extensive testing in a sandboxed environment, with appropriate guardrails.
The first process in four weeks
Before week 1: ask whether the problem needs AI. NIST points out that an AI system may not be the right solution for the task. The federal government's Guia Unificado de Inteligência Artificial starts with a prospecting phase, in which the challenge is identified and prioritized before any experiment. Sometimes a simple rule or an integration between systems solves it.
Week 1. Choose and measure. Choose a process with high volume, known rules and an owner in the business area. Time 30 to 50 tasks the current way and record volume, time, cost and error rate. Write down the target and the shutdown criterion, for example: "if the cost per accepted task does not fall below half of the current one, we stop".
Week 2. Define authority and tests. Decide the initial authority level (A1 or A2), who reviews and what counts as an accepted task. Set aside 50 to 100 real cases with the correct answer known. Every change to the system goes through them before reaching users.
Week 3. Put it in the flow. The solution reads from and writes to the systems where the work happens. In the first weeks, review 100% of outputs and record each one as accepted, corrected or redone. Each error found becomes a new test case.
Week 4. Run the numbers and decide. Fill in the nine-field spreadsheet and compare with the baseline. The possible decisions are expand (more volume, less review or more authority), adjust (go back to week 2 with the errors mapped) or switch off, if the week 1 criterion was not met.
Switching off a project that did not meet the criterion is also a result: in four weeks you know that process does not pay off. Without a written criterion, the same discovery usually takes a year and surfaces in a board meeting.
This cycle of planning, executing, checking and correcting is the same that ISO/IEC 42001, the AI management system standard, adopts as continual improvement.
Where to go next
The SEI, in the maturity model published with Accenture in 2026, defines AI maturity as the ability to build reliable and resilient systems, with rigorous engineering practice and governance tied to business results.
Run the diagnostic with your leadership team, pick a process and do the math in the spreadsheet. If you want help with the first process, Catech Forward puts engineers inside your operation, with a baseline, tests and cost per accepted task from the first week.
Method note
The five levels and the six dimensions are a CatechLabs synthesis. The SEI uses eight dimensions and Microsoft uses five pillars. We condensed these structures to fit a board-level diagnostic for a mid-sized company. The mapping between levels is approximate, because each model measures slightly different things:
| CatechLabs level | SEI / Accenture (2026) | Microsoft (2026) | Gartner | AWS |
|---|---|---|---|---|
| Exploration | Exploratory | 100 Initial | Awareness | Envision |
| Experimentation | Implemented | 200 Repeatable | Active | Experiment |
| Operation | Aligned | 300 Defined | Operational | Launch |
| Scale | Scaled | 400 Capable | Systemic | Scale |
| Transformation | Future-Ready AI | 500 Efficient | Transformational | (none) |
The authority levels (A0 to A4), the nine-field spreadsheet and the four-week plan come from our practice. The reference numbers (samples of 30 to 50 tasks, 50 to 100 test cases) are practical starting points, without a statistical basis.
References
Carnegie Mellon SEI and Accenture, AI Adoption Maturity Model v1.0 (2026)
Five-level, eight-dimension model aimed at taking organizations from experimentation to predictable results.
Microsoft, Agentic AI Adoption Maturity Model (2026)
Five levels and five pillars, with guidance to assess each pillar separately and based on evidence.
Gartner, AI Maturity Model Toolkit
Five-level scale, from Awareness to Transformational.
AWS, Generative AI Maturity Model
Four levels (Envision, Experiment, Launch, Scale) assessed by aspect.
NIST, AI Risk Management Framework 1.0 (2023)
AI risk management in four functions: govern, map, measure and manage.
ISO/IEC 42001:2023 and 23894:2023
AI management system and guidance on AI risk management.
Secretaria de Governo Digital, Guia Unificado de Inteligência Artificial para o Setor Público
AI project cycle in seven phases, from prospecting to decommissioning.
OECD, BCG and INSEAD, The Adoption of Artificial Intelligence in Firms (2025)
Survey of companies in the G7 and Brazil on AI adoption and barriers, including uncertainty about return.
Stanford HAI, AI Index 2026, economy chapter
Data on adoption of AI, generative AI and agents in organizations.
McKinsey, The state of AI: How organizations are rewiring to capture value (2025) and The state of AI in 2025
Workflow redesign as the attribute most associated with financial impact, and data on agents at scale.
Brynjolfsson, Li and Raymond, Generative AI at Work, NBER and Quarterly Journal of Economics
Causal effect of AI assistance on the productivity of 5,179 support agents.
METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (2025) and update (2026)
Controlled experiment in which the perceived gain diverged from the measured performance.
Hamel Husain and Shreya Shankar, LLM Evals FAQ
Practical reference on how to test the quality of AI systems.
Anthropic, Building effective agents (2024)
Recommendations from people who built agents with dozens of teams: start with the simplest solution, test in a sandboxed environment and define guardrails.
The Register, Vibe coding service Replit deleted user's production database and Replit makes vibe-y promise to stop its AI agents making vibe coding disasters (2025)
The incident in which an agent deleted a production database during a code freeze, and Replit's response.
Klarna, Klarna AI assistant handles two-thirds of customer service chats in its first month (2024), and Bloomberg, Klarna Turns From AI to Real Person Customer Service (2025)
The announcement of the assistant's results and the company's own later review of its strategy.
Civil Resolution Tribunal of British Columbia, Moffatt v. Air Canada, 2024 BCCRT 149
Decision that held the company responsible for the wrong information given by its virtual assistant.
Frequently asked questions
- How do you implement AI in a company?
- Pick a process with volume, an owner and a known cost. Measure how it works today, define what the AI can do on its own and what needs approval, put the solution inside the real workflow, and compare before and after using the same indicators. Handing out licenses is the start of usage. Implementation starts when a process changes and the change can be measured.
- Where should an AI project start?
- Start with the process. The choice of model comes later. The best first case has high volume, known rules, a clear owner and a current cost that someone already knows how to calculate. Before building, ask whether the problem needs AI at all: NIST points out that an AI system may not be the right solution for the task.
- What is AI maturity?
- It is the ability of an organization to turn artificial intelligence into business results in a way that is repeatable, measurable and controlled. It depends on strategy, processes, people, data and technology, governance and measurement. The number of tools purchased says little about it.
- What are the levels of AI maturity?
- Five: exploration (individual use), experimentation (isolated pilots), operation (AI inside a real process, with an owner and metrics), scale (proven cases replicated on a common base) and transformation (processes redesigned around what people, software and AI each do best). Assess each dimension separately, without a single score for the company.
- How do you measure the ROI of AI?
- Compare the total cost per accepted task with the cost of the same task before AI. Total cost includes the model, infrastructure, maintenance and the time of the people who review and correct the output. Then check whether the savings or freed-up capacity showed up in a business indicator such as lead time, margin, revenue or risk.
- Which KPIs should you use for artificial intelligence?
- Organize them in four layers: usage (who uses it and how often), output (how many tasks were completed), performance (time, cost, acceptance rate and rework per task) and results (lead time, operating cost, revenue, margin, risk). Usage and output indicate adoption. Only performance and results indicate value.
- How do you know if a company is ready for AI agents?
- An agent takes actions, so the question is whether the company can say what it may do, check what it did and stop it when it fails. That requires a documented process, controlled access to systems, a set of test cases with known answers, an owner for the outcome and a baseline to compare against.
- How do you scale an AI pilot to production?
- A pilot becomes operation when it gets an owner, integration with the systems where the work happens, human review rules, quality tests repeated at every change and metrics tracked every month. Scaling means repeating this in other processes while reusing what was already built: integrations, access controls, tests and dashboards.