How to Measure AI Impact: A Practical Guide to Department-Specific KPIs and Metrics

Column
How to Measure AI Impact: A Practical Guide to Department-Specific KPIs and Metrics

Introduction

This article is aimed at heads of digital transformation, corporate planning, and individual business units who are uncertain about how to measure the outcomes of AI initiatives. It sets out how to design AI performance metrics on a department-by-department basis, and how to connect efficiency indicators, quality indicators, and revenue contribution into a coherent framework. The goal is to move beyond vague, qualitative assessments such as "it was useful" — and instead establish a position where leading and lagging indicators enable informed, ongoing decisions about whether to continue, adjust, or scale AI adoption.

Tatsuya Ito

Tatsuya Ito

Artificial Intelligence Consultant

company-icon

Third Scope Inc.

Born in 1985 and originally from Mie Prefecture, Japan. In 2012, he joined an AR startup in Hong Kong as an engineer. Since then, he has been involved in new business development and AI service launches at several AI startups. In 2018, he founded the current Third Scope Inc. by taking over an AI service and its development team. He now supports companies in adopting and utilizing AI, with a focus on AI-driven business development, operational transformation, and product development. He has also been involved in AI research as a Project Researcher at the University of Tokyo. Today, he continues to work at the forefront of AI project development, providing practical consulting from both technical and business perspectives.

“In the end, I can’t explain, department by department, what AI actually improved.”

Ms Saeki (a pseudonym), who leads DX promotion within the corporate planning department of a B2B company with around 1,200 employees, said this at the monthly investment-return review meeting.

The sales department reported that “proposal preparation has got faster”, while HR said “the burden of handling enquiries has eased”, and the IT department had been accumulating usage logs. However, as of six months earlier, most of these results were expressed only in terms of user impressions or usage frequency, which was not enough to inform decisions on whether to continue AI investment or how to set next year’s budget.

The company therefore linked AI chat and AI summarisation tools to specific tasks within each department, and set KPIs such as first-draft creation time for proposals in the sales department, first-response rate to enquiries in HR, and time spent on verification tasks in administrative departments.

According to internal figures, when comparing first drafts of proposals produced by the sales planning department over the three months from January to March 2026, average creation time fell from 90 minutes to 45 minutes. However, assessing this result properly requires checking the number of cases involved, the difficulty of each project, the measurement method, and whether review and revision time was included.

This article explains how to design department-level KPIs for DX leads, corporate planning staff and department heads who are struggling to measure the results of AI use. The aim is to combine efficiency, quality, adoption, business contribution and risk, moving away from qualitative impressions such as “it was useful” towards an evaluation that can genuinely inform decisions about continued investment.

Why it’s hard to explain the results of AI use

画像待ち 6-1-en - Why it's hard to explain the results of AI useの挿絵

At companies that have introduced generative AI, even when users respond positively, it can be difficult to explain the results at management meetings.

On the ground, AI can be used for a wide range of tasks: drafting emails, producing meeting minutes, rewriting documents, handling enquiries, organising notes from client meetings, and more. Meanwhile, management meetings and departmental reviews demand answers to questions such as the following.

  • Which departments and tasks are seeing results
  • How much have working time and volumes handled changed
  • Has the quality of outputs or responses improved
  • How much has this contributed to revenue, customer service and employee experience
  • Can investment be sustained while managing risk

Data on “how many times AI was used” alone cannot answer these questions. Usage frequency is a useful leading indicator of adoption, but it is not, in itself, a measure of business outcomes.

For example, if a salesperson uses an AI chat tool 100 times a month, that alone doesn’t tell you whether proposal quality has improved, or whether it simply took more attempts to get the desired response. Likewise, if HR’s use of AI summarisation increases, separate indicators are still needed to confirm whether enquiry handling time, repeat-enquiry rates, or the accuracy of responses have actually improved.

When measuring the results of AI use, it helps to distinguish at least the following three things.

Usage
Who used it, for which tasks, and how much.
Business outcomes
How time, volume and quality have changed.
Impact on the business and organisation
How revenue, customer experience, employee experience and risk have been affected.

The OECD’s review of research on generative AI and productivity points out that the productivity impact of generative AI is not uniform, and varies according to the type of task, the capability of the user, and how it is implemented.

For this reason, rather than looking only at company-wide usage counts, results need to be measured at the level of specific tasks.

What to decide before designing department-level KPIs

画像待ち 6-1-en - What to decide before designing department-level KPIsの挿絵

When designing department-level KPIs, rather than starting by drawing up a list of metrics, first decide on the target tasks, the baseline, and the timing of evaluation.

Define the target tasks specifically

First, clarify the scope of tasks in which AI will be used.

A definition such as “the sales department uses AI” is far too broad. Within sales alone, there are quite different types of tasks: preparing for client meetings, researching customers, drafting proposals, organising meeting notes, and writing follow-up emails.

For example, break tasks down to this level of specificity.

  • Sales department: drafting the first version of proposals
  • Marketing department: drafting article outlines
  • HR department: drafting first responses to internal enquiries
  • Accounting department: organising items to check during monthly processing
  • IT department: classifying help-desk enquiries

If the target tasks remain vague, the scope of usage logs, what is measured, and the definition of “results” will all stay vague too.

The International Labour Organization’s “Generative AI and Jobs: A 2025 Update” similarly analyses the impact of generative AI not at the level of whole occupations, but at the level of the individual tasks that make up an occupation.

When companies design KPIs, it likewise helps to define not just “which department used it” but “which task, and which step of that task, it was used for” — this makes results much easier to compare.

Set a baseline before implementation

Next, measure the state of things before AI is introduced. This pre-implementation reference point is called the baseline.

Where AI is used for drafting proposals, capture the average creation time and number of review rounds before implementation. Where it is used for handling enquiries, check the time to first response, the volume handled, and the repeat-enquiry rate.

Candidate baseline measures include the following.

  • Working time per item
  • Monthly volume handled
  • Number of review rounds
  • Number of reworks or returns
  • Time to first response to enquiries
  • Repeat-enquiry rate
  • Customer or employee satisfaction
  • Self-reported workload from staff

You don’t necessarily need to gather data on every single case from the outset. You can start with a sample covering one or two weeks, but in that case you should record the number of people involved, the number of cases, the timing of measurement, and how busy or quiet that period was.

Time from starting work to completing a first draft was measured for 30 proposals produced by five salespeople during a normal (non-peak) period.

Defining things this way makes it easier to compare results under the same conditions after implementation.

Separate leading and lagging indicators

KPIs for AI use should be designed by separating leading indicators from lagging indicators.

Leading indicators
Indicators that change before the final results appear. These include the proportion of users, the usage rate for target tasks, template usage rate, and training completion rate.
Lagging indicators
Indicators closer to the actual results for the business or task. These include working time, rework rate, first-time resolution rate, conversion-to-opportunity rate, and satisfaction.

If you evaluate purely on revenue or ROI immediately after implementation, the results can be hard to see. It can help to change what you evaluate depending on the stage of rollout, as follows.

Key indicators to check at different stages after AI implementation
Timing Main focus of evaluation
Immediately after implementation Usage environment, training status, review arrangements
1 to 3 months after implementation Working time, volume handled, quality
Quarterly or half-yearly Business contribution, risk, return on investment

These timeframes are only a guide. Tasks that occur only a few times a month, or that see large swings between busy and quiet periods, need a longer measurement period.

Grouping AI evaluation metrics into five categories

画像待ち 6-1-en - Grouping AI evaluation metrics into five categoriesの挿絵

When designing department-level KPIs, splitting metrics into five categories — efficiency, quality, adoption, business contribution and risk — helps avoid over-reliance on usage counts or time savings alone.

Efficiency metrics

Efficiency metrics look at how working time or processing speed have changed before and after AI use.

  • Working time per item
  • Time to produce a first draft
  • Time to first response to enquiries
  • Volume handled within a given period
  • Time from the end of a meeting to sharing the minutes
  • Lead time for producing documents

Time savings are relatively easy to measure and useful for justifying AI investment. However, even if AI speeds up first-draft production, if checking and revising then takes longer, that doesn’t necessarily mean the task as a whole has become more efficient.

It is important to measure “time to produce the first draft” and “total completion time including review and revision” separately.

Research on workplace tasks by the UK’s AI Security Institute similarly measures productivity gains from AI by looking at output quality alongside processing speed, rather than speed alone.

Quality metrics

Quality metrics look at how the quality of AI-assisted outputs or responses has changed.

  • Number of issues flagged during review
  • Rework rate
  • Rate of incorrect responses
  • Number of omissions in information
  • Rate of repeat questions from customers or employees
  • Number of items returned during internal checks
  • Fulfilment rate for mandatory items
  • Proportion of responses where evidence or sources could be confirmed

For example, if the time to produce meeting minutes falls from 30 minutes to 5 minutes, but decisions, owners and deadlines are missing, it cannot be said that business outcomes have genuinely improved.

Efficiency and quality metrics should, as a rule, be evaluated together.

Adoption metrics

Adoption metrics look at whether target staff are continuing to use AI for the designated tasks.

  • Number of monthly active users
  • Proportion of target staff who are actual users
  • Repeat usage rate
  • AI usage rate for the target task
  • Usage rate of approved templates
  • Training completion rate
  • Rate at which AI output is reviewed

Simple usage counts should be treated as a sign of adoption rather than a result in themselves. Higher usage counts don’t necessarily mean bigger results — they may just reflect more attempts being needed to complete the same piece of work.

Conversely, a low usage rate does not necessarily mean the tool itself is at fault. It may be due to the target task being poorly defined, limited opportunities to use it, insufficient training, or a complicated approval process.

Business contribution metrics

Business contribution metrics look at how process improvements involving AI relate to revenue, customer experience, employee experience and similar outcomes.

  • Lead time to submitting a proposal
  • Number of proposals per salesperson
  • Rate of follow-up after client meetings
  • Conversion-to-opportunity rate
  • Email response rate
  • Number of content pieces published
  • Cost per lead acquired
  • Customer satisfaction

However, revenue and conversion-to-opportunity rates are influenced by many factors, including market conditions, pricing, salespeople’s experience, customer budgets, advertising activity and the competitive landscape.

Rather than stating directly that “AI increased revenue”, it is better to compare AI users against non-users, compare before and after implementation, and compare similar projects against one another. Even then, if other factors cannot be sufficiently ruled out, it is more appropriate to say that “improvement was observed following process improvements that included the use of AI”.

Risk management metrics

Using AI requires measuring not only efficiency and revenue, but also safety and governance.

  • Number of instances where factually incorrect output was found
  • Number of instances used without human review
  • Number of instances where restricted or confidential information was entered
  • Number of instances that deviated from the approval process
  • Number of instances where information with unverifiable sources was used
  • Rate of pre-submission review before external release
  • Number and severity of incidents
  • Time from discovering a problem to correcting it

Risk metrics should be managed by severity as well as by count. For example, treating a typo in an internal memo and an error in contract terms sent to a customer as equally “one incident” would give a misleading picture of the actual situation.

The OECD’s “Generative AI and the SME Workforce” reports operational benefits from using generative AI, while also raising concerns about copyright, legal liability and regulatory compliance.

Result measurement should therefore look not only at “how much faster things became” but also at “whether use stayed within an acceptable level of risk”.

Examples of AI KPI design by department

画像待ち 6-1-en - Examples of AI KPI design by departmentの挿絵

The KPIs actually adopted shouldn’t be decided by department name alone, but narrowed down according to the target task and the step in the process where AI is used.

Sales department

In the sales department, it helps to separate “efficiency of proposal preparation” from “quality of proposals and follow-up”.

  • Time to produce the first draft of a proposal
  • Time to complete a proposal, including review and revision
  • Time spent on customer research
  • Time spent organising notes from client meetings
  • Number of review comments per proposal
  • Rate of follow-up after client meetings
  • Number of proposals per salesperson
  • Conversion-to-opportunity rate

When reporting measurement results, state the period, scope, number of cases, and method of calculation clearly.

From 1 January to 31 March 2026, we compared creation time before and after AI implementation for 30 proposal first drafts produced by the sales planning department. The average time from starting work to completing the first draft fell from 90 minutes before implementation to 45 minutes afterwards. This does not include review and revision time.

Writing it up this way makes clear exactly what the figures measure.

Where projects vary widely in difficulty or proposal length, report the median and figures broken down by project category, not just the simple average. A simple average alone can be skewed by a small number of exceptionally time-consuming cases.

Marketing department

In marketing, it helps to evaluate “production efficiency”, “quality before and after publication”, and “business outcomes” separately.

  • Time to produce an article outline
  • Time to produce a whitepaper first draft
  • Number of social media post drafts produced
  • Number of content pieces published
  • Number of times returned during review
  • Total working time until editing is complete
  • Number of corrections made after publication
  • Search traffic
  • Conversion rate
  • Number of leads acquired and cost per lead

Tracking only the number of items published risks missing duplicated content, factual errors, and inconsistencies in brand tone.

Alongside editing workload and the number of corrections before publication, also check the number of corrections after publication, the contribution to enquiries, and the contribution to sales opportunities. However, since search traffic and lead volumes are also affected by topic, advertising, publishing volume and seasonality, it is important not to attribute these solely to the effect of AI.

HR and general affairs department

In HR and general affairs, the main targets are handling internal enquiries, maintaining FAQs, and producing training materials.

  • First-response rate for internal enquiries
  • Time to first response
  • Rate of hand-off to a staff member
  • Self-resolution rate via FAQ
  • Recurrence rate of the same question
  • Number of incorrect answers regarding company policies
  • Time to produce training materials
  • User satisfaction
  • Number of enquiries handled per staff member per month

For metrics such as “first-response rate”, where interpretation can vary, define the calculation method clearly.

For example, define the first-response rate as “the proportion of enquiries received in the period that were resolved by the first response alone, without needing further handling by a staff member”.

Even if response time falls, it cannot be called an improvement if incorrect guidance on policy or repeat enquiries are increasing. Alongside efficiency, check the accuracy and clarity of responses and the repeat-enquiry rate.

Accounting and administration department

In accounting and administration, AI can support verification tasks, monthly processing, and organising documents.

  • Lead time for monthly processing
  • Time spent organising items to check
  • Time to respond to internal enquiries
  • Number of items returned
  • Number of input errors
  • Fulfilment rate for mandatory checks
  • Rate of pre-approval checks carried out
  • Number of exception cases detected

In accounting and administrative work, accuracy and approval accountability matter more than processing speed alone. One approach is to limit AI’s role to “identifying points to check”, “organising the issues”, and “drafting explanatory text”, rather than “making the final decision”.

Rather than KPIs based purely on “the proportion automated by AI”, it may better reflect actual working practice to measure “whether staff could identify all the points needing review, thoroughly and quickly”.

IT department

In the IT department, the main targets are help-desk support, enquiry classification, knowledge search, and FAQ creation.

  • Automatic classification rate for enquiries
  • Accuracy rate of classification results
  • Time to first response
  • First-time resolution rate
  • Self-resolution rate via FAQ
  • Recurrence rate for similar enquiries
  • Escalation rate
  • Success rate of knowledge searches
  • Number of responses that referenced outdated information

The “automatic classification rate” is the proportion of enquiries for which AI was able to present a classification. The “accuracy rate of classification results”, on the other hand, is the proportion of those classifications that matched a human reviewer’s assessment.

Without distinguishing the two, you risk missing a situation where AI is classifying a large number of enquiries but misclassification is also increasing.

Where response accuracy is low, the cause is not necessarily the AI model alone. It may be that the internal documents being referenced are out of date, that necessary information hasn’t been registered, or that no one has been assigned responsibility for managing the documents. Keep track of when knowledge content was last updated and how often it is updated.

A framework for explaining AI ROI

画像待ち 6-1-en - A framework for explaining AI ROIの挿絵

When explaining the return on AI investment, calculating it purely as “time saved x labour cost” can produce a misleading picture.

For example, if 100 sets of meeting minutes a month are each produced 25 minutes faster, the calculated time saving works out as follows.

100 items x 25 minutes = 2,500 minutes = approximately 41.7 hours

However, this 41.7 hours does not necessarily translate directly into reduced labour costs.

Unless headcount, overtime, or outsourcing costs have actually been reduced, it is more appropriate to think of this not as an accounting cost saving but as “time created” that can be reallocated to other work.

The effects of AI are easier to explain when split into the following three layers.

Time-saving effects

  • Time to produce meeting minutes has fallen
  • First drafts of proposals are produced faster
  • Time spent classifying enquiries has fallen
  • Time from starting research to organising the information has fallen

These are relatively easy to measure early in implementation. However, also check total working time, including review and revision.

Quality-improvement effects

  • Fewer omissions
  • Fewer issues flagged in review
  • Less variation in responses
  • Smaller quality differences between individual staff
  • Easier reuse of past knowledge

Quality improvements are hard to convert into monetary terms, but they affect operations through reduced rework and repeated effort.

When assessing that “quality has improved”, specific indicators need to be shown, such as the number of review comments, the rate of items returned, and the repeat-enquiry rate.

Business contribution effects

  • Number of proposals has increased
  • Follow-up after client meetings has got faster
  • Number of content pieces published has increased
  • Response time to customers has got shorter
  • Customer or employee satisfaction has improved

When explaining business contribution, avoid stating it as an effect of AI alone.

Note the conditions and limitations of the evaluation, for example: “improvement was confirmed following the implementation of process-improvement measures that included AI” or “changes to the sales structure were also made during the same period, so the contribution of AI alone cannot be calculated”.

Common pitfalls in KPI design

画像待ち 6-1-en - Common pitfalls in KPI designの挿絵

Treating usage counts alone as the result

Even if monthly usage counts are rising, if working time, quality and staff workload haven’t improved, that is not enough to demonstrate results.

Treat usage counts as a leading indicator, and combine them with efficiency and quality metrics.

Measuring solely with company-wide KPIs from the start

What is expected of AI differs by department and task.

Preparation time may matter most in sales, first-response rate in HR, first-time resolution rate in IT, and verification time in administrative departments.

A practical approach is to first design task-level KPIs, then separately select indicators that can be compared company-wide, such as investment amount, user proportion, and the number of serious incidents.

Setting unsubstantiated company-wide targets

Targets such as “AI-enable 30% of all tasks” or “raise company-wide productivity by 20% within six months” are easy to understand, but cannot be verified unless the target tasks and measurement method are defined.

In the early stages, limit the scope, for example, to the following.

  • Drafting proposal first drafts in the sales department
  • Organising notes from client meetings
  • Drafting responses to internal enquiries

Establishing a measurement method at a small scale before broadening the scope makes it easier to compare results before and after implementation.

Not setting any quality metrics

Focusing only on time savings risks missing an increase in misinformation or rework.

Particularly where external-facing documents, contracts, internal policies, customer interactions, or financial information are involved, include human review, verification of evidence, and approval records in the KPIs.

Overloading staff with measurement work

Adding too many detailed metrics can turn data entry and aggregation itself into a burden on the ground.

In the early stages, one approach is to start with around three metrics per high-priority task, as follows.

  • Efficiency metric: completion time per item
  • Quality metric: rework rate or error rate
  • Adoption/risk metric: usage rate for the target task or rate of review carried out

Separate logs that can be captured automatically from data that requires manual input by users, and keep manual-entry items to a minimum.

Using Kanata to run department-level KPIs

画像待ち 6-1-en - Using Kanata to run department-level KPIsの挿絵

There are several options for running department-level KPIs: general-purpose generative AI services, AI features built into business applications, and internal AI platforms.

When choosing between them, you need to check the following points, not just response quality.

  • Whether it has the functionality needed for the target task
  • Whether it can properly reference internal documents
  • Whether usage environments can be separated by department or project
  • Whether permissions management and usage log review are possible
  • Whether approved prompts can be shared
  • Whether data storage location and purpose of use can be confirmed
  • Whether it can be operated together with training and adoption support

Kanata, our own product, is characterised by managing prompts and reference information at the task level alongside AI chat and AI summarisation, and by combining this with training features such as e-learning. The latest available features and contract terms should be confirmed in the official product materials.

As such, it can be one option for companies aiming not just for one-off AI use, but for an operating model in which “target tasks are defined by department, usage is standardised, and training and result measurement continue over time”.

That said, introducing Kanata does not automatically mean results will be measured. The types of logs that can be captured, retention periods, the units used for departmental aggregation, and the scope of integration with external systems may vary depending on the contract and configuration.

Before implementation, it is important to check the KPIs you want to measure against the data that can actually be captured.

Decide on the target task for each department

Start by limiting each department to one or two tasks.

  • Sales department: proposal first drafts, organising notes from client meetings
  • Marketing department: article outlines, social media post drafts
  • HR department: enquiry response drafts, producing training materials
  • IT department: enquiry classification, FAQ creation
  • Administration department: organising items to check, summarising monthly reports

Rather than “using AI for everything”, define “who uses it, when, and for what purpose”.

Prepare prompts and reference information

Next, prepare the prompts and reference materials to be used in each department.

For example, if HR is drafting responses to enquiries, organise employment regulations, FAQs, application procedures, and past enquiry examples. For sales, candidates might include proposal templates, case studies, product specifications, and approved competitor comparison materials.

Also, the term “training data” can easily be confused between data used to train the AI model itself and documents that are searched or referenced at the time of response. Check how the product actually uses data, and, where necessary, use distinct terms such as “reference data” or “internal knowledge” as appropriate.

Check leading indicators in the early stages of rollout

In the early stages of rollout, check leading indicators such as the following on a monthly basis.

  • Proportion of target staff who are actual users
  • Usage rate for the target task
  • Usage rate of approved templates
  • Training completion rate
  • Rate at which human review is carried out
  • Issues reported by users
  • Number of updates to reference information

If the usage rate is low, this is not necessarily due to AI performance alone. There may be several causes: the target task is poorly defined, people don’t know how to operate it, approved prompts are hard to use, or reference information is insufficient.

Trying to simply push up usage counts without checking the underlying cause risks encouraging use that isn’t actually needed for the work.

Check lagging indicators once enough cases have accumulated

Once operations are underway and enough cases have built up for comparison, check the following lagging indicators.

  • Has total working time fallen
  • Has rework or the error rate increased or decreased
  • Has response quality stabilised
  • Has response time to customers or employees improved
  • Have the number of proposals or pieces of content changed
  • Has the sense of workload on the ground changed
  • Have any risk events occurred

Checking on a quarterly basis helps smooth out month-to-month variation. However, if case numbers are low, switch to a longer period, such as every six months, to ensure a sufficient number of cases.

An operating cycle for embedding AI result measurement

画像待ち 6-1-en - An operating cycle for embedding AI result measurementの挿絵

Measuring the results of AI use isn’t a one-off exercise that ends once KPIs are set. As the target tasks, users’ proficiency, AI functionality, and internal policies change, the indicators worth watching change too.

Measure the current state before implementation

  • How many minutes does each item take
  • How many items are processed in a given period
  • Where does rework occur
  • Who is carrying the heaviest workload
  • Which information takes a long time to find

Record what is measured, the period, the number of cases, and the aggregation method, and compare under the same conditions as far as possible after implementation.

Check the usage environment immediately after implementation

  • Are target staff able to use it
  • Are the situations for using it clear
  • Are templates and procedures easy to use
  • Is output being reviewed
  • Is confidential or personal information being handled appropriately
  • Can users report concerns or problems

Rushing to see results at this stage risks making “increasing the usage count” an end in itself.

After implementation, check efficiency and quality together

Check whether review comments or revision time have increased even as working time has fallen. If enquiry handling has got faster, also check whether repeat enquiries or incorrect responses have increased.

Where both efficiency and quality have improved and risk stays within an acceptable range, it becomes easier to judge that genuine business results have been achieved.

Evaluate the return on investment after a set period

  • Which tasks has the time created been reallocated to
  • Have volumes handled or proposal numbers increased
  • Has the quality of customer or employee interactions changed
  • Do the benefits justify the implementation, running, training and verification costs
  • Can you explain the decision to continue, revise, or stop
  • Can you identify the next task to target

Even where a particular task hasn’t shown results, this doesn’t have to be judged as a failure of AI overall. Check whether the cause lies in how the task was chosen, the reference information, the prompts, or the training method.

On the other hand, if effects still can’t be confirmed despite ongoing improvements, and operational burden or risk outweighs the benefits, a decision to stop use for that task is also necessary.

Summary: measure AI’s results at the level of individual tasks

画像待ち 6-1-en - Summary: measure AI's results at the level of individual tasksの挿絵

What matters when measuring the results of AI use is not judging things by uniform, company-wide figures alone.

Sales, HR, marketing, IT and administrative departments each expect different things from AI. Some departments want to shorten the time it takes to produce proposals; others want to reduce variation in enquiry responses; others want to cut the time spent searching for internal information.

When designing department-level KPIs, proceed in the following order.

  1. Decide on the target task and the step where AI will be used
  2. Measure the baseline before implementation
  3. Choose indicators for efficiency, quality, adoption, business contribution and risk
  4. Define the measurement period, number of cases, and calculation method
  5. Judge whether to continue, revise or stop based on the results

There is no need to calculate a precise, company-wide ROI from the outset. Start with a single department, a single task, and a small number of indicators, and broaden the scope as you confirm the validity of your measurement approach.

The results of AI use cannot be adequately explained by impressions such as “it was useful” or by usage counts alone. At the same time, it isn’t appropriate to attribute improvements in revenue or productivity solely to AI either.

Setting KPIs that fit the context of the task, and checking efficiency, quality, business contribution and risk together, brings you closer to an evaluation that both the front line and management can verify.

KPIs are not fixed once and for all. As AI functionality, target tasks, user proficiency and internal rules change, it is important to revisit their definitions and targets, and keep updating them into evaluation measures that genuinely support your company’s decision-making.

Q&A on measuring AI results

What should we measure first after implementing AI?

What should be measured first is not simply the usage rate for the target task, but basic indicators that let you compare before and after implementation.

For proposal drafting, for example, start with three metrics: “first-draft creation time”, “completion time including review and revision”, and “number of review comments”. Because this lets you check efficiency and quality at the same time, it prevents time savings alone from being treated as the result.

If usage counts are rising, does that mean AI use is a success?

A rising usage count alone cannot be taken as a sign of success.

Usage count is a leading indicator that shows adoption. It needs to be assessed together with business outcomes such as working time, volume handled, rework rate, and rate of incorrect responses.

It’s also possible that a rise in usage count reflects more trial and error being needed to get a usable response.

How should we convert AI time savings into a monetary figure?

First, separate time saved from actual cost reductions.

Even if working time falls by 40 hours a month, if headcount, overtime, or outsourcing costs haven’t actually been reduced, that 40 hours is not a direct cost saving — it is time created that can be used for other work.

If you do convert it into a monetary figure, also check which tasks the freed-up time was reallocated to, and what results that produced.

Is it a problem if KPIs differ between departments?

It isn’t a problem. In fact, setting different KPIs according to the target task makes it easier to understand the actual situation.

Sales may prioritise proposal-drafting time, HR the first-response rate, and IT the first-time resolution rate — the results that matter most differ by department.

At the same time, setting separate indicators that are checked consistently across the whole company — such as investment amount, the proportion of target staff who are users, and the number of serious incidents — makes the data more useful for management decisions.

If we can’t confirm results, should we stop using AI?

There’s no need to stop all use immediately. First, check whether there are problems with the target task, the reference information, the prompts, training, or the review process.

If sufficient benefit still isn’t achieved after making improvements, and operating costs or risk outweigh the results, it is reasonable to scale back or stop use for that particular task.

Rather than treating the continued use of AI as an end in itself, it is important to judge whether to continue, revise, or stop on a task-by-task basis.