We’re collecting the logs, yet we can’t work out what to fix next.
That was the remark of Mr Morita (a pseudonym), from the DX promotion team, in a meeting room at a manufacturing firm. Laid out on the table were the number of AI chat sessions, access figures broken down by department, and a list of prompt logs. The numbers were certainly there. The charts had been drawn up. And yet the mood in the room remained somehow heavy.
The IT team reported that “usage is climbing.” Meanwhile, a manager from one of the operational departments noted that “on my team, plenty of people still aren’t sure how to use it.” From sales came the observation that “it’s handy for a first draft of a proposal, but the quality of the output varies.”
Watching all this, I sensed that AI adoption had entered its next phase. In the early days, the challenge is simply getting people to use the thing. But once a certain volume of AI usage logs begins to accumulate, the next requirement is to learn from how it is actually being used.
Six months earlier, the company had rolled out KANATA AI’s AI chat and AI summarisation features, and had been collecting AI usage logs and prompt logs. At the monthly meeting, however, what they reviewed was chiefly the trend in usage counts. Even when voices were raised — “sales are using it, but it hasn’t taken hold in quality control,” or “the same failure cases keep getting shared on Slack” — none of it was being translated into concrete improvement measures.
So, over the most recent 90 days, we examined 1,842 usage logs across 12 target departments, sorting them by the nature of the question, failure patterns, user behaviour, and departmental segment. As a result, it became possible to see which tasks attracted a great many regenerated answers, which prompts saw users drop out partway through, and what the fast-growing teams had in common. The company is now in a position where, within a monthly AI improvement cycle, it can distil all this into prioritised improvement measures.
This article sets out how to use AI log analysis, log-visualisation AI, and benchmarks, and how to feed all of it into a PDCA cycle for AI operational improvement. The aim is not merely to gaze at the logs, but to reflect what you find back into the next round of training, prompt refinement, and process design. That said, log analysis alone will not solve everything: only when it is combined with on-the-ground interviews and a review of operating rules does it lead to genuinely repeatable improvement.
AI usage logs are the raw material for improvement, not the result of it
At companies that have introduced generative AI, the usage rate is the first thing to draw attention.
How many people used it. How many questions were asked. Which departments use it most. These are important indicators for grasping how far AI adoption has spread.
But looking at usage counts alone tells you nothing about whether the business AI is genuinely proving useful.
What I often see on the ground when supporting AI adoption is the case where the report reads “usage is up, so all is well,” and yet satisfaction among the people actually doing the work is rather low. On paper it is being used. In practice, however, people may be asking the same thing over and over, or finding the answers unusable as they stand and reworking them heavily by hand.
For instance, even if a particular department records a high number of AI uses, you cannot tell from the figure alone whether that is because “it’s so handy people keep coming back to it” or because “it keeps failing to give a good answer, so they keep asking again.” Only a deeper look at the logs will settle the question.
Conversely, even a department with few uses may be deriving a substantial benefit from those handful of sessions. Even if something is used only a few times a month, if those occasions involve drafting an important proposal or marshalling the materials for a board meeting, the business impact is hardly trivial.
AI usage logs are not simply figures for reporting on how widely the tool is used. They are the raw material for finding where people stumble, where the work has been designed badly, where training is lacking, and where there is room to improve prompts.
What matters is not only “how much it was used,” but “where people got stuck,” “why they didn’t keep using it,” and “which ways of using it led to results.”
The basic lenses to apply in AI log analysis
When you begin AI log analysis, there is no need to scrutinise everything in fine detail from the outset. Faced with a pile of logs, one is tempted to analyse the lot. But looking too closely, too soon, can actually obscure the next move.
To start with, sorting the logs into five views — usage frequency, prompt logs, failure cases, user behaviour, and segment analysis — makes the improvement points easier to spot.
Usage frequency: who is using it, and for what work
The first thing to look at is how often the AI is used.
“How many times was it used across the whole company,” however, is not enough on its own. By breaking the figures down by department, by role, and by task, the light and shade of AI adoption begins to show.
For example, while the sales department may be using it to knock out first drafts of proposals, the administrative department may be using it to check regulations and handle enquiries. The development team might use it to organise specifications, and marketing to draft article outlines or ad copy.
The important thing here is not to write off a heavy-using department as simply “the able one.” How frequently AI is used is also shaped by the nature of a department’s work. It is easy to reach for in departments that do a lot of writing, while in departments centred on hands-on operations the occasions to use it can be harder to see.
The first step towards improvement is to consider the reasons for heavy use and light use separately. Departments that use it heavily have success patterns. Departments that use it little may have an issue somewhere — in the workflow, in training, or in a psychological barrier.
Prompt logs: which questions are leading to results
Next to examine are the prompt logs. By prompt logs I mean the record of the instructions and questions users have put to the AI.
Prompt logs preserve what sort of requests staff are making of the AI. The everyday frustrations and the quirks of the work show through here in their unvarnished form.
For example, where the following kinds of question are common, the underlying issue differs in each case.
- Tidy this text up so it reads more clearly
- Summarise these minutes
- Draft a proposal for this client
- Answer in line with our internal regulations
- Boil this down briefly for my manager
At first glance these all look like fairly ordinary AI use. Yet once you tally up the logs, you begin to see whether usage leans towards writing, towards summarising, or towards searching internal knowledge.
When reading prompt logs, look not only at the questions that went badly but also at the ones that went well. Prompts that led to a good result tend to have features in common. The clearer the purpose, the more context supplied, and the more the output format is specified, the more likely a prompt is to yield output that is genuinely usable in practice.
For instance, “write up the minutes” is far less effective than “please take the meeting notes below and organise them, for internal circulation, into decisions, to-dos, open items, and points to check next time” — the latter is far more likely to land close to what the user actually wanted.
In an environment such as Kanata, where frequently used instructions can be saved and shared as a prompt library, it becomes easier to stop a prompt that produced results from ending its life as one person’s private knack and to turn it into a shared asset for the organisation. The same can be done with other AI tools, provided they offer template management or knowledge sharing.
Failure cases: where is answer quality slipping
What matters especially in AI operational improvement is the analysis of failure cases.
A failure case is not merely an instance of “the AI gave a wrong answer.” Logs of the following kinds, for example, are also candidates for improvement.
- The same question is being rephrased again and again
- The answer is regenerated immediately
- The user leaves without copying the output
- Follow-up instructions such as “that’s not it” or “no, rather…” are added
- The matter is, in the end, checked with a person
These are signs that the AI’s answer may not be reaching what the user hoped for.
I take the view that failure logs are precisely the logs worth having. The reason is that failure logs tend to reveal plainly where the misalignment lies — in the tool, the work, the training, or the data design.
That said, it is hasty to look at a failure log and conclude on the spot that “the AI just isn’t accurate.” Often the cause lies not in the AI model itself but in the vagueness of the prompt, a shortage of reference data, operating rules that have not been pinned down, or insufficient user training.
If, for example, you see many failures of the kind “I asked the AI about an internal regulation but got back only generalities,” it may be that the regulation documents are not within the AI’s reference scope, or that no instruction has been set to make it cite its sources when answering.
With Kanata, you can organise internal documents, FAQs, and operating rules as training data and make them easy to reference from AI chat and AI summarisation. But registering documents does not, of itself, make everything fall into place. You need to settle questions such as whether outdated or contradictory material has crept in, and who holds responsibility for keeping it up to date.
User behaviour: what happens after the answer
When reading AI usage logs, look not only at the content of the question but also at what the user does after the answer.
If the user copies the AI’s answer, that answer probably had a degree of practical use. On the other hand, if the user immediately moves to a different question, or repeatedly regenerates, the response may not be reaching the hoped-for answer.
The behaviours worth watching are these.
- Did they copy the answer
- Did they go on to ask another question
- Did they regenerate
- Did they abandon the conversation partway
- Did they open a separate chat on the same theme
- What times of day does usage cluster around
In B2B business AI especially, “did the next action in the work follow on” matters more than “was it used.” By seeing whether the answer was copied, shared internally, or carried over into minutes or a proposal, you learn where the AI sits within the flow of the work.
For example, where an AI summary log shows a high proportion of outputs being copied afterwards, that summary format may suit the people doing the work. Conversely, where summaries are generated but not copied, the heading structure, the level of detail, or the tone may not fit practice.
Read the logs as numbers, with a cool head. But on the far side of those numbers there is always human behaviour. In log-analysis meetings I make a point of asking, again and again, “behind this number, what were the people on the ground actually doing?”
Segment analysis: seeing the differences by department, role, and task
Looking at AI usage logs through the company-wide average alone makes it easy to overlook the improvement points.
Suppose the company-wide usage rate is 30%. In sales it might be 70%, in administration 15%, and on the factory floor 5%. In that case, judging from the company-wide average alone that “adoption still has a way to go” will not lead to the right measures.
Sales may need successful patterns rolled out more widely. Administration may need internal regulations and FAQs prepared as training data. The factory floor may need a route in built around smartphones and standard forms rather than AI use that presumes a PC.
In segment analysis, look at, at the very least, the following cuts.
- By department
- By role
- By task
- By length of use
- By usage frequency
- Whether a successful prompt exists
- Whether training has been attended
Looking at things divided up this way makes it clear “to whom, and which improvement measure, should be delivered.”
Even on the ground at the firms I support, where the same AI tool has been introduced, the results can differ entirely from one department to the next. That gap arises not from a difference in the tool but, more often than not, from how the tool is woven into the work, the involvement of the manager, the sharing of success stories, and the clarity of the input rules.
Log-visualisation AI becomes a common language for improvement decisions
AI usage logs are hard to put to use for improvement if you merely stare at a CSV or an admin-screen list. This is where log-visualisation AI and dashboards earn their keep.
That said, visualising the data does not automatically surface the improvement points. Producing a tidy chart and being able to make an improvement decision are two different things. I have seen, more than once, the situation where “there’s a dashboard, but nobody can decide on the next move.”
In log visualisation, the first thing that matters is to narrow the indicators. Trying to look at every log from the start tends, on the contrary, to leave you unsure what to improve.
The five basic indicators worth keeping in view, common to most cases, are these.
| Indicator | What it tells you |
|---|---|
| Number of users | How widely AI adoption has spread |
| Number of uses | Which tasks and departments it is used in |
| Regeneration rate | Answer quality and any misalignment in prompts |
| Continued-use rate | Whether those who tried it once have stuck with it |
| Low-rating / drop-out logs | Finding failure cases and candidates for improvement |
The important thing here is not to judge on a single number alone.
Even with a high number of uses, a high regeneration rate may mean users are quietly dissatisfied. Even with a low number of uses, a high copy rate or continued-use rate may mean that, for certain tasks, the benefit is already there.
Log-visualisation AI is not there to make the numbers look pretty. It is a common language for the people on the ground to talk through what they ought to fix next.
In an environment such as Kanata, where chat, summarisation, training data, and prompts can all be handled within the same operational platform, it becomes easier to feed the issues surfaced by the logs back into improvement measures. For instance: turn the prompts that get regenerated a lot into templates; build up the training data in areas where answers tend to be generic; and design use-case training for the departments where usage is low. Connecting the logs to improvement measures in this way is what counts.
Don’t let the AI improvement cycle end at the analysis
A common failing in AI log analysis is to produce an analysis report and feel the job is done.
Checking a dashboard monthly and reporting that “sales use it a lot,” “use of the summary feature is rising,” and “in some departments usage is low” will not, on its own, lead to AI operational improvement.
What matters is connecting log analysis to a PDCA cycle. PDCA is a management method that repeats plan, do, check, and act. To my mind, when it comes to embedding AI in everyday work, this unglamorous PDCA matters more than any flashy initiative.
Decide the purpose of the analysis
First, decide what you want to improve.
The objectives might be divided up, for example, as follows.
- Raise the usage rate
- Stabilise answer quality
- Increase the take-up rate in a particular department
- Reduce failure cases
- Roll out successful prompts more widely
- Cut the time spent handling enquiries
Look at the logs with a vague objective, and you will end up merely gazing at the numbers. Setting the improvement theme first narrows down which logs to look at.
“We want to push AI adoption further,” for instance, is too broad. If you can say “we want to lower the regeneration rate on the AI chat used for drafting proposals in the sales department,” then both the logs to look at and the measures to take become concrete.
Decide the target period and target logs
Next, decide the target period and the target logs.
What you can see changes depending on whether it is “the whole company’s logs over the most recent 30 days,” “the sales department’s logs over the most recent 90 days,” or “a two-month comparison either side of a training session.”
Whenever you handle figures, always make the period, the target department, and the number of cases explicit.
- Poor example
- AI usage has gone up.
- Good example
- Over the 90 days from 1 January to 31 March 2026, the sales department’s AI chat usage rose from 412 uses a month to 689.
Making the conditions explicit in this way is what lets you verify the improvement work.
In AI log analysis I place more weight on “being in a state you can compare” than on “the numbers themselves.” Without a basis for comparison, you cannot judge whether something has risen, fallen, improved, or worsened.
Classify the failure patterns
Next, classify the failure logs.
| Failure pattern | Likely cause | Improvement measure |
|---|---|---|
| Answers turn generic | Insufficient training data | Register internal documents |
| Answers are too long | Output format not specified | Create prompt templates |
| Repeated re-asking | Vague premises in the question | Teach input examples |
| Large usage gaps between departments | Doesn’t fit the workflow | Design department-specific use cases |
| Many identical questions | Not turned into an FAQ | Build out the knowledge base |
Once you can classify in this way, you stop thinking of AI improvement purely as “a model problem.” You can separate out whether the thing to improve is training, the prompt, the training data, or the workflow.
What is important here is not to reduce failure to an individual’s lack of skill. Stop at “staff aren’t asking well” and the organisation cannot improve. Why can’t they ask well? Is there no template? Are there no worked examples from the business? Is the line on what information may be entered unclear? You need to dig that far.
Set priorities
It is important not to try to carry out every improvement measure at once.
When setting priorities, the following two axes make things easier to sort out.
- Is the improvement impact large
- Is it easy to carry out
For example, a measure to improve the minutes-summary prompt used company-wide is comparatively easy to carry out, and its reach may be broad. Linking up with core systems or a large-scale redesign of access permissions, by contrast, is high in difficulty to execute even where the effect would be great.
The realistic course is to start with “measures whose effect is easy to see and which are easy to carry out.”
On the ground with AI adoption, I put more weight on “a succession of small improvements” than on “grand reform.” Fix one prompt. Update one piece of training data. Share the way of using it with one department. That accumulation makes a substantial difference a few months down the line.
Verify with the following month’s logs
Finally, once an improvement measure has been carried out, always verify it with the following month’s logs.
If you improved the minutes prompt, for example, the indicators to watch are as follows.
- Did use of the minutes summary increase
- Did the regeneration rate fall
- Did the copy rate rise
- Did the range of users’ departments widen
- Did the complaints on Slack and in meetings decrease
Only once you have looked this far can you say the AI improvement cycle is turning. Log analysis exists not to produce an analysis report but to inform the next improvement.
Common failure cases and the direction of improvement
Analyse AI usage logs and you find similar failure cases at a great many companies. Here are the representative examples worth keeping in mind, common to most.
Usage is high, but so is regeneration
Look at a heavy-using department and, at first glance, AI adoption seems to be going well. But a high regeneration rate calls for caution.
Where users are reworking answers again and again, the first answer may be missing the mark.
The improvement measure in this case is to build out prompt templates.
Let users enter not only “what they want produced” but “who it is for,” “in what format,” “roughly how many words,” and “which information to draw on.” With Kanata, registering the prompts that led to results in a library, in a form reusable across the company, helps reduce the differences in how individuals phrase their requests and makes the quality of output easier to keep steady.
Only a handful of people use it
It is also common for AI adoption to skew towards a few staff.
In this case, taking those who don’t use it to task will not make it stick. You need to look at the logs and break down “why it isn’t being used.”
The reasons are various.
- They aren’t sure which tasks to use it for
- They have no time to use it
- They are nervous about judging what information may be entered
- They can’t trust the AI’s answers
- Their manager or team isn’t using it
Here you need department-specific use cases presented, and clear rules on what information may and may not be entered. Embedding AI adoption calls not only for introducing the tool but for operating rules that let people use it with confidence.
In my experience, the main reason people don’t use AI is rarely “lack of interest.” It is the unease of “I don’t want to use it and fail,” “I’m not sure what I’m allowed to put in,” and “I don’t know how my manager will judge it.” That is precisely why looking at the psychological barriers on the ground matters every bit as much as the log analysis.
Successful prompts aren’t being shared
One member of staff uses the AI well; the prompts that lead to results have common features. The context supplied is concrete, the output format is clear, and the conditions for judgement are written out.
When you find a successful prompt, don’t let it end as one person’s knack — make it a team template. Templating the recurring tasks first — minutes, proposals, enquiry handling, weekly reports, one-to-one preparation — tends to pay off most readily.
With Kanata, you can put frequently used instructions into a form the team can reuse, and keep prompts organised by task. It is useful when you want to move from a state where only those comfortable with AI get results to one where the whole team can produce a consistent level of quality.
The AI only ever returns generalities
“Whatever I ask the AI, all I get back are generalities” is another common refrain.
In this case, the problem lies not with the AI itself but, very likely, with a shortage of internal data it can refer to.
Organise internal regulations, product materials, FAQs, past proposals, minutes, and training materials as training data, and the AI finds it easier to answer in keeping with your own organisation’s context.
That said, it is not a case of registering everything. Throw in outdated material, contradictory material, or highly confidential material, and answer quality can actually drop. Training data needs regular stocktaking and updating.
I regard preparing training data not as “the task of feeding documents to an AI” but as “the task of putting the company’s knowledge in order.” The areas where AI can’t answer are, more often than not, areas where the information hasn’t been organised within the company in the first place.
With Kanata, logs and improvement measures are easy to handle in one flow
When pressing ahead with AI operational improvement, there are any number of tool options. You might combine general-purpose generative AI services, an internal chatbot, a BI tool, a log-analysis platform, and a knowledge-management tool.
Among these, the distinguishing feature of using Kanata is how readily AI chat, AI summarisation, prompt management, training-data management, and learning content can be joined up within everyday operations.
For example, if the AI chat logs reveal that “requests to summarise minutes are common,” the next thing to do is to prepare a prompt for minutes. You then register that prompt in the prompt library so that everyone can use it in the same form.
And where questions about internal rules are common, you can organise the regulations and FAQs as training data and shift to an operation that shows the basis when answering.
Further, where the same failure case keeps recurring, you can reflect it in a short piece of learning content or an internal study session, rolling it out as a brief course for staff.
In other words, the results of log analysis can be connected to the following four improvements.
- Prompt improvement
- Training-data preparation
- Workflow review
- Staff education
By not letting it end at looking at the logs, and instead feeding it back into the AI environment used in daily work, AI operational improvement becomes easier to sustain.
What I look for in a business AI platform such as Kanata is not simply being able to use AI. It is being able to find issues from the logs on the ground, fix the prompts, tidy the training data, link it to education, and verify it again with the logs — to keep that loop turning within a single operational platform.
AI adoption is not something that ends once introduced. If anything, it is only after introduction that the real improvement begins.
To make AI log analysis stick, a monthly review works well
Carry out AI log analysis just once and the effect is unlikely to be great. To turn it into operational improvement, you need to embed it as a monthly review.
For a monthly review, the following flow is realistic.
- Review the previous month’s usage logs
- See which tasks saw more use and which saw less
- Check the logs with high regeneration or drop-out
- Share successful prompts and failure cases
- Narrow the improvement measures to no more than three
- Decide the verification indicators for the following month
What matters here is not to keep it within the IT or DX promotion team alone. The numbers in the logs sometimes leave you unable to tell what is happening on the ground.
For instance, suppose usage fell in a particular department. AI may not have become unnecessary; people may simply have had no room to use it during a busy spell. Conversely, even where usage rose, the state of affairs on the ground may be “we’re using it because we have to, but we’re not happy with the accuracy.”
Only by looking at the numbers and the voices on the ground together does log analysis become genuine improvement work.
The three questions I most often use in a monthly review are these.
- From these logs, what troubles on the ground can we see
- What one thing will we change by next month
- Which log will we use to confirm whether that improvement worked
If you can answer these three, AI log analysis becomes not a reporting document but a tool for operational improvement.
When handling AI usage logs, aim at improvement, not surveillance
AI usage logs are useful, but they call for care in how they are handled.
First, don’t use the logs to keep staff under surveillance. Pore over them to point out who used it how many times, or whose prompts are immature, and staff will simply stop using AI. The purpose of log analysis is not to appraise individuals but to improve the work and the environment.
Second, don’t judge the ground from the logs alone. The logs preserve the traces of behaviour, but not the circumstances behind them. Whether low usage reflects dissatisfaction with the tool, a low business need, a busy period, or psychological unease — that has to be asked of the people on the ground.
Care is also needed over how the information contained in the logs is treated. Prompt logs may contain client names, deal information, and the contents of internal documents. When analysing, you need to make access rights, the scope of who may view them, anonymisation, and retention periods explicit. Authoritative frameworks set out much the same expectation: the UK ICO’s guidance on AI and data protection stresses lawful, fair, and transparent handling of the personal data that AI systems process. The NIST AI Risk Management Framework and the OECD AI Principles, too, point to the importance of AI trustworthiness, risk management, and human oversight.
AI operational improvement cannot do without striking a balance between convenience and safety. Settling in advance the scope of what you look at in the logs, the purpose of looking, and how it will be used for improvement also matters for keeping the trust of your staff.
Summary: AI usage logs become a map for improvement on the ground
AI usage logs are not a mere record of how much the tool was used.
In which tasks is AI being used. Where are people getting stuck. Which prompts are leading to results. Which departments need training. Which training data ought to be prepared.
They are a map for finding all of this — a map for improvement on the ground.
That said, looking at the logs alone yields no answers. The logs are no more than clues for finding the improvement points. Only when combined with on-the-ground interviews, process design, prompt improvement, training-data preparation, and education measures does the AI improvement cycle begin to turn.
On the ground with AI adoption I always put more weight on “learning from how it is used” than on “making people use it.” AI is not finished the moment it is introduced. It is by being used on the ground, failing, being fixed, and being used again that it finally settles into the work.
Don’t leave the logs simply accruing; create an occasion, even once a month, to review them. Decide on a small improvement measure. Watch for the change in the following month’s logs. It is this repetition that turns AI from a passing tool into something embedded in the work.
Q&A: Common questions on AI log analysis and operational improvement
Q1. With AI usage logs, what should I look at first?
To begin with, looking at five things is realistic: number of users, number of uses, regeneration rate, continued-use rate, and low-rating / drop-out logs. Trying to analyse every log in fine detail from the outset makes it hard to translate into improvement measures. First, grasping “which departments are using it” and “where people are getting stuck” is what counts.
Q2. If the number of uses is rising, can I say AI adoption is a success?
A rise in the number of uses is an encouraging sign, but it cannot on its own be judged a success. Where the regeneration rate is high, outputs aren’t being copied, or the same question is repeated, there may be an issue with answer quality or the workflow. The number of uses needs to be read in combination with other indicators.
Q3. What should I be careful of when reading prompt logs?
Prompt logs may contain client names, deal information, the contents of internal documents, and the like. For that reason, it is important to settle viewing rights, anonymisation, retention periods, and the purpose of analysis in advance. They should also be handled with the aim of finding what the organisation should improve, not of appraising how individuals use the tool.
Q4. When the AI’s answers tend to be generic, what should I improve?
First, check whether the internal documents, FAQs, and operating rules the AI can refer to are in good order. Where training data is lacking, the AI is prone to answering in generalities. Alongside this, it helps to make the output conditions explicit on the prompt side too — “based on internal regulations,” “where there is no basis, write that it needs checking,” and so on.
Q5. How often should I run AI log analysis to keep it going?
At most companies, starting with a monthly review is realistic. Look at the previous month’s logs, narrow the improvement measures to no more than three, and confirm the effect with the following month’s logs. Make the cadence too frequent and the operational load grows, making it hard to sustain. I would recommend starting small, on a monthly basis.