So, whose way of doing this is actually the correct one again?
That was the remark Mr Yamaoka from the sales planning team let slip in a Monday-morning meeting. On the sites where I am asked to advise on rolling out AI agents, I hear something very close to this line again and again. Companies want to bring AI in. They want to make enquiry handling more efficient. They want to automate routine work. Yet behind all that hope there often lies a rather awkward truth: even among the human staff, the judgements were never aligned in the first place.
This is the story of a business-design team at a BtoB company that was trying to introduce an AI agent for internal enquiry handling. HR, IT and frontline leaders gathered to run trials, but six months earlier each person had interpreted the operating manual differently, and both the standard operating procedures and the exception-handling branches were left vague. As a result, the AI agent could answer the routine cases well enough, but it stalled on enquiries with unusual contract terms or on requests where the approver was away, and Slack filled up with posts saying, in effect, “this one is probably best handed back to a person.”
The team then took stock of the 12 target tasks, working through 82 enquiries from the most recent three weeks, and tidied up the SOPs, the escalation conditions and the fallback wording. SOP here is short for Standard Operating Procedure: a standard set of steps that helps anyone handling the task arrive at broadly the same judgement. Where you use a business-support platform such as Kanata, which lets you handle AI chat, AI summarisation, e-learning and the like on a per-project basis, one advantage is that standardised procedures are easy to connect to training data and learning content. That said, a tool is only ever one option among several. What matters is getting your own operating rules down to a level of detail an AI can actually refer to.
In this article, written for those wanting to take business standardisation with AI forward, I will set out how to design AI-ready operations so that an AI agent can act without hesitation, how to think about exception handling, and how to put guardrails in place for unforeseen scenarios. The aim is a state in which the criteria for a decision do not wobble whoever happens to be on duty, and AI and people run the work to the same set of rules. Even so, simply drawing up standard operating procedures does not solve everything by magic. Only with continual review, frontline education and reviews by a responsible owner do you edge towards genuinely stable operation.
The first stumbling block when introducing an AI agent is the vagueness of the work itself
When companies consider bringing in an AI agent, most of them think about which AI to use, how far they can automate, and how much human working time they might save.
The performance of the model and its ability to connect to external systems matter, of course. In my own work supporting AI adoption and product development, I check model selection, security requirements, data integration and the operating setup. Yet on real adoption sites, the first thing to cause trouble is, more often than not, the vagueness on the business side rather than the AI itself.
Suppose, for instance, you try to hand internal enquiry handling to an AI agent. Expense claims, contract checks, sales-material requests, confirming internal regulations: the operating manuals may well exist. And yet, once you get onto the floor, you may find a state of affairs like this.
- Ms A checks on Slack before approving.
- Mr B looks at past emails to decide.
- Mr C sends it back once before asking a superior.
- Ms D reckons that “if it is urgent, you may push it through as an exception.”
Between people, this vagueness can be papered over in conversation. The work keeps moving on tacit knowledge: “we did it this way last time,” “if this person is asking, it must be urgent,” “let it through for now and check later.” I have no wish to dismiss that frontline flexibility in itself; many companies keep their daily work turning precisely because of it.
An AI agent, however, does not run stably on the assumption of tacit knowledge. Unless it is settled which information to consult, on what conditions to answer, and from what point to hand over to a person, your AI-ready design will be unstable.
In short, the precondition for putting an AI agent to good use is business standardisation.
Business standardisation with AI does not simply mean turning the operating manual into a PDF, or feeding existing documents into an AI. It means organising the flow of the work, the decision conditions, the exception handling and the escalation routes down to a unit the AI can actually act on.
When I am asked to advise on AI adoption, I make a point of not opening with “so, what shall we automate?” I am more inclined to ask, “is the judgement on that task currently changing from person to person?” Bring in the AI alone while skirting this question, and the floor may find itself chasing extra checks rather than enjoying any new convenience.
Simply “having an operating manual” is not enough for an AI agent to act
Most companies already have some sort of operating manual. The form varies: an internal wiki, a spreadsheet, a PDF, Notion, a pinned post in Teams.
For an AI agent, however, what matters is not whether a manual exists, but whether it contains information on which a decision can actually be made.
Suppose, for example, you have an operating manual like this.
Check the contents of the application and, where necessary, confirm with a superior.
To a human, this more or less makes sense. To an AI agent, though, the information needed to decide is missing.
- What exactly is to be checked under “the contents of the application”?
- Under what conditions does “where necessary” actually apply?
- Who is “a superior”: the applicant’s direct line manager, or the head of the department that owns the process?
- After checking, is it to be approved, sent back, or put on hold?
The trouble with using an operating manual for AI, then, is this: the prose may exist, but it has not been broken down to a grain the AI can use.
When I come in to support a process improvement, before feeding the existing manual straight into an AI, I first mark up the points where a judgement arises. More often than not, the manual turns out to be peppered with phrases such as “as appropriate,” “as required,” “in principle” and “according to the situation.” These are handy expressions that lean on human experience. To an AI agent, though, they are also undecidable, woolly words.
For the standard operating procedure you hand to an AI agent, you need, at the very least, to organise the following information.
| Item | What to organise |
|---|---|
| Input | The request, files, forms and chat text the AI receives |
| Reference information | Regulations, FAQs, past responses, product information, customer information and so on |
| Decision conditions | The conditions for approval, answering, sending back, holding, or requesting confirmation |
| Output | The answer text the AI returns, the recorded content, the notification text, the next action |
| Exception handling | Cases the AI does not handle; cases that are passed to a person |
| Responsible owner | The final checker, the escalation route, the operations manager |
A standard operating procedure is not a write-up of the order in which tasks are done. It is a shared language for the work, one that helps anyone looking at it arrive at much the same judgement.
Introduce an AI agent without this shared language, and the AI simply mirrors the vagueness of the floor straight back at you. Get the shared language in order, on the other hand, and the AI becomes far more likely to act as a practical, workmanlike assistant to the business.
An SOP for an AI agent should be designed by separating “routine handling” from “exception handling”
In an SOP for putting an AI agent to work, it is important to design routine handling and exception handling as two separate things.
Routine handling means the work an AI agent may process along a fixed set of rules: answering questions clearly set out in internal regulations, summarising to a set format, handling enquiries based on an existing FAQ, and the like.
Exception handling, by contrast, means the work that is not to be completed by the AI alone: cases with unusual contract terms, cases with a large impact on the customer, cases needing a legal or HR check, cases where information is missing, and so on.
Leave this line blurred while the AI agent is running, and two problems follow.
- The AI answers things it ought not to answer.
- The AI bounces matters back to a person more than it needs to, so the work fails to move on after all.
On the sites I have supported, both of these were happening at the early trial stage. For one enquiry the AI would overreach with its answer; for another it would return even content it could perfectly well have handled with a curt “please check with the person in charge.” From the floor’s point of view, that is a state of “handy, but I am not ready to trust it yet.”
For that reason, in designing exception handling for AI, you make the branches explicit, as follows.
| Branch | The AI’s response | Example |
|---|---|---|
| Can answer clearly | The AI answers | Stating an application deadline that is explicitly set out in the regulations |
| Information is missing | Ask a follow-up question | Confirming the period covered, the amount, and the application category |
| Authority to decide is required | Escalate to a person | Approver away, exceptional approval, contract change |
| Risk is high | Do not answer; direct to the relevant department | Legal judgement, personal data, undisclosed information |
| No basis | Reply that “confirmation is needed” | No corresponding entry in the regulations or FAQ |
The point is not to make the AI “answer as much as it possibly can.” It is to draw a clear line between the range it may answer and the range it must not.
Japan’s AI Guidelines for Business (METI) similarly set out an approach in which businesses involved with AI recognise risk across the whole lifecycle and take the necessary measures. This way of thinking matters for the autonomous operation of an AI agent, too. Autonomy does not mean removing human checks; it means making clear the points at which a person ought to check. Japan’s AI Guidelines for Business (METI)
Exception handling is not about eliminating the unforeseen, but about deciding how to receive it
A common misconception when introducing an AI agent is that “if you simply list out every unforeseen scenario, things will be stable.”
In real work, however, you cannot enumerate every exception in advance. Departmental habits, customer-specific contract terms, rush requests, reorganisations, absent staff: exceptions arise all the time.
I have stood in plenty of rooms, on product development and business-reform projects, where someone has said “we hadn’t anticipated this.” It is not in the specification, but it happens on the floor. It is not on the workflow chart, but it is needed when dealing with a customer. That very margin is the reality of frontline work.
The aim of exception handling, therefore, is not to drive the unforeseen down to zero. It is to create a state in which, when the unforeseen does occur, the AI agent does not run amok, stops appropriately, and can hand over to a person.
What matters here are guardrails and fallbacks.
- Guardrail
- A boundary the AI must not cross. For example, rules such as “do not make judgements on changing contract terms,” “do not answer requests containing personal data,” and “do not speculate on content not covered by the regulations.”
- Fallback
- Where the AI falls back to when it cannot handle something autonomously. Rather than merely replying “please check with the relevant department,” it sets out which department, with what information attached, and how the confirmation should be made.
A poor fallback looks like this.
Please check this matter with the person in charge.
It seems harmless at a glance, but the requester has no idea what to do next. The upshot is a return to person-dependent ways of working: asking someone else on Slack, DMing whoever handled it last time, confirming verbally in a meeting.
A more practical fallback takes a form like this.
This application does not fall under the standard rules, so the AI cannot make a judgement. Please confirm with the approver in the sales planning team, attaching the following three points.
- The name of the customer the application concerns
- How it differs from the standard terms
- The desired deadline for a response
A fallback, then, is not merely a form of words for “stopping.” It is a design that lets the work keep moving forward even after it has been handed to a person.
When I design exception handling, I often put it as “let us decide how the AI should behave when it gets stuck.” More important, in practice, than the AI answering perfectly is that it stops correctly when it does not know.
In AI-ready design, decide what NOT to delegate to AI before deciding what to delegate
Discussions about introducing an AI agent tend to revolve around “what shall we hand to the AI?” From the point of view of stable operation, however, the first thing to settle is “what shall we not hand to the AI?”
While the work that is off-limits remains vague, the floor cannot tell how far to trust the AI’s answers. If people end up second-guessing every answer the AI gives, the checking effort only grows. Place too much faith in the AI, on the other hand, and you risk a mistaken judgement working its way into the business.
Examples of work not to delegate to AI include the following.
- Final contract decisions
- Answers requiring specialist judgement, such as legal, labour or tax matters
- Firm commitments to customers on changes of terms
- Decisions on HR appraisal or disciplinary action
- Processing involving personal or sensitive information
- Communications carrying the company’s official position
- Anything touching undisclosed information or management decisions
This does not mean the AI can play no part at all. There are supporting roles it can take: organising the points at issue, drafting, building comparison tables, drawing out the items to be checked.
That said, the final judgement rests with a person. It is important to write this dividing line into the standard operating procedure.
The OECD AI Principles set out a way of thinking about the development and use of trustworthy AI, and they are drawn on in shaping AI risk frameworks in various countries. In designing AI-ready operations within a company, too, you need to keep “trustworthiness,” “explainability” and “where responsibility lies” front of mind in your operating design.
When I support AI adoption, I place more weight on agreeing first on “the boundary the AI must not cross” than on widening “the territory the AI is handed.” It is precisely because that boundary exists that the floor can use the AI with confidence. AI use without a boundary may look convenient, but it tends to stall in operation.
Standard operating procedures fit reality better when built from the floor’s own logs
When drawing up a standard operating procedure, trying to sketch the ideal workflow from the outset tends to end in failure, because it fails to reflect the exceptions and the hesitations that actually occur on the floor.
What I would recommend instead is to take stock of the work as it really is, drawing on past enquiry logs, Slack exchanges, emails, tickets and meeting notes.
When I go onto a site myself, I rarely start by drawing a tidy workflow diagram. I look first at the exchanges that actually took place. Where did things stall? Who is being consulted? Which wording brings out the hesitation? Which enquiries keep coming round again and again? That is where the points to settle before introducing an AI agent reveal themselves.
Look through 82 enquiries over three weeks, for example, and you can sort them like this.
| Category | Content | Direction for AI handling |
|---|---|---|
| Standard answer | The answer is in the FAQ or regulations | The AI answers |
| Missing information | Required items are lacking | The AI asks a follow-up |
| Approval decision | A superior’s or owner’s judgement is needed | Escalation |
| Exceptional query | Does not fit the standard rules | Pass to a person |
| Manual not in place | No basis for an answer exists | Subject for knowledge development |
Sort things this way, and the work that ought to be handed to the AI, and the work that ought to be standardised first, both come into view.
Especially important are the cases the AI could not answer well. Hidden there are gaps in the operating manual, vagueness in the decision criteria, and a lack of clarity over where responsibility sits within the organisation.
Introducing an AI agent is not merely an automation project. It is also a project to make the vagueness of the work visible and to set it down as a standard operating procedure.
I think of this work as “the process of turning the floor’s hesitations into an asset.” As the hesitations are recorded, classified and reflected into the procedures, the response when the same case next arises grows steadily more stable.
If you use a tool, design business standardisation and training together
To run an AI agent steadily, you need more than system settings; you also need training that brings the floor’s ways of using it into line.
Here it matters how you combine tools such as AI chat, AI summarisation, e-learning, knowledge management and workflow management. There are several options. You might make use of an existing groupware or internal wiki, or you might use a business-support platform with AI features built in.
If you use Kanata, you can organise functions such as AI chat, AI summarisation and e-learning on a per-project basis, which makes it easier to connect the SOPs and exception-handling rules produced through business standardisation to learning data, prompts and training content. In Kanata, AI chat, AI summarisation and e-learning can be added as individual apps, and the design lets you organise AI settings, prompts and learning data within a project library.
You might, for example, operate it like this.
- Build a standard operating procedure from the work logs
- Register the SOP as learning data
- Turn the decision flows you use often into prompts
- Create training content for the floor
- Handle day-to-day enquiries through AI chat
- Review the cases the AI could not answer on a monthly basis
- Update the SOPs, prompts and training content
Set up this loop, and business standardisation stops being a one-off bit of document writing.
I take the view that the success or failure of AI adoption is decided less by the “initial setup” than by “the mechanism for returning to operation.” However splendid an operating manual you produce, it means nothing if the floor never looks at it. However careful the training, it will not take root unless its use is refreshed in the course of daily work.
In training for putting AI agents to work, and in advanced reskilling, what matters is not only teaching how to operate the tool. It is enabling the floor to judge which work may be handed to the AI, in which cases to return it to a person, and how to check which outputs.
In running an AI agent, keep growing the rules by reading the logs
An AI agent is not a set-it-and-forget-it affair. The work, the organisation, customer conditions and internal regulations all change. The standard operating procedure you first drew up will, in time, drift away from the floor.
For that reason, running an AI agent calls for a mechanism that keeps growing the rules by reading the logs.
The logs worth reviewing fall mainly into four kinds.
| Log | What to look for |
|---|---|
| Logs where the AI answered | Any wrong answers, over-confident assertions, or lack of basis |
| Logs where the AI handed over to a person | Whether there are too many escalations |
| Logs the floor corrected | Whether the SOP or prompts have gaps |
| Unresolved logs | Whether new exception patterns are arising |
It is realistic to set the frequency of this review according to volume and risk: weekly in the early days of adoption, then monthly thereafter, say. In areas where a wrong answer carries heavy consequences, though, such as legal, labour, finance, healthcare or safety management, you need to check on a shorter cycle.
What to watch here is that you do not judge the AI on answer accuracy alone. You need to separate out whether the cause of an AI mistake lies in the prompt, in the learning data, or in the business rules.
A failure I often see is one in which everything gets filed away under “the AI is not accurate enough.” There may, of course, be room to improve the AI’s output quality. In reality, though, the cause may equally be stale reference data, business rules that were never settled, vague approval authority, or an undecided owning department.
While the business rules stay vague, no amount of tuning will make the AI stable. Conversely, the further business standardisation progresses, the clearer the range the AI agent can cover becomes.
Growing an AI agent does not mean improving the AI alone. It means reviewing the work itself and, little by little, meshing human judgement with the AI’s processing.
Whatever your role, there are points everyone should have in common
The thinking on business standardisation and exception handling needed to put AI agents to work is not a matter for one department alone.
For the executive, it is a question of investment and risk: how far across the business to entrust the AI agent. For the head of IT, it is a question of system integration, access management, log auditing and security design. For the head of marketing and sales, it is a question of standardising customer handling, deal support, knowledge sharing and proposal quality.
The points to watch differ by role, but the foundation they share is the same.
- Separate the work to delegate to AI from the work not to delegate
- Get standard operating procedures down to a decidable grain
- Decide the exception handling and escalation conditions
- Design the guardrails and fallbacks
- Review continually on the basis of frontline logs
- Run not only operational training but training on business design
I am wary of talking up an AI agent as “something that quietly clears the work away for you.” Such expectations may give adoption a push in the short term. But to build an AI that keeps being used on the floor, you need a rather more humdrum, practical sort of design.
An AI agent is not something that magically tidies the work away on its own. If anything, it is something that mirrors the vagueness of the work back at you.
That is precisely why introducing an AI agent is also a fine opportunity for business standardisation. You put into words the judgements you have been running on experience and gut feel, and you bring them closer to the same quality whoever is on duty. Only with that foundation in place does an AI agent become able to support the work reliably.
In summary: an AI agent shows its strength on top of standardised work
What putting AI agents to work calls for is not advanced technology alone. If anything, what you need first is to break the floor’s work down carefully and to set the standard operating procedures and exception handling in order.
While the work stays person-dependent, an AI agent cannot run steadily. Even with a manual, if the decision conditions are vague, the AI is left either to guess or to bounce things back to a person.
On the other hand, once the SOPs, the branches, the escalation, the guardrails and the fallbacks are in order, the range the AI can cover and the range a person should judge both become clear.
The key is not to try to make everything autonomous from the start. Begin with work whose decision conditions are easy to organise, such as enquiry handling, minute-taking, an internal FAQ, or pre-submission checks. Then, on the basis of the logs you gather there, update the SOPs and the exception handling. It is this accumulation that leads to the stable operation of an AI agent.
On AI adoption sites, I am sometimes told, “this is more low-key than I expected.” It is true: reading the work logs, classifying the exceptions and tidying up the fallback wording is hardly flashy work. And yet it is precisely that low-key design that holds up an AI which keeps being used on the floor.
Business standardisation with AI is not a preliminary step before AI adoption; it is the central step that lets AI take root in the work.
Create a state in which the AI agent acts without hesitation and people can concentrate on the work they ought to be judging. The first step is to put your own work into words at “a grain the AI can understand.”
Q&A
Where should I begin with business standardisation for AI?
The first thing to do is to gather the logs for the target work. Look at enquiry histories, Slack exchanges, emails, tickets and meeting notes, and see where the judgements diverge. Rather than building the ideal flow straight off, building the SOP from the hesitations and exceptions that actually occur on the floor fits reality far better.
If we have an operating manual, can the AI agent be used straight away?
Not necessarily. If the manual is full of vague expressions such as “as required,” “as appropriate” and “in principle,” the AI agent will hesitate over the judgement. To let an AI use it, you need to make the input, the reference information, the decision conditions, the output, the exception handling and the responsible owner explicit.
What is the most important thing in exception-handling design for AI?
It is to separate the range the AI may answer from the range it must not. In particular, legal judgements, personal data, contract terms, undisclosed information and firm commitments to customers need a design that does not let the AI complete them on its own. Rather than forcing it to answer when it does not know, you provide a fallback that lets it stop appropriately and hand over to a person.
How often should we review an AI agent’s logs?
A workable approach is weekly at the start of adoption and monthly thereafter. For work where a wrong answer has a large impact, check on a shorter cycle. The points to look at are not only the AI’s wrong answers: check too whether there are too many escalations, where the floor is making corrections, and whether unresolved exceptions are on the rise.
In what situations is Kanata easy to consider?
It is an easy option to consider where you want to combine AI chat, AI summarisation, e-learning and a project library, and to advance business standardisation and frontline training as one. It suits operations where, for instance, you organise SOPs as learning data, turn frequently used decision flows into prompts, and roll these out into training content for the floor. That said, an existing internal wiki or workflow system may be sufficient, so it is important to weigh it up against your own information management, access design and the range of departments using it.