AI Risk Monitoring After Deployment: How to Design Incident Response and Key Safeguards

Column
AI Risk Monitoring After Deployment: How to Design Incident Response and Key Safeguards

Introduction

This article sets out how to design AI risk monitoring and embed AI incident response into day-to-day operations. The goal is not to halt AI adoption on the grounds of risk, but rather to detect anomalies at the earliest opportunity and ensure that all stakeholders can act swiftly and without hesitation when issues arise.

Tatsuya Ito

Tatsuya Ito

Artificial Intelligence Consultant

company-icon

Third Scope Inc.

Born in 1985 and originally from Mie Prefecture, Japan. In 2012, he joined an AR startup in Hong Kong as an engineer. Since then, he has been involved in new business development and AI service launches at several AI startups. In 2018, he founded the current Third Scope Inc. by taking over an AI service and its development team. He now supports companies in adopting and utilizing AI, with a focus on AI-driven business development, operational transformation, and product development. He has also been involved in AI research as a Project Researcher at the University of Tokyo. Today, he continues to work at the forefront of AI project development, providing practical consulting from both technical and business perspectives.

“Who checked this answer? Couldn’t it have been stopped before it went out to the customer?”

That question was raised at a Monday morning meeting by Morita (a pseudonym) of the Information Systems Department at Aoba Manufacturing Co., Ltd., who is responsible for the rollout of generative AI.

At the company, as more and more departments began using generative AI, the way each department handled incorrect outputs, checked usage logs, managed access rights, and reported problems when they arose all differed. The Legal Department was worried about compliance risks, while the DX Promotion Department wanted to avoid excessive controls that would stifle adoption. Voices from the Sales Department also raised concerns, saying, “We don’t know what kind of event should be reported as an incident.”

The company now runs its operations using a combination of usage logs on Kanata and internal rules. Looking at the 47 generative-AI-related enquiries received in the three months after this approach was introduced, the average time from receiving an enquiry to making an initial decision is said to have fallen from 2.5 working days before the new approach to 0.8 working days.

This article sets out how to design AI risk monitoring and build AI incident response into everyday operations. The aim is not to halt AI adoption for the sake of risk. It is to create a state in which warning signs are spotted early and everyone involved can act without hesitation.

Preparing a classification table or a tool alone will not prevent every risk. Organisations need to combine training, access management, record-keeping and regular reviews, and continuously improve their own AI governance.

The risks after introducing AI are not limited to “wrong answers”

画像待ち 6-9-en - The risks after introducing AI are not limited to

When people think of the risks of generative AI, hallucination is often the first thing that comes to mind.

Hallucination refers to a phenomenon in which generative AI presents information that is not based on fact as though it were factual. Examples include citing internal regulations or customer case studies that do not exist, or confidently stating figures that are not found in any source document.

However, the risks that a company needs to manage are not limited to errors in the content of answers. The following issues are also worth considering.

  • The AI is able to reference documents that users should not, in principle, have access to
  • Customer information or personal data is entered into prompts more extensively than necessary
  • AI output is used in customer-facing documents or emails without being checked by a person
  • Answers are based on outdated regulations or old price lists that predate a revision
  • Usage rules differ from department to department, so it is unclear where to report a problem
  • Usage logs are kept, but it has not been decided who checks them or against what criteria
  • Access rights remain in place for staff who have changed roles or left the company
  • Administrators are not aware of the full scope of data shared with external services

In other words, the risks that arise after adopting AI are not simply a matter of “the AI getting things wrong” — they also stem from how the organisation introduces, uses and oversees the AI.

For this reason, AI risk monitoring needs to cover not just the content of outputs, but also users, input data, reference data, access rights, where outputs are used, approval procedures, and reporting routes.

the “AI Business Operator Guidelines (Version 1.2)” jointly issued by the Ministry of Economy, Trade and Industry (METI) and the Ministry of Internal Affairs and Communications (MIC) also sets out the principle that organisations should recognise the risks of using AI and take the necessary measures across the entire lifecycle. It also calls for measures such as documenting and retaining relevant information and improving AI literacy, in order to ensure transparency and accountability.

AI governance is not a mechanism for halting adoption

画像待ち 6-9-en - AI governance is not a mechanism for halting adoptionの挿絵

AI governance is a framework through which an organisation sets out policies, responsibilities, procedures and oversight methods for using AI, so that it can achieve its objectives while managing risk.

It is not something that only concerns the executive board or the audit department.

Even in everyday work — for example, a salesperson having AI draft the first version of a proposal, an HR staff member drafting an answer to an FAQ about internal regulations, or a marketing staff member discussing the structure of an article with AI — the following judgement calls are needed.

  • How much can be left to the AI
  • At what stage human checking is required
  • What information may be entered
  • Who to report to if a problem is noticed
  • Whose approval is needed before sending anything outside the organisation

Turning these decision criteria into something that can actually be carried out on the ground is also part of AI governance.

If the rules are vague, the more cautious departments may simply avoid using AI altogether. On the other hand, if approval procedures and checklists become too detailed, the operational burden increases and the rules risk becoming a mere formality.

What is needed is not to restrict all uses uniformly, but to vary the level of control depending on the purpose, the sensitivity of the information, and the potential impact on external parties.

For example, the level of checking required differs between using AI to organise one’s own ideas and using it to draft a contract-related document to be sent to a customer. Organisations need to design a “workable operation” that classifies risks, detects them, carries out an initial response, and feeds lessons back in to prevent recurrence.

First, decide the definition of an “AI incident”

画像待ち 6-9-en - First, decide the definition of an

When setting up AI incident response, the first thing to decide is “what kind of event counts as an incident”.

If operations begin without a definition in place, staff on the ground will run into uncertainty of the following kind.

  • Should even a simple wrong answer from the AI be reported
  • Is a report unnecessary if the document was only used internally
  • Does it count if the error was caught before being sent to a customer
  • If it was shared on an internal chat, who needs to be told
  • What should be done if no actual harm occurred, but the incident could have led to an information leak

the OECD’s “Defining AI Incidents and Related Terms” distinguishes between events in which AI has actually caused harm and hazardous events that could potentially lead to harm. Within a company, too, it is useful to record not only “incidents that caused actual harm” but also “near misses” — such as errors caught before publication or information entered incorrectly — as this makes them easier to use for preventing future accidents.

When classifying events, at least the following points should be checked.

  • Whether the information reached anyone outside the organisation
  • Whether personal or confidential information was involved
  • Whether it affected any internal or external decision-making
  • Whether it can be withdrawn or corrected
  • Whether it may breach the law, a contract, or internal regulations
  • Whether the same type of event has occurred repeatedly
Example severity classification for AI incidents
Level Classification Main assessment criteria Approach to response
Level 1 Minor error / near miss Discovered during internal use, no external impact, can be corrected on the spot Record it and use it to improve prompts or reference data
Level 2 Potential impact on internal business decisions Could affect decisions relating to regulations, contracts, revenue, HR, etc. Check with the responsible department and make an initial decision within a set deadline
Level 3 Has affected, or is about to affect, an external party Reached a customer or business partner, or was about to be published/sent Check the scope of distribution and consider correction, retraction, or suspending publication
Level 4 Possible information leak or serious breach of law or contract Personal data, trade secrets, credentials, etc. may have been handled inappropriately Escalate into the existing information security incident response process

Level 1: Minor error / near miss

An event discovered during internal use, with no external impact, that can be corrected on the spot.

Examples include the following.

  • Using outdated internal terminology
  • Misreading the context in part of a summary
  • Generating a heading that was not in the original text
  • An inappropriate expression that was corrected before publication

At this stage, rather than treating it as a serious incident, it is better to record it as a near miss and use it as material for improving prompts and reference data.

That said, if the same type of error keeps recurring, or if it occurs in a critical area such as finance, HR, or contracts, it may need to be treated at a higher level.

Level 2: Events that could affect internal business decisions

An event in which the AI’s output could affect internal decisions or business processes.

  • Misinterpreting internal regulations
  • Overlooking an important point when reviewing a contract
  • Making an unfounded, definitive statement about sales forecasts or customer response policy
  • Being used for decisions with a significant impact on people, such as recruitment, performance evaluation, or credit assessment

Rather than the user reaching a conclusion alone, the matter should be checked with the relevant department — Legal, HR, Finance, or IT — depending on the subject.

As well as deciding who to check with, it also helps to set a deadline — for example, “an initial decision within X hours or X working days” — to keep the response from stalling.

Level 3: Events that have affected, or are about to affect, an external party

An event in which the AI’s output reached a customer, business partner, job candidate, the media, or similar, or was about to be sent or published.

  • Sending incorrect product specifications to a customer
  • Including a non-existent case study in a proposal document
  • Publishing unverified figures in an advertisement or on a website
  • A response to one customer containing information about a different customer
  • Text generated by AI potentially causing a copyright issue or an advertising-display compliance issue

At this stage, before analysing why the AI made the mistake, check “who received what information, when, and by what route”.

As needed, report to the sales manager, PR, Legal, information security staff, and senior management, and consider correction, retraction, suspending publication, or explaining the matter to the customer.

Level 4: Possible information leak or serious breach of law or contract

An event in which personal data, trade secrets, unpublished information, credentials, or information subject to contractual restrictions may have been handled inappropriately through the use of AI.

In this case, rather than treating it as a routine operational improvement matter, it should be escalated into the existing information security incident or personal data breach response procedure.

  • The information that was entered or generated
  • The user and the date and time of use
  • The service, app, or project used
  • Where the data was sent and stored
  • Who it was shared with, and the scope of sharing
  • The possibility of external access
  • Whether the data can be deleted or its use suspended
  • Whether there is a reporting obligation under law, contract, or internal regulations

Whether there is a legal reporting obligation, or a need to notify the individuals concerned, differs depending on the type of information, the scope of impact, contractual terms, and the country or region in which the business operates. Legal, personal data protection, and information security staff need to make these judgements on a case-by-case basis.

Narrowing down what to monitor

画像待ち 6-9-en - Narrowing down what to monitorの挿絵

The term “AI risk monitoring” might make some people imagine a system that constantly watches every conversation.

However, having a person check every single input and output becomes harder to sustain as usage grows. Consideration also needs to be given to notifying employees, privacy, labour management, and who has access to the logs.

In practice, what to check and how often should be decided according to the purpose and the sensitivity of the information.

  • Focus checks on high-risk departments or highly confidential projects
  • Detect signs of information being entered that is banned, such as personal data or credentials
  • Carry out regular sampling reviews of outputs intended for external use
  • Check for unusual behaviour, such as a sudden surge in usage or large volumes of data being entered
  • Review, after the fact, any conversations that led to an incident or enquiry
  • Increase the frequency of checks for a set period after launching a new use case

When designing monitoring, it is important to decide not only “what to check” but also who checks it, how often, and against what criteria.

Input data

The first thing to check is the information entered into the AI.

Check whether customer names, personal names, contract amounts, unpublished management information, passwords or API keys, bank account details, health information, and the like are being entered unnecessarily.

Particular care is needed when handling contracts, meeting minutes, sales-meeting notes, HR documents, and enquiry histories.

Input rules should be simple enough for users to apply while carrying out their day-to-day work.

  • Do not enter personally identifiable information unless it is genuinely needed for the task
  • Anonymise or mask information where necessary
  • Do not enter credentials such as passwords, private keys, or API keys
  • Do not enter unpublished financial or management information outside approved environments
  • Handle confidential customer information only after checking the contractual terms and the environment being used
  • If in doubt, do not enter the information and check with the responsible department

The principle of “if in doubt, don’t enter it” is easy to understand, but if there is nowhere to turn for advice, work will grind to a halt. A point of contact and a response deadline should also be set.

Reference data

Next, check the internal documents and knowledge that the AI draws on to generate answers.

This includes internal regulations, sales materials, FAQs, product manuals, contract templates, and past meeting minutes.

A common problem is that outdated regulations, old price lists, and old organisation charts are left in place instead of being removed. If users cannot check the source, they may accept an answer based on outdated information as correct.

Reference data should carry at least the following management information.

  • Document name
  • Responsible department
  • Version or revision number
  • Date last updated
  • Confidentiality classification
  • Departments or roles permitted to use it
  • Person responsible for updates
  • Next review date
  • Expiry date
  • Where the original is stored

Simply registering a document does not complete the management process. A person responsible for updates and a review frequency need to be set, and expired documents must be reliably removed.

Users and access rights

The risk posed by AI varies greatly depending on who can access which information.

Information available to the whole company needs to be separated from information handled only by specific departments. Bringing together customer information from Sales, evaluation data from HR, contract information from Legal, and unpublished materials belonging to senior management into the same environment can lead to unintended sharing of information.

The following principles should form the basis of access management.

  • Divide projects based on “who may view the same information”
  • Grant users only the access rights they need
  • Keep the number of people with administrator rights to a minimum
  • Set deadlines for changing access rights when someone changes role, takes a leave of absence, or leaves the company
  • Avoid using shared accounts as a rule
  • Restrict highly confidential documents to a defined department, purpose, and period of use
  • Carry out regular reviews of access rights
  • Set expiry dates for temporary access rights

The convenience of consolidating information and the security provided by restricting access need to be designed together, as a pair.

Where outputs are used

Where an AI’s output is used is also an important thing to monitor.

The same output requires a different level of checking depending on whether it is used as a personal note or published as a customer-facing document.

  • Personal thinking aids or drafts
  • Reference material within a department
  • Internal decision-making materials
  • Communications to customers, business partners, or job candidates
  • Advertisements, websites, press releases
  • Decisions relating to contracts, legal matters, finance, HR, or safety

As a rule, any output that could go outside the organisation should be reviewed by a person. In particular, there needs to be a rule that figures, proper nouns, quotations, sources, legal judgements, and content relating to medicine, finance, or safety should never be used on the basis of AI output alone.

Reviewers should be given specific checkpoints, rather than simply being told to “check the content”.

  • Does it match the primary source
  • Are the figures, dates, product names, and people’s names correct
  • Are there any unfounded, definitive statements
  • Does it contain confidential information about a customer or third party
  • Are there any discriminatory, offensive, or misleading expressions
  • Are there any copyright, advertising-display, or contractual issues
  • Has it been approved by the necessary person in charge

Think of the initial response as “stop, keep, and tell”

画像待ち 6-9-en - Think of the initial response as

When an AI-related problem is suspected, the first thing that leaves people on the ground unsure is “where to start”.

If the initial response is made too complicated, reporting may be delayed. It becomes easier to manage if it is organised into three stages: “stop”, “keep”, and “tell”.

  1. Stop

    As soon as the problem is noticed, stop the impact from spreading further.

    • Stop using the incorrect answer
    • If it has not yet been sent to the customer, stop it from being sent
    • If it has already been published, check whether it needs to be taken down or corrected
    • If it is an internally shared document, temporarily restrict who it is shared with
    • If there is a risk of an information leak, suspend the relevant account or integration
    • Temporarily suspend any work that uses the same prompt or reference data
  2. Keep

    Next, keep a record so that the facts can be verified.

    • The date and time the problem was noticed
    • The date and time the AI was used
    • The user
    • The AI service, app, or project used
    • A summary of what was entered
    • The output
    • The reason it was judged to be a problem
    • The data or documents referenced
    • Who it was shared with, and the scope of sharing
    • Whether there was any impact on customers or external parties
    • Any action taken at the time
    • Where the related logs or screenshots are stored

    Care should also be taken not to needlessly copy personal or confidential information into the incident report. Detailed information should be stored somewhere with restricted access, and the report form itself should contain only the minimum information necessary.

    The purpose of record-keeping is not to blame the person who reported the issue, but to understand the impact, analyse the cause, and prevent recurrence. If a system penalises people for reporting problems quickly, incidents and near misses are less likely to come to light.

  3. Tell

    Pass on what has been recorded to the pre-defined point of contact.

    Example reporting destinations by AI incident classification
    Event Main reporting destination
    Minor error Project manager, business owner
    Error affecting a business decision Responsible department, department head
    Event with external impact Sales manager, PR, Legal, senior management
    Possible information leak Information security, personal data protection, Legal
    Event relating to HR or recruitment HR manager, Legal, senior management if necessary

    In addition to the reporting destination, a deadline for the initial decision, an out-of-hours contact, and the person ultimately responsible for the response should also be decided in advance.

    Simply telling a manager verbally, or posting in an internal chat, can leave it unclear who is responsible, and the matter may fall through the cracks. Where possible, issue a ticket number through an enquiry management system so that progress, deadlines, and the person responsible can be tracked.

Don’t let recurrence prevention stop at “users being more careful”

画像待ち 6-9-en - Don't let recurrence prevention stop at

One thing to avoid in incident response is concluding simply with “we will be more careful in future”.

While users being careful is necessary, it is difficult to prevent the same problem recurring in day-to-day work through vigilance alone.

When preventing recurrence, break the cause down into at least the following four areas.

Problems with the prompt

Check whether the instructions given to the AI were vague, leading to incorrect answers or overly definitive statements.

For example, an instruction such as “please review this contract” alone may not align the scope of the check or the assessment criteria with what the user actually intended.

“Please organise the clauses that could be disadvantageous to our company, from the perspective of liability for damages, confidentiality, contract termination, intellectual property, and governing law. Flag any points you cannot judge as ‘needs confirmation’, and do not provide a final legal opinion.”

Making the checkpoints and output conditions explicit in this way helps reduce variation in the output.

That said, improving the prompt does not guarantee the accuracy of the output. For contract or legal matters, checking by a specialist should always be assumed.

For frequently used tasks, rather than leaving it to individuals, the responsible department should manage a reviewed prompt as a template.

Problems with reference data

Check whether the documents referenced by the AI had any of the following problems.

  • The information is out of date
  • The scope of application is not clearly stated
  • The source or original document is unknown
  • The content is duplicated or contradictory
  • Pre-revision versions of the document remain in place
  • Access rights are too broad
  • Documents that the AI should not use to formulate answers are included

As countermeasures, replace with the latest version, expire old versions, standardise naming conventions, set update dates, and clearly state the responsible department.

Problems with access design

Check whether users who should not, in principle, have access were able to use a particular AI app, project, or reference data.

Particularly for cross-departmental projects, the guiding principle should be “include only those who may view the same information”, rather than “include everyone because it’s convenient”.

Beyond individual access errors, also analyse whether there were problems in the overall process of requesting, approving, granting, reviewing, and revoking access rights.

Problems with training and operations

Check whether users understood the following.

  • What information may and may not be entered
  • How to check the AI’s output
  • The approval procedure for external use
  • Who to report to if a problem is noticed
  • How to stop use in an emergency
  • How usage logs and records are handled

Training is not something to be done once and forgotten.

Update the training materials and rules whenever a new use case is launched, a service’s specifications change, the organisation is restructured, or an incident occurs.

It is also important to check not just completion rates, but whether users can actually make sound judgements in practice, through case exercises and comprehension checks.

AI risk management should be run across departments

画像待ち 6-9-en - AI risk management should be run across departmentsの挿絵

AI risk management cannot be handled by the IT department alone. It is equally difficult to run it sustainably through Legal alone, the DX Promotion Department alone, or front-line staff alone.

Responsibilities should be shared across several departments, while keeping accountability clear.

Senior management

Senior management sets the risk appetite and priorities for AI adoption.

  • Tasks where AI should be actively used
  • Tasks where AI use is banned or restricted
  • Information that should not, in principle, be handled by AI
  • The threshold of incident severity that requires reporting to management
  • The people and budget invested in risk measures
  • Who is ultimately accountable for adoption and control

Senior management does not need to check every individual AI response, but it does need to set out the level of risk it is willing to accept and the scope of accountability.

Chief information officer / IT department

The CIO or IT department manages the system environment, accounts, logs, access rights, and external integrations.

  • Who can use which functions
  • Which documents are registered in each project
  • Which logs are kept, and for how long
  • Who can access the logs
  • How often the logs are checked
  • When access rights are changed following a role change or resignation
  • What data is sent to external services

Simply storing logs does not constitute monitoring. The conditions for checking them, who is responsible, and the criteria for escalation all need to be decided.

Legal / risk management department

Legal and risk management departments establish the judgement criteria and the conditions under which specialist departments should be consulted.

In areas such as contracts, personal data, copyright, advertising compliance, industry regulation, discrimination, and recruitment or performance evaluation, decisions cannot be made on the basis of AI output alone.

They should make clear in which cases a specialist department’s review is required, and set up a point of contact that is easy for front-line staff to consult.

DX Promotion Department / business units

The DX Promotion Department and business units embed AI adoption into day-to-day operations.

  • Creating templates that are easy for staff on the ground to use
  • Sharing not only success stories but also failures and near misses
  • Setting up review checklists for each task
  • Feeding enquiries from users back into improvements
  • Regularly evaluating both the benefits and the risks

Their role is not just to create rules, but to verify that they can actually be carried out on the ground.

Front-line users

Front-line users are at the very front of AI governance.

  • Not taking AI output at face value as fact
  • Checking figures, proper nouns, quotations, and dates against primary sources
  • Getting a review before sending anything to a customer or outside the organisation
  • Not ignoring a sense that something is off
  • Reporting problems quickly once noticed
  • Checking the sensitivity of information before entering it

However, responsibility should not be concentrated solely on front-line users. The organisation itself needs to provide the rules, points of contact, training, and system-level controls that allow users to make sound judgements.

An example operational design when using Kanata

画像待ち 6-9-en - An example operational design when using Kanataの挿絵

There are several options for managing AI use in a business setting — an internal portal, individual generative AI services, a knowledge management system, an enquiry management tool, and so on.

What matters is not simply which tool is chosen. It is whether users, purposes, reference data, access rights, and those responsible for management can all be managed as a single, consistent unit.

Where projects, AI apps, and libraries can be used within Kanata, one approach is to use these as the units for managing operations and access rights.

Example of dividing Kanata projects by function
Example project Example information registered and managed
Company-wide shared project Public information, common FAQs, AI usage rules
Sales project Proposal templates, approved customer-response knowledge
HR and general affairs project Internal regulations, enquiry FAQs, training materials
Legal project Contract templates, review criteria, checking workflow
Management project KPI definitions and management materials. Whether to register highly confidential information should be decided only after checking the environment and access rights

The criteria for dividing projects should not simply follow the organisation chart, but rest on the following three points.

  • Whether the same people may view the same information
  • Whether it is used for the same purpose
  • Whether the same person can be responsible for managing it

Where prompts and reference data can be managed as a library, build in a review at the time of registration and periodic housekeeping.

  • Data with no clearly assigned responsible department
  • Documents whose update deadline has passed
  • Duplicate documents
  • Prompts that are not being used
  • Templates that have led to incorrect answers
  • Data left without an owner because its creator has left the company

Leaving this kind of information unattended can become a cause of incorrect answers or unauthorised access.

Where usage logs can be checked, look not just at simple usage counts, but at whether the items needed for an incident investigation can actually be retrieved. Specifically, this includes the user, the date and time of use, the app or project used, the reference data, the operation history, and the history of changes made by administrators.

AI risk monitoring checklist

画像待ち 6-9-en - AI risk monitoring checklistの挿絵

Items to check at the time of adoption

  • The AI usage policy and its scope are documented
  • Information that is banned from being entered, or that requires approval, is defined
  • Definitions of an AI incident and a near miss exist
  • Severity classifications and reporting destinations are defined
  • The deadline for the initial decision and who is responsible are defined
  • Projects and access rights match actual working practice
  • The department responsible for reference data, and who is in charge of updating it, are defined
  • There are review criteria for use before sharing externally
  • The scope, retention period, and viewing rights for logs are defined
  • Users have been notified and trained

Items to check on a regular basis

  • There are no unusual usage patterns
  • AI output is not being used unchecked in high-risk tasks
  • No reference data past its review deadline remains in use
  • Access rights have been updated for staff who have changed role, taken leave, or left the company
  • Near misses are being reported from the front line
  • The same type of incorrect answer is not recurring
  • Prompt templates are being improved
  • The time taken for the initial decision on incidents is not getting worse
  • No cases remain outstanding past their response deadline
  • The scope and frequency of monitoring match the level of risk

Not every item needs to be checked every month. Depending on usage volume, the sensitivity of the information, and the size of the organisation, checks can be spread across monthly, quarterly, or half-yearly cycles.

Items to check after an incident has occurred

  • The facts and any assumptions were recorded separately
  • Whether there was any external impact, and its scope, was confirmed
  • It was shared with the relevant departments and those responsible
  • Logs and records serving as evidence were preserved
  • The cause was analysed by breaking it down into prompts, reference data, access rights, training, and systems
  • Measures to prevent recurrence were reflected in rules, settings, templates, and training materials
  • It was checked whether the same risk exists in other departments or projects
  • The appropriateness and speed of the response were reviewed
  • It was considered whether the classification criteria or reporting routes needed to change

Summary: building risk response into everyday operations

画像待ち 6-9-en - Summary: building risk response into everyday operationsの挿絵

Monitoring risk and responding to incidents after adopting AI is not something that only specialists should be responsible for.

There are areas — usage logs, access rights, checking legal and contractual compliance — that specialist departments should own. But it is often front-line staff who use AI on a daily basis who are the first to see its output.

Noticing when something feels wrong, stopping it before it reaches the customer, and reporting it promptly are all important roles played by people on the ground.

That is precisely why AI risk management needs to be thought of not just as “regulations for policing violations”, but as “operational design that allows people to use AI with confidence”.

It will be difficult to get everything — handling incorrect outputs, incident classification, initial response, recurrence prevention, usage logs, and access rights — perfectly in place from the outset.

One approach is to start with the following three points.

  1. Decide what should be reported
  2. Decide who to report to when a problem is noticed
  3. Review the operation monthly or quarterly

From there, update your classifications, training, access rights, and templates based on the enquiries and near misses that actually occur.

AI can make mistakes, and if the context or reference data provided is insufficient, it may generate inaccurate answers. What matters is making clear which parts require human checking, and building a mechanism through which the organisation keeps learning and improving.

Building risk response into everyday operations, so that AI adoption is not halted — that is the next operational challenge for any company that has introduced AI.

Q&A on AI risk monitoring

If the AI gives a wrong answer, should it always be reported as an incident?

Not everything needs to be treated as a serious incident.

Minor errors that have no external impact and can be corrected on the spot can instead be recorded as near misses. On the other hand, if the output was used for a business decision, sent to a customer, or contains personal or confidential information, it needs to be reported to the relevant department.

Rather than leaving the decision to the individual user, it is important to have a classification based on criteria such as the scope of impact, the sensitivity of the information, and whether it can be corrected.

Do all AI usage logs need to be checked by a person?

Not necessarily — every single log does not need to be checked by a person.

Where usage volume is high, options include focusing on high-risk tasks or highly confidential projects, extracting usage that meets certain conditions, or sampling a fixed proportion of cases.

That said, simply storing the logs is not enough. It is necessary to decide who checks them, how often, the criteria for judgement, and who to report to if something unusual is found.

Should an error found before it was sent to a customer still be recorded?

Yes, it is recommended to record it even if it had no external impact.

Events discovered before publication or sending, which did not turn into a serious incident, can be treated as near misses and used as material for improving prompts, reference data, and review procedures.

In particular, if the same error occurs multiple times, it may indicate a systemic issue rather than an individual’s checking mistake.

To what extent should AI output be checked by a person?

The level of checking required depends on how the output will be used.

A simple check may be enough for personal note-taking or idea organisation, but content relating to customer-facing materials, advertising, contracts, finance, HR, or safety requires checking by the responsible department.

At the very least, figures, dates, proper nouns, quotations, and sources should be checked against primary sources, and the content should be reviewed for unfounded, definitive statements or confidential information.

Where should we start when setting up AI incident response?

There is no need to introduce a complex monitoring system from the outset.

Start by deciding the following.

What events need to be reported Severity classifications The initial response to be taken on the ground The reporting destination for each classification The deadline for the initial decision What information to record

From there, review your classifications and reporting procedures based on actual enquiries and near misses. Rather than aiming for a finished system from day one, it is more realistic to improve it through operation.