I thought this answer was perfectly fine, but the team next door sent it back for revisions.
That was Sato (a pseudonym), who drives generative AI adoption in the sales planning department, speaking in a meeting room on a Monday morning. Draft answers for the internal FAQ, summaries of proposals, first drafts of customer emails. The occasions for checking the quality of AI output had grown, yet the judgements of the people on the front line, the quality-management team, and the DX office were not aligned.
I have sat in on scenes like this many times. The mood in the room is by no means hostile. Quite the opposite: everyone is earnestly trying to protect quality. The trouble is that one person is looking at whether the writing reads naturally, another at whether the facts are correct, and yet another at whether it would feel right to put in front of a customer. Everyone is saying something valid, and still the judgements fail to line up.
In the past, comments tended to be impressionistic, along the lines of “easy to read”, “a bit worrying”, or “doesn’t sound like us”, and the standards for hallucination checks, fact-checking, and tone of voice differed from person to person. These days, things are shifting towards an arrangement in which we keep a proper AI evaluation rubric that scores accuracy, completeness, expression, risk, and reproducibility on a five-point scale when reviewing deliverables produced with AI chat or AI summarisation, and in which a weekly sampling review is used to iron out discrepancies in judgement.
This article explains how to organise the problem when reviews of AI output become person-dependent and judgements of AI output quality vary from one reviewer to the next. The aim is a state in which, whoever does the checking, the points to inspect and the acceptable thresholds are shared, and in which the operation of your quality-management AI keeps improving over time.
That said, simply drawing up review criteria will not solve everything. It is only when education, operating rules, and regular reviews are added that the criteria become something usable on the ground. If any of this sounds familiar, do read on with your own review meetings in mind.
Why reviews of AI output quality so easily become person-dependent
As the use of generative AI advances, the first thing that happens at many companies is that the volume of output increases.
Draft emails, minutes, FAQs, proposals, research memos, advertising copy, internal notices. Deliverables that people once built from scratch can now be drafted by AI in a fraction of the time.
How that output is to be evaluated, on the other hand, tends to be put off.
On the ground, conversations like the following arise.
- The writing is natural enough, but the content is a touch shallow.
- I am uneasy because the basis for the figures is unclear.
- It is a little too casual to send to a customer.
- Is this wording all right from a legal standpoint?
- It was fine last time, so why is it a no this time?
Every one of these is an important point. But when the angles of the comments are not aligned, the AI review criteria end up resting on individual instinct.
One person prioritises readability, another prioritises fact-checking. One department scrutinises tone of voice strictly, another puts speed first. Let this carry on and the management of AI output quality becomes unstable.
What I often notice when supporting teams is that the companies whose reviews descend into confusion are not the ones with low quality awareness. If anything, it is the reverse. It is precisely because they care about quality that each person inspects things closely, drawing on their own experience. The snag is that this experience has never been put into words, so every review raises the question of whose standard the judgement should follow.
The problem is not the performance of the AI alone. It is that, on the human side, there is no shared language for what counts as good output.
AI review criteria are about aligning the basis for judgement, not about voicing impressions
The purpose of drawing up AI review criteria is not to silence reviewers’ impressions.
Rather, the purpose is to convert the front line’s instincts and experience into a basis for judgement that others can understand too.
For instance, the comment “this writing makes me uneasy” on its own leaves the author with no idea what to fix. Break it down as follows, however, and it leads to improvement.
- It contains figures that need fact-checking.
- The source is not stated.
- It has assertive phrasing that customers could easily misread.
- The expression is stronger than our own tone of voice.
- It is short on explanation of technical terms for the intended reader.
Seen this way, review criteria are not a tool for deciding “good or bad” by feel; they are a tool for sharing which angle to judge from, and how far is acceptable.
I think of the work of drawing up review criteria as the work of translating the front line’s tacit knowledge. Turning the points a seasoned hand glances at without thinking into words that land with juniors and other departments. Turning the pass mark a quality lead holds in their head into checklist items the whole team can use. That accumulation raises the reproducibility of AI use.
In running business AI in particular, drawing up review criteria makes the following three things easier to achieve.
- The reproducibility of judgements improves.
- Revision instructions become concrete.
- Reasons for sending work back accumulate, feeding improvements to prompts and operating rules.
AI review criteria are not there to shackle the AI. They are a shared language that lets people use AI with confidence.
Start by sorting AI output by purpose
There is something to settle first, before you build an AI evaluation rubric.
That is which kind of AI output you are going to evaluate.
AI output is not all of a piece; the quality demanded varies with the purpose. An internal memo and a proposal sent to a customer need not be reviewed with the same severity.
Build your review criteria without making this distinction and they become awkward to use on the ground. Demand rigorous fact-checking of every single output and the pace of AI use slows. Conversely, wave even customer-facing documents through with a cursory check and the risks mount.
When I provide support, I usually start by sorting output into the following four categories.
AI output for internal use
Summaries of internal meetings, sorting out the issues, idea generation, rough first cuts for research, and the like.
In this territory, the absence of gaps between fact and conjecture matters more than flawless prose. If the material is for internal deliberation, having it clearly marked “to be verified” is itself part of the quality.
For example, when AI is used to sort out the issues ahead of a planning meeting, that output is not the final document. It is material for people to discuss. Here, broad coverage of the points worth considering matters more than the odd rough phrase.
Customer-facing AI output
Emails, proposals, FAQ answers, sales materials, marketing content, and so on.
In this territory, on top of accuracy, tone of voice, resistance to misreading, brand expression, and risk language all matter. Even if the writing is natural, any expression that raises undue expectations in the customer needs to be stopped at review.
In sales and marketing especially, text the AI has produced can sail through simply because it “sounds about right”. But documents used at the customer touchpoint bear directly on the impression of the company. Here you need to check not only how well the writing reads but also the bounds of what you may promise and what you may assert.
AI output used for decision-making
Output bearing on board materials, business plans, KPI analysis, investment decisions, hiring decisions, and the like.
In this territory, the basis for figures, the underlying assumptions, dissenting views, risks, and the treatment of alternatives all matter. AI output should not be fed straight into a decision; it must be treated as material for people to verify.
With output bearing on decisions, I weight the “assumptions” above the “conclusion”. What information is it resting on? Which conditions, if they changed, would change the conclusion? Seen from the opposing side, where is it weak? Output that gives no sight of these points is precarious as decision material, however polished it may be.
AI output bearing on legal, HR, and security matters
Contract reviews, answers about employment regulations, documents containing personal data, answers bearing on security policy, and the like.
In this territory, the premise is that the AI does not have the final word. The review criteria, too, need to spell out the conditions under which a specialist department must check the work.
AI is no substitute for a legal adviser, an HR lead, or an information-security lead. It is useful for preliminary research and for sorting out the issues, but the final judgement should always be made by someone with the specialist knowledge. Writing this line into your review criteria is part of what protects the organisation.
The angles to build into an AI evaluation rubric
When drawing up AI review criteria, there is no need to over-complicate things from the outset.
To begin with, being able to evaluate against the following six angles will serve you across a great many tasks.
The AI evaluation rubric meant here is a table that sets out the angles and the scale for evaluating AI output. By putting into words “how many points for accuracy” and “what state warrants sending it back”, you reduce the discrepancies in judgement between reviewers.
I recommend making your first rubric “short enough that the front line can read it as they work”. Too many items and merely reading the checklist becomes a burden. Starting simple and adding the missing angles as you operate tends to bed in more readily.
Accuracy
Accuracy is the foundation of AI output quality.
The things to check with particular care are proper nouns, figures, dates, quotations, names of schemes, product names, contract terms, and the like.
AI is good at producing plausible-sounding writing. For that very reason, the more natural the prose, the harder the errors can be to spot. Therein lies the difficulty of the hallucination check.
By hallucination here I mean the phenomenon whereby AI generates information that has no basis in fact, dressed up to sound plausible.
For an accuracy review, set standards such as the following.
- Are fact and conjecture kept apart?
- Is there a basis for the figures and dates?
- Do claims that require a source carry one?
- Does it contain non-existent scheme names or product names?
- Does it contradict internal materials or primary sources?
What matters is getting to a state of “can be verified”, not “looks right”.
For instance, if the AI writes “many of the companies that adopted it have felt the benefit”, you should not wave it through as it stands. We cannot tell how many “many” is, over what period, or which survey it rests on. If it cannot be verified, the wording needs amending to match the facts, to something like “some of the companies that adopted it have confirmed a benefit”.
Completeness
Completeness is the standard for whether any necessary angle is missing.
For a customer-facing FAQ, for instance, you check not only the body of the answer but whether the eligibility conditions, the exceptions, the contact point, and the caveats are all included. For a proposal summary, you look at whether the issue, the proposed solution, the benefits, the cost, the rollout steps, and the risks are laid out with nothing missing and nothing surplus.
For a completeness review, questions like the following are useful.
- Is the information the reader will want next included?
- Are any important assumptions missing?
- Have exceptions or caveats been left out?
- Is the reasoning that leads to the conclusion explained?
- Is the material the stakeholders need in order to decide all present?
Push completeness too far, though, and the output grows long. So it helps the operation to separate, by purpose, the “items that must always be included” from the “items to include where needed”.
A method I often use on the ground is to separate “required items” from “supplementary items”. For a customer-facing FAQ, say, the conclusion, eligibility conditions, caveats, and contact point are required; background and a detailed account of the scheme are supplementary. Simply sorting them this way speeds up the review.
Contextual fit
Contextual fit is the standard for whether the output suits the situation in which it will be used.
The same content calls for different phrasing and different amounts of information depending on whether it is aimed at the board, at front-line staff, at customers, or at an internal Slack channel.
For executives, for example, you need to lead with the conclusion and the decision points. For an information-technology lead, system integration, access management, logging, security, and operational load matter. For a marketing or sales lead, the customer experience, conversion to opportunities, brand expression, and the effect on lead generation carry weight.
For a contextual-fit review, check points such as the following.
- Is the amount of information suited to the intended reader?
- Is the conclusion that is needed in the moment placed up front?
- Are technical terms explained appropriately?
- Is it clear what the reader should do next?
- For the situation of use, is the expression neither too heavy nor too light?
Taken on its own, AI output can look well put together. But if it does not suit the audience or the occasion, it falls short as business quality.
When I look at contextual fit, I ask myself “will the person who receives this be able to act on it without hesitation?”. A piece of writing that leaves the reader unsure what to do, however beautiful, does not function in a business setting.
Tone of voice
Tone of voice is the standard for getting the impression of the writing right.
It matters especially for customer-facing emails, advertising copy, sales materials, recruitment communications, and internal notices.
Left without instruction, AI tends to produce generic, inoffensive prose. Yet every company differs in its “house style”, its “expressions to avoid”, and the distance it keeps from customers.
It is worth building items like the following into the review criteria.
- Does it read in our own voice?
- Is there any overly categorical phrasing?
- Does it stray into scaremongering?
- Does it raise excessive expectations in the customer?
- Is the honorific register or the sentence ending unnatural?
- Does it follow the brand and service style rules?
Sorting out tone of voice is also the work of crafting “the company’s voice”. The more you have AI write, the more you need to have put your company’s character into words. Otherwise, whoever the company, the writing ends up looking much the same.
Risk
Risk assessment is especially important among the AI review criteria.
The risks meant here include legal risk, compliance risk, personal-data risk, security risk, brand-damage risk, and customer-misunderstanding risk.
For example, output like the following warrants caution.
- Expressions that appear to guarantee an outcome
- Expressions that read as a legal judgement
- Expressions containing personal or confidential information
- Expressions that assert internally unverified figures
- Expressions that unfairly disparage other firms or competitors
- Strongly categorical expressions in medicine, finance, law, or safety
For a risk review, it is important to separate “things a fix will make usable” from “things that need a specialist department to check”.
Where a simple turn of phrase will do the job, the front line can handle it. Where contract terms, regulation, personal data, or security are involved, on the other hand, you need to make the route to legal, IT, or HR for confirmation explicit.
What I want to stress is this: do not use the risk review “to ban things because they frighten you”. Pile on the prohibitions and the front line simply stops using AI. What matters is making clear how far the front line may judge on its own, and from where on a specialist department must be consulted. Once the safe range is clear, the front line in fact finds it easier to use AI with confidence.
Reproducibility
Reproducibility is the angle of whether you can keep obtaining output of much the same quality.
In AI use, one good output is not enough for a business operation. It matters that a steady standard of output is obtained whoever uses it and on whichever day.
When looking at reproducibility, check the following.
- Is the prompt saved?
- Is the format of the input information consistent?
- Are the review criteria written down?
- Are the reasons for sending work back recorded?
- Are improved prompts shared across the team?
- Is the operation being revised on the back of the evaluation results?
To stabilise AI output quality, rather than giving the AI a one-off instruction, you need to cultivate the prompts, the training data, the review criteria, and the operating rules as a set.
In AI use, I regard “it happened to go well” as not yet a result but a sign. To turn it into a result, you need a state in which someone else can reproduce the same quality. That is what calls for review criteria and a rubric.
Make the review criteria concrete with five-point scoring
When drawing up review criteria, merely listing the angles is not enough.
For each angle, you need to decide what state passes and what state needs fixing.
A five-point AI evaluation rubric is effective for this.
For accuracy, for example, you might define it as follows.
| Score | State | Handling in operation |
|---|---|---|
| 5 | Facts, figures, dates, and proper nouns are verified, and the sources are clear | Usable as is |
| 4 | The main facts are verified, but a few minor items remain to be confirmed | Usable after a minor fix |
| 3 | No major errors, but several expressions rest on an unclear basis | Re-check after the author has amended it |
| 2 | Important fact-checking is lacking, and using it as is could mislead | Send back for review |
| 1 | Contains clear misinformation, non-existent information, or unfounded assertions | Not usable. Revisit the prompt or the input conditions |
Putting the scale into words this way makes it easier for reviewers’ judgements to line up.
When I build a rubric on the ground, I weight “what to do next” above the score itself. At 3, who fixes it? At 2, on which angle do we send it back? At 1, do we revisit the prompt, or the input data? Decide that far ahead and the review is less likely to stall.
Treat the hallucination check as a standalone item
When drawing up AI review criteria, I recommend treating the hallucination check as a standalone item.
The reason is that the naturalness of the prose and the correctness of the facts are two different things.
AI output can be misstated on the facts even while reading smoothly and looking well put together. The information to watch particularly is the following.
- Non-existent scheme names
- Superseded, out-of-date specifications
- Non-existent survey data
- Market sizes with no basis
- Wrong company names or job titles
- Fabricated quotations
- Answers at odds with internal rules
In a hallucination check, there is no need to verify everything with equal weight. Start with the items that have the greatest impact on the business.
For a customer-facing document, that means product specifications, prices, contract terms, delivery dates, and claimed benefits. For an answer about internal regulations, the eligible parties, the application deadline, the approval flow, and the exception conditions. For board material, the figures, the comparison basis, the assumptions, and the effect on the decision.
Into the review criteria, write rules such as “information that cannot be verified is marked as to be verified”, “claims that require a source must carry one”, and “judgements in specialist areas are referred to the relevant department”.
Rather than aiming solely to stop the AI getting things wrong, it is important to build an operation in which mistakes can be caught.
I see hallucination countermeasures not as “blaming the AI for its flaws” but as “building a verification step into the business process”. We check figures and proper nouns even in documents people have written. With AI-produced output, all the more reason to keep it in a form that is easy to verify.
Bed it in on the ground with a sampling review
Reviewing every piece of AI output in fine detail, every time, is not realistic.
This is where a sampling review proves useful.
A sampling review is an arrangement in which you periodically pull a portion of the deliverables the AI has generated and check them against the review criteria.
Once a week, say, you pick five sets of minutes produced by AI summarisation and check them on the angles of accuracy, completeness, expression, and risk. Or you pick ten draft customer emails and check them for tone of voice and misreading risk.
What matters here is that you are not reviewing in order to blame individuals.
A sampling review has three aims.
- To spot discrepancies in the review criteria
- To grasp the tendencies of the AI output
- To feed improvements to the operation
If there is an item where the scores of different assessors diverge, you revisit how the criterion is worded. If much the same error keeps appearing, the cause may lie in the prompt or the training data. Categorise the reasons for sending work back and the points to teach, and the check items worth turning into a template, come into view.
Reasons for sending work back can be classified as follows, for instance.
- Insufficient fact-checking
- No basis for figures
- Tone-of-voice mismatch
- Too little information for the reader
- Risky expressions present
- Specialist-department check required
- Insufficient prompt instruction
- Insufficient input information
As this classification accumulates, the AI review becomes not mere inspection but an improvement cycle for running business AI.
In a sampling-review session, I try to look at “why the output came out as it did” rather than “who produced it”. Was the input information lacking? Was the prompt vague? Were the materials that should have been referenced never organised? Tracing the cause back to the system rather than the individual leads to improvement for the whole team.
Building the environment to put review criteria into operation
Manage AI output quality only through personal notes and local files and it readily becomes a hollow formality.
What matters is getting the review criteria, the prompts, the training data, and the output results into a state the team can handle together.
There are several ways to do this. An internal wiki, a knowledge-management tool, a document-management system, an AI platform: choose to suit the environment you already have and you will be fine.
On top of that, if you want to organise AI use by task, an environment such as Kanata, which lets you combine AI chat, AI summarisation, and a project library, is one option. You can split projects by task, for example sales-material reviews, internal FAQ reviews, marketing-copy reviews, and minutes reviews, and share the prompts and reference materials you use most across the team.
For a review prompt, registering an instruction like the following makes it handy in practice.
Evaluate the AI output below on a five-point scale across the six angles of accuracy, completeness, contextual fit, tone of voice, risk, and reproducibility. For each item, give the reason a fix is needed and a proposed fix. Where information is uncertain, do not assert it; classify it as “to be verified”.
It also helps to keep tone-of-voice materials, a list of prohibited expressions, the FAQ, product specifications, internal regulations, and past OK and NG examples in a place the team can reference, so reviewers need not explain the criteria from scratch every time.
The point is not to hand the review over to the AI entirely. The realistic approach is to use AI for the first-pass check and for sorting out the angles, while people make the final judgement.
In AI use on the ground, I regard separating “the work to leave to AI” from “the work people take responsibility for” as the single most important thing. Drafting, summarising, classifying, and sorting out the angles are easy to leave to AI. The final judgement, the customer relationship, and accountable explanation, on the other hand, should be borne by people. Leave that line blurred and AI use stays a mix of convenience and unease.
How to stop the review criteria becoming a dead letter
Review criteria are not finished the moment you draw them up.
If anything, it is only once they are in operation that the gaps come to light.
Try to build a perfect rubric from the start and the items grow so numerous that the front line stops using it. Begin instead by choosing one key deliverable and starting small across the six angles.
For example, narrow the first target to “AI drafts of customer emails”. Run it for a month and gather the reasons for sending work back. If tone-of-voice discrepancies turn out to be common, add expression standards; if insufficient fact-checking is common, strengthen the hallucination check.
To bed the review criteria in, the following operating habits are also needed.
- Appoint a review lead
- Revisit the criteria monthly
- Add OK and NG examples
- Update the items when a new risk appears
- Run short training for reviewers
- Feed the reasons for sending work back into prompt improvement
Managing AI output quality is not something a single design settles once and for all. It is something to update in step with the nature of the work, customer interactions, internal rules, and changes in the AI models.
What I particularly want to avoid is the review criteria becoming a “document you draw up and then forget”. Nothing is more of a waste to an organisation than a checklist no one on the ground uses. Review criteria are not a document to read out in a meeting; they are a tool to make day-to-day judgements a little easier.
For that, it is important to write the criteria in the language of the front line. Rather than abstractions like “ensure appropriateness”, pin them to words you can actually check against, such as “is there any categorical wording customers could misread?” or “are any figures used without a basis?”.
In summary
When introducing AI to a business, attention tends to settle on “how much the AI can do”.
But what gets tested in actual operation is how people and the organisation handle what the AI produces.
With no review criteria in place, the quality of AI output rests on the reviewer’s experience and instinct. Change the reviewer and the boundary between OK and NG shifts too. As a result, the front line feels uneasy about using AI, and the quality lead ends up unable to keep it under control.
With an AI evaluation rubric, by contrast, what to look at, how far is acceptable, and the cases in which a specialist department must be consulted are all shared.
Of course, review criteria are no panacea. Even with criteria, you cannot prevent every hallucination. Nor can you foresee every risk in advance.
Even so, by aligning the angles of judgement, accumulating the review results, and revisiting the operation, AI output quality grows steadier little by little.
What it takes to bed business AI in is not a binary of “trust the AI or doubt it”.
It is deciding by which criteria you check the AI’s output, within what range you use it, and from where on people take responsibility.
As that shared language, AI review criteria and an AI evaluation rubric become an indispensable foundation for the quality-management AI of the years ahead.
In using AI in business development and operational improvement myself, the more I use it the more I feel that what matters most in the end is not “the technology itself” alone. Technology advances quickly. The features available keep growing. Whether it can keep being used on the ground, however, hinges on whether you have rendered it into a form people can judge.
Review criteria for assessing the quality of AI output are the bridge to that. Carrying the possibilities of AI onto the ground while protecting business responsibility, the trust of customers, and internal safety. To strike that balance, review criteria will only grow more important.
Q&A
Which department should draw up the AI review criteria?
To begin with, it is realistic for the front-line department that actually uses the AI and the department that manages quality and risk to draw them up jointly. Drawn up by the front line alone, they lean too far towards practicality; drawn up by the management department alone, they risk being awkward to use. For sales materials, that means sales planning together with quality management; for internal use, the operating department together with the IT department.
Does all AI output need to be reviewed?
Not necessarily all of it. Low-risk output, such as internal memos or idea generation, may be adequately handled by the author’s own self-check. Output bearing on customer-facing materials, contract and legal matters, HR, security, and management judgement, on the other hand, should have human review or a specialist-department check built in. It is important to use full reviews and sampling reviews selectively, according to the level of risk.
What should be prioritized during a hallucination check?
Prioritize information that could have a significant business impact. This includes proper nouns, numerical values, dates, prices, contract terms, names of policies or programs, product specifications, quotations, and internal company rules. In customer-facing documents and decision-making materials in particular, it is important not to present information as fact when its source cannot be verified. Clearly mark uncertain information as “Needs verification” and, when necessary, confirm it using primary sources or with the relevant department.
How many items should the evaluation rubric start with?
Starting with around five to six items is recommended. The six angles of accuracy, completeness, contextual fit, tone of voice, risk, and reproducibility will cover most tasks. Building it out in too much detail from the start can mean it goes unused on the ground. It beds in more readily if you start by focusing on one deliverable and add items as you watch the reasons work gets sent back.
What should you do if the review criteria go unused on the ground even after you draw them up?
Likely causes are that the criteria are too abstract, that there are too many items, that the situation of use is undecided, or that there is no one accountable. First narrow the target to one, and for example apply it only to “AI drafts of customer emails”. Then, lining up OK and NG examples and using it in real reviews, cultivate it as you go.