How to Build an AI Operating Model: Continuous Improvement, Prompt Management, and Shared Knowledge Systems

Column
How to Build an AI Operating Model: Continuous Improvement, Prompt Management, and Shared Knowledge Systems

Introduction

A practical look at how to run prompt operations, knowledge improvement and workflow improvement as a proper AI operating model rather than leaving them to individuals. It sets out roles, review meetings, improvement KPIs, guidelines and how to think about tool selection.

Tatsuya Ito

Tatsuya Ito

Artificial Intelligence Consultant

company-icon

Third Scope Ltd.

Born in 1985 and originally from Mie Prefecture, Japan. In 2012, he joined an AR startup in Hong Kong as an engineer. Since then, he has been involved in new business development and AI service launches at several AI startups. In 2018, he founded the current ThirdScope Inc. by taking over an AI service and its development team. He now supports companies in adopting and utilizing AI, with a focus on AI-driven business development, operational transformation, and product development. He has also been involved in AI research as a Project Researcher at the University of Tokyo. Today, he continues to work at the forefront of AI project development, providing practical consulting from both technical and business perspectives.

So who, exactly, is holding the latest version of this prompt?

That was the question Mori (a pseudonym) from the DX office put to the room during the Monday-morning sales planning meeting. Three months earlier, the company had certainly spread its use of generative AI, but sales, customer success and the back office were each building their own prompts, and internal documents were scattered across personal folders and Slack. When the quality of the AI’s answers slipped, nobody could tell whether the culprit was the prompt, the knowledge it was drawing on, or the workflow itself.

Today the firm manages its minutes templates, FAQs and meeting-note tidying prompts across departments, and reviews what needs improving at a fortnightly review meeting. Over the most recent eight weeks, looking at thirty recurring tasks, the team managed to consolidate nine of the twelve most-amended prompts into a single current version. That said, these figures come from observing a slice of one company’s work, and they are no guarantee that every organisation will see the same effect.

This article sets out how to run prompt operations, knowledge improvement and workflow improvement as an AI operating model rather than leaving each to its own devices. The aim is a state in which someone’s good idea does not vanish the moment it is used, but accumulates as organisational knowledge. Mind you, simply installing a tool will not drive continuous improvement on its own. You also need to design the roles, the guidelines, the review meeting and the improvement KPIs together.

AI adoption tends to degrade after the launch

AI adoption tends to degrade after the launch

The period straight after a generative AI rollout is the one most likely to show the front line a clear, tangible change. Minutes take less time to write, a first draft of an email appears in moments, internal documents are summarised more quickly. These early wins create the genuine sense that “this is something we can actually use at work”.

After a few weeks or months, however, a different sort of problem comes into view.

The prompt that worked so well at first is still sitting in someone’s personal notes or chat history. The document being used as an FAQ has gone out of date. One department uses something to good effect while another repeats exactly the same trial and error from scratch. In this state, AI use may look as though it is spreading, yet the organisation is accumulating no real know-how.

In a sales team, for instance, you may find several prompts for tidying up meeting notes. One person organises them in BANT format; another structures them around the customer’s issues and the next action. Both are useful in practice, but if the team has not standardised on one, the entries that reach the CRM and the quality of any handover will vary.

In customer success, the knowledge fed to the AI for handling enquiries may mix documents from before and after a pricing change. Here the trouble is less that the AI’s answer is wrong and more that the source information has been poorly managed.

In the back office, a team may build an AI chat that answers staff questions from internal regulations, only for the AI to give a flatly definitive answer even to questions that call for an exception or an individual judgement. The upshot is often more work, as someone has to go back and check afterwards.

So it goes: in running business AI, tidying up only one of prompts, knowledge or workflow is never quite enough. Unless all three are managed together, AI use gradually becomes one person’s job, quality wobbles, and the front line’s trust is easily lost.

The things to keep improving are prompts, knowledge and workflow

The things to keep improving are prompts, knowledge and workflow

When thinking about an AI operating model, the first thing worth sorting out is what to keep improving.

At many companies, improving AI use brings to mind chiefly how to write the prompt. Prompt operations certainly matter. But the AI’s output is not decided by the prompt alone; it is also shaped by the quality of the knowledge it draws on, how it is used within the workflow, and how the output is reviewed afterwards.

There are broadly three things to keep improving.

Prompts
The instructions that tell the AI what you want, on what assumptions, and in what format. If an individual is only going to use one once, a little roughness rarely matters; but if it is to be reused across an organisation, it needs tidying so that anyone using it gets a consistent level of quality.

Knowledge
The internal documents, FAQs, regulations, past minutes, proposals and operating manuals the AI refers to. Knowledge that is out of date, duplicated, of unknown provenance, or managed separately by each department will leave the AI’s answer quality anything but stable.

Workflow
In which task, and at what point, the AI is used; who checks the output; whether the AI’s answer can be used as is or whether a human judgement is needed. Hand out the AI without settling these, and the decisions are left to each part of the front line.

Ask merely for “please write the minutes”, for example, and the format and level of detail wander every time. Settle on headings such as meeting overview, decisions, to-dos, points of discussion and items to confirm next time, and insist that an owner and a deadline always go in, and you end up with minutes that are genuinely usable downstream.

In an AI operating model, the important thing is to see prompts, knowledge and workflow not as separate matters but as a single cycle of improvement.

With prompt operations, “maintaining” matters more than “creating”

With prompt operations,

The phrase “prompt operations” may bring to mind nothing more than how to write a prompt. Yet what an organisation needs, rather more than the writing of a good prompt, is a mechanism for maintaining, updating and sharing the good ones.

Suppose a department improves its prompt for drafting minutes. Decisions now always carry a “who” and a “by when”, and confirming the tasks after a meeting becomes that much easier. If that improvement merely lives on in one person’s chat history, it is no improvement for the organisation. Only once it has been added to a library, given a name, described by its purpose and shared as the current version does it become something others can use too.

There are at least five things worth settling for prompt operations.

  • Which prompt to use as the standard
  • Who owns the prompt
  • What naming convention to manage it under
  • Where to keep the change history
  • How to retire prompts that are no longer used

The naming convention need not be elaborate from the outset. Something like “purpose_target task_version” that means something to anyone who reads it will do nicely.

Examples of a prompt naming convention
Prompt name Intended use
minutes_dept_regular_v1 Drafting minutes for the regular departmental meeting
meeting_notes_new_sales_v2 Tidying meeting notes for new-business sales
faq_answer_hr_admin_v1 Answering internal HR and general-affairs FAQs

Once you can tell what a prompt is for, which task it serves and which version it is, even minimal management becomes a good deal easier.

Let names such as “latest version”, “test”, “made by Yamada-san” and “the nice one” proliferate, on the other hand, and things become hard to find in remarkably short order. With prompt operations, it is not only the quality of the contents that matters, but how easy they are to find, to update and to retire.

Knowledge improvement is the foundation that underpins answer quality

Knowledge improvement is the foundation that underpins answer quality

When the AI’s answers are unsteady, there is a tendency to reach only for the prompt. In practice, though, the cause not infrequently lies on the knowledge side.

Consider an internal FAQ that the AI consults. If a member of staff asks “what is the daily allowance for a business trip?” and the AI consults an outdated travel-expense regulation, the guidance will, naturally enough, be wrong. However carefully you craft the prompt, an out-of-date source will not yield a correct answer.

The point of knowledge improvement is not to add information but to keep it in a usable state.

Registering a vast pile of internal documents does not make the AI any cleverer. Mix in old documents, duplicates, drafts that were never finalised and material of unknown provenance, and answer quality is, if anything, liable to fall.

In managing knowledge, it is worth checking the following.

  • Is this document the current version?
  • Are its source and originating department clear?
  • Is it actually referred to in the work?
  • Do several documents of the same content exist?
  • Is this information you are content for the AI to consult?
  • Are the update date and any expiry clear?

Above all, it matters to decide who is responsible for the knowledge.

HR regulations belong with the HR department, the price list with sales planning or corporate management, product specifications with the product owner or IT department; the owner needs to be someone who can judge whether the information is correct.

Try to tidy the knowledge with the DX office alone, and you may sort out the categories and the storage, but you will struggle to vouch for the correctness of the contents. A realistic shape for an AI operating model is one in which the DX office or IT department builds the mechanism while each business department takes responsibility for what is inside.

Building AI in as a workflow improvement

Building AI in as a workflow improvement

Even with prompts and knowledge in good order, AI use will not bed in unless it is built into the workflow.

What commonly happens on the front line is a state of “only those who want to, use it”. In the early days that is fine. But to give the work repeatability, you need to decide at which moments the AI is used.

Take running a meeting, which divides up rather neatly. Before the meeting, the AI drafts an agenda from the previous minutes and progress notes. After it, it drafts the minutes from a recording or notes. The chair then confirms the decisions, owners and deadlines, amends where necessary, and shares.

An example of dividing roles between AI and people in running meetings
Stage What the AI handles What people handle
Before the meeting Draft an agenda from the previous minutes and progress notes Adjust it to the meeting’s purpose, priorities and attendees
After the meeting Draft the minutes from the recording or notes Confirm and share the decisions, owners and deadlines

The same holds for handling enquiries. The AI drafts an answer to a staff question by consulting internal regulations and the FAQ. Anything not set out plainly in the regulations, or any matter calling for an exception, is routed to the relevant department to confirm. Deciding in advance the areas in which the AI must not be definitive is important.

In tidying meeting notes, too, the AI can extract the key points and propose candidate next actions. The reading of the customer’s mood, the priority of a proposal and any decision on discounts, on the other hand, are for people to do.

The reason an AI goes unused is not only that it is short of features. “I cannot use it because I do not know how far I am allowed to trust it” is just as real. Workflow improvement is not about adding places to use the AI; it is about making clear the respective roles of AI and people.

The roles an AI operating model needs

The roles an AI operating model needs

To keep continuous improvement turning over, a division of roles is indispensable.

In the early days a keen individual may carry the whole thing alone. Let that state persist, though, and the operation grinds to a halt the moment that person gets busy. Updating prompts, tidying knowledge, the review meeting itself, all become “something that so-and-so does when they have a spare moment”.

In an AI operating model, it is worth settling at least the following roles.

Business owner
Defines which business outcome this use of AI is meant to serve. For a sales team that might be shorter deal-preparation time or more consistent proposal quality; for customer success, better first-response quality on enquiries; for HR, a higher self-resolution rate on internal queries.

Prompt manager
Tidies, standardises and version-controls the prompts used on the front line. Gathers the front line’s voices, “it would be better to add this phrasing”, “this output format is hard to use”, and folds the suggested improvements in.

Knowledge manager
Manages the freshness, accuracy and provenance of the material the AI consults. The role includes not only adding material but deleting the old, merging duplicates, and separating what should be consulted from what should not.

Front-line user
While using the AI’s output, feeds back what worked and what did not. Without the voice of the people actually using it, neither prompts nor knowledge can be improved.

Operations lead
Runs the review meeting, checks the improvement KPIs, coordinates across departments and updates the guidelines. They need not do every task themselves, but they keep an eye on the whole so the operation does not stall.

The crucial thing is not to make AI operations the job of one department alone. Since this is AI woven into the work, ownership by the business departments is essential.

Design the review meeting as a place to grow your AI use

Design the review meeting as a place to grow your AI use

To actually keep continuous improvement turning, a review meeting earns its keep.

There is no need to pile on grand meetings, mind. A fortnightly half hour will do to begin with. What matters is having a regular place to check how the AI is being used, the quality of its output, and what is troubling the front line.

The themes a review meeting takes up can be organised like this.

  • The prompts in heavy use
  • The outputs that get amended a lot
  • The questions the AI could not answer
  • The bottlenecks in the workflow

First, look at the prompts in heavy use. A frequently used prompt is one the front line is likely to value. By the same token, precisely because it is used so often, any wobble or error in its output carries a large effect.

Next, look at the outputs that get amended a lot. Where people make the same correction every time to the AI’s draft, that correction can often be folded back into the prompt. If the refrain is “the conclusion always comes last, so I keep moving it to the front”, say, you can add a conclusion-first condition to the output.

Next, look at the questions the AI could not answer. An unanswered question may be a sign of a gap in the knowledge. Work out whether there is information to add to the FAQ, whether the material exists but has not been registered, or whether it is a question the AI ought not to answer at all.

Lastly, look at the bottlenecks in the workflow. Where the AI’s output is perfectly good yet goes unused on the front line, the trouble may be that the input takes too much effort, that no one has been named to check the output, or that hopping back and forth with existing tools is a chore.

A review meeting is not a place to grade the AI but a place to improve the work. Rather than stopping at “the AI got it wrong”, it asks “why did the output come out that way?”, “can the prompt fix it?”, “should we update the knowledge?”, “should we hand this back to a workflow where a person decides?”.

Do not judge improvement KPIs on usage alone

Do not judge improvement KPIs on usage alone

When you come to evaluate an AI operating model, the easiest thing to see first is usage. How many people use it, how many times it is used, in which departments. These are important measures.

Usage alone, however, will not tell you whether the AI is leading to better work. It may be used a great deal yet need substantial amendment every time, in which case the real effect is rather limited. Conversely, even with low usage, if it is used on heavy work such as monthly reports or board-meeting materials, the value may well be large.

Improvement KPIs are easier to organise if you split them across the three of prompts, knowledge and workflow.

An example classification of improvement KPIs
Target Example KPIs
Prompt operations Number of standard prompts registered, times used, times improved, consolidations into the current version, requests for amendment from the front line
Knowledge improvement Number of updates, duplicates removed, old documents tidied away, and unanswered questions turned into FAQ entries
Workflow improvement Task time, number of rework cases, review time, first-response time on enquiries, and time to share the minutes after a meeting

For drafting minutes, say, you can measure it as “the time to draft the minutes of a sixty-minute meeting was, on average, thirty minutes before adoption and, including the checking, ten minutes after”. When you put out figures of that sort, though, you need to be clear about the period, the number of meetings and the measurement conditions.

There is no need to strain to quantify everything. Qualitative changes such as “sharing after a meeting got quicker”, “new joiners find it easier to dig out past minutes”, and “the variation in answers between individuals on enquiries has narrowed” are themselves important results of operational improvement. Recording the changes in what people say and do, and sharing them at the review meeting, raises the front line’s sense of conviction.

Guidelines should show “how to use it”, not only “what is forbidden”

Guidelines should show

An AI operating model needs guidelines.

But if the guidelines become nothing but a list of prohibitions, the front line finds them hard to work with. “Do not put this in”, “be careful of this”, “confirm this”, and that alone leaves people none the wiser as to which task to use it on, and how.

Guidelines that earn their keep in practice need at least these three things.

Rules on handling information
Define what information may and may not be put into the AI, such as personal data, customer data, contract data and undisclosed financial information. Where appropriate, set out the method of masking too.

Rules on where it is used
Set out in which tasks the AI may be used, in which it must not, and in which a human check is mandatory. For instance, it may be used to draft an email or summarise minutes, but not for the final judgement on a contract or a performance review.

Rules on review
Decide who checks what before the AI’s output leaves the building. Figures, dates, proper nouns, quotations and anything touching legal, financial or HR matters need particular care.

Guidelines are not a one-and-done affair. Use throws up exceptions and borderline cases, and each time you need to update them and share with the front line.

Ideally, the guidelines themselves become a target of continuous improvement. Gather the voices, “this case left me unsure”, “this wording does not land with the front line”, “this task could probably be made AI-permitted”, and grow them into a shape that fits the work.

Start small, and widen from a range you can actually operate

Start small, and widen from a range you can actually operate

Say “AI operating model” and there is a temptation to build a company-wide mechanism from the very start. In most cases, though, the bigger you start, the heavier the operation becomes.

The recommendation is to begin with one department, one task and one set of prompts.

Take only the tidying of meeting notes in the sales department, for example. After a deal, the AI tidies the notes and outputs the customer’s issues, the next action, points of concern and a summary for the CRM. Narrowing to this one task, you standardise the prompt, put the knowledge it consults in order, and improve it at the review meeting.

Or take only the handling of enquiries in HR and general affairs. You put the work rules, expenses regulations and attendance rules in order as training data, and the AI drafts answers to staff’s frequently asked questions, with any exception routed to the owner to confirm.

Narrowing the range makes the points for improvement easier to see. It makes clear who the owner is. And it makes the KPIs easier to measure.

Once an operation begun small is turning over, you can roll it out to other departments. At that point the naming convention, the way of running the review meeting, the improvement KPIs and the guidelines built in the first department become a working draft.

What it takes for AI use to bed in is not building a perfect design from the outset. It is building the smallest unit you can keep improving, and then spreading it sideways.

In selecting a tool, look at how easy it is to manage and improve

In selecting a tool, look at how easy it is to manage and improve

Which tool you use also matters for continuously improving business AI. That said, bringing in a particular tool will not automatically solve your operational challenges. What to look at is whether it is easy for the front line to use, easy for managers to improve, and easy to keep control over how information is handled.

For example, managing prompts and reference material individually, person by person, makes organisation-wide improvement difficult. What matters is being able to organise prompts by department or task, manage the training data, separate users and permissions, and easily trace the history of improvements.

A service such as Kanata, which lets you handle AI chat, AI summaries, a prompt library and a training-data library on a per-project basis, is one option well suited to this sort of operation. The sales department might gather its meeting-note tidying prompts and proposal materials, while HR gathers its internal regulations and FAQs, making it easier to keep the uses apart.

Before bringing in a tool, however, you need to settle which task you start from, who becomes the owner, which material the AI consults, and which outputs a person checks. A tool supports the operation, but it does not stand in for the operational design itself.

Summary: business AI is not something you adopt, but something you grow

Summary: business AI is not something you adopt, but something you grow

Adopting generative AI does, to be sure, make part of the work quicker. But to stop the early win from being a one-off, you need an operating model that keeps reviewing prompts, knowledge and workflow.

Prompts must not be left to end with one person’s ingenuity, but tidied into a shape the organisation can reuse. Knowledge must not merely be registered, but kept fresh, accurate and properly sourced. And workflow must separate the tasks left to the AI from the judgements for which people take responsibility.

For that, a division of roles between business owner, prompt manager, knowledge manager and front-line user is indispensable. Through the review meeting and the improvement KPIs, it matters to build AI use into a cycle of business improvement.

Business AI is not something you adopt and have done with. You grow the prompts. You grow the knowledge. You grow the workflow. It is this accumulation that turns one person’s ingenuity into organisational knowledge and builds an AI operating model that can keep improving.

Q&A: common questions on prompt operations, knowledge improvement and the AI operating model

With prompt operations, where should we start?

Start by gathering the prompts in heavy use on the front line. Then separate out the duplicates, the ones with steady output quality, and the ones that get amended a lot. Rather than building a company-wide standard from the outset, the realistic route is to settle a standard prompt in one department and one task, and improve it as you review.

With knowledge improvement, which material should we tidy first?

Tidy first the material that is used most and most likely to bear on business judgements, such as internal regulations, FAQs, price lists, product specifications, proposal materials and enquiry histories. Since a mix of old or duplicate documents affects answer quality, it matters to confirm the current version, the originating department, the update date and whether it may be used, before registering it.

How far should we trust the AI’s answers?

The AI’s answers are best used as a draft, a summary, a way of organising the points of discussion, or a way of surfacing angles. Anything that leaves the building, anything touching legal, financial or HR matters, and any formal reply to a customer should be run on the premise that a person always checks it. Even when the AI’s answer looks natural, the figures, dates, proper nouns and sources need confirming.

How should we set the improvement KPIs?

Combine measures that reveal the effect on the work, not just the number of uses. For prompt operations: times used, times improved, consolidations into the current version. For knowledge improvement: number of updates, duplicates removed, unanswered questions turned into FAQ entries. For workflow improvement: task time, rework cases, review time and the like are candidates. Where quantifying is hard, record the front line’s words and changes in behaviour as well.

Which department should lead the AI operating model?

The overall design is often led by the DX office or IT department, but the business departments need to take responsibility for the correctness of the content. HR regulations are confirmed by HR, the price list by sales planning or corporate management, product specifications by the product owner. It matters not to confine AI operations to the technical department, but to share the work with the business departments.

Share this article