How to Move AI from Proof of Concept to Production: Scaling Implementation Projects Company-Wide

Column
How to Move AI from Proof of Concept to Production: Scaling Implementation Projects Company-Wide

Introduction

We set out why AI PoCs stall at the trial stage and lay out how to run an implementation project that carries through to production, operational design, governance and building AI capability in-house.

Tatsuya Ito

Tatsuya Ito

Artificial Intelligence Consultant

company-icon

Third Scope Ltd.

Born in 1985 and originally from Mie Prefecture, Japan. In 2012, he joined an AR startup in Hong Kong as an engineer. Since then, he has been involved in new business development and AI service launches at several AI startups. In 2018, he founded the current ThirdScope Inc. by taking over an AI service and its development team. He now supports companies in adopting and utilizing AI, with a focus on AI-driven business development, operational transformation, and product development. He has also been involved in AI research as a Project Researcher at the University of Tokyo. Today, he continues to work at the forefront of AI project development, providing practical consulting from both technical and business perspectives.

We saw clear gains in the PoC, yet here we are, stuck once more on the question of going to production.

Those were the words of Saeki-san (a pseudonym), who leads digital transformation at a manufacturing firm we’ll call Company A, one Monday morning in the meeting room. The regular review brought together three groups, the operational leads from the business division, the IT systems department and corporate planning, and the agenda was the results of a PoC using generative AI to support internal enquiry handling. Over three months and 120 enquiries, the time spent drafting replies fell from twelve minutes on average to five. And yet, six months earlier, even a PoC that delivered results would not move on to a production decision at the company; the pile of trials simply kept growing.

These days the firm settles the criteria for taking AI into production, along with user acceptance, operational design, governance and the division of responsibilities for building AI capability in-house, right at the start of the PoC. With BtoB AI platforms such as Kanata, which let you handle AI chat, summarisation and training-data management on a per-task basis, among the options on the table, the company is steadily turning the PoC from a one-off exercise into a mechanism that keeps testing business value.

In this article I’ll set out why an AI PoC fails to reach production, and lay out a way of thinking that carries an AI implementation project through to company-wide rollout. That said, simply adopting a tool will not solve everything. AI tends to stick only once decision-making, frontline training and a habit of operational improvement are all in place.

Why a ‘successful’ AI PoC can still look like a failure

Why a 'successful' AI PoC can still look like a failure

When people hear the phrase ‘failed AI PoC’, they’re inclined to picture technical shortcomings, the accuracy wasn’t good enough, or the answers weren’t what they’d hoped for. Model performance, the quality of the training data and the awkwardness of system integration can certainly be to blame.

In a corporate AI implementation project, however, it’s far from rare for the PoC itself to show a respectable result while the evidence needed to move into production is simply not assembled.

Consider, for instance, the following situation.

  • The PoC’s outcome stalls at impressions, ‘that was handy’, ‘it looked usable’.
  • No one has settled the business workflow for a production version.
  • It’s vague who checks the AI’s output and who makes the final call.
  • Sign-off from the IT, legal and security functions keeps being put off.
  • There’s no design for the training and support needed to keep the frontline using it.
  • There are no metrics that explain the value to senior management in business terms.

In this state, even a good PoC result tends to stall at ‘a useful trial’. In other words, a failed AI PoC isn’t caused by inadequate AI performance alone. It also happens when the business design and decision-making process needed to move from PoC to production are lacking.

Shift the purpose of the PoC from ‘trying it out’ to ‘deciding on production’

Shift the purpose of the PoC from 'trying it out' to 'deciding on production'

When firms set out on a PoC, many think, ‘let’s just give it a go’. A small start is sound enough in itself. Assume a company-wide rollout from the outset and the cast of stakeholders swells while the questions multiply.

A small start and an aimless trial, though, are two different things.

The purpose of a PoC in an AI implementation project is not merely to confirm whether generative AI works. Properly understood, it is to gather the material needed to judge whether embedding that use of AI into production work is worth doing.

To that end, you need to settle at least the following three things before the PoC begins.

Metrics for business value
Define how much you want to cut working time, how you’ll gauge the quality of enquiry handling, and how you’ll assess the burden of preparing for sales meetings. Make clear not the use of AI in itself, but what change it brings to the work and to the business.

Metrics for operational viability
Check whether frontline staff can use it without strain, whether it fits into existing workflows, and whether administrators can update the training data and prompts. However high the accuracy, a setup that needs preparation every time it’s used will struggle to take hold.

Conditions for acceptable risk
Decide what data may be entered, how far the AI’s output may inform business judgements, and who checks things when a wrong answer appears. Leave governance to the end and the decision tends to stall just before production.

It matters to design the PoC not as ‘a place to look at the results of a trial’ but as ‘a place to gather the material for deciding whether to go to production’.

In an AI implementation project, redesign the business process first

In an AI implementation project, redesign the business process first

A common failure when adopting AI is trying to slot AI straight into one part of the existing work, unchanged.

Take enquiry handling: if you think only in terms of ‘feed the enquiry to the AI and have it produce a reply’, the benefit is limited. The actual work runs through a whole sequence, classifying the enquiry, referring to past answers, drafting a reply, having a staff member review it, sending it, and updating the knowledge base.

In an AI implementation project, you need to break this sequence down and then separate the steps you hand to AI from the steps people own.

Steps well suited to AI, and steps people should own
Category Main steps
Steps well suited to AI
  • Extracting the key points from a large body of information
  • Drafting answers on the basis of past material
  • Tidying up the tone of a piece of writing
  • Pulling action points out of minutes and meeting notes
  • Organising knowledge into a form that’s easy to search
Steps people should own
  • Making the final decision
  • Responding in light of a customer’s or an employee’s circumstances
  • Checking high-risk judgements such as legal, contractual and HR matters
  • Deciding how to handle exceptions
  • Assessing whether the AI’s output fits the business purpose

AI is not there to replace the work wholesale. It lightens the load of organising information and producing drafts, so that people can concentrate on the areas where their judgement is needed.

Press on with a PoC while this division stays vague, and you’ll readily provoke the reaction, ‘if a person ends up checking everything anyway, what’s the point?’. Sort out the roles of people and AI from the start, by contrast, and the frontline can begin using it with rather more confidence.

Design user acceptance during the PoC, not just before production

Design user acceptance during the PoC, not just before production

User acceptance is easily overlooked when taking AI into production. By user acceptance I mean the process of confirming that the intended users can use it in their actual work, and of gathering any operational worries and points for improvement.

In a PoC, the users tend to be the transformation team and a handful of keen members, so the response can look rosier than it is. In production, though, the users also include staff who aren’t well versed in generative AI, frontline people pressed for time, and managers comfortable with the existing way of doing things.

So you need to decide, from the PoC stage, exactly ‘who will use it in production’.

The points worth checking are as follows.

  • Are the users staff in a business division, administrators, or every employee
  • At which point in the work will the AI be used
  • Who teaches people how to use it
  • How the output is to be checked
  • Where to turn with a query when it proves hard to use
  • How requests for improvement are to be collected

User acceptance is not merely a how-to briefing. It is about reaching a state in which the frontline can judge, ‘yes, I can use this in my own work’.

For example, having an environment where you can split projects by department or by task and organise the AI chat, AI summaries and the training data each one refers to makes it easier to design operations that suit the work at hand. Kanata, a per-project AI platform of that sort, is one option to consider when weighing up such an arrangement.

The important thing is not to dwell solely on explaining the tool’s features. You need to design ‘when, what and how’ it will be used, to fit the user’s actual work.

Governance is there not to halt AI use but to widen it

Governance is there not to halt AI use but to widen it

Try to push ahead with a company-wide AI rollout and the governance debate is sure to surface. By governance I mean the rules, permissions, responsibilities and review mechanisms for using AI safely and on a lasting basis.

There are, for instance, questions such as these.

  • May confidential information be entered
  • May the AI’s answers be used as they stand
  • When a wrong answer appears, who bears responsibility
  • How far through the organisation use is permitted
  • How usage logs and the record of improvements are to be managed

Put such questions off and the brakes come on hard at the company-wide rollout stage. Settle the basic rules first, by contrast, and the frontline can take it up with more confidence.

There are, broadly, four things to settle in AI governance.

The range of information that may be entered
Classify public information, general internal information, customer information, personal data and confidential information, and decide how far each may be handled.

How the output is to be treated
Make clear whether the AI’s answer is used as a draft, whether it may be shared internally, or whether a person must always review it before anything goes outside the company.

Usage permissions
Separate the features every employee may use from those only certain departments may use. Where customer or contract information is involved in particular, it can be advisable to manage permissions on a per-project basis.

A mechanism for improvement and audit
Review regularly which uses are most common, where wrong answers occurred, and which prompts or training data ought to be updated.

Governance is not solely a means of restricting AI use. It is the foundation for widening the scope of use with confidence.

Building AI in-house does not mean ‘making everything yourself’

Building AI in-house does not mean 'making everything yourself'

Hearing the phrase ‘building AI in-house’, you might picture developing your own models or hiring specialist engineers in droves.

For most firms, though, a realistic in-house AI capability isn’t about building the AI platform itself from scratch. It’s about reaching a state where you can find use cases that fit your own work, prepare the prompts and training data, and improve operations over time.

For that, a clear division of roles across departments is indispensable.

Senior management
They decide which business challenge the use of AI is tied to. Whether it’s plain efficiency, a better customer experience, or the search for new business, the priorities shift accordingly.

The transformation and corporate-planning functions
They create a template the whole company can use, designing how PoCs are run, the evaluation metrics, governance, training content, and the lateral spread of success stories.

The IT systems department
They guarantee safety and connectivity, owning account management, permissions, logging, integration with existing systems, and security checks.

The business divisions
They bring in the real operational challenges and use cases, pinning down which tasks are a burden, which judgements take time, and which knowledge is locked up in particular individuals.

Building AI in-house is hard to advance through a single department alone. It needs to be designed as a project that joins up technology, the business and management judgement.

Steps for moving from PoC to production

Steps for moving from PoC to production

To keep an AI implementation project from ending at the PoC, it helps to proceed in the following five steps.

Narrow it to a single business challenge

There’s no need to aim from the outset at a grand, company-wide theme. If anything, you should narrow it to a specific task at first.

‘Improving sales productivity’, for instance, is too broad. Narrow it to ‘tidying up notes after a meeting and drafting the CRM entry’ and the steps where AI can help come into view.

Decide the criteria for going to production

Before the PoC begins, decide which conditions, once met, will prompt you to consider production.

The criteria might, for instance, be as follows.

  • The working time for the task in question is cut by a meaningful margin
  • A good share of users say they’d like to keep using it
  • Even when a wrong answer occurs, a human review can catch it
  • It can be folded into the existing workflow without much added burden
  • It clears the points to be checked on security, legal and information management

If you set a numerical target, make the period, the number of cases and the basis for comparison clear, for instance ‘three months before and after adoption’, ‘120 enquiries covered’, or ‘the average time to draft a reply’.

Map the roles of people and AI onto the workflow

Set down in writing what you hand to AI and what a person checks.

What matters here is not just ‘who looks at the AI’s output’. It’s designing the whole thing, from the information entered before the AI is used, through the check after the AI produces its answer, to transcribing into the business system and reflecting it in the knowledge base.

Run an operational test with a small group of users

Too few subjects for a PoC is a problem, and so is too many. Too few, and it leans on one person’s skill. Too many, and you can’t keep up with the queries and improvement requests.

It’s realistic to begin with a small number of people who know the task well and are keen to improve it. The headcount varies with the size of the organisation and the nature of the work, but with, say, five to ten people it’s easier to keep a handle on usage and improvement requests.

On that footing, check usage and issues weekly or fortnightly, and improve the prompts, training data and operating rules.

Get it into a reproducible form before the company-wide rollout

Once a PoC delivers results, rather than rolling it out company-wide straight away, organise it into a reusable form.

Concretely, that means things such as the following.

  • A business-flow diagram
  • Rules of use
  • Standard prompts
  • How the training data is managed
  • Frequently asked questions
  • Failure patterns
  • A record of production decisions
  • A manual for users

Only once these are all in place can another department try it as readily. Company-wide AI rollout is not about spreading a single success story unchanged. It’s about reaching a state where a successful way of working can be reproduced in other tasks too.

A way of thinking when you bring in Kanata

A way of thinking when you bring in Kanata

Kanata, a BtoB AI platform of that kind, makes it easier to design the move from PoC to production on a per-project basis.

If you take enquiry-handling support as the theme, for example, you might proceed as follows.

  1. Create a project for the department concerned. Into it, organise as training data the FAQs, operational manuals and past sample answers referred to in enquiry handling.
  2. Create an AI chat app so staff can use it to draft questions and proposed answers. Where needed, register the tone of replies and the points to check as prompts.
  3. Use AI summarisation to tidy up meeting notes and a record of responses, and draw out the FAQs and knowledge worth improving. In this way, rather than using AI as a one-off chat tool, you make it easier to fold it into a cycle of operational improvement.

That said, with Kanata as with any tool, it matters not to make adoption itself the goal. You need to make clear which business process you’ll change, whose burden you’ll lighten, and which business value you’ll tie it to, and use it on that basis.

A checklist for not stopping at the PoC

A checklist for not stopping at the PoC

Before taking an AI implementation project into production, please check the following items.

Checking business value

  • You can describe in a single sentence the business challenge AI will solve
  • You have a current baseline metric for the task in question
  • You’ve settled the metric you’ll compare against after the PoC
  • There’s value you can explain to senior management or the business owner
  • You’ve considered not just efficiency but the effect on quality and the customer experience

Checking the business process

  • You can see the workflow before and after AI is used
  • The steps you hand to AI and the steps people judge are separated
  • The person who checks the output has been decided
  • There are rules for handling exceptions
  • You’ve checked how it connects to existing systems and existing forms

Checking user acceptance

  • The users are clearly identified
  • You can explain not just how to operate it but how to use it in each business situation
  • The point of contact for queries has been decided
  • There’s a mechanism for collecting improvement requests from users
  • Frontline managers understand the purpose of using it

Checking governance

  • It’s settled what information may be entered and what is forbidden
  • There are conditions for reviewing AI output before it goes outside the company
  • You can manage permissions by department or project
  • You can review logs and usage
  • There’s a regular point for revisiting things

Checking in-house capability

  • The roles of the business divisions, transformation, IT systems and senior management are settled
  • There’s someone able to update the prompts and training data
  • There’s a mechanism for spreading success stories to other departments
  • Rather than leaving it all to an outside vendor, operational know-how stays within the company
  • There’s a forum or review cycle that keeps small improvements going

Summary

Summary

The further along a firm is with AI, the less it treats raising the sheer number of PoCs as the goal in itself. What matters is not how many trials you ran, but how many tasks remain in production and what business value they keep generating.

PoCs are necessary. Roll out grandly from the start and the impact when things go wrong is grand as well. Beginning with a small start, improving in an agile fashion, and settling frontline acceptance and governance as you go is a realistic way to approach AI implementation.

Before you begin a PoC, though, you need to settle the conditions for going to production. Without those conditions you can’t decide, even when results come in. Without operational design it won’t take hold on the frontline. Without governance the company-wide rollout decision becomes hard. Without a division of roles for in-house capability the improvements won’t continue.

An AI implementation project that doesn’t stop at the PoC is not about testing an AI tool; it’s an effort that designs the business process, the decision-making and frontline operation all at once.

Adopting generative AI is no cure-all. But by breaking the work down, rethinking the roles of people and AI, and building a habit of continuous improvement, the PoC turns from a passing experiment into an operation that keeps testing business value.

Q&A

What are the main reasons an AI PoC fails to reach production?

Beyond a shortfall in AI accuracy, the chief reasons are that the criteria for production, the workflow, the division of responsibility, governance and frontline training have not been settled. You need to assess a PoC’s outcome by business value and operational viability, rather than letting it end at ‘that was handy’.

What should be settled before a PoC begins?

It’s important to settle the metrics for business value, the metrics for operational viability, and the conditions for acceptable risk. You might, for instance, sort out in advance the working time, the number of cases covered, user satisfaction, the review arrangements, and the range of information that may be entered.

How should we split the work given to AI from the work people own?

AI readily takes on summarisation, classification, drafting, organising information and helping with knowledge search. People, on the other hand, should own the final decision, exception handling, regard for customer feeling, and high-risk judgements in areas such as legal, HR and contracts.

Why is governance necessary for a company-wide AI rollout?

While it stays vague what information may be entered, how AI output is treated, who has permission to use it and who is responsible for review, the risks mount at the company-wide rollout stage. Governance is there not to halt AI use but as the foundation for widening the scope of use with confidence.

In what situations might an AI platform such as Kanata be worth considering?

It’s an option worth considering where you want to run things while organising AI chat, summarisation, training data and prompts by department or by task. That said, don’t make adoption itself the goal; you need to make clear first which business process you’ll change and which business value you’ll tie it to.

Share this article