We’ve got more people using AI now, yet when it comes to appraisals we’re rather at a loss as to what we should actually be looking at.
It’s a remark that tends to surface from HR teams and front-line managers after a generative-AI training session. There are certainly people who now draft documents more quickly, who tidy up meeting notes more deftly, who can knock together a first cut of a client proposal in short order. What rather a lot of organisations have yet to pin down, however, is how to assess that as an “AI skill” and how to feed it into grading systems and development plans.
In the past, AI use was left to individual ingenuity, and even when people went through reskilling courses it wasn’t joined up properly with day-to-day goal setting, one-to-ones or appraisals. These days the prevailing move is to stop asking simply whether someone “used AI” and instead break it down into behaviours such as work design, prompt design, review and knowledge sharing, then reflect those in skill maps and development plans. When you’re thinking about developing people and making skills visible, frameworks for AI skills and workforce policy such as the OECD AI Principles are a useful point of reference.
In this article we set out how to design “AI use in HR appraisal” and “AI skill assessment”, and how to connect reskilling assessment to people development. The aim is not a state in which only a handful of AI-fluent people deliver results, but one in which each department defines the behavioural indicators it needs and people can grow their AI skills along a clear career path. That said, simply bolting on a few extra appraisal items won’t do the trick. Only with operating rules, assessor understanding and regular review in place does it actually function as a system.
What to think through before putting AI skills into HR appraisal
As generative AI spreads through an organisation, the next question that arises is not “who is using it and how much” but “who is turning it into business results”.
The number of times an AI tool is used, or the number of prompts entered, is rather thin as a basis for appraisal. Someone might use AI every single day and yet, if they pass the output along without checking it, the risks to quality and compliance only grow. Equally, there are people who don’t use it especially often but who apply it adroitly to meeting preparation, client proposals, process improvement and knowledge sharing, and so contribute to the results of the whole team.
For that reason, when you fold AI skills into HR appraisal, you first need to keep two things apart: “being able to use the tool” and “being able to improve the work”.
What deserves to be assessed is not the bare fact that someone used AI. It is whether, by using AI, they rethought how the work is done, raised the quality of the deliverable, and shared it in a form others can reproduce.
Take the same task of putting a document together: with the following two people, what merits assessment is quite different.
- A case that stops at personal efficiency
- Someone who has AI write the text, tidies it up a little, and submits it. They may well have shaved off some time, but it hasn’t become a way of working or a piece of team know-how.
- A case that leads to organisational improvement
- Someone who breaks the document-creation process down, decides where to use AI across “drafting the outline”, “organising the client’s issues”, “adjusting the wording” and “the final review”, and shares that procedure with the team. Here AI use goes beyond personal time-saving and becomes a behaviour that is easy to credit as improving the organisation’s productivity.
If you’re going to build AI use into HR appraisal, it’s important not to overlook this distinction.
Four perspectives worth examining in AI skill assessment
When you bring AI skill assessment into a system, building in finely detailed appraisal items from the very start tends to leave the front line unable to operate it. It’s easier to start by narrowing things down to four perspectives common to every job type.
Work-design capability
Work-design capability is the ability to separate the work you hand to AI from the work a person ought to judge.
Generative AI is well suited to drafting text, summarising, organising the key points, generating ideas, and producing comparison tables. Final judgements, reaching agreement with clients, HR appraisal, and decisions touching on legal or financial matters, on the other hand, are things a person needs to take responsibility for.
People with strong AI skills don’t hand everything to AI; they think through “which steps of this task can AI handle” and “from where a person ought to check”.
Put as behavioural indicators, you might define it like this:
- Can break down their own work process and explain which steps AI can be used for
- Clearly distinguishes the part handed to AI from the part a person checks
- Can explain how AI use changed the time taken, the quality of the deliverable, and what was shared with those around them
What matters here is not simply writing “made it more efficient”. When you use it for appraisal, you need to record the task in question, the comparison period, the deliverable and the checking method as a set — something along the lines of “used AI for the first draft of the monthly meeting materials and organised the process into three stages: outline, body text and review”.
Prompt-design capability
Prompt-design capability is the ability to convey to AI, clearly, the purpose, the assumptions, the output format and the constraints.
An instruction along the lines of “please sum this up nicely” won’t produce output of any consistent quality. By conveying who the text is for, what it is meant to achieve, what information it should assume, and what format to produce, you make AI’s answers genuinely usable in the work.
This capability is not merely an operating knack. It is also the ability to put your own work purpose into words and to organise the conditions you need.
As behavioural indicators, you might consider the following:
- Can give instructions that include the purpose, the intended reader, the background information and the constraints
- Doesn’t settle for a single round of output, but can improve it with follow-up instructions
- Organises frequently used instructions into a form the team can reuse
For instance, where you use an environment such as Kanata in which prompts and learning data can be accumulated across a team, it becomes easier to turn individual ingenuity into a shared asset. That said, simply introducing the tool isn’t enough. You also need to decide which prompt is treated as the official version, who updates it, and how older prompts are reviewed.
Review capability
Review capability is the ability to check what AI has produced and judge whether it is fit to be used.
AI’s answers, however natural and persuasive they may appear, can contain factual errors, out-of-date information, over-confident assertions and suggestions that don’t fit the context. Figures, proper nouns, dates and anything touching on legal, HR or financial matters in particular need to be checked against the original sources or internal rules.
Leaving this review capability out when you assess AI use is asking for trouble. It means only those who produce output quickly get the credit, while those who check carefully fade from view.
As behavioural indicators, you might frame it like this:
- Checks AI output by separating fact, inference and the points that remain uncertain
- Builds a human review into anything submitted externally or any important material
- When they spot mistaken output, shares the cause and the corrective instruction
In AI skill assessment, you need to look at “the ability to doubt correctly” every bit as much as “the ability to produce quickly”. This matters from an AI-governance standpoint too. For literacy, accountability and verifiability in AI use, the NIST AI Risk Management Framework is a useful reference.
Sharing behaviour
Sharing behaviour is the ability to spread your own AI know-how across the team.
In an organisation where AI use rests on individual effort, a “only the capable few can do it” state persists. Left like that, even if you bring reskilling assessment into a system, it won’t lift the whole organisation.
What matters is turning the ways of using AI that worked into a form others can reproduce.
As behavioural indicators, you might consider the following:
- Shares frequently used prompts and work procedures with the team
- Shares not only the AI successes but also the failures and the points to watch
- Gives concrete support so that juniors and colleagues can use AI
Bringing in this perspective turns AI use from an individual knack into organisational learning.
Using a skill map to make each department’s expectations visible
To assess AI skills, drawing up a skill map is an effective approach. A skill map is a table that lists the skills required for each job type and grade and makes the current position and the development gaps visible.
That said, applying the same items uniformly to every employee creates a mismatch with the front line.
Sales, marketing, HR, accounting, IT and corporate planning all use AI in different situations. So you need to design common skills and department-specific skills separately.
| Category | Main items | Approach to designing the assessment |
|---|---|---|
| Common skills | Work-design capability, prompt-design capability, review capability, sharing behaviour | Defined as the foundational AI-use behaviours expected across every job type. |
| Department-specific skills | Behaviours suited to a particular role, such as sales, HR, marketing or IT | Defined as behaviours tied to each function’s business results. |
Department-specific skills are defined by tying them to each function’s business results. For sales, that might cover organising sales-call notes, drafting proposals and framing hypotheses about client issues. For marketing, article outlines, ad copy, persona work and campaign retrospectives. For HR, training design, structuring appraisal comments, maintaining FAQs and producing onboarding materials.
What matters here is expressing AI use not as “operating a tool” but as “a behaviour within the job”.
For instance, making a skill-map item read “can use the AI chat” leaves the assessment shallow. Define it instead as “can use AI to organise, from sales-call notes, the points needed for the next proposal, and put it to a manager’s review”, and you get an assessment far closer to the actual work.
The same thinking applies when you build AI skills into a grading system. Of juniors you expect “can use AI for routine tasks”; of mid-level staff, “can build AI into the work process”; of senior staff, “can roll out a model of AI use across a team or department”.
How to think about building AI skills into a grading system
When you reflect AI skills in a grading system, it’s important not to place the same standard on every grade.
The AI use you expect of a junior employee differs from what you expect of a manager. Of junior employees you ask for efficiency in their own work and improvement of their deliverables. Of managers, by contrast, you ask for the role of reviewing the team’s whole work process and embedding AI use across the organisation.
By way of example, the expectations for each grade might be organised like this:
| Grade / role | Expected AI-use behaviours |
|---|---|
| Member level |
|
| Leader level |
|
| Manager level |
|
Varying the behavioural indicators in line with each grade’s role like this connects AI skills to the career path.
For a company that has adopted a job-based HR system, one option is to add the expected AI-use behaviours to the job description. A job-based system is an approach to design in which the placement, appraisal and reward of people are considered on the basis of clearly defined job content and expected outcomes. That said, writing the AI-use descriptions in fine detail from the outset risks leaving them unable to keep pace with changes in tools and work.
A realistic way to run it is to trial it first with the main job types and review it each quarter or half-year.
Connecting reskilling assessment to development plans
A common failing in reskilling assessment is to assess the mere fact of having attended the training.
Attending the course is, of course, important. But if you’re folding AI into people development, you need to look at what they put into practice afterwards, which work they applied it to, and what improvement it led to.
For that reason, design AI training as a set of four parts: “attend”, “practise”, “reflect” and “the next development theme”.
- Attend a company-wide common AI-literacy course through e-learning or classroom training
- Tackle the practical tasks of each department — sales, HR, planning and so on
- Reflect on the results of using AI in one-to-ones and team meetings
- Feed what comes out of that reflection into the next development plan and goal setting
For sales, building a proposal hypothesis from sales-call notes; for HR, producing training notices and FAQs; for planning, drafting the outline of meeting materials — that sort of thing.
In the reflection, check the following points:
- Which work AI was used for
- Which steps were shortened
- What was watched for when checking the output
- How it might be improved next time
- Whether there’s a model that can be shared with other members
Feed this reflection into the development plan and AI use becomes not a one-off course but a continuing theme for growth.
In appraisal meetings too, it matters to ask not “did you attend the AI training” but “which work did you apply it to after the course, and what did you learn”. On developing people for the AI era, the OECD AI Principles are also worth consulting.
How to bring AI use into goal setting
When you build AI skills into HR appraisal, the connection to goal setting matters too.
That said, a goal such as “use AI every week” isn’t one I’d recommend. Make frequency of use the sole goal and you risk an increase in aimless use.
Set goals by tying them to business results and changes in behaviour.
For example, goals along these lines:
- Apply AI to the process of creating the monthly meeting materials and rethink the steps from drafting the outline to producing the first draft
- Apply AI to organising the key points before a client proposal, making it easier to set out the items to check at a manager’s review
- Share two reusable prompts a month within the department, contributing to standardising the team’s work
- After attending the course, organise the AI-use procedure for three of their tasks
Here too, what matters is not leaning too heavily on the numbers alone. AI use bears not only on short-term time-saving but on the quality of thinking, the precision of review and knowledge sharing.
For that reason, it’s wise to combine quantitative and qualitative goals.
If, for instance, you write “shorten the time taken to produce the first draft of the monthly materials”, make the comparison period and the measurement method clear. Set a condition such as “record the average production time over the month before adoption and compare it with the month after” and it becomes easy to verify later.
Time-saving alone, on the other hand, says nothing about quality. Add a qualitative goal alongside it, such as “rather than using AI output as is, set out in writing the procedure for checking the evidence and adjusting the wording”, and you can assess in a way that takes in quality and reproducibility too.
Making the development of AI skills a systematic process
To build AI skills into a system, you need to run training, practice, sharing and reflection within the same loop.
First, prepare company-wide foundational AI training and department-specific AI-use training. The training covers not only the basics of using generative AI but the handling of information, checking of output, and which work is fine to use it for and which calls for caution.
Next, tackle practical tasks. For the sales department, summarising sales-call notes and structuring proposals; for HR, training design and organising appraisal comments; for planning, drafting meeting agendas and decision-making materials.
Then share the prompts and work procedures that went well across the team. The means of sharing — an internal wiki, a chat tool, an LMS, a knowledge-management tool — can be chosen to fit the existing environment. Where you use a service such as Kanata, which handles AI chat, e-learning and the management of prompts and learning data within one environment, the point in its favour is that it’s easy to join training, practice and sharing into a single flow.
Whichever tool you use, though, what matters is the operation. You need to decide who updates the materials, which prompt is treated as the official version, and how the handling of confidential information is checked.
Managers check what was put into practice with AI in one-to-ones and appraisal meetings. The point, however, is to treat it as material for development rather than to monitor usage logs.
You might, for instance, use questions like these:
- Which work did you use AI for?
- Which parts of the output did you amend?
- What looks as though it could be improved next?
- Is there a model that could be shared with other members?
Through questions of this sort, AI skills are not merely assessed — they grow.
Common pitfalls when getting started
When you build AI skills into an appraisal system, there are a few patterns of failure.
Assessing the number of uses alone
The thing most to be avoided is making the number of times AI is used the sole assessment metric.
The number of uses can serve as background information, but it shows nothing of results or quality. If anything, an increase in aimless use brings the risk of more checking work and of mistaken information spreading.
In appraisal, rather than the number of uses, you need to look at how the work process and the deliverables have changed.
Assessment items that are too abstract
An item that reads only “actively makes use of AI” leaves judgement split across assessors.
One manager rates highly the person who uses it every day; another may set store by the quality of the deliverable. That way you can’t maintain fairness in appraisal.
For that reason, assessment items need to be brought down to behavioural indicators.
Rather than “makes use of AI”, for instance, express it as “organises which steps of their work AI can be used for, and shares the results of putting it into practice with the team”.
Managers who can’t assess it
To assess AI skills, a certain amount of understanding is needed on the manager’s side too.
While assessors don’t know how to use AI, they can’t properly see what their reports are putting into practice. The upshot is that appraisal comes to rest on self-report, or the loudest voices get rated highly.
So you need not only training for members but assessor training for managers. Assessors need a viewpoint that takes in work design, review and sharing behaviour, not merely how slick the prompts are.
Ignoring departmental differences and assessing uniformly
How AI is used differs from one department to another.
The IT department might use AI for organising requirements and handling enquiries. The marketing department would use it for content planning and ad copy. The leadership might use it to organise the key points of a decision and compare scenarios.
A company-wide common standard is needed, but on its own it isn’t enough. It matters to separate common items from department-specific items and tie them to each function’s results.
Five steps to embedding AI skills as a system
The procedure for building AI skills into HR appraisal and development plans is easier to take forward if you think of it in the following five stages.
-
Decide the purpose of AI use
First, make clear what you are assessing AI skills for. Build the assessment items with a vague purpose and it risks being taken as “the company just wants us to use AI”.
The purpose should sit with growth support, not control.
- Raise the reproducibility of process improvement
- Connect reskilling to the actual work
- Make each department’s level of AI use visible
- Reflect the new skill requirements in the career path
Explaining this purpose at the outset raises the front line’s sense of buy-in.
-
Define the common skills
Next, define the AI-use skills common to every employee.
My recommendation is the following four categories:
- Work-design capability
- Prompt-design capability
- Review capability
- Sharing behaviour
What matters is not making them too detailed from the start, but putting them in words anyone can understand.
-
Bring them down to department-specific behavioural indicators
Once the common skills are defined, build behavioural indicators to suit each department’s work.
Tie them to proposal activity for sales, to training and appraisal for HR, to enquiry handling and requirements work for IT, and to content production and analysis for marketing.
At this point it’s important not to have HR draw them up on its own, but to bring the front-line managers in. Assessment items not defined in the language of the front line tend not to get used.
-
Pair training with practical tasks
People don’t grow on skill definitions alone.
Build a flow in which they learn the basics in training, try them on practical tasks, and reflect in one-to-ones. Combine AI chat, e-learning, an internal wiki and knowledge-sharing tools so that the training content and the actual work don’t come apart.
-
Align the standards across assessors
Finally, line the standards up across assessors.
If the same behaviour is rated high by one assessor and low by another, trust in the system falls. You need to align them using concrete examples, within appraisal meetings or manager training.
You might, for instance, prepare cases like the following:
- Someone who uses AI often but is lax about checking the output
- Someone who doesn’t use it especially often but shares how to use it with the team
- Someone who shows a time-saving effect but whose deliverables vary in quality
- Someone who improves not only their own work but the procedures of the whole department
Talking things through on the basis of examples like these sharpens the resolution of the assessment.
In summary: the purpose of assessing AI use is growth support, not control
Hearing that AI skills are being built into HR appraisal, people sometimes take it as “are they going to monitor us” or “will our rating drop if we don’t use AI”.
For that reason, the message matters enormously in designing the system.
The purpose of assessing AI use is not to control employees. It is to make clear the skills the new ways of working require, reflect them in development plans, and connect them to the career path.
The gap between those who use AI and those who don’t may well widen further from here. But leave that gap to individual effort alone and the organisation’s learning won’t progress.
By joining up HR appraisal, the grading system, reskilling and development plans, AI use becomes easier to treat not as the skill of a few but as the capability of the whole organisation.
That said, a system isn’t finished once you’ve built it. The AI tools, the work and the skills required all keep changing. Rather than aiming for a perfect appraisal system from the start, the wise posture is to begin small and review each quarter or half-year.
Start with one department, one job type and one development theme. The practical examples you gain there become the foundation for rolling it out to the next department.
Q&A
Should AI skills go straight into HR appraisal?
Rather than tying them straight to an appraisal score, it’s more realistic at first to treat them as part of the development items or goal setting. Even when you do bring them into appraisal, you need to look at behaviours — process improvement, deliverable quality, review, knowledge sharing — rather than “the number of times AI was used”.
Which items should I look at in AI skill assessment?
It’s easier to operate if you start by organising things into four: work-design capability, prompt-design capability, review capability and sharing behaviour. More than fine points of tool operation, what matters is whether they’re turning AI into business results and sharing it in a form others can reproduce.
How should I assess the results of attending reskilling training?
Assessing on attendance alone is best avoided. Check which work they used it for after the course, what improvement there was, and what they’ll make the next development theme. Joining up training, practical tasks, one-to-ones and appraisal meetings helps reskilling take root in the actual work.
Are there things to watch when putting AI skills into a grading system?
It’s important to vary the behaviours you expect by grade. It’s easier to organise if you expect the member level to use it safely in their own work, the leader level to spread its use within the team, and the manager level to build it into work processes and development plans.
Won’t employees feel they’re being monitored if AI use is assessed?
That concern is real. So in designing the system you need to convey clearly that the purpose is “growth support”, not “monitoring”. Rather than looking at usage logs alone, the preferable way to run it is to check the person’s own learning and improvement behaviours through one-to-ones and reflection.