If you haven’t already checked out the first in this two-part article, “Start with the work, not the tool: Choosing your first AI use case” – head over and read that now.
So, you’ve identified a real problem or friction point, used it to choose your sensible and bounded first AI use case, and can articulate clearly what outcomes you’re expecting AI to improve. You’re already ahead of most!
But this is also the point where the most promising early ideas can start to evolve into something a bit less straightforward. This is where you find out whether your use case is actually… useful!
Remember, the aim of a first AI pilot isn’t to prove that AI works. It’s to understand where it can help your organisation and your teams, where it will struggle, and what needs to be in place before you rely on it more widely.
Let’s return to the example from our previous article: using AI to help tutors prepare for progress reviews.
On paper, the use case sounds fairly simple: Bring together relevant information and produce a first summary for the tutor to review.
But if you ask five experienced tutors to tell you in precise detail what they each actually do before their reviews, you’ll usually uncover five variations of preparation – much more than each of them following the same straightforward information gathering exercise.
Some may check whether a missed activity is particularly unusual for that learner. Another may like to compare recent employer feedback with earlier conversations. They may know that that one specific overdue task is an administrative issue, as opposed to a sign of disengagement.
So before deciding what a ‘good’ AI-generated summary looks like, do involve the people who prepare and carry out the reviews. Ask them what they look for, what they currently have to search for, what they would never want the system to assume and what would make them trust or distrust the output.
This is not just key for making sure the technology can support the process, but also helps you secure team buy-in.
We’ve discussed before that not every AI use case needs the same level of control (see The agentic shift: why funded learning's future is intelligent, connected, and governed).
Using AI to tidy up the wording of an email is very different from using it to flag a learner as at risk, recommend an intervention or influence a decision about programme progress.
Before you test your first AI use case, be clear about what could happen if the output is incomplete, misleading or simply wrong.
For a lower-risk task, a straightforward human review is likely enough. For more involved use cases, you may need restricted permissions, a clear escalation route, an audit record and explicit approval before any action is taken.
As we explored in our article on why AI governance is the foundation of innovation in funded skills, these controls are not something to bolt on later. They should shape the use case from the beginning.
Test the awkward examples, not just the tidy ones
Vendors demonstrating new AI functionality tend to use clean information and a carefully chosen example. But as we know, real delivery environments are rarely that accommodating.
You’re likely to have some incomplete learner records, and even with platforms that offer structured workflows like Bud, there’s a chance that different team members may record similar information in slightly different ways. A learner may have changed employer, returned from a break in learning or have support needs that change how the situation should be understood.
So we do recommend you test your use case with a few examples that are most likely to expose its limitations and give it a challenge!
For our progress-review summary example, that might include:
We’re not trying to ‘catch the AI out’ for the sake of it, but we do need to understand the conditions in which its output remains useful and the point at which a person will need to step in. This is also how confidence is built – team members are much more likely to trust it when they have seen both what it does well and where its limits sit.
One of the trickiest things about reviewing AI output is that it can sound very, VERY convincing. Even when it is absolutely factually correct.
A summary can be clear, fluent and professionally written while still missing the most important piece of information. A recommendation can sound perfectly reasonable while being based on only part of the learner’s record.
So the most vital piece of your testing puzzle is to go beyond asking ‘Is this accurate?’.
Also carefully review:
This links directly to the point we made in why AI assistants are only as useful as the information they can access. A polished response is not necessarily a reliable one. The quality, consistency and context of the underlying information are vital.
Once your pilot is running, it’s easy to become distracted by the technology itself.
Your team members are likely to start discussing the quality of the summaries, how often they use the new capability, if they like it or not. All of those things matter, but they are not the original outcome.
If the problem was that tutors were spending too much time preparing for reviews, has that actually changed?
Also look out for unintended consequences.
Has the new process created another administrative step? Are staff spending the time they saved checking weak output? Has a standardised summary made people less likely to look more deeply at the learner’s individual circumstances?
The AI will produce something; that’s not in question. But is the work genuinely better as a result of it?
Not every pilot should move straight into wider use – and don’t be disheartened if this is the case with your first AI test.
You might discover that while the idea is sound, the AI doesn’t yet have access to the right information to make the output valuable enough. You may discover that team members need clearer guidance on following the process. You may also determine through your testing that the risk is higher than you originally expected.
None of these point to a failed pilot – they point to a pilot that’s surfaced valuable data.
A useful pilot gives you evidence about what AI technology can do inside the current reality of your organisation. You may well determine that your organisation needs to lay stronger platform foundations for AI before moving into further initiatives (check out our article on the AI Readiness maturity curve here).
So, when the output is consistently useful, the controls are understood and the people doing the work believe it helps them – scale it!
When the value is clear but the information, workflow or guidance needs work – improve it!
When the benefit is uncertain or the risks cannot yet be managed properly – pause it!
Your first AI use case doesn’t need to be perfect, and it doesn’t need to feel overwhelmingly ‘transformative’ for the whole organisation. It just needs to solve a recognisable problem well enough to earn you and your team’s confidence.
That means involving the people who understand the work, testing the complicated cases as well as the easy ones, checking what the AI misses or leaves out and measuring whether the original friction has actually improved.