RELACTIS BlogMeasurement5 min read

What AI Skills Tests Can't Measure

Score your people on AI ability and you will usually pick the wrong person for the job. Here is what a skills test actually measures, and what it quietly leaves out.

"Let's start by measuring everyone's AI proficiency." Once an AI rollout gets concrete, somebody says this. It does not sound like a bad idea. You cannot plan without knowing where you are, you cannot show training worked without a before and after, and a number makes the proposal easier to get approved.

So a test gets built. Multiple choice on prompting, a few questions on tool operation, maybe a scenario or two. Scores come back. Departmental averages come back.

Then you sit down with the results to decide who should do what, and the numbers turn out to be strangely unhelpful.

What the test is actually measuring

A skills test measures real things: whether someone knows the tools, knows the vocabulary, and can recall a handful of prompting patterns. Run it before and after a training session and it will tell you honestly whether the training landed.

The trouble starts past that line. A high scorer is someone who knows, which is not the same as someone who does. Those two come apart more often than you would expect. Plenty of people can define every term and have never once used AI on their own work, and plenty of people who could not define a single term have it wired into their Tuesday morning.

There is also a structural problem. The moment people know a score is attached, they start managing the score. That is not dishonesty, it is a normal reaction to being graded, and it lands hardest exactly where you need candour most. The people who are not using AI, and the people who are unconvinced by it, are the ones whose behaviour decides whether the rollout works. They are also the ones with the strongest reason to give you the answer they think you want.

Who this is for

Two kinds of reader, and it is worth saying which. If you are the person driving AI adoption inside a company and someone has just proposed starting with a proficiency test, this piece is about why that instinctive hesitation of yours is well founded.

If you are here because you enjoy assessments, the interesting part is the design question. Every assessment has to decide whether it is producing a rank or a shape, and that decision determines what you can do with the result.

Direction, not amount

Picture four colleagues. One keeps finding uses nobody asked for. One turns other people's tricks into templates the rest of the team can run. One checks the source on every output before it goes anywhere. One barely touches the tools but knows precisely who to ask.

Now rank them. You cannot, at least not usefully. They are not four points on one line, they are four directions. A skills test is a single line by construction, so it flattens the four into an order, and the flattening destroys the information you actually wanted.

This is why RELACTIS measures seven factors instead of one score: Creation, Crossover, Autonomy, Advocacy, Systemization, Verification, and Engagement. High is not better on any of them, and the low end of each carries real value. A separate piece walks through what each one asks.

Three decisions a score will get wrong

That may sound abstract, so here are three places where teams feel it.

Choosing who pilots a new tool. The temptation is to hand it to your highest scorer. The person you actually want is the one who enjoys unfinished things: undeterred by rough edges, back in three days with uses nobody had thought of. That is The Pioneer. Give the same assignment to someone whose strength is following a documented process and you will get a careful review of the onboarding flow. Both people did their job. Only one answered the question you were asking.

Choosing who owns quality. Sooner or later an AI workflow needs someone whose reflex is "what's the source for this?" That is The Verifier, and skills tests do not surface this person at all. Because they tend to voice objections, they often get filed as an obstacle to the rollout instead. Read the scores alone and you will sideline the person the project most needed.

Choosing who teaches. This is the easiest one to get wrong, because being good at AI and being good at teaching AI are separate talents. The teacher you want is The Evangelist, who translates a feature into the language of someone else's job. Your most advanced user gives a demo that impresses the room. The Evangelist gives a nudge that is still being used a week later. Pick by score and you will get the demo.

The common thread is that all three decisions need to know which way a person leans, and a test only reports how much they know. Staff a rollout on the second kind of information and you find out it was wrong after the roles are already handed out.

Where skills tests do belong

In fairness, there are places where a test is the right instrument.

If you have run training on a specific tool, testing comprehension before and after is reasonable. Checking that people understand policy, that confidential material must not be pasted into a public model, for instance, suits a test format well. Wherever there is a standard everyone genuinely should reach, a test is how you check they reached it.

What a test cannot do is tell you who to put where. That question wants a shape, and a test only returns a height.

So what should you measure?

Our answer is a typology rather than a score. Twenty-four questions produce a reading on seven factors and place you near one of fifteen types. Nothing gets ranked.

Refusing to rank has a practical payoff. When no answer can lose, people stop managing their response, which means the non-users and the sceptics answer honestly. As noted above, those two groups are the ones you most needed to hear from.

Laid out across a team, the same results show who you have and which role is sitting vacant, which this piece covers in detail. You do not need to go that far to see the difference, though. One person's result is enough to show that it returns something a score does not.

Next time someone proposes measuring AI proficiency, the useful reply is probably this: measuring is fine, let's just be careful about what we measure. It is free, takes about four minutes, and needs no signup, so take it yourself first and judge from there.

Which of the 15 types are you?

24 questions · about 4 minutes · free · anonymous · no sign-up

Take the free assessment →

Or meet all 15 AI usage types first →

Read next