
26 August 2026
Last time, this series looked at the framework underneath — the taxonomy that quietly decides what an organisation is able to see about its own capability. But a vocabulary only becomes data when somebody applies it to a person. Someone sits down, reads a skill description, thinks about a colleague or about themselves, and picks a number. That moment — small, routine, repeated thousands of times a year — is where every dashboard, every gap analysis and every workforce plan ultimately comes from.
It is also the least examined part of the whole apparatus. Organisations will spend a quarter designing a framework and an afternoon deciding how it gets rated, then treat what comes out as measurement rather than judgement. This post is about that moment: what a skills rating actually is, the two directions it reliably drifts in, and what makes one honest enough to plan on.
A Rating Is a Claim, Not a Fact
A skills rating looks like a measurement. It arrives as a number on a five-point scale, it averages neatly, and it renders beautifully in a bar chart. But it is not a measurement in any sense a scientist would recognise. It is a claim — one person’s judgement about capability, made at a particular moment, on whatever evidence happened to be within reach that week. The number is a compression of that judgement, and like every compression it discards the part that mattered most: the reason.
This is not an argument against rating people. Judgement is unavoidable and often rather good, and organisations have run on it for a century. It is an argument against forgetting what the number is made of. The moment a rating is treated as a fact, it stops being questioned — and an unquestioned claim will sit happily in a skills analytics view for two years, shaping succession shortlists and hiring briefs, long after the person it described has moved on. Everything downstream depends on remembering that each cell in the grid was once a sentence somebody said out loud.

Generous, Frightened, or Both
Ratings drift in two directions, and most organisations manage both at once. The first is generosity. Marking a colleague down is unpleasant, the conversation is awkward, and a three feels like a criticism where a four feels like encouragement. Nobody sets out to inflate anything; each individual four is defensible and kindly meant. Two cycles later, ninety per cent of the workforce is proficient or better in almost everything, the gaps have vanished from the report, and the only honest conclusion available from the data is that no development is required anywhere.
The second direction is fear, and it usually shows up in self-assessment. If people suspect a rating will be read as a verdict — that admitting to a two will cost them a project, a promotion or a place on the list — they will rate themselves defensively, and the safest answer is always the middle. What makes a skill-will conversation work is precisely the opposite reflex: the willingness to say plainly that you are not there yet and would like to be. That candour is a function of consequence, not of scale design. Nobody tells the truth about a gap to an audience they expect to be judged by.
Evidence Is What Turns a Number Into a Sentence
The most reliable cure for drift is not a better scale. Five points, four points, behavioural anchors, elaborate rubrics — organisations rework these endlessly and the ratings move very little, because the problem was never the granularity. The problem is that a number arrived without anything attached to it. Ask for one piece of evidence alongside every rating, a specific thing the person did and roughly when, and the whole exercise changes character. It becomes much harder to award a casual four when a sentence has to follow it.
This is also the argument for keeping assessment close to work rather than to the calendar. Structured evaluations and feedback gathered near the moment something happened are worth more than a retrospective sweep in November, when the only evidence anyone can recall is whatever occurred in October. The same applies at the other end: a rating that produces a specific training goal is being used, and things that get used get corrected. Ratings that go nowhere are never wrong, because nothing ever tests them.

Calibration Is a Conversation, Not a Committee
Different managers mean different things by the same number, and no amount of guidance in a handbook fixes that. What does fix it, slowly, is managers comparing notes out loud. Half an hour with three or four peers, each explaining what a four in one specific skill looks like in their team, does more for consistency than a page of definitions — mostly because it surfaces the disagreement rather than averaging it away. The point of the session is not to force agreement on every rating. It is to ensure that when two people say four, they are describing roughly the same thing.
It is worth being clear about what this is not. Calibration in the skills sense is a shared-language exercise, not a forced distribution and not a ranking; the moment it starts allocating a fixed quota of fives it has become a performance mechanism and the honesty drains out of it within one cycle. Done well, it is a quiet, recurring habit attached to the development review cycle, and its main output is not a corrected spreadsheet but a group of managers who now describe capability in compatible terms.
Honest Beats Complete
There is a useful diagnostic available to anyone with access to the reporting. Look at the distribution rather than the average. A healthy skills dataset has spread — ones and twos in places, genuine variation between teams, some skills where the organisation is visibly weak. A dataset where almost everything clusters at three-and-a-half is not describing a uniformly competent workforce; it is describing a population of people who have worked out what the safe answer is. The shape of the data is usually more informative than anything in it.
Which leads somewhere slightly uncomfortable. A partial dataset that people believe is worth considerably more than a complete one they quietly discount, and coverage is the wrong first target. Better to have forty skills rated candidly, with evidence, by managers who have compared notes, than four hundred rated to fill in a form. The framework tells you what can be seen. The rating tells you whether what you are seeing is real — and no amount of downstream sophistication will rescue a number that nobody meant.
