Performance calibration

Performance calibration is a meeting where managers review each other's proposed ratings before they are final, so that the same rating means roughly the same thing across teams.

Home HR glossary for HR teams

Performance calibration

Performance calibration is a working session in which a group of managers, usually with HR facilitating, review the ratings they intend to give before those ratings become final. The aim is not to change anyone's opinion of their own team. It is to make sure that a rating of exceeds expectations means something comparable whether it was awarded in sales or in finance.

The need for it comes from a structural fact about performance reviews: each manager rates their own small group in isolation, using their own internal scale. Some are strict, some are generous, most drift over time, and none of them has visibility into how their peers are scoring. Without a step that compares across groups, the rating on the form describes the manager at least as much as the employee.

Why it matters more than it sounds

Ratings are rarely just feedback. They feed promotion decisions, performance-based pay, bonus pools and, in a downturn, decisions about who stays. If one department's manager rates generously and another's rates strictly, the generous manager's team receives more money and more opportunity for identical work. That is a fairness problem with a legal dimension attached, and it accumulates quietly year after year.

Calibration also catches the biases a single rater cannot see in themselves: the halo effect, recency, the tendency to rate everyone in the safe middle, and patterns of unconscious bias that only become visible when you look at a whole population at once rather than one review at a time.

How a session actually runs

A typical session covers one level or one function at a time, runs with five to ten managers, and takes two to three hours. The facilitator opens with the distribution as it currently stands, by team and by manager. Then attention goes to the outliers rather than to everyone: the highest and lowest ratings, any manager whose spread looks unlike the rest, and any employee whose rating changed sharply from the previous cycle.

For each of those, the manager gives evidence. The questions that do the work are simple. What did this person do that someone rated a level lower did not do. Would you rate them the same if they worked in another team. What would have had to be different for the next rating up. A rating survives if the evidence holds and moves if it does not, and the change is recorded with a reason so the manager can explain it to the employee in their own words rather than reporting that a committee decided something.

Forced distribution is a separate question

Calibration is often confused with forced distribution, where a fixed share of employees must land in each rating band. They are not the same thing. Calibration compares evidence across managers; forced distribution imposes a shape on the result regardless of the evidence.

The case for a distribution guideline is that without one, ratings inflate until almost everyone is above average and the scale stops carrying information. The case against is that real teams are not distributed identically: a small team that genuinely performed well is forced to demote someone, which damages trust and makes people avoid the strongest teams. Most companies that have tried strict forced ranking have moved away from it. The workable middle is to publish an expected distribution as a reference, require a written explanation when a team departs from it substantially, and let the explanation win when it is good.

What separates a useful session from a negotiation

  • Evidence rules, not advocacy. Without a rule that every claim is backed by a specific example, the outcome tracks which manager argues best.
  • Shared rating definitions written before the cycle. If exceeds expectations is not defined behaviourally per level, the session is people comparing private definitions.
  • A facilitator with no team in the room. Usually HR, whose actual job is to ask the uncomfortable question and to watch for patterns across gender, tenure, working pattern and location.
  • Calibrate before results are shared. A rating that reaches the employee and then changes costs more trust than the inconsistency it fixed.
  • Write down the reason for every change. The manager owns the conversation afterwards and cannot have it without knowing why.

The failure modes are the mirror image. Senior managers protecting their own people and juniors staying quiet. Trading, where one manager accepts a downgrade in exchange for support elsewhere. Sessions so long that everything after the first hour is decided by fatigue. And the version where ratings are quietly adjusted after the meeting by someone who was not in it, which is the fastest way to make managers stop taking the whole process seriously.

Preparing a calibration session with PeopleForce

Calibration happens in a room, but it only works if the numbers on the screen are comparable in the first place. That is what the scoring setup in PeoplePerform is for: review questions are tied to defined competencies rather than free text, each competency in the cycle carries an explicit weight, and the weights have to total 100 before the cycle can move on. Every manager is therefore scoring the same competencies with the same weighting, which removes the most common reason two ratings cannot be compared at all.

The inputs a session needs are in the same place. Each employee's results break down by reviewer group, with a radar chart per competency, so a manager rating that sits well above the peer and upward assessments of the same person is visible before anyone has to ask about it. Sharing is controlled at cycle level and can be held manually, which is what makes it possible to calibrate before results reach employees rather than after. Cycles have collaborators alongside the owner, so HR can work inside the cycle during the review, and a cycle can be locked while that work is going on. For the distribution view across departments and managers, HR analytics is where the pattern shows up, including the cross-cycle comparison that tells you whether a manager is consistently generous or had one unusual quarter.

Let us show you what's possible

From Core HR to advanced workforce analytics — see the platform saving 80 hours a month for teams just like yours. Fully tailored to your workflow.