Skip to main content

Performance Management

Performance Calibration: Making Ratings Fair Across Teams

Performance Management9 min read

Performance calibration is a process where managers compare their proposed ratings before finalising them, so that a rating means the same thing regardless of who assigned it. Without it, a generous manager’s “exceeds expectations” is a strict manager’s “meets” — and the data becomes meaningless across teams.

Calibration matters most in multi-site organisations, where the same role is assessed by different managers at different locations with no shared reference point.

Why ratings drift without calibration

Rating drift is not a sign of bad managers. It is a predictable result of asking people to make subjective judgements independently.

Three well-documented patterns produce it:

PatternWhat happensEffect on ratings
Leniency biasManagers avoid conflict by rating generouslyEveryone is above average
Severity biasManagers signal high standards through low ratingsStrong performers under-rated
Central tendencyManagers avoid both extremesEveryone lands in the middle; no information
Reference group effectManagers rate against their own team, not the organisationA weak team’s best is rated like a strong team’s best

The last one is the most damaging in multi-site businesses and the least discussed. A venue manager rating their strongest bartender “exceeds expectations” is comparing them to the other seven bartenders at that venue — not to the 180 bartenders across the group.

When calibration is worth doing

Calibration adds real effort. It is worth it when ratings drive a decision, and largely not worth it when they do not.

SituationCalibrate?
Ratings feed remuneration or bonus decisionsYes — essential
Ratings inform promotion or successionYes
Multiple managers assess the same role across sitesYes
Reviews are purely developmental, no ratingNo — not applicable
A single manager assesses the entire teamNo — nothing to compare against
Fewer than about 20 employees in totalUsually not worth the effort
Expert tip

If you are not sure whether to calibrate, ask what the rating is for. A rating that serves no decision creates anxiety without adding information — and the right fix is to remove the rating, not to calibrate it.

How to run a calibration session

Managers submit provisional ratings. With a specific example supporting each one. No example, no rating.
HR distributes the spread beforehand. Each manager sees the anonymised distribution across all teams before the session, which surfaces the outliers without anyone having to name them.
Group by role, not by manager. Discuss all venue supervisors together, then all educators. This is what makes comparison meaningful.
Discuss the edges first. The highest and lowest ratings. The middle rarely needs debate and consumes the whole session if you let it.
Require evidence, not advocacy. “She’s great” is advocacy. “She trained four new starters and covered two weeks of close shifts” is evidence.
Agree adjustments and record why. If a rating changes, the manager needs to be able to explain the change to the employee.
Finalise before any employee conversation. Never tell someone their rating and then revise it.

A session covering 40–60 employees typically takes 90 minutes to two hours. Longer than that and attention degrades; shorter and the edges do not get proper discussion.

Who should be in the room

All managers assessing the same role level, their shared manager, and an HR facilitator. The facilitator’s job is to keep the discussion on evidence and stop the loudest manager from setting the standard for everyone.

The most common structural error is calibrating within a site rather than across sites. That addresses drift between individual managers at one venue but does nothing about the gap between venues — which in multi-site businesses is the larger problem.

Prosper Performance Management See every rating side by side, across every site Prosper shows review outcomes by manager, team and location together — so the distribution is visible before the calibration session rather than assembled by hand from spreadsheets afterwards. See how Prosper supports review cycles

Should you use forced distribution?

Forced distribution requires a fixed percentage of employees in each rating band — commonly 20% high, 70% middle, 10% low.

The argument for: it eliminates leniency drift entirely and forces genuine differentiation.

The argument against, which is stronger: it assumes performance is normally distributed within every team, which is often false. A high-performing team of six is required to produce a low rating that nobody has earned. It also makes managers compete rather than collaborate, and employees notice quickly that the outcome was structural rather than personal.

Most large organisations that adopted forced ranking in the 2000s have since abandoned it. For frontline businesses with small site-level teams it is particularly ill-suited — a venue with five staff cannot meaningfully be divided into three performance bands.

Common mistake

Applying a distribution target to individual sites rather than to the organisation. A single venue with six people cannot produce a normal distribution. If you use guidance at all, apply it at the level of the whole role population and treat it as a sense-check, not a quota.

Calibration in frontline and multi-site organisations

Three adaptations that make calibration workable when your managers run venues rather than sit in offices.

Calibrate by role across sites, not by site. All duty managers together. All room leaders together. This is the entire point and the thing most commonly done backwards.

Do it remotely and keep it short. Ninety minutes on a video call at a time that suits shift patterns beats a head office day that half the managers cannot attend.

Account for context. A supervisor at a struggling venue with high turnover faces a different job to one at a stable site. Calibration should account for that without lowering the standard — the question is what someone did with the situation they had.

What to do after calibration

Two things, both often skipped.

Managers must own the outcome. A manager who tells an employee “I rated you higher but HR changed it” has destroyed the credibility of the entire process. If a rating changed, the manager needs to understand why well enough to explain it in their own words.

Review the distribution afterwards. If one site consistently rates higher every cycle, that is worth a separate conversation — it may be a genuinely stronger team, or it may be a manager who avoids difficult ratings.

Common mistakes

Calibrating within sites instead of across them. Solves the smaller problem and misses the larger one.

Allowing ratings without evidence. If a manager cannot name a specific example, the rating is an impression.

Spending the session on the middle. The edges are where the disagreement and the information are.

Changing a rating after telling the employee. Finalise everything before any conversation happens.

Calibrating ratings that serve no decision. If the rating does not affect pay, promotion or development, remove the rating rather than calibrating it.

See it in practice Review cycles that produce comparable data Configurable templates by role or site, self and manager assessment in one place, and completion and outcomes visible across every location — which is what makes calibration possible in the first place. Explore Prosper Performance Management

Frequently asked questions

A process where managers compare their proposed performance ratings before finalising them, so that a rating means the same thing regardless of who assigned it.

Because independent subjective judgements drift predictably — through leniency, severity, central tendency, and managers rating against their own team rather than the organisation. Without calibration, ratings are not comparable across teams.

Managers submit provisional ratings with supporting examples, HR distributes the anonymised spread beforehand, the group discusses by role rather than by manager, focuses on the highest and lowest ratings first, and finalises everything before any employee conversation.

Around 90 minutes to two hours for 40 to 60 employees. Longer and attention degrades; shorter and the outlying ratings do not get proper discussion.

Generally no. It assumes performance is normally distributed within every team, which is often false, and it is particularly ill-suited to frontline businesses where site-level teams are small. Most organisations that adopted it have since abandoned it.

All managers assessing the same role level, their shared manager, and an HR facilitator whose job is to keep the discussion on evidence.

By role, across sites. Calibrating within a site addresses drift between managers at one location but does nothing about the gap between locations, which is usually the larger problem.

The discussion happens in the session, not afterwards. Once finalised, the manager must own the outcome — telling an employee that HR changed their rating destroys the credibility of the whole process.

See it in practice

Every rating, side by side, across every site

Review outcomes by manager, team and location in one view — so the distribution is visible before the calibration session, not assembled by hand afterwards.

Ratings you can compareReview outcomes across every manager and site
See how it works