Interview calibration is a session where interviewers score the same real anonymised scorecards.

Interviewers score independently, then reveal their ratings at the same time.

Interviewers then argue about the differences. Quarterly is enough for a stable team.

One measurement tells you it worked: the spread of ratings on the same candidate.

Two interviewers, same candidate, same 45 minutes. One writes “strong hire”. The other writes “no”.

Two honest interviewers disagreeing is a measurement problem. Your loop’s output partly depends on who was free that week.

Interviewer calibration is the fix. It costs about 2 hours a quarter.

What an interview calibration session produces: a shared rubric, a worked example per rating, and the disagreements written down where the panel can see them
An interviewer calibration session, 2 hours

What is interview calibration?

Interview calibration is a session where interviewers score the same real evidence independently.

Compare afterwards. Then argue about the differences until the standard is shared.

Interviewer calibration is neither a lecture on interviewing nor a document.

The arguing is the mechanism. Skip it and you have a written bar nobody applies the same way.

How do you run an interview calibration session?

Pick 3 anonymised scorecards from real past interviews. Strip the names.

Include at least 1 where the panel disagreed at the time.

Everyone scores alone. Reveal all ratings simultaneously. Start with the widest gap.

Separate disagreement about the evidence from disagreement about the standard. Reveal the actual outcome last.

An interview calibration session is 2 hours and 6 to 10 interviewers. 10 silent minutes per scorecard.

Ratings go on the board at the same time. Sequential reveals anchor the room on whoever spoke first.

Afterwards, write down what changed. Usually 2 or 3 sentences about what a given rating requires.

The written change is the actual output, and it is worth more than the meeting.

Four findings from interviewer calibration: people score different things, the bar differs per interviewer, some scorecards hold no evidence, and the senior standard was never stated
The 4 findings interviewer calibration produces

What comes out of it?

Four findings recur in almost every interview calibration.

People are scoring different things. One scores whether the candidate reached the answer, another how they got there.

“Senior” means different things per team. The most senior person’s standard is the one everyone else has been guessing at.

Some interviewers never say no. A round that never rejects carries no information, and its owner should know.

The evidence in scorecards is thin. Half of a typical scorecard is adjectives and conclusions with no observation attached.

What does a usable scorecard line look like?

Three parts. What you asked, what the candidate did, and what you concluded.

A line with all 3 survives a debrief and an interview calibration session. A line with 1 part is an adjective.

The format itself is in write the scorecard before you write the job ad.

Chose a queue over direct calls without prompting, and named the failure it protects against. When I doubled the load he identified the consumer as the bottleneck before I asked.

Could not explain why the retry count was 3. Asked twice, in different ways. Suggests the code was produced without the decision being made.

Compare with what usually gets written:

Good communication skills. Solved the problem. Seems senior. Would work well with the team.

Nothing in the second set can be argued with or checked. It does not belong in a decision.

A usable scorecard line names the observation, the conclusion and the test, against a line that names only impressions
A scorecard line with all 3 parts, and one without

How do you know it worked?

Track the spread of ratings on the same candidate across the panel. Interview calibration worked if it narrows.

Wide disagreement on 30 to 40% of candidates is common before calibration. Disagreement should fall afterwards.

Track it per interviewer too. The person always 1 rating above everyone else is using a different standard.

The remaining disagreements should be about genuinely borderline people. Borderline candidates are what disagreement is for.

Quarterly is enough for a stable team. Monthly for one hiring heavily.

Everyone who interviews attends, the 5-year veterans most of all. New interviewers calibrate before their first real round.

Why does interview calibration get skipped?

Interviewer calibration costs 2 hours of senior engineering time a quarter. It produces no candidates.

Compare it against what 1 bad hire costs. 12 months of salary, and your best engineers’ time.

The arithmetic is in the real cost of a bad engineering hire.

Interviewer calibration costs about 8 hours of senior time a year against 12 months of salary for one bad hire
What interviewer calibration costs against 1 bad hire

Common questions

How often should a panel calibrate?

Quarterly for a stable team, monthly while hiring heavily or onboarding new interviewers. Two sessions a year keeps a panel roughly aligned. Less than that and the standard drifts back within about 2 quarters. Every interviewer recalibrates against whoever they last sat in a debrief with.

What material should you calibrate on?

Three anonymised scorecards from real past interviews, with names and companies stripped. Include at least 1 where the panel disagreed at the time. Invented material produces polite agreement. Real disagreement is the thing you are there to examine. Pick the cases that were genuinely hard.

Why does interviewer calibration get skipped?

Interviewer calibration costs 2 hours of senior engineering time per quarter. It produces no candidates. Compare it against what 1 bad hire costs. 12 months of salary, and your best engineers’ time. And a resignation the team blames on the bar. That comparison is what gets it scheduled.

What if one interviewer is always harsher than the rest?

A harsh interviewer is using a different standard. The session exists to find out which one. Ask what evidence would have moved them to the next rating up. Often the answer reveals they are scoring a level above the role. That is a fixable misunderstanding rather than a personality.

How are interviews scored?

Each interviewer rates the signals their round owns. The scale is agreed by the team in advance. Every rating carries a quote or an action observed. Scores are submitted before the debrief and locked. Without interviewer calibration those scales drift. One candidate then collects a strong hire and a no hire.

The short version

Interviewer calibration is a session where interviewers score real anonymised scorecards independently.

The argument about the differences follows.

Reveal ratings simultaneously. Start with the widest gap.

Separate disagreement about evidence from disagreement about the standard.

Write down what a rating requires. Track the spread. Watch for the interviewer who never says no.

I run this session for teams, either once or across a hiring quarter.

Shadowing and calibration from 5,000 euro, or a workshop sprint at 2,500.

The first call is 20 minutes and free.

The debrief that follows the interviews is its own meeting.

Want this fixed in your loop?

Interview training for hiring managers and their panels: one workshop on your open roles, scorecards written for them, and calibration until the panel agrees. It starts with a free 20-minute call.

No pitch if the loop is already fine.