Best Practices8 min read

Outcome-Calibrated Recruiting: How to Measure What Actually Predicts a Good Hire

Time-to-fill measures how fast you closed, not whether you were right. How to build recruiting metrics that link interview scores to real outcomes.

Andreas Amann

Outcome-calibrated recruiting means judging your hiring decisions against what actually happened to the people you hired, then correcting the criteria you score candidates on. The measures that predict a good hire are retention and ramp at 30, 60 and 90 days, hiring manager satisfaction asked identically every time, and the link from those outcomes back to the interview scores. Time-to-fill tells you how fast you closed a role, not whether you were right.

I ran a recruitment agency before I built software, and I reported time-to-fill in every client review I ever did, because it was the only number I could produce from two timestamps. That is why it dominates dashboards: cheap to compute, not informative.

Time-to-fill is the most reported and least useful metric in recruiting

It is a throughput metric wearing the costume of a quality metric.

First, it is gameable in the worst possible direction. The fastest way to improve time-to-fill is to lower the bar: approve the second candidate instead of waiting for the fifth, skip the reference call. Every metric that only measures speed rewards saying yes, and saying yes is what produces expensive hires.

Second, it is silent on the thing you care about. Two roles closed in 28 days each. One person is promoted within a year, the other is gone in four months. The dashboard shows two identical wins.

None of that makes speed worthless. An open role has a real weekly cost, and dragging a process out is its own way of losing good candidates, which is why reducing time-to-hire is still worth doing. Treat speed as a cost metric and stop treating it as a correctness metric. Reported next to 90-day retention it is informative; reported alone it instructs your team to be less careful.

What to measure instead

Retention and ramp at 30, 60 and 90 days

Retention at 90 days is blunt and honest. A departure that early is almost always a hiring failure or an onboarding failure, and both are yours to fix. Ramp is the more useful half: did the person hit the first milestone the role was supposed to produce? That only works if the milestone was written down at offer time. A milestone invented at review time is a description of what happened, not a measurement.

Hiring manager satisfaction at 90 days, asked the same way every time

One question, same wording, same scale, same day count, every hire: "Knowing what you know now, would you make this hire again?" Plus one free-text follow-up: "What surprised you?" That is the whole survey.

The discipline is in the sameness. Change the wording and the comparison across the year is gone. Ask at 30 days and you measure relief that the seat is filled; ask at twelve months and you measure a working relationship. Ninety days is late enough for reality to arrive and early enough that the interview is still the reason.

Quality-of-hire proxies that are honest about being proxies

There is no direct measurement of quality of hire. Every number anyone has shown you is a composite someone chose, which is fine. Hiding the recipe inside a single figure quoted as though it were measured is not.

Publish the components. Mine would be three: twelve-month retention, the would-hire-again answer, the ramp milestone. Version the formula and never compare across versions.

This is the part almost nobody does, and it makes the rest worth collecting. You have a score per criterion for everyone you hired, and now you have their outcomes. Join them.

Vanity metric versus calibrated metric

Commonly reportedWhat it actually measuresCalibrated replacement
Time-to-fillHow fast you closedTime-to-fill segmented by 90-day retention
Candidates screenedSourcing volumeShare of shortlisted candidates the manager rates as genuinely interviewable
Scorecard completion rateComplianceWhether the recorded scores separate strong performers from weak ones
Cost per hireSpendCost per hire still in seat at twelve months
Quality-of-hire scoreA composite someone choseThe same composite, published with its components and its version

The calibration loop, concretely

  1. Fix the criteria and the scale. Five to seven criteria per role family, scored one to four. Four points, not five, because a midpoint is where an interviewer parks an opinion they will not defend. If you are still designing these, interview scorecards done properly covers the structure.
  2. Record the score at the moment of decision. A score reconstructed three days later is a rationalisation of the outcome, not evidence about the candidate.
  3. Record the outcome. This is where teams quietly fail. It exists in HR records, in payroll, in the manager's head, and it is never joined back to the requisition it came from.
  4. Compare the tails. For each criterion, take the hires who scored four and the hires who scored two, and compare their outcomes.
  5. Act on the answer. If the fours and the twos perform the same, the criterion is not predictive. Drop it or redefine it.
  6. Re-run quarterly and version the criteria so you can tell later which definition produced which result.

Step five is the one that changes how you hire. Suppose you have scored "communication skills" on forty hires and the fours are indistinguishable at 90 days from the twos. Three explanations, all actionable: the criterion is not observable in an interview and you are scoring likeability; it has no variance, because almost every candidate scores four and it works as a formality rather than a filter; or it genuinely does not matter for this role and you have been rejecting people over it for years.

The fix is usually to replace the abstraction with something observable. Not "communication skills" but "explained a past failure clearly enough that someone outside their function understood what went wrong". That one you can score consistently.

Be honest about what breaks this

It takes months. My rule of thumb: do not act below about twenty recorded outcomes, and treat twenty to forty as directional. If you hire eight people a year this will take years, so start recording now and rely on structure meanwhile. Asking every candidate the same questions helps whether or not you run the analysis.

Survivorship bias is the real ceiling. You only see outcomes for people you hired. The strong performer you scored a two and rejected never enters the data, so the loop is good at proving a criterion useless and weak at proving one essential.

Nobody wants to record a bad hire. Outcomes have to be captured on a schedule, by someone whose job it is, or they get captured only for the hires that went well. And managers who fought for a candidate rate that candidate higher, which inflates the would-hire-again number.

Calibrate per role family, not per company. Averaging a sales hire and a backend engineer destroys the signal. Agencies get the same split one level up, per client as well as per role. Either way it means fewer data points per bucket, which collides with the first problem. There is no clean way out of that tension.

The minimum viable version

Whether you are an in-house team or an agency desk, if you place fewer than fifty people a year, do this and nothing more: five criteria, a one-to-four scale, one 90-day question with fixed wording, one spreadsheet with a row per hire holding the scores and the outcome. No dashboard, no vendor, no project. The value is in the joining, which is trivial once both halves exist.

First check whether you are recording enough to run it. The free recruiting audit is an eight-question wizard, about two minutes and no signup, plus a deeper read-only connection to your existing system that reports scorecard and interview compliance across your real hiring history. Compliance is the precondition: if half your interviews never produced a recorded score, there is nothing to calibrate against, and that is worth knowing before you design a measurement programme.

Where Pickr fits

Pickr is the AI-native recruiting platform that ties what happened to the people you hired back to the signals you scored them on, so criteria that do not predict performance get corrected instead of repeated. Scores get recorded at the moment of the interview, because scorecards arrive pre-filled with evidence mapped to each criterion rather than as an empty form. Undocumented decisions get challenged while the reasoning is fresh. And outcomes feed back into how the next candidates are evaluated, the longer argument in every hire should make your next hire better. All of it is candidate data, and Pickr hosts it in Germany.

The limit is the same one above: none of this works on day one, because a system that learns from outcomes needs outcomes, and a new account has none. Pickr imports your past hiring records when you switch systems, which shortens the wait if you have history worth importing. Anyone promising calibrated criteria in week one is describing a demo, not a measurement.

The decision is not which dashboard to buy. It is whether you will record, for every hire, both the score you gave and what happened afterwards, then accept an answer you will not like about criteria you have used for years. Most teams will not. The ones that do stop repeating the same hiring mistake at a slightly faster time-to-fill each quarter.

Frequently Asked Questions

What is outcome-calibrated recruiting?

Outcome-calibrated recruiting means judging your hiring decisions against what actually happened to the people you hired, and then correcting the criteria you score candidates on. You record a score per criterion at the moment of the interview, record the outcome at 30, 60 and 90 days, and join the two. Criteria where high scorers and low scorers produce the same outcomes are not predictive and should be dropped or redefined.

What are the best recruiting metrics to track?

Track retention and ramp at 30, 60 and 90 days, hiring manager satisfaction measured at 90 days with identical wording every time, and a quality-of-hire proxy whose components are published rather than hidden inside a single score. Then measure the predictive power of each interview criterion against those outcomes, which is the number almost no hiring team keeps. Speed and volume metrics such as time-to-fill and candidates screened are cost and throughput measures, useful for capacity planning but silent on whether the hire was right.

How do you measure quality of hire?

You do not measure it directly. Every quality-of-hire number is a composite proxy, usually some combination of retention at twelve months, whether the hiring manager would make the same decision again, and whether the new hire hit a ramp milestone that was written down before they started. The honest approach is to publish exactly what your proxy is made of, version it, and never compare numbers across two different versions of the formula.

How long before outcome calibration produces useful answers?

Outcome calibration starts producing useful answers in months, not weeks. My working rule of thumb is that a criterion needs roughly twenty recorded outcomes before a difference means anything, and forty before I would act on it with confidence. A team hiring fifty people a year gets its first real answers after two to three quarters; a team hiring eight a year should start recording now and expect the loop to become useful in a couple of years. The recording has to start before the analysis is possible.

Why is time-to-fill a misleading recruiting metric?

Time-to-fill is misleading because it measures how fast you closed a role, not whether the person you closed on was the right one, and it is trivially improved by lowering the bar. It remains a legitimate cost metric, since an open role has a real weekly price, but it should never be the headline number a hiring team is judged on. Reported alongside 90-day retention it becomes useful; reported alone it rewards saying yes.

Free recruiting audit · 2 minutes

Find out what your hiring process is actually costing you.

Answer eight questions, or connect your current system read-only, and get a report on where your funnel loses candidates and which changes are worth making. No signup, no API key stored, data stays in the EU.

A

Written by Andreas Amann

Founder of Pickr. Former operator at startups in Berlin and Silicon Valley, where he helped scale companies from 40 to 200+ people. Built Pickr after years of using every major ATS as a recruitment agency owner at ScalingPPL.

Read more