Best Practices7 min read

Recruiting Analytics Nobody Measures (But Should)

A metric earns dashboard space only if you can name the decision it changes. Which recruiting numbers to keep, which to delete, and what each one is for.

Andreas Amann

The recruiting numbers nobody measures are the ones that arrive late: source quality read at 6 and 12 months instead of at offer, whether a given interviewer's scores predict anything, pass-through rate read one stage at a time, and why offers actually get declined. The numbers everyone measures — averaged time-to-fill, source volume, candidates in pipeline — arrive immediately and change nothing. This article is the keep list, the delete list, and the reason for each.

I ran a recruitment agency before I built Pickr, and the monthly report I sent clients had a dozen charts on it. Going back through a year of those reports, I could not find one instance where a chart, rather than a conversation, made anyone do something differently the following week. That is the only test a metric has to pass, and most of mine failed it.

Why time-to-fill cannot be acted on

Time-to-fill compresses brief quality, market scarcity, client responsiveness, interviewer availability and your own execution into a single figure. When it moves from 38 days to 51 days you cannot tell which of those five changed, so there is nothing to fix. A number you cannot attribute is a number you cannot act on. That it also rewards lowering the bar is the separate argument about hiring outcomes; the point here is narrower. Even when nobody games it, the average tells you nothing.

The average is the second problem. Most role portfolios are bimodal: a cluster that closes in about 3 weeks because the market is liquid, and a tail that takes 3 months or dies quietly. The mean lands in the gap where no actual search lives. Reporting it means reporting a role nobody has ever worked on.

Cycle time does matter, because candidates disappear while you deliberate. But the version you can act on is median time in stage, and the interventions themselves are covered in how to reduce time to hire.

How to measure source quality: 6 and 12 months, not at offer

Every applicant tracking system reports source at the point of hire. Say 40% from one job board, 30% from referrals, the rest split across outbound and agency. That is a count of volume being read as a measure of quality.

The question worth answering is which source produced the hires who are still there and still good a year later. Two figures per channel: the share of hires from that source still employed at 12 months, and the share their manager marks as would-hire-again at 90 days.

The decision this changes is where the sourcing budget goes, and it frequently points the opposite way from the volume chart. A board that supplies 40% of your hires and 15% of your 12-month survivors is not a channel, it is a cost centre with good throughput. Referrals tend to look strong on survival and thin on volume, which is an argument for spending on the referral programme rather than on more inbound.

Be honest about the constraints. The first complete reading takes roughly 18 months, because the earliest cohort has to age. Below about 10 hires per source, the difference between two channels is noise. And the comparison is confounded whenever channels feed different role families, so split by role family before drawing any conclusion. Whatever tool you run it in, this is the figure that deserves the top of the recruiting analytics view, and in most teams it is not on the screen at all.

How to measure interviewer calibration

This is the metric with the highest ratio of value to effort, and I have almost never seen a team run it.

Take each interviewer, every score they have given, and what happened to the candidates who were hired. You want two things. First, spread: does this person use the scale, or does everyone come out at 4 out of 5? Second, prediction: do their above-median and below-median calls line up with the 90-day and 12-month verdict on the people who got the job?

You will find three groups. A few predictors whose scores track outcomes. A larger group producing noise, usually because they rate almost everything the same. And occasionally an inverse case, someone whose enthusiasm turns out to be a mild negative signal. That last one is rare, uncomfortable, and worth knowing.

The decision this changes is who sits on which panel and how much weight a debrief gives their view. It also identifies who needs an hour of coaching rather than a lecture to the whole team. An interviewer who rejects nearly everyone is not being rigorous either; zero variance carries zero information.

The honest caveat is a real one. You only ever observe the outcome for candidates who were actually hired. Everybody an interviewer correctly rejected is invisible to the analysis, so what you are measuring is the top slice of their judgement rather than all of it. It is still the best signal available, but it is not a clean accuracy score and nobody should present it as one.

Two conditions make this possible at all. Scores have to exist, written close enough to the interview to be real rather than reconstructed, which is the practical case for structured scorecards. And every interviewer has to be in the system. The moment interviewer seats cost money, someone caps the licences and half the panel's opinions arrive as chat messages that never reach a record.

Stage pass-through: which hiring stages to delete

Breaking the funnel into stages with a pass-through rate and a median time in stage is standard advice, and reading it for bottlenecks is covered in why your pipeline board misleads you. The use almost nobody makes of it is deletion.

Read pass-through in the direction people forget. A stage that passes more than 90% of candidates is not a good stage, it is a ceremony. If your first screen passes 9 of every 10, it is not a filter, it is a scheduling cost with a calendar invite attached. Give it a real bar or remove it.

The worse pattern is high pass-through plus a long median. Everyone gets through, 11 days later. That is pure loss: no decision made, and enough elapsed time for a competing offer to land.

This is the only recruiting analysis I know that reliably ends with something being removed. A 5-stage process where two of the stages pass 95% is a 3-stage process that has not admitted it yet.

Why candidates decline offers, and how to code the reasons

Almost nobody records why an offer was declined beyond the word "comp", which is what candidates say when they do not want to explain.

Code every decline into 5 or 6 buckets and count them: base compensation, a competing offer that moved faster, the manager or the team, location or remote policy, scope that differed from the description, and a counter-offer from the current employer. Each bucket points somewhere completely different. Compensation means your bands are wrong for this market, and now you have evidence for that conversation instead of an anecdote. Lost to a faster process is the one place cycle time genuinely costs money, and it is a narrow fix. Scope mismatch means the brief was wrong, which is a recruiting failure and an entirely solvable one. Manager or team is not a recruiting problem at all, and pretending otherwise wastes a quarter.

The sample is small. A mid-sized team might see 5 or 6 declines in a year, which is not enough to run a trend line on. It is still worth doing, because 5 declines with a reason attached beat 50 without one.

Which recruiting metrics to keep and which to delete

MetricVerdictThe decision it changes
Time-to-fill, averagedDeleteNone
Median time in stage, per stageKeepWhich stage to own, shorten or remove
Source volume at hireDeleteNone, it is a vanity count
Source survival at 6 and 12 monthsKeepWhere the sourcing budget goes
Cost per hire, blendedDemote to annualBudgeting only
Interviewer score spread and predictionKeepWho sits on which panel
Offer acceptance rate, uncodedDeleteNone without the reason attached
Offer declines by coded reasonKeepPay bands, brief quality, process speed
Candidates in pipelineDeleteNone, it rewards hoarding
Scorecard completion rateKeepWhether any of the above is trustworthy

Two items on that list depend on plumbing rather than on analysis. Pickr is the AI-native recruiting platform that scores candidates on evidence of skills rather than keyword matches, including adjacent and transferable ones, and feeds what happened to the people a company actually hired back into how the next candidates are evaluated. Scores only exist if writing them is easy, so interviews in Pickr are transcribed and the scorecard arrives pre-filled with evidence mapped to each criterion, meaning the interviewer edits a draft instead of facing an empty form days later. The whole panel has to be in the system for calibration to mean anything, so interviewer and hiring-manager seats are free. For agencies, placement tracking is native rather than a permission workaround, so a placed candidate stays on the record instead of disappearing the day the invoice goes out. Candidate data is hosted in Frankfurt, Germany.

None of this is fast, and that is the part worth saying plainly. Interviewer calibration needs roughly 15 to 20 scored interviews per person before it says anything, and source survival needs about 18 months of history before the first cohort is even complete. A dashboard that becomes meaningful in month one is a dashboard measuring the wrong things.

What to put on the dashboard first

A metric earns its place only if you can name the decision it changes. Time-to-fill fails that test and survives on habit. Source survival, interviewer calibration, per-stage pass-through and coded decline reasons each pass it, and all four are harder to collect than the number they replace. Start with the cheapest one: code your next 10 offer declines. It takes about 2 hours a year, and it is the only one of the four that pays off before the quarter is out.

Frequently Asked Questions

Which recruiting metrics should you delete from your dashboard?

Averaged time-to-fill, source volume counted at the point of hire, uncoded offer-acceptance rate and the total number of candidates in pipeline. Each one is either unattributable, gameable, or a count of activity rather than a result, and none of them names a decision you would make differently next week if the number moved 10 points. Deleting them costs nothing, because nobody was acting on them in the first place.

What recruiting metrics actually change decisions?

Source quality measured by how many hires from each channel are still there and rated well at 6 and 12 months rather than counted at offer; interviewer calibration, showing whose scores predict outcomes and whose are noise; pass-through rate paired with median time in stage, read one stage at a time; and offer declines coded into 5 or 6 reason buckets. Each of the four maps to a specific decision, which is the test a metric has to pass to earn space on a dashboard.

How do you measure interviewer calibration?

Take every interviewer, the scores they gave, and what happened to the candidates who were hired. Two numbers per interviewer matter: the spread of their scores, and whether their above-average and below-average calls line up with the 90-day and 12-month verdict. An interviewer who rates almost everyone the same is not being rigorous, they are producing no information. Expect to need roughly 15 to 20 scored interviews per person before the reading means anything, and remember you only observe outcomes for the people who were hired.

How long before source quality data is useful?

Longer than most vendors admit. A 12-month survival figure needs at least 18 months of history before the first cohort is even complete, and you want around 10 hires per source before the gap between two channels is worth acting on. Below that, treat it as an anecdote that suggests where to look rather than a number that settles an argument.

Should you still track cost per hire?

Yes, but annually rather than monthly, and never as a single blended average across every role. Cost per hire is a budgeting number, not an operating one. It moves too slowly to guide a weekly decision, and the blended figure hides the fact that one expensive channel is quietly subsidised by several cheap ones that produce worse hires.

Free recruiting audit · 2 minutes

Find out what your hiring process is actually costing you.

Answer eight questions, or connect your current system read-only, and get a report on where your funnel loses candidates and which changes are worth making. No signup, no API key stored, data stays in the EU.

A

Written by Andreas Amann

Founder of Pickr. Former operator at startups in Berlin and Silicon Valley, where he helped scale companies from 40 to 200+ people. Built Pickr after years of using every major ATS as a recruitment agency owner at ScalingPPL.

Read more