AI in Recruiting7 min read

The Match Quality Gap: Why Good Candidates Get Rejected

Keyword screening rejects qualified people daily and nobody notices, because the rejected never reach a report. Why it happens and how to measure it.

Andreas Amann

Good candidates get rejected at the CV screen because the screen compares vocabulary, not capability. The three it drops most reliably are the person who did the work under a different job title, the person whose CV names the framework but never the language underneath it, and the career changer whose real evidence sits in a projects section no filter weights. Nobody notices, because a rejected candidate never reaches a report anyone reads: time to hire, conversion by stage, offer acceptance and quality of hire are all calculated over the survivors of the filter. The false negative is the one error in recruiting with no feedback loop attached to it, which is why it has outlived 30 years of recruiting software.

I ran a recruitment agency before I built Pickr. The placements I am least proud of are not the ones that went wrong. They are the shortlists I sent where a better candidate was sitting in my own database, already rejected, because my Boolean string and their vocabulary did not overlap. I never found out at the time. Nobody does.

Why false negatives in hiring stay invisible

A bad hire announces itself. Within 90 days somebody is having an uncomfortable conversation, and within 6 months there is a cost attached that finance can see. The error is loud, so teams build process around preventing it.

A wrongly rejected candidate announces nothing. They take a job somewhere else, do it well for 4 years, and no signal ever travels back to the system that screened them out. There is no incident report for a person you did not meet.

Three structural reasons the gap stays hidden:

The denominator is always the survivors. Your funnel report starts at "applications received" and immediately narrows to the ones a human read. Everything upstream of that is a single number with no quality attached to it.

CV-stage rejection is usually unattributed. Later stages produce reasons, however thin. The CV screen produces a status change and nothing else, often without a person involved at all.

The only test that would expose it is one nobody schedules. Finding out means re-reading your own rejections, which feels like auditing yourself for fun on a Thursday afternoon.

So the filter runs for years with no measured error rate. If a supplier shipped you parts on that basis, you would not accept it.

Three ways a keyword filter drops a qualified candidate

They did the work under a different title. Title is the laziest available proxy for capability, and it is the field most filters weight most heavily. An operations manager who owns all the reporting for a business is an analyst. A team lead at a 30-person company is doing what a head of department does at a 300-person one. The work is comparable and the noun is not.

The CV names the framework but not the language. Few people write "JavaScript" when their last 4 years were React and Next.js. Few write "SQL" when they have spent that time building data models in dbt. Practitioners write the layer they actually touch; job ads are written one layer down. The filter reads the gap as absence.

The career changer's evidence is in the projects section. Someone who retrained has thin job history and thick proof: shipped side projects, a public repository, a portfolio, contract work. All of it sits at the bottom of page two in a section that carries almost no weight, under an experience block that says "teacher".

None of these people are edge cases. They are the ordinary strong candidate, because people who are better at doing the work than at describing it are, on average, better at doing the work.

A worked example: the CV that scores zero and meets every requirement

This one is constructed, but it is assembled from CVs I screened for real analytics roles.

The brief says senior data analyst: SQL, Python, Tableau, 3 years in an analyst role.

The CV says operations manager, regional logistics, 4 years. The bullets underneath say: owns weekly demand and depot performance reporting across the network; rebuilt that reporting in dbt and Looker; wrote pandas notebooks for forecast variance; cut reporting turnaround from 3 days to about 4 hours; trained 6 regional leads to self-serve.

The words SQL, Python, Tableau and analyst appear nowhere. Keyword match: zero of four. Rejected in under a second, by nobody.

RequirementWhat the filter readWhat was actually there
SQLNot present4 years of dbt models, which are SQL with a build step
PythonNot presentpandas notebooks in production use
TableauNot presentLooker, a different vendor for the same job, plausibly 2 weeks to transfer
3 years as an analystNot present4 years doing analysis under an operations title

Every requirement was met. Not one was matched.

The second error is quieter and lands on your calendar rather than in your blind spot: the same ranking promotes the CV that contains all four words, whatever sits underneath them. Both errors have one cause, which is that a string comparison is being asked to do a judgement job. Only one of the two is ever measured. This is the concrete version of the broader argument for evaluating evidence of skills rather than keywords.

Why more sourcing does not fix a broken screen

The usual response to a weak shortlist is more supply: another job board, more outbound, a second agency. It is the wrong lever. If your screen drops 20% to 30% of the people who could do the job, doubling the top of the funnel doubles the number of good people you throw away. You widen the pipe and the same grate is still sitting in it.

That is not an argument against sourcing. Sourcing that scores every profile it surfaces earns its keep precisely because the thing downstream of it can tell a strong candidate from a well-worded one. Run in the other order, extra supply is expensive noise, and the cost lands on whoever has to read it.

What accurate candidate matching does differently

Three things, and the third one matters most.

It reads for demonstrated capability, not word overlap. The question is what the person owned, at what scope, for how long, and what happened when it broke. 4 years of dbt is evidence of SQL. Capability gets derived from the work described rather than matched against a string.

It scores adjacent and transferable skills as partial evidence. Keyword logic is binary: the word is there or it is not. Capability is continuous. Looker instead of Tableau is a near miss with a short transfer cost, and it should be scored as a near miss rather than a zero. Score the near misses as zeros and they vanish from the shortlist entirely, which is how a role with plenty of applicants still runs another 6 weeks with nobody to send.

It shows the reasoning so a human can overrule it. A score with no evidence behind it is a keyword filter with extra steps and worse accountability. Every score needs the lines of the CV it was built from sitting next to it, so a recruiter can look at a 62 and say "no, that project is exactly what we need" — and have the disagreement recorded rather than lost.

Pickr is the AI-native recruiting platform built around that idea: it scores every candidate on evidence of skills rather than keyword matches, adjacent and transferable skills included. What happened to the people a company actually hired feeds back into how the next candidates are evaluated, so a strong match for a role stops being a definition someone guessed at before meeting a single candidate. That is what evidence-based candidate matching means in practice.

Now the limitation. An evidence-based score is still a judgement, and it will be wrong about individual people in both directions. It will put candidates in front of a hiring manager who reads them as a stretch, and sometimes the manager is right. On a role you have never hired for, there is no outcome history to learn from yet, so the score is only as good as the brief it was given. What changes is not that the error disappears. It is that the error is visible and arguable, where a keyword rejection is neither.

How to measure your own false negative rate this week

You do not need to buy anything to find out how bad yours is.

  1. Pull 50 CV-stage rejections from last quarter.
  2. Give 20 of them to someone who was not involved, with the skills list and without the rejection reasons.
  3. Count how many they would take to a phone call. 2 out of 20 is a 10% false negative rate at that stage.
  4. Check how many of the 50 carry any recorded reason at all. That number tells you whether you have a filter or a habit.
  5. Search your own database for people you rejected who now hold the same title elsewhere.

About 2 hours of work, and it is the cheapest diagnostic in recruiting. Do it with a role you fill repeatedly, so the result is a rate rather than an anecdote. If you would rather start from your real funnel numbers, the recruiting audit connects to your current ATS read-only and reads your existing hiring history, so you can see where the drop-off actually sits before changing anything.

The takeaway

The candidates worth worrying about are not the ones in your pipeline. They are the ones who were qualified, applied, and were removed by a string comparison before a human was involved, and who will never appear in any number you report to anyone. Measure that one stage once. Whatever you find, you have been paying for it for years without ever seeing the invoice.

Frequently Asked Questions

What is a false negative in recruiting?

A false negative is a candidate who could have done the job well but was rejected before anyone looked properly, almost always at the CV screen. It is the error with no feedback loop attached, because a rejected candidate never appears in your funnel report, your quality-of-hire numbers or your post-mortems. You find out about bad hires within about 90 days. You never find out about good rejections.

Why do keyword filters reject qualified candidates?

Because a keyword filter compares vocabulary, not capability. It drops the person who did the work under a different job title, the person whose CV names a framework without ever naming the language underneath it, and the career changer whose evidence sits in a projects section that carries almost no weight. None of these people lacked the skill. They described it in words the search string did not contain.

How do I measure my own false negative rate?

Pull 50 CV-stage rejections from last quarter and give 20 of them to someone who was not involved, with the skills list and without the original rejection reasons. Count how many they would take to a phone call. 2 out of 20 is a 10% false negative rate at that one stage. The exercise takes about 2 hours, needs no new software, and it is the only number in hiring that nobody reports.

Is a false negative worse than a bad hire?

Not per candidate, but in aggregate it is usually larger and always less visible. A bad hire costs you one salary, some management time and a rehire, and everybody involved knows it happened. A screen rejecting 20% of the people who could do the job costs you a share of every shortlist you have ever sent, spread thinly enough that no single decision looks wrong and no report ever shows it.

What does evidence-based matching do differently from keyword matching?

It reads a CV for demonstrated capability rather than word overlap, so 4 years of building data models counts as evidence of SQL even when those three letters never appear. It scores adjacent and transferable skills as partial evidence instead of treating them as absence, and it shows the lines it based each score on, so a recruiter can disagree with the score and have that disagreement recorded rather than lost in someone's inbox.

Free recruiting audit · 2 minutes

Find out what your hiring process is actually costing you.

Answer eight questions, or connect your current system read-only, and get a report on where your funnel loses candidates and which changes are worth making. No signup, no API key stored, data stays in the EU.

A

Written by Andreas Amann

Founder of Pickr. Former operator at startups in Berlin and Silicon Valley, where he helped scale companies from 40 to 200+ people. Built Pickr after years of using every major ATS as a recruitment agency owner at ScalingPPL.

Read more