Guide
How to evaluate the accuracy and ROI of AI resume screening
Last updated Reviewed by Hab Business Solutions
Every vendor in this category quotes an accuracy percentage. Almost none of them will tell you what it was measured against. Here is how to run the test yourself, and how to work out whether the result is worth paying for.
Short answer
How do you evaluate the accuracy of an AI resume screening tool?
Why a published accuracy percentage means almost nothing
Accuracy is not a property of a screening tool. It is a property of a tool, a job family, a candidate pool, and a definition of what counts as a correct answer. Change any one of those and the number moves. A figure measured on software engineering applicants in one market tells you very little about maintenance technicians in another.
There is a second problem, which is who did the measuring. When a vendor publishes its own accuracy figure, the vendor chose the test set, the baseline, and the success criterion. That is not fraud, but it is not independent evidence either, and a careful buyer should treat it as a claim rather than a finding.
Hab Business Solutions does not publish an accuracy percentage for MinMaxHR. We have not run an independent benchmark that would justify one. What we publish instead is the scoring method, the eight dimensions behind every score, the evidence under each dimension, and the test below, which you can run on us and on anyone we are competing with.
- Ask which job family the figure was measured on
- Ask what the comparison baseline was: a recruiter, a keyword search, or nothing
- Ask who defined a correct shortlist, and whether they saw the tool's output first
- Ask for the size of the test pool, and whether it was a real pool or a constructed one
- Ask whether the study was run by anyone other than the vendor
The test that actually settles it
This is a back-test, and it is the only vendor evaluation we think is worth your time. It uses candidates you already have, so it costs nothing but the effort of assembling the file, and it produces a result nobody can argue with because the outcomes were known before the tool saw the data.
Assemble a set of past applicants for one role you understand well. Include people your hiring managers rated highly, people your recruiters shortlisted and the manager rejected, people who interviewed and failed, people you hired who worked out, and people you hired who did not. Strip the outcomes. Give each vendor the same job description your recruiters actually worked from, and ask for a ranked list.
Then compare the top of each ranked list against the candidates you already know were strong. Run the same test for a second and third role, because a tool that is good at one job family may be poor at another, and one role is a story rather than a result.
- How many of your known-strong candidates appear in the top ranks
- How many candidates your managers rejected appear there too
- How many strong candidates the tool buried, which is the failure nobody measures
- Whether the evidence shown for each ranking survives a hiring manager reading it
- Whether the same inputs produce the same order on a second run
What to check beyond the ranking
A ranking you cannot interrogate is a ranking you cannot defend, so the explanation matters as much as the order. Open the top ten and the bottom ten, and read what the system says it found. If the reasons are generic, the score is generic. If the reasons point at specific lines in the resume, you have something a recruiter can take into a hiring manager conversation.
Check the failure behaviour too, because that is where the real risk sits. Ask what happens to a resume the system cannot read properly, or a job description that is vague. A system that routes poor and ambiguous input to human review is behaving correctly. A system that scores it confidently anyway is producing numbers that mean nothing, and you will not be able to tell which ones.
- Whether any candidate can be rejected without a person making that decision
- Whether employment gaps reduce a score automatically, or are surfaced for review
- Whether the same job description produces the same ranking on a repeat run
- What happens to unreadable, scanned, or low-quality documents
- Whether criteria can be marked mandatory rather than merely weighted
The four cost lines nobody quotes you
Vendors price the licence. The licence is rarely the largest number. Budget in four parts, and get the other three on the table before you sign, because they are what determines whether the project delivers anything at all.
The licence covers the software. Integration covers connecting it to the systems you already run, which is where most of the delay lives. Training covers the people who will operate it daily, and skipping it is the most common reason a working system goes unused. Governance and measurement cover the baseline, the reporting, and the audit trail that let you prove the thing worked, which is the part that gets cut first and missed most.
- Licence: what the vendor invoices
- Integration: connecting to your ATS and the systems around it
- Training: the recruiters who will use it every day, in their own language
- Governance and measurement: baseline capture, reporting, and the audit trail
Working out ROI without inventing a number
ROI on screening is measurable, but only if you measure the process before you change it. Capture the baseline first: how many hours screening consumes on a given requisition, how long it takes to produce a shortlist, how much of the pool is genuinely read, and how often a shortlist survives the hiring manager. Those four figures are your comparison set.
Then run a bounded pilot on one process, capture the same four figures on the same definitions, and compare. If the shortlist arrives faster but the hiring manager rejects more of it, that is not a gain, and the acceptance figure is what stops you mistaking one for the other. Speed with quality held constant is the only version of this that is worth anything.
This is the sequence Hab Business Solutions uses on every engagement: baseline, pilot, measure, compare, decision. It is slower than quoting you a percentage, and it produces a number you own rather than one you have to take on trust. MinMaxHR and CandidRanker are deployed inside that sequence, never ahead of it.
| What you are shown | What it proves | What to ask next |
|---|---|---|
| A published accuracy percentage | That the vendor ran a test they designed | Which job family, which baseline, and who defined a correct shortlist |
| Resumes processed, or hours saved | That the tool operated at volume | Whether shortlist quality was measured at the same time |
| A named customer result | That it worked for someone else's process | Whether that customer will take a call, and how close their roles are to yours |
| A back-test on your own candidates | How the tool ranks people whose outcomes you already know | Nothing. This is the evidence. Run it for a second role. |
Frequently asked
What accuracy percentage does MinMaxHR publish?
None. We have not run an independent benchmark that would justify one, and a vendor-reported figure is a claim rather than a finding. We publish the scoring method, the eight named dimensions, and the evidence behind each score, and we will run a back-test on your own historical candidates so the number you get is one you measured.
How many candidates do I need for a useful back-test?
Enough that the result is not luck, and drawn from a role you understand well. Include strong candidates, rejected candidates, and hires that did not work out, because a test made only of good candidates cannot tell you anything about what the tool buries.
Should I compare AI screening against my recruiters or against nothing?
Against your recruiters, on the same requisition. Comparing against nothing tells you the tool is faster than not screening, which was never in doubt. Comparing against your current process tells you whether it is better, which is the question you are actually paying to answer.
How long before ROI on AI resume screening shows up?
As soon as the pilot has run against a baseline, which is why the baseline has to be captured first. Without it there is no before to compare the after against, and every figure produced afterwards is an assertion.
Does faster screening mean worse shortlists?
It can, which is exactly why shortlist acceptance is tracked alongside speed throughout a pilot. If acceptance falls while time-to-shortlist improves, the speed gain has been paid for out of quality and the pilot has told you something useful.
Want this running on your next role?
Bring one job description and its applicants. You will see the ranked pool and the evidence behind it before the call ends.
No retainers to start · Pilot-first · Human-in-the-loop governance
