Buyer guide

How to evaluate AI hiring tools without taking the demo on trust

Published Reviewed by Hab Business Solutions

A buyer-side checklist for AI recruitment software: what to ask about scoring, explainability, data handling, recruiter control and rollout effort, and which answers should end the conversation.

Short answer

What should you look for when choosing AI recruitment software?

Five things, in this order. Whether scoring is deterministic and reproducible on a repeat run. Whether every score opens into the specific evidence that produced it. Whether a candidate can be rejected without a named human deciding it, which should be impossible. Where the data is stored and processed, and under which law. And what the rollout actually costs beyond the licence, meaning integration, training, and the measurement that proves the thing worked. Run every vendor through the same five, using one of your own job descriptions and one of your own applicant pools.

Bring your own data, always

Vendor sample data is chosen to flatter the product. Bring one real job description your recruiters actually worked from, and one real batch of applications for that role. Anything you learn from a curated demo set is a fact about the demo set.

Better still, bring candidates whose outcomes you already know: people your managers rated highly, people who were rejected, people who were hired and then struggled. Hide the outcomes and see where the tool ranks them. That is the only evaluation in this category worth the calendar time.

Scoring: reproducible, or an opinion

Ask the vendor to run the same batch twice in front of you and compare the order. If a generative model is producing the verdict, the second run can differ, which is acceptable when drafting an email and not acceptable in a decision you may have to explain a year later.

Then ask how the score decomposes. A single mysterious percentage cannot be defended to a hiring manager, a client, or anyone reviewing the process afterwards. Named dimensions can. Ask why the third-ranked candidate is not first, and listen for whether the answer points at specific lines in the resume or restates the score.

  • Same inputs, same order, on a repeat run
  • Score decomposes into named dimensions, not one figure
  • A zero on a dimension can be explained specifically
  • Equivalent terms are matched through aliases, so wording is not a silent penalty
  • Mandatory requirements can be marked mandatory, not merely weighted heavily

Control: who is actually deciding

The most important question in the room is who is recorded as the decision maker when a candidate is rejected. If the answer involves the system, the process is not reviewable and it will not survive procurement or a complaint.

Ask to see the audit trail. Changes to weights, filters, skills and terminology should be traceable with person, time, reason, and the before and after values. Ask what happens when a recruiter disagrees with the ranking, and whether that disagreement is recorded anywhere.

Data: where it lives, and who can reach it

Resumes are personal data. Under India's DPDP Act that brings purpose limitation, notice, retention limits and processor obligations with it, and under GDPR it brings rules on automated decision making. Ask the vendor to draw the data flow rather than describe it, and note any hesitation.

Concrete questions get concrete answers: which region stores and processes the data, whether uploads are malware scanned, how API keys are held, whether support staff can read customer data and for how long, and how tenants are isolated from each other.

  • Region of storage and processing, named
  • Retention period, and who can trigger deletion
  • Support access: scope, and how quickly it expires
  • Tenant isolation between workspaces
  • Sub-processors listed, not summarised

Fit with what you already run

Most teams do not need to replace their applicant tracking system, and the projects that tried usually spent their budget on migration rather than on hiring quality. Ask whether the tool works alongside the ATS, and how the ranked output gets back to the people who work in it.

Ask about the formats too, because this is where day-two friction lives. Can the pool arrive as PDFs, Word files, scans and phone photos. Can the ranking leave as PDF, HTML, CSV or JSON. Is there an API, and on which plan.

Cost: the licence is not the number

Budget in four parts. The licence is what the vendor invoices. Integration is connecting to the systems you already run, which is where most of the delay lives. Training is the recruiters who will use it daily, and skipping it is the most common reason a working system goes unused. Governance and measurement cover the baseline, the reporting and the audit trail, which get cut first and missed most.

Get all four on the table before signing. A cheap licence attached to a six-month integration is not a cheap system.

Answers that should end the conversation

Some responses are disqualifying, and recognising them early saves a quarter.

  • The ranking cannot be explained beyond the score itself
  • The same batch produces a different order on a second run, and this is described as learning
  • Candidates below a threshold are removed automatically, with no human record
  • The data flow cannot be drawn
  • There is no answer to what would make them tell you not to buy
The five questions, and what a good answer sounds like
What to askA weak answerA usable answer
Run the same batch twice, do we get the same order?The model keeps improving, so results evolveYes, scoring is deterministic. Here it is, run twice.
Why is this candidate ranked third?The overall match score is lowerTwo mandatory skills are unevidenced, here are the lines we read
Who rejected this candidate?The system filtered them outA named recruiter, at this time, with this written reason
Where does our data sit?Securely, in the cloudA named region, with retention, scanning and support access stated
What does year one cost in total?The licence priceLicence, integration, training, and measurement, itemised

Frequently asked

What is the difference between AI recruitment software and an ATS?

An applicant tracking system is a system of record: it stores applications and moves candidates through stages. AI recruitment software of this kind is a system of judgement: it scores each applicant against the specific role and orders the pool with the evidence attached. Most teams run both.

How long should an evaluation take?

A working session with your own job description and applicant pool, then a back-test on candidates whose outcomes you already know, then a bounded pilot on one role with a baseline captured first. Weeks, not quarters, and each step produces evidence rather than an impression.

Should we trust a published accuracy percentage?

Treat it as a claim rather than a finding. Accuracy depends on the job family, the pool, the baseline and the definition of a correct shortlist, all of which the vendor chose. Hab Business Solutions does not publish one for MinMaxHR, and publishes the scoring method and the back-test instead.

Can we try it before committing?

MinMaxHR has a free tier covering 100 resumes, 10 job descriptions and 10 hiring decisions with the full ranking engine and no card required, which is enough to run a real back-test on one role.

What if our recruiters resist it?

Usually a signal that the tool is deciding rather than advising. Recruiters adopt ranking when it shortens reading and leaves the judgement with them, and resist it when it overrules them without explanation.

Want this running on your next role?

Bring one job description and its applicants. You will see the ranked pool and the evidence behind it before the call ends.

No retainers to start · Pilot-first · Human-in-the-loop governance