Screening

Before any score exists: what a resume pool goes through first

Published Reviewed by Hab Business Solutions

Most conversations about AI hiring start at the score. The work that decides whether the score means anything happens earlier, in the intake stage nobody demos.

Short answer

What happens to resumes before an AI scores them?

Four things, in order. The file is parsed, including PDFs, Word documents, scanned images and phone photos through OCR. The document type is identified and checked to see whether the content is usable at all. Duplicates are detected by content hash before any AI runs, so the same person applying twice does not occupy two shortlist places. Only then is the pool evaluated in one batch pass against the job description. Skip any of those steps and the ranking is confidently wrong.

A pool is messier than a demo dataset

Real applications arrive as exported PDFs, Word files a recruiter reformatted, screenshots, and photographs of printed pages taken on a phone. A tool that only handles clean text quietly drops the rest, and the pool you evaluate is not the pool that applied.

MinMaxHR parses PDFs, Word documents, plain text, HTML, and image and scanned formats through OCR. That is an intake requirement, not a feature. A candidate should not lose a role because of the file format their laptop produced.

Duplicates are removed before the AI runs, not after

The same candidate applies through the careers page, then again through a job board, then a recruiter forwards the same file. Deduplicating after scoring wastes the compute and, worse, gives one person several positions in a ranked list.

Detection happens by content hash at intake, before any model sees the pool. It costs nothing and it keeps the shortlist honest.

Usability is checked and reported, not assumed

Some documents are not resumes. Some are resumes a scanner mangled. The system identifies the document type and checks whether the content can support an evaluation, and says so when it cannot.

This matters for the number recruiters actually care about. A pool where nine percent of files failed to parse is a different pool from one where everything read cleanly, and you should be told which one you have.

Then, and only then, the batch runs

With a clean, deduplicated, readable pool, CandidRanker evaluates every candidate against that specific job description across eight named dimensions in a single pass. Same inputs, same outputs, every run.

Hab builds the same discipline into work outside hiring. In document processing and approvals workflows the intake stage decides the quality of everything downstream, and it is where most implementations are won or lost.

Frequently asked

Does OCR affect the score?

It affects what the score is based on. A poorly scanned document yields less evidence, which is why the match score and the document quality score are reported separately rather than blended into one number.

What happens to a file that cannot be read?

It is surfaced rather than silently discarded. A file that fails intake is a recruiter task, not a rejection.

Want this running on your next role?

Bring one job description and its applicants. You will see the ranked pool and the evidence behind it before the call ends.

No retainers to start · Pilot-first · Human-in-the-loop governance