How AI CV screening works at agencies (and where it quietly fails)

The letters AI and a question mark written in black marker on a whiteboard

Recruitment agencies have a sales pitch right now: AI cuts shortlist time from days to hours. That part is true. What the pitch skips is the part where good candidates get quietly auto-rejected and nobody ever checks.

If you’re hiring through an agency or running one, here is what the system actually does, where it earns its keep, and where the liability is sitting unnoticed in most small agency operations today.

What agencies mean by “AI screening”

When an agency says it uses AI to screen candidates, four things are usually happening in sequence:

  1. Parsing: Pulling structured data out of an unformatted CV, job titles, dates, skills, qualifications, location.
  2. Matching: Scoring that data against a job spec, weighted for must-haves versus nice-to-haves.
  3. Conversational screening: A chatbot asking three to six pre-set questions before a human ever sees the application.
  4. Ranking: Sorting hundreds of applicants into a shortlist a consultant can work through in a morning instead of a week.

Applicant tracking systems have used keyword matching since the 2000s. What changed between 2023 and 2025 is that the matching got smarter, large language models reading CVs in context rather than just counting keywords, and the conversational layer got good enough that candidates often can’t tell they’re talking to software for the first two or three exchanges.

️ The tools doing the work

Most UK and US agencies run one of a handful of platforms, or a combination of them:

  • Bullhorn and Manatal for core ATS and parsing
  • Textkernel for CV parsing specifically (it’s the engine behind many white-label ATS features you’d never know were Textkernel)
  • Paradox’s Olivia for conversational screening
  • HireVue for video interview analysis
  • Eightfold or Beamery for enterprise-side talent matching

Smaller agencies increasingly bolt a custom GPT-based screening layer onto whatever ATS they already have. Building a simple screening chatbot is now a two-week job, not a six-month one.

robot and human hands reaching toward ai text

A real deployment: 40-person Manchester agency, 1,200 placements a year

A logistics and warehouse recruitment agency with roughly 40 staff ran this process manually before automation. Each consultant was reading every incoming CV against six or seven open roles simultaneously. Their own numbers: an experienced consultant could get through 15 to 20 CVs an hour if being careful. A busy Monday with 300 new applications took the best part of two days just to triage.

A parsing and ranking layer went in front of their ATS. The AI extracted relevant experience, forklift licence, shift patterns, distance from site, and scored each CV against live job specs. The routing logic ran like this:

  • Above 70 percent match: straight to a consultant’s queue
  • 40 to 70 percent match: automated screening chat, three questions (right to work, licence expiry, notice period), then routed to a consultant
  • Below 40 percent: auto-archived with a rejection email

After eight weeks, average time to shortlist dropped from 11 days to just under 2. Consultants were spending their time on calls and site visits. That’s the stat every case study stops at.

⚠️ What the case studies don’t tell you

A manual re-read of 200 auto-archived CVs from that first eight weeks turned up 14 candidates who should have made it through.

One had worked as a “materials handler” rather than “warehouse operative.” The parser didn’t map the job title, despite the candidate holding exactly the forklift certification the role required. Another had a two-year gap for parental leave. The AI didn’t penalise the gap directly, but the gap reduced the volume of recent “relevant experience” in the scoring window the system used, which pushed the score below the threshold.

Nobody at the agency had checked before the audit. Not because they didn’t care, but because the entire point of the system was that they no longer had to look. That’s the part nobody selling AI screening tools puts on the homepage: the efficiency gain and the quality risk come from the exact same mechanism. The system is fast precisely because a human isn’t checking its judgement. And a human not checking its judgement is exactly how good candidates disappear without anyone noticing.

How the pipeline runs, step by step

  1. Ingestion: CVs arrive via job boards, agency website forms, or referrals and pull automatically into the ATS, regardless of format (PDF, Word, LinkedIn export).
  2. Parsing: The AI extracts structured fields and normalises them against the agency’s own taxonomy.
  3. Scoring: Each candidate gets a match score against the specific job spec, weighted for must-haves versus nice-to-haves.
  4. Tiering: Candidates split into bands. Top tier goes to a recruiter, mid tier goes to automated screening, bottom tier gets auto-rejected or held in a talent pool.
  5. Conversational screening: Mid-tier candidates get a chatbot or SMS screen asking three to six qualifying questions, usually under five minutes for the candidate.
  6. Human review: A consultant reviews the shortlist, listens to any recorded video answers, and picks who to put forward to the client.
  7. Feedback loop: The best agencies feed placement outcomes back into the scoring model. In practice, most agencies skip this step entirely because it’s slower and less exciting than switching on the next feature.

Where this earns its keep (and where it doesn’t)

For high-volume, low-differentiation roles, warehouse, call centre, hospitality, driving, this works well. The qualifying criteria are objective (licence type, clearance level, shift availability), there are hundreds of applicants per role, and speed matters more than nuance.

For senior, niche, or client-facing roles, it earns its keep far less. A CFO search or a specialist engineering placement usually draws fewer than 30 applicants, and the differentiators are things a parser can’t read: how someone handled a specific board conflict, whether their leadership style suits a founder-led business. Agencies applying the same automated logic across every role level are the ones losing placements to boutique competitors who still read every application by hand for their top-tier roles.

Woman working at desk with coffee

⚖️ The bias and legal exposure most agencies ignore

In the UK, the Equality Act 2010 applies to automated decisions exactly as it applies to human ones. An AI system rejecting candidates on a pattern correlated with age, disability, or maternity leave is still discrimination, even if no human made the individual call. The Information Commissioner’s Office has published guidance specifically on AI and automated decision-making in recruitment, and it puts the burden squarely on the employer or agency, not the software vendor, to prove the system isn’t discriminating.

Most small and mid-size agencies have never run a bias audit on their screening tool. They’ve bought a product, switched it on, and trusted the vendor’s marketing that it’s “fair by design.” That’s a genuine liability sitting quietly in a lot of agency operations right now.

Three habits the agencies getting this right share

  • They keep a human reviewing a random sample of auto-rejected candidates every month, not just the ones who complain.
  • They set automation thresholds differently by role seniority rather than using one blanket cutoff score for every vacancy.
  • They tell candidates plainly that part of the process is automated. Counterintuitively, this tends to increase completion rates on screening chats rather than putting people off. Candidates would rather know than guess.

What job seekers should do right now

If you’re applying through an agency, assume your CV is read by software first. That means mirroring the exact job title language used in the advert (not a close synonym), spelling out qualifications in full rather than abbreviations the parser might not recognise, and keeping your most recent and most relevant experience near the top rather than buried in a strict chronological format.

It’s not gaming the system. It’s writing for the actual reader, which happens to be software before it’s a person.

The honest verdict

AI screening is a real efficiency gain for volume roles, and the time savings are not exaggerated. An 11-day shortlist down to under 2 days is genuinely useful. But the efficiency and the risk are the same mechanism. Agencies that adopt this without building in a manual audit step aren’t being efficient. They’re being unlucky in a way they haven’t discovered yet.

The fix isn’t complicated: sample your reject pile monthly, tier your thresholds by role seniority, and tell candidates what’s happening. Most agencies skip all three because switching on the next feature is more interesting than auditing the last one.

Stay on top of AI & Automation with BizStack Newsletter
BizStack  —  Entrepreneur’s Business Stack
Logo