ATS demo questions that expose fake AI
Five demo questions that separate real AI matching from keyword filters with a new label: live reasoning, moving scores, training data, audit trails.
On this page
- Why do demos hide this so well?
- Question 1: show me a match, with the reasoning
- Question 2: change one requirement and rescore
- Question 3: what data trained or grounds this?
- Question 4: show me the audit trail
- Question 5: how would a candidate contest a score?
- What does a passing demo look like?
- Where does Recruitifly fit?
The fastest way to expose fake AI in an ATS demo is to ask for a match with its reasoning, then change one requirement live and watch whether the scores move intelligently. Follow with three more questions: what data the model was trained or grounded on, where the audit trail lives, and how a candidate would contest a score. A vendor with real matching answers all five on screen. A keyword filter with a new label stalls within minutes: filters count terms; they cannot explain, re-reason, or account for themselves.
Why do demos hide this so well?
Because demos are rehearsed and matching looks identical from a distance. A list sorted by keyword density and a list sorted by evidence-weighted reasoning look the same on a slide. The difference only appears under interaction: when you ask the system to justify a specific score, or when the inputs change and the output has to change with them.
Vendors have every incentive to blur this line. “AI matching” commands a premium, and the gap between legacy systems and AI-native ones is now the main axis they compete on. Some closed that gap with engineering, others with a rename. Your demo hour is where you find out which. This post is the interrogation script for that hour; it assumes you know roughly how candidate matching works under the hood, because here the goal is purely forensic: you are in a demo and want the truth before the contract.
Question 1: show me a match, with the reasoning
Bring a real job and, if the vendor allows it, two or three anonymised CVs from a recent search. Ask for a match score on one specific candidate, then ask the only question that matters: why this number, for this person?
Real matching produces an explanation grounded in the documents: scored well because of four years on the exact stack you listed, scored down because the leadership requirement is only weakly evidenced. The explanation references the candidate, not a template.
Fake AI produces one of three tells:
- A generic explanation. “Strong skills alignment” that would fit any candidate above 70.
- A keyword receipt. “Matched 7 of 9 terms,” a filter describing itself honestly while the marketing says otherwise.
- No explanation at all. “The algorithm is proprietary.” Proprietary is fine; unexplainable is not. A vendor who cannot explain a score in a demo cannot help you explain it to a candidate or a regulator later. Our ATS selection checklist treats explainability as a hard requirement for exactly this reason.
Question 2: change one requirement and rescore
This is the question that ends most fake demos, so do not let it be skipped. Take the job you brought and change exactly one requirement, live: drop “5 years experience” to 3, swap one must-have skill for an adjacent one, or promote a nice-to-have. Then ask for the same candidates to be rescored. What you are watching for:
| Behaviour after the change | What it tells you |
|---|---|
| Scores shift, and the per-candidate reasoning updates to reference the new requirement | Real matching: the system re-evaluated evidence |
| Nothing changes | The scores were precomputed or cosmetic; the “match” is decoration on a list |
| One candidate swings from 85 to 30 | Term lookup: a single keyword stopped matching, so the candidate vanished |
| The vendor offers to “configure that and follow up” | The model cannot do it interactively, which means it cannot do it at all |
A genuinely useful matcher treats a requirement change the way a recruiter would: it re-weighs the same evidence against a new bar. A filter treats it as a different query against the same index. The behaviours diverge the moment you force the comparison, which is why unrehearsed input is the one thing demo scripts are built to avoid.
Question 3: what data trained or grounds this?
Ask plainly: was a model trained on hiring outcome data, and if so, whose? Is it a foundation model grounded on the job and CV text at request time? Something else?
Several architectures are legitimate; you are testing whether the vendor can answer at all, and whether the answer survives two follow-ups:
- If trained on outcomes: whose outcomes, and what was done about the bias already sitting in them? Models trained on historical hiring decisions inherit historical hiring preferences, the core problem you will eventually face when auditing AI scoring for bias.
- If grounded on your documents at request time: does candidate data leave the platform to reach the model, where does it go, and does it stay in the EU? “We use a leading AI provider” is not an answer to where.
A vendor who hesitates here has usually never been asked, which itself tells you what their other customers did not check.
Question 4: show me the audit trail
Scores influence decisions, and decisions about candidates need a paper trail. Ask to see, on screen, the record for one score: when it was generated, against which version of the job requirements, what evidence it cited, who viewed it, and what action a human took afterwards.
Under the EU AI Act, recruitment AI is high-risk, which brings logging, documentation and human-oversight duties, and the deployer (you) carries them alongside the vendor. If the system cannot reconstruct how a score came to be, you cannot demonstrate oversight, investigate a bias complaint, or answer a candidate’s access request honestly. A vendor selling “AI” without an audit trail is selling you their compliance gap.
Question 5: how would a candidate contest a score?
The quiet killer. GDPR gives candidates the right not to be subject to purely automated decisions with significant effects, and rejection is such a decision. So ask: a candidate believes your system scored them unfairly and writes to us. Walk me through what happens.
A serious vendor has a real answer: the score is advisory, a human made the decision and the trail proves it, and here is how to retrieve the reasoning to review the case. A relabelled filter has never considered the question, because filters were never supposed to be decisions. Watch the salesperson’s face on this one; it is the most honest moment of the demo.
What does a passing demo look like?
All five questions answered on screen, in the session, with your job and your edits. Reasoning per candidate, scores that move sensibly under a live requirement change, a straight answer on data, a visible trail, a thought-through contest process. Plenty of established vendors can show real substance on several of these; the point of the script is not to catch a particular vendor but to make every vendor demonstrate rather than assert. Anyone who answers “let me get back to you” more than once has answered already.
Where does Recruitifly fit?
We publish this script knowing it will be used on us, which is rather the point. Recruitifly’s matching runs through Fly, one assistant across the whole platform: ask it to score and rank candidates against a job and it returns the ranking with its reasoning, and it compares candidates side by side with the evidence per dimension. Every action Fly takes is propose-then-confirm, so a human approves each change before it happens, which is the oversight model the EU AI Act expects from hiring AI, not a limitation we apologise for. Recruitifly is built in the EU, with GDPR tooling as standard and a Compliance Engine add-on for heavier obligations.
We are in private beta. If you want to run these five questions against a live system with your own role and your own edits, talk to us and we will sit the demo, scores moving and all.
Frequently asked questions
How can I tell if an ATS really uses AI matching?
Ask for a live match with the reasoning attached, then change one requirement during the demo and watch whether the scores move in a sensible way. Real matching re-evaluates evidence and explains itself per candidate. A keyword filter relabelled as AI returns generic percentages, cannot say why a specific candidate scored 82, and either ignores your edited requirement or swings wildly because one term changed.
What questions should I ask an ATS vendor about their AI?
Five hold up well: show me a match with the reasoning for one candidate, change this requirement and rescore, what data was the model trained or grounded on, show me the audit trail for a score, and how would a candidate contest a decision influenced by it. Vendors with substance answer all five on screen in minutes. Vendors with a renamed filter ask to follow up by email.
Why does the audit trail matter for AI scoring?
Under the EU AI Act, AI used in hiring is high-risk, which brings documentation and human-oversight duties, and GDPR already restricts purely automated decisions with significant effects. If the system cannot show who saw a score, what evidence produced it, and which human acted on it, you carry that compliance gap, not the vendor. An audit trail is also how you investigate bias before a regulator or candidate asks.
Is keyword matching in an ATS always bad?
No. Keyword and boolean search are honest, useful tools, and recruiters rely on them daily. The problem is mislabelling: a keyword filter sold as AI matching promises judgment it cannot deliver, ranks candidates on term frequency rather than evidence, and silently rejects strong people who used different vocabulary. Buy keyword search as keyword search. Just refuse to pay an AI premium for it.
Recruitifly Editorial
Editorial
Related reading
Adding LLM screening to your ATS without creating duplicate records
Layer LLM screening over your ATS without splitting your candidate data: one source of truth, stable ID sync, scores written back as fields, not copies.
Sourcing with adjacent job titles and skills
Searching one job title misses most of the market. A worked SRE example plus a repeatable method for mapping adjacent titles and skills for any role.
AI Act candidate disclosure: notice template
What to tell applicants when automated screening is used, under the EU AI Act and GDPR Articles 13, 14 and 22, plus a copy-paste disclosure notice template.
Want to see how this looks on your own data?
No hard promises. Just a straight conversation about exports, stages, and your current stack.
Contact us