How to evaluate AI recruitment tools before you buy

Evaluate AI recruitment software by testing it against a realistic hiring workflow: inspect the evidence reviewers receive, verify that humans control decisions, exercise failure paths, review data handling and traceability, and require support for every material vendor claim.

A practical vendor checklist covering workflow fit, visible evidence, human control, traceability, data handling, and credible regulatory claims.

Vendor checklist · 11 min readAuthored by: Skilltage OÜPublished: Updated: Reviewed by: Janus JektvikFacts reviewed:

Begin with the problem, not the AI label

Before comparing vendors, write down the operational problem: too many CVs, inconsistent criteria, slow hiring-manager feedback, poor evidence visibility, or weak auditability. A product should improve that workflow without creating a more serious decision or data-governance problem.

Ask the vendor to demonstrate your real workflow rather than a polished generic example.

Question 1: What can the system decide or change?

Require an exact map of every automated output and every state-changing action. Ask whether the product can automatically:

  • reject or progress a candidate;
  • rank or shortlist candidates;
  • change workflow status;
  • send candidate communication;
  • learn from recruiter actions.

Request a precise explanation of where AI assistance ends and authenticated human action begins. “Human in the loop” is not enough without product-level detail.

Human action boundary

A useful demonstration should distinguish AI-supported evidence from the authenticated reviewer action that changes status or sends communication.

Skilltage cards separating AI-supported evidence from reviewer actions such as marking for interview or requesting clarification
Skilltage app UI capture with fictional candidate data.

Question 2: What does the reviewer actually see?

Require a live view that connects each assessment to inspectable candidate evidence. Ask to see:

  • approved role requirements;
  • candidate evidence mapped to each requirement;
  • source locations in the submitted material;
  • missing, weak, or contradictory evidence;
  • uncertainty and processing failures.

If the primary output is a score or ordered list, ask how a reviewer can challenge the result without rereading every document from scratch.

Evidence visibility

Requirement-level evidence, assessment state, source material, and challenge controls should be visible in the reviewer workflow.

Skilltage reviewer evidence rows with requirements, assessment status, visible source, and inspect or challenge actions
Skilltage app UI capture with fictional candidate data.

Question 3: How are criteria created and changed?

Criteria should be human-approved, role-specific, and versioned. Determine who can define hard requirements, preferences, and irrelevant information. Ask whether a change can silently alter previously generated candidate assessments.

Good review tooling should make the approved basis of comparison visible.

Question 4: What happens when the system is uncertain?

The system should expose uncertainty and route failures to review instead of manufacturing confidence. Use difficult examples during the demonstration: unusual job titles, equivalent experience from another industry, scanned PDFs, incomplete employment dates, and missing information.

Look for explicit manual-review states. A system that always produces a confident answer is hiding failure rather than managing it.

Failure states and retry controls

Processing queues should expose parsing, failed, and ready states so reviewers can retry, remove, or route material to manual review.

Skilltage candidate document states showing analyzed, missing, and failed-analysis source material
Skilltage app UI capture with fictional candidate data.

Question 5: What data is processed, where, and for how long?

Require written, contract-aligned answers for the complete data flow. Ask for concrete detail covering:

  • storage and processing regions;
  • subprocessors and model providers;
  • whether customer or candidate data trains models;
  • retention after case closure or deletion;
  • access controls, export, deletion, and incident support;
  • telemetry and logs that may contain personal data.

The EDPB opinion on AI models and personal data illustrates why personal-data questions depend on the facts and safeguards of a deployment. These answers should align with contractual and privacy documentation, not only sales statements.

Question 6: What traceability is available?

Require proportionate traceability that reconstructs the review without retaining everything indefinitely. Ask whether the organisation can determine which criteria and system version were used, when processing occurred, what output was presented, and what human action followed.

Storing every prompt and document forever is not the same as useful traceability.

Question 7: Which regulatory claims are substantiated?

Treat every broad regulatory claim as a request for specific evidence. The EU AI Act lists AI intended to analyse or filter job applications and evaluate candidates among the employment uses that may be high-risk. Article 6(3) contains limited, fact-specific exceptions; systems that profile natural persons remain high-risk. The European Commission's current guidance places the high-risk employment rules from 2 December 2027. Ask the vendor to classify the actual intended use and separate:

  • implemented controls and available documentation;
  • current legal assumptions;
  • planned work;
  • completed conformity or registration steps;
  • external certification, if any.

Treat broad claims such as “AI Act compliant”, “bias free”, or “legally safe” as claims requiring specific evidence.

Which evidence and warning signs should you compare?

Skilltage guidance: use the same evidence request and warning-sign definition for every vendor:

CriterionEvidence to requestWarning sign
Workflow fitLive completion of your role-to-review workflow, including handoffsA polished generic demo replaces your real process
Reviewer evidenceRequirement-level source excerpts, uncertainty, and challenge controlsOnly a score, rank, or summary is visible
Human controlProduct demonstration of authenticated progression, rejection, and communication actions“Human in the loop” is asserted without showing the action boundary
Failure handlingA failed parse, missing information, retry state, and manual-review routeEvery input produces a confident output or failures disappear
Data handlingContract, privacy documentation, subprocessors, regions, retention, deletion, and access controlsSales answers conflict with written documentation or omit part of the flow
TraceabilityCriteria version, processing record, presented output, and subsequent human actionThe vendor offers either no reconstruction or indefinite indiscriminate logging
Material claimsNamed evidence for each regulatory, certification, security, or performance statementBroad claims are unsupported, ambiguous, or described as future work

How should you run a realistic vendor test?

Skilltage guidance: run the same numbered procedure with every shortlisted vendor:

  1. Lock one realistic role. Provide five approved requirements, including one hard requirement and one preference.
  2. Use one representative candidate document. Include relevant evidence under terminology that differs from the role description.
  3. Introduce one failure case. Remove a key detail, use an unsupported or damaged copy, or create a controlled source contradiction.
  4. Observe the complete workflow. Record visible source evidence, uncertainty, status controls, communication boundaries, data handling, and traceability.
  5. Score only what you observed. Attach the demonstration note or document supporting every rating and mark unproven claims as unverified.

Use a scored comparison only after the questions

Create a short decision table with workflow fit, reviewer evidence, human control, failure handling, data governance, traceability, implementation effort, and vendor credibility. Record evidence from the demonstration beside each rating.

The score should summarize your findings, not replace them.

Vendor comparison after evidence review

A buyer-side scored comparison is useful only when each rating is backed by notes from the demonstration, visible evidence, and documented claim boundaries.

Screenshot of a Skilltage AI recruitment vendor comparison table with evidence visibility, human action boundary, uncertainty handling, data region, traceability, and unsupported claims criteria
Synthetic buyer-side evaluation artifact with fictional vendor data.

Limitations and decision boundaries

This checklist does not determine whether a tool or deployment is lawful, suitable, or conformant. That depends on product behavior, configuration, organisational use, jurisdiction, and professional advice.

Skilltage recommends evaluating every vendor—including Skilltage—against the same operational and evidence standard.

Use structured screening requirements as the test input, the human-review boundary for AI-assisted screening as the decision-control baseline, and the manufacturing CV-screening workflow for a production-specific vendor test.

Next step

Choose one active role and one representative candidate document. Ask each shortlisted vendor to demonstrate the complete path from approved requirements to human-reviewed candidate action, including a failure case.

References and provenance

Related resources

This is practical information, not legal advice. Skilltage supports human-reviewed decision support, not automated hiring decisions.

Want to see the workflow?

Bring your current screening workflow and evaluate Skilltage against the same evidence, control, and data-handling questions.

Use the checklist in a Skilltage walkthrough