Begin with the problem, not the AI label
Before comparing vendors, write down the operational problem: too many CVs, inconsistent criteria, slow hiring-manager feedback, poor evidence visibility, or weak auditability. A product should improve that workflow without creating a more serious decision or data-governance problem.
Ask the vendor to demonstrate your real workflow rather than a polished generic example.
Question 1: What can the system decide or change?
Require an exact map of every automated output and every state-changing action. Ask whether the product can automatically:
- reject or progress a candidate;
- rank or shortlist candidates;
- change workflow status;
- send candidate communication;
- learn from recruiter actions.
Request a precise explanation of where AI assistance ends and authenticated human action begins. “Human in the loop” is not enough without product-level detail.
Human action boundary
A useful demonstration should distinguish AI-supported evidence from the authenticated reviewer action that changes status or sends communication.

Question 2: What does the reviewer actually see?
Require a live view that connects each assessment to inspectable candidate evidence. Ask to see:
- approved role requirements;
- candidate evidence mapped to each requirement;
- source locations in the submitted material;
- missing, weak, or contradictory evidence;
- uncertainty and processing failures.
If the primary output is a score or ordered list, ask how a reviewer can challenge the result without rereading every document from scratch.
Evidence visibility
Requirement-level evidence, assessment state, source material, and challenge controls should be visible in the reviewer workflow.

Question 3: How are criteria created and changed?
Criteria should be human-approved, role-specific, and versioned. Determine who can define hard requirements, preferences, and irrelevant information. Ask whether a change can silently alter previously generated candidate assessments.
Good review tooling should make the approved basis of comparison visible.
Question 4: What happens when the system is uncertain?
The system should expose uncertainty and route failures to review instead of manufacturing confidence. Use difficult examples during the demonstration: unusual job titles, equivalent experience from another industry, scanned PDFs, incomplete employment dates, and missing information.
Look for explicit manual-review states. A system that always produces a confident answer is hiding failure rather than managing it.
Failure states and retry controls
Processing queues should expose parsing, failed, and ready states so reviewers can retry, remove, or route material to manual review.

Question 5: What data is processed, where, and for how long?
Require written, contract-aligned answers for the complete data flow. Ask for concrete detail covering:
- storage and processing regions;
- subprocessors and model providers;
- whether customer or candidate data trains models;
- retention after case closure or deletion;
- access controls, export, deletion, and incident support;
- telemetry and logs that may contain personal data.
The EDPB opinion on AI models and personal data illustrates why personal-data questions depend on the facts and safeguards of a deployment. These answers should align with contractual and privacy documentation, not only sales statements.
Question 6: What traceability is available?
Require proportionate traceability that reconstructs the review without retaining everything indefinitely. Ask whether the organisation can determine which criteria and system version were used, when processing occurred, what output was presented, and what human action followed.
Storing every prompt and document forever is not the same as useful traceability.
Question 7: Which regulatory claims are substantiated?
Treat every broad regulatory claim as a request for specific evidence. The EU AI Act lists AI intended to analyse or filter job applications and evaluate candidates among the employment uses that may be high-risk. Article 6(3) contains limited, fact-specific exceptions; systems that profile natural persons remain high-risk. The European Commission's current guidance places the high-risk employment rules from 2 December 2027. Ask the vendor to classify the actual intended use and separate:
- implemented controls and available documentation;
- current legal assumptions;
- planned work;
- completed conformity or registration steps;
- external certification, if any.
Treat broad claims such as “AI Act compliant”, “bias free”, or “legally safe” as claims requiring specific evidence.
Which evidence and warning signs should you compare?
Skilltage guidance: use the same evidence request and warning-sign definition for every vendor:
| Criterion | Evidence to request | Warning sign |
|---|---|---|
| Workflow fit | Live completion of your role-to-review workflow, including handoffs | A polished generic demo replaces your real process |
| Reviewer evidence | Requirement-level source excerpts, uncertainty, and challenge controls | Only a score, rank, or summary is visible |
| Human control | Product demonstration of authenticated progression, rejection, and communication actions | “Human in the loop” is asserted without showing the action boundary |
| Failure handling | A failed parse, missing information, retry state, and manual-review route | Every input produces a confident output or failures disappear |
| Data handling | Contract, privacy documentation, subprocessors, regions, retention, deletion, and access controls | Sales answers conflict with written documentation or omit part of the flow |
| Traceability | Criteria version, processing record, presented output, and subsequent human action | The vendor offers either no reconstruction or indefinite indiscriminate logging |
| Material claims | Named evidence for each regulatory, certification, security, or performance statement | Broad claims are unsupported, ambiguous, or described as future work |
How should you run a realistic vendor test?
Skilltage guidance: run the same numbered procedure with every shortlisted vendor:
- Lock one realistic role. Provide five approved requirements, including one hard requirement and one preference.
- Use one representative candidate document. Include relevant evidence under terminology that differs from the role description.
- Introduce one failure case. Remove a key detail, use an unsupported or damaged copy, or create a controlled source contradiction.
- Observe the complete workflow. Record visible source evidence, uncertainty, status controls, communication boundaries, data handling, and traceability.
- Score only what you observed. Attach the demonstration note or document supporting every rating and mark unproven claims as unverified.
Use a scored comparison only after the questions
Create a short decision table with workflow fit, reviewer evidence, human control, failure handling, data governance, traceability, implementation effort, and vendor credibility. Record evidence from the demonstration beside each rating.
The score should summarize your findings, not replace them.
Vendor comparison after evidence review
A buyer-side scored comparison is useful only when each rating is backed by notes from the demonstration, visible evidence, and documented claim boundaries.

Limitations and decision boundaries
This checklist does not determine whether a tool or deployment is lawful, suitable, or conformant. That depends on product behavior, configuration, organisational use, jurisdiction, and professional advice.
Skilltage recommends evaluating every vendor—including Skilltage—against the same operational and evidence standard.
Use structured screening requirements as the test input, the human-review boundary for AI-assisted screening as the decision-control baseline, and the manufacturing CV-screening workflow for a production-specific vendor test.
Next step
Choose one active role and one representative candidate document. Ask each shortlisted vendor to demonstrate the complete path from approved requirements to human-reviewed candidate action, including a failure case.