What does better CV screening mean?
Better screening is not one number. A useful review asks whether candidates progress through a controlled process, whether reviewers can complete the work, whether decisions use visible evidence consistently, and whether failures remain actionable. A faster screen may be worse if it hides missing evidence, parsing failures, or human reversals.
Acas guidance recommends applying a consistent process against the job description and person specification, while CIPD guidance describes structured selection methods and their evidence base. Neither source supplies a universal scorecard for every role. The team still has to define what it is measuring and retain the denominator.
The EU AI Act provides context on transparency, human oversight, accuracy, robustness, and monitoring for covered systems. The ICO recruitment audit work records concerns found in particular audited recruitment tools and practices. These sources do not prescribe the worksheet below, prove that a process is compliant, or establish that the measures are fair or valid.
Set the baseline before comparing periods
Define one role, one approved criterion version, one applicant source scope, and one screening-decision rule. Skilltage guidance: use the first 25 completed reviews or four weeks, whichever is later, as the baseline. Review every four weeks after that. Independently double-review 20% of the cohort, with 10 to 30 cases per window. If fewer than 10 cases are available, show the raw cases and do not claim a trend.
Do not combine materially different roles, criterion versions, languages, or applicant sources into one rate. If one changes, start a new stratum or state the break in the comparison.
Outcome measure
Screening-to-interview progression
- Numerator: candidates in the cohort for whom an interview was scheduled after screening.
- Denominator: candidates in the same cohort who reached a completed human screening decision.
- Cohort and window: one role, criterion version, applicant-source scope, and four-week window.
- Caveat: progression reflects applicant mix, hiring capacity, and threshold choices; it is not proof of screening quality or a better hire.
- Owner: recruitment lead.
- Review threshold: investigate an absolute change of 10 percentage points from baseline or any material cohort change.
- Action: inspect the criteria, source mix, capacity, and decision records; do not automatically raise or lower the progression target.
- Next review: the next four-week scorecard review, recorded in the worksheet.
Process measures
Screening throughput
- Numerator: completed human screening decisions.
- Denominator: reviewer hours spent on those decisions; also retain the raw decision count and median elapsed time separately.
- Cohort and window: the same role cohort and four-week window as the outcome measure.
- Caveat: a higher rate can reflect less scrutiny, simpler candidates, or better preparation; it is not a quality result by itself.
- Owner: screening operations owner.
- Review threshold: investigate a change of 20% or more from baseline in either direction.
- Action: sample decision records and failure cases before changing staffing or workflow.
- Next review: next four-week review.
Clarification frequency
- Numerator: reviewed candidates sent at least one role-related clarification request.
- Denominator: all candidates with a completed screen in the cohort.
- Cohort and window: the same role and four-week window; count a candidate once even if several questions were sent.
- Caveat: a clarification can be healthy evidence gathering; fewer requests are not automatically better.
- Owner: lead reviewer.
- Review threshold: investigate an absolute change of 10 percentage points or any repeated question affecting at least three candidates.
- Action: decide whether the criterion, application question, job ad, or reviewer guidance needs revision.
- Next review: next four-week review, or earlier after a criterion revision.
Quality measures
Reviewer agreement
- Numerator: independently double-reviewed cases receiving the same initial human review status.
- Denominator: all cases independently double-reviewed before discussion in the window.
- Cohort and window: the stratified 20% sample for the role, using the same criterion version; record evidence-state agreement separately.
- Caveat: agreement does not prove that a criterion is relevant, fair, lawful, or correct.
- Owner: calibration owner.
- Review threshold: any disagreement on a hard requirement, or a fall of 10 percentage points from baseline.
- Action: compare source evidence and wording, calibrate reviewers, and version any approved criterion change.
- Next review: immediately after calibration and again in the next four-week sample.
Evidence coverage
- Numerator: sampled criterion-candidate checks containing traceable source evidence or an explicit missing, uncertain, contradictory, or failed-analysis state.
- Denominator: all in-scope criterion-candidate checks in the sample.
- Cohort and window: all criteria for the double-reviewed sample in the four-week window.
- Caveat: a recorded state improves traceability; it does not prove the evidence is sufficient or the criterion is valid.
- Owner: evidence-review owner.
- Review threshold: any decisive criterion without a traceable source/state, or total coverage below 95%.
- Action: recover the source, correct the state, or pause the decision before relying on the result.
- Next review: after remediation and in the next scheduled sample.
Calibration drift
- Numerator: audited decisions that do not follow the locked criterion wording and evidence interpretation.
- Denominator: all decisions audited against that exact criterion version.
- Cohort and window: the double-reviewed sample plus any escalated case in the four-week window.
- Caveat: a changed interpretation may reveal a genuine role change; do not silently rewrite the baseline around a preferred candidate.
- Owner: hiring manager and recruitment lead jointly.
- Review threshold: any threshold movement during live review, or drift in more than 10% of audited cases.
- Action: pause, document the reason, approve and version the criterion if warranted, then restart comparison under the new version.
- Next review: before further live screening under the changed criterion and at the next four-week review.
Safety and control measures
Parsing failure frequency
- Numerator: submitted documents with a failed parse or a material omission found during manual source recovery.
- Denominator: all documents entering the defined screening cohort; show blocked, queued, failed, and ready states separately.
- Cohort and window: documents for the role received in the four-week window.
- Caveat: document layout and language affect the rate; a failure is unavailable input, not negative candidate evidence.
- Owner: intake owner.
- Review threshold: any failure that influences a decision without source recovery, or a failure rate at least twice the baseline.
- Action: pause the affected decision, inspect the original, record recovery, and escalate repeated failure patterns.
- Next review: after recovery and at the next four-week review.
Reopen or reversal frequency
- Numerator: candidates whose first recorded closed review state was later reopened or changed.
- Denominator: candidates in the cohort with an initial shortlisted or rejected state.
- Cohort and window: decisions first closed during the four-week period, followed through the next review date.
- Caveat: a reversal can be appropriate when new evidence arrives; the reason matters more than minimizing the rate.
- Owner: recruitment lead.
- Review threshold: any reversal without a named reason and owner, or a rate above 10%.
- Action: classify the cause as new evidence, criterion drift, source failure, reviewer error, or process change and address that cause.
- Next review: at the case closeout and next four-week scorecard review.
Human override frequency
- Numerator: sampled cases where a reviewer declined or changed a visible tool suggestion.
- Denominator: sampled cases where such a suggestion was actually present; record “not captured” when it was not.
- Cohort and window: the double-reviewed sample in the four-week window.
- Caveat: a high or low override rate alone says nothing about quality, and Skilltage does not currently provide this scorecard as an analytics dashboard.
- Owner: human-oversight owner.
- Review threshold: any override without a recorded evidence-based reason, or an absolute change of 10 percentage points from baseline.
- Action: inspect the source, suggestion, reviewer reason, and criterion; correct the workflow rather than targeting an override rate.
- Next review: after any unsupported override and in the next four-week review.
Baseline and periodic-review worksheet
Use one row per measure. Never enter a percentage without its numerator and denominator.
| Field | What to record |
|---|---|
| Role and cohort | Role, criterion version, applicant sources, languages, inclusion and exclusion rules |
| Window | Baseline or periodic-review start and end dates |
| Measure | Exact measure name and dimension: outcome, process, quality, or safety |
| Numerator and denominator | Raw counts, rate if appropriate, and treatment of unavailable data |
| Comparison | Baseline value, current value, and any break in cohort comparability |
| Caveat | Small sample, mix change, source failure, capacity change, or other limitation |
| Owner | Named person responsible for review and follow-up |
| Threshold | The pre-recorded trigger used for this measure |
| Action and decision | Investigate, recover, calibrate, revise, pause, or retain; include the reason |
| Next review | Date and the criterion/cohort version to use |
Worked example
For a four-week baseline of 40 completed screens, the team records 10 interviews scheduled (10/40), 8 candidates receiving clarification (8/40), 10 matching initial statuses across 12 double-reviews (10/12), and 188 explicit evidence states across 200 criterion-candidate checks (188/200). It also records 3 documents with parsing failure or material omission out of 40 received (3/40), 2 reopened decisions among 34 initially closed states (2/34), and 4 overrides among 15 sampled cases where a suggestion was visible (4/15).
These figures are a baseline, not a verdict. At the next review, the owner compares like with like, applies the recorded thresholds, explains denominator changes, and records the action. A total or average score must not hide a zero on source recovery or human control.
What can Skilltage make visible today?
Skilltage guidance: map the worksheet to the real workflow without claiming an implemented analytics product. Skilltage separates define, intake, review, and interactions; exposes document processing states; presents requirement-level evidence and gaps; keeps candidate review status human-selected; and records bounded human and AI trace events. Existing usage accounting concerns product consumption such as model tokens, not the quality scorecard above. The team must calculate, interpret, and own these measures outside the product unless a separate analytics capability is implemented and validated.
Limitations and next step
This scorecard cannot prove fairness, legal compliance, validity, accuracy, or better hires. Small samples fluctuate, agreement can preserve a poor rule, progression changes with applicant mix, and recorded evidence can still be misinterpreted. Protected-group impact and jurisdiction-specific obligations require appropriately governed analysis and qualified advice.
Start by setting human-control boundaries for AI-assisted CV screening. If this is a product trial, define the measures in a reversible AI-assisted screening pilot. When agreement drifts, recalibrate the screening criteria; when source extraction fails, use the manual CV parsing fallback.