Business Functions · Use-case guide
AI for HR and Talent: Recruiting, Performance, Retention, and Compliance
AI now touches every stage of the employee lifecycle: screening candidates, drafting performance reviews, predicting attrition, and recommending internal moves. The technology is mature enough to buy; the evidence base and the law are what should shape how you buy it. This guide maps the four HR AI workloads, the evaluation questions that matter for each, and the compliance layer that spans them all.
Share of US firms in which workers use AI in work-related tasks — 23% of firms, 41% employment-weighted, so the typical employee is far more exposed than the typical firm.[^census-wp-26-25]
US Census Bureau, CES Working Paper 26-25
US workers who quit their jobs in June 2026 alone — a 2.0 percent quits rate for the month, and the scale of the problem attrition-prediction vendors sell against.[^bls-jolts-june-2026]
BLS Job Openings and Labor Turnover Survey, June 2026
The date from which high-risk AI systems — a category the EU AI Act applies to "employment, management of workers and access to self-employment" — must meet strict pre-market obligations in the EU.[^ec-ai-act-framework]
European Commission, AI Act regulatory framework
HR is where enterprise AI meets its most heavily regulated subject: decisions about people's livelihoods. Screening, performance, retention, and mobility tools all work well enough to deploy today — and all four sit inside the strictest AI rules now coming into force. So the buying decision is less about model quality than about evidence, auditability, and legal posture. This guide maps the four workloads and the compliance architecture that spans them.
Read those numbers together and the shape of the decision appears. AI exposure in the workforce is already broad, and broadest at large employers, where HR processes are most industrialized.[1] Churn remains a multi-million-worker monthly event, which keeps budget flowing toward hiring speed and retention analytics.[2] And the EU has put employment AI on its high-risk list by name, while US cities and states layer bias-audit and disclosure duties on top of anti-discrimination law.[3] Whatever you deploy here will eventually be inspected — by a regulator, a plaintiff's counsel, or your own works council.
One budget line, four different workloads
"AI for HR" is not one purchase. It is four workloads with different data, failure modes, and legal exposure — and vendors increasingly bundle all four into one talent suite, convenient for procurement and terrible for risk analysis. A screening model that rejects candidates carries disparate-impact exposure. A performance-summary feature creates documents that will be read in litigation. An attrition model makes some of the most sensitive inferences an employer can make about a person. Treating them as one line item means governing them at the level of the least risky one.
| Workload | What the AI does | Decision it touches | Dominant risk |
|---|---|---|---|
| Recruiting and screening | Parses, ranks, and scores applicants; assesses interviews and tests | Who gets considered and rejected | Disparate impact; regulated as high-risk in the EU[^ec-ai-act-framework] |
| Performance management | Drafts review summaries; extracts themes and sentiment from feedback | Ratings, promotion, compensation | Unowned drafts becoming the record; discoverability |
| Retention and mobility | Predicts attrition risk; maps skills; recommends internal moves | Who gets attention, offers, and paths | Sensitive inference; self-fulfilling scores; monitoring creep |
| Compliance overlay | Bias testing, audit trails, human-oversight workflow | Whether the other three are defensible | Doing it as paperwork instead of architecture |
One adjacent workload is deliberately out of scope: HR service delivery — agents that answer policy questions, route cases, and handle requests. That is an operations problem with a different risk profile, covered in /use-cases/employee-service-agents-guide. This guide stays on the workloads where AI output changes someone's employment outcome.
Recruiting and screening: the workload under the microscope
The screening market has settled into three archetypes. Talent-intelligence platforms — Eightfold is the best-known example — score candidate-role fit and rediscover past applicants. Assessment and interview vendors, with HireVue the most prominent, evaluate candidates through structured tests, games, or analysis of recorded interviews. High-volume screening tools — the category Ideal occupied — rank incoming resumes into a prioritized queue for recruiters. The archetypes matter more than the brands, because each concentrates a different risk: matching platforms depend on the historical hiring data they learn from, assessment vendors on the validity of what they claim to measure, and volume-ranking tools on what never surfaces for human review at all.
| Screening archetype | Typical pitch | What to demand before buying |
|---|---|---|
| Talent-intelligence matching | Score candidate-role fit; rediscover past applicants; see the whole talent pool | The prediction target in writing; the training population; how fit scores behave across demographic groups; whether your own hiring history will retrain the model |
| Assessment and interview analysis | Measure capability and potential, not resume pedigree | Validation evidence tying the assessment to job performance for your roles; which candidate signals are used and which are excluded; accommodation paths for disabled candidates |
| High-volume resume ranking | Cut screening time; surface the best of thousands | What happens to the bottom of the ranking; sampling audits of never-reviewed candidates; recruiter override rates and logs |
The reason to insist on written answers is that public evidence about vendor practice is thin. The most systematic academic look at this market — Raghavan, Barocas, Kleinberg, and Levy's audit of algorithmic pre-employment assessment vendors — worked from what vendors had publicly disclosed about "their development and validation procedures," with a focus on "efforts to detect and mitigate bias."[4] The premise is itself the finding: growing interest in algorithmic hiring as a bias remedy, built on very little public knowledge of how the systems work. The authors flag the two design choices deserving the most scrutiny — data collection and the prediction target — and note that de-biasing techniques "interface with, and create challenges for, antidiscrimination law."[4]
"Yet, to date, little is known about how these methods are used in practice."[^arxiv-1906-09208]
The prediction target is where most screening claims quietly fail. A model must be trained to predict something, and in hiring that something is usually a proxy: who past recruiters advanced, who past managers rated highly, who stayed two years. Each label encodes the judgment — and the bias — of the process that produced it. A vendor who says the model "predicts quality of hire" is making a claim about a label; ask to see the label. If the honest answer is "probability that your historical process would have advanced this candidate," the tool accelerates your current process, including whatever is wrong with it.
Predictive of what?
Before any screening pilot, get three things in writing: the exact outcome the model predicts, the population it was trained and validated on, and the subgroup outcome data the vendor will export on demand. A vendor who cannot answer the first question precisely is selling you your own historical hiring behavior at machine speed.
One more screening trap for multinationals: fairness engineering does not travel. Most screening tools are built against US legal concepts — selection-rate comparisons, US demographic categories. Researchers who examined prominent automated hiring systems from a UK legal standpoint concluded that systems developed for the US frame "may obscure rather than improve systemic discrimination in the workplace" when transplanted into a different regime.[5] If you hire in the EU or UK, a vendor's US-style bias report is not evidence of compliance there — the categories, the lawful bases for processing, and the EU AI Act's conformity requirements all differ.[3]
Performance management: drafts, sentiment, and the record
The AI features shipping in performance suites — Lattice and 15Five are the recognizable mid-market names, and the large HCM platforms bundle equivalents — do a narrower, more defensible job than screening tools. They draft review summaries from accumulated feedback and check-ins, extract themes and sentiment from open text, and prompt managers with coaching suggestions. This attacks a real weakness of performance programs: thin, recency-biased reviews written by managers for whom synthesizing a year of notes is tedious. An AI draft that surfaces the March incident alongside the October win is doing genuine work.
The risks are specific. First, summarization flattens: an AI summary reads as authoritative even when it over-generalizes, and a manager under time pressure will sign it unread. Second, sentiment extraction degrades quietly on the inputs real workforces produce — non-English feedback, sarcasm, terse engineering prose — and the failure is invisible because the output still looks fluent. Third, every generated summary is a document: if an AI-drafted review later feeds a termination, the draft, the edits, and the prompt history are part of the story a court reconstructs. None of that argues against the feature; it argues for deciding what is retained, what is a draft, and who owns the final text.
- What languages and input styles has summarization been evaluated on, and can you test it on your own review corpus before rollout?
- Can a generated summary be traced to the source comments it drew from, so a manager can verify rather than trust?
- What is the measured edit rate — how often managers materially change the draft — and is that metric exposed to you?
- Where do drafts, prompts, and intermediate outputs live, under whose retention policy — discoverable records or ephemeral state?
- Does employee feedback flow into model training, and is there a contractual bar on your people data improving the vendor's product for others?
- Can an employee or manager opt out, and does the review process still function when they do?
The summary is a draft, the manager is the author
Set one bright-line policy before enabling AI review features: the manager is the author of record for every performance document, no AI draft becomes final without human edit and sign-off, and edit rates are monitored. A record nobody actually authored is indefensible before an employee, an arbitrator, or a judge.
Retention and internal mobility: prediction is the cheap part
The retention pitch writes itself against the churn numbers: with 3.2 million US workers quitting in June 2026 alone, and total separations at 5.4 million for the month, even a small improvement in regretted attrition is worth real money.[2] Attrition models promise exactly that — a per-employee risk score built from tenure, compensation, engagement signals, and manager relationships, so HR can intervene before the resignation letter.
The honest technical picture is harder. Attrition is a rare event in any scoring window, so models fight base rates, and accuracy claims mean little without the time horizon and the definition of a positive. Historical labels conflate people who quit, who were managed out, and who left for reasons no feature captures. And the most predictive features are often the most legally loaded — leave usage, commute distance, age-correlated tenure — proxies that can reintroduce protected characteristics into a score that drives who gets a retention bonus. The evaluation question is therefore not "how accurate is the model," but "what does the score cause to happen, and is that defensible for the people it selects and skips."
The score also changes the system it observes. A manager told that a report is a flight risk behaves differently — sometimes with a counteroffer, sometimes by routing work away, which produces the predicted resignation. So the product you are actually buying is the intervention workflow: who sees scores, what actions they trigger, and how you measure whether intervened-on employees stay longer than a comparable un-intervened group. A vendor whose demo is a dashboard of red names, with no intervention design and no lift measurement, is selling anxiety, not retention.
Internal mobility is the constructive twin of the same data. Skill-mapping systems infer a live inventory of capabilities from work artifacts, learning records, and role histories, then recommend lateral moves, gigs, and promotion paths — attacking the perceived-stagnation driver of attrition rather than its symptoms. Two things determine whether this works. First, the skills taxonomy: if it is stale or generic, every recommendation is noise, and building it is inseparable from the training investments in /guides/ai-workforce-upskilling-guide. Second, trust: skills inference runs over employee-generated data, so the pipeline powering career recommendations is, architecturally, a monitoring system. Employees can tell a tool that opens doors from one that watches them work — and so can regulators.
Score nothing you will not act on
For each score a people-analytics system produces, name the funded action it triggers and the metric that will show the action worked. A score with no attached intervention is pure liability — knowledge you are accountable for, value you never collect.
The compliance layer spans all four
Employment is where AI bias stops being an ethics-deck topic and becomes litigation exposure. In the US, disparate-impact doctrine treats a facially neutral practice as unlawful discrimination when it disproportionately screens out a protected group and cannot be justified as job-related — no discriminatory intent required. The doctrine long predates AI and applies to a ranking model exactly as it applied to a paper test. US law also brings specific selection-rate tests and thresholds, deliberately not stated here: those belong to employment counsel, applied to your data in your jurisdictions, not hard-coded from a vendor whitepaper into a pipeline. The general practice of fairness measurement and bias testing is covered in /guides/responsible-ai-in-practice.
One framework from that guide is worth restating because HR programs habitually get it wrong. NIST SP 1270 "identifies three categories of bias in AI — systemic, statistical, and human — and describes how and where they contribute to harms."[6] Most HR AI diligence tests only the middle category — the statistical properties of the model. But hiring is the canonical home of the other two: systemic bias arrives through historical hiring data and the institutional practices that generated it; human bias arrives through the recruiters and managers who label outcomes, override scores, and act on flags. A bias program that audits the model alone is auditing a third of the problem.[6]
On top of the federal doctrine, US jurisdictions have begun regulating the tools directly, and the pattern is consistent even where details differ: recurring independent bias audits, publication or disclosure of results, advance notice to candidates that an automated tool will assess them, and consent or alternative-process rights. New York City's audit-and-notice regime for automated employment decision tools is the most prominent; Illinois imposes consent and disclosure duties on AI analysis of video interviews. Track specifics with counsel; for stack planning the lesson is structural: assume any screening tool will eventually need an audit trail, subgroup outcome reporting, and a candidate-notice workflow, and buy only tools that can feed those.
The EU has gone furthest. The AI Act's high-risk list includes "AI tools for employment, management of workers and access to self-employment" — the European Commission's own illustrative example is "CV-sorting software for recruitment."[3] Starting December 2, 2027, high-risk systems face strict obligations before they can be put on the market: risk assessment and mitigation, training-data quality controls aimed at reducing discriminatory outcomes, activity logging for traceability, detailed documentation, clear information to the deployer, appropriate human oversight, and a high level of robustness, cybersecurity, and accuracy.[3] That list is close to the diligence a careful buyer runs anyway — and deployers carry duties too: for EU candidates, "the vendor handles compliance" is not a position, it is a contract clause to negotiate and verify.
Monitoring and termination sit at the sharp end. Productivity monitoring, communication analysis, and sentiment inference are employee-data processing first and AI features second — the lawful-basis, minimization, and retention questions are worked through in /guides/personal-data-protection-ai, and they bind harder for employees because consent is rarely meaningful where a paycheck depends on it. For terminations, the structural rule is jurisdiction-proof: an identifiable human decision-maker of record, informed by AI output but able to explain the decision without it, with the AI's role documented and the audit trail retained. A termination explainable only as "the model said so" is a losing position everywhere.
Where the legal line sits
Nothing here is legal advice, and the omissions are deliberate: selection-rate thresholds, audit methodologies, and jurisdiction-specific tests belong to employment counsel. The stack side owns capability — the logging, subgroup data export, notice workflows, and override paths that make whatever counsel requires implementable. Buy those up front; they cannot be retrofitted onto a closed vendor pipeline.
Honest objections
"Human hiring is worse, and nobody audits it." The strongest argument for the technology, and substantially right: unstructured resume review by tired humans is biased, inconsistent, and leaves no log. An algorithmic process is at least testable — run the same ten thousand candidates through it twice and inspect the outcomes, which no interview panel allows. But the advantage is conditional: auditability arrives only if the vendor exposes logs, scores, and subgroup outcomes — and public disclosure of validation practice in this market has been thin.[4] Consistency without transparency is just bias with better throughput. Buy the auditable version of the technology, and the objection becomes the business case.
"Bias audits are compliance theater." Also partly right. An annual point-in-time audit, on data the audited party assembles, against a test it helps choose, proves little — and check-the-box auditors predictably appear where audits are mandated. But the mandated audit is the floor, not the control. The control is continuous outcome monitoring under your own governance: selection-rate and error-rate tracking by subgroup, reviewed on a cadence, with a named owner and an escalation path when drift appears. Do that, and the annual audit becomes a formality you pass with data you already have.
"Regulation will make employment AI more trouble than it is worth." For some deployments, genuinely yes — that is what a high-risk classification is for, and a marginal screening feature that cannot justify its documentation burden should die in review. But look at what the obligations require: risk assessment, data quality, logging, documentation, human oversight, robustness.[3] That is a description of well-engineered enterprise software. The compliance cost is real but mostly front-loaded into buying better tools and writing better contracts — and deploying decision-making AI on your workforce without those properties was never a defensible position, regulation or not.
The read
Buy the four workloads on four different theories. Screening is a throughput purchase with a hard legal perimeter: take the efficiency, but only from vendors who put the prediction target, validation evidence, and subgroup data export in the contract. Performance AI is a drafting assistant: valuable as long as humans remain the authors of record and edit rates are watched. Retention and mobility are intervention programs wearing a prediction costume: fund the actions and the skills taxonomy, or skip the scores. Compliance is not a workload you buy at all — it is the architecture requirement that decides which vendors are eligible for the other three.
The portable takeaway is a data-boundary decision. Every talent-suite vendor wants the full lifecycle — hiring, performance, engagement, and attrition data — because the workloads genuinely compound on shared data. That compounding is also concentration: one vendor, one model lineage, one failure mode behind hiring, rating, and termination decisions, inside the most regulated AI category there is. Decide how much of the lifecycle one platform may see, demand exit rights over your people data, and keep the audit capability — logs, outcomes, overrides — in systems you control. The vendors will call that friction. Your counsel will call it defensibility.
How to apply this
- Inventory every AI touchpoint in the employee lifecycle — including features inside suites you already own — and classify each against the four workloads.
- Map the inventory to your jurisdictions with counsel: EU AI Act high-risk scope, US bias-audit and notice laws, disparate-impact exposure.
- For every screening tool, get the prediction target, training and validation populations, and bias-testing methodology in writing.
- Contract for audit capability: activity logs, subgroup outcome export, candidate-notice support, and advance notice of model changes.
- Name a human decision-maker of record for every rejection, rating, and termination an AI system touches; document the AI's role in each.
- Stand up continuous selection-outcome monitoring under your own governance, beyond mandated point-in-time audits.
- Pilot performance-summary features with edit-rate measurement and a written author-of-record policy before rollout.
- Fund retention interventions and lift measurement before enabling attrition scores; no score gets an audience without an attached action.
- Route employee-data flows — especially skills and sentiment inference — through the review in /guides/personal-data-protection-ai.
- Revisit the inventory annually; both the vendor features and the law are moving.
Sources
Every quantitative or attributed claim above is linked to a primary source. Last verified at publication.
- [1]The Microstructure of AI Diffusion: Evidence from Firms, Business Functions, and Worker Tasks (CES Working Paper 26-25)US Census Bureau, Center for Economic Studies · · accessed
- [2]Job Openings and Labor Turnover Summary — June 2026US Bureau of Labor Statistics · · accessed
- [3]AI Act — Shaping Europe's digital future (regulatory framework for artificial intelligence)European Commission · accessed
- [4]Mitigating Bias in Algorithmic Hiring: Evaluating Claims and PracticesarXiv (Raghavan, Barocas, Kleinberg, Levy) · · accessed
- [5]What does it mean to 'solve' the problem of discrimination in hiring? Social, technical and legal perspectives from the UK on automated hiring systemsarXiv (Sanchez-Monedero, Dencik, Edwards) · · accessed
- [6]