What the research actually shows about employee turnover

Full-paper syntheses for people leaders who want the finding, the context and the limitation in the same place.

Layered research papers, tracing sheets, a pen and a magnifying lens
Full paperMethods, sample and limitations reviewed beyond the abstract.
Primary linksStable arXiv record plus the version of record when available.
Status visiblePreprints and synthetic datasets stay clearly labeled.

Start with the question

Finding, source, and boundary travel together.

Where the numbers came from

Widely quoted figures, followed back to the document that produced them.

Review method and bibliography

Employee turnover research is full of impressive model scores and confident recommendations. The harder question is whether a result came from a real workforce, a public benchmark or a synthetic teaching dataset, and whether it predicts an outcome or actually changes one.

This library reads the full papers and keeps those distinctions visible. The first release covers nine papers across prediction, network effects, job embeddedness, explainability and employee data rights.

How we review the evidence

Each synthesis starts with the complete paper, not its abstract. We record the population, outcome, evaluation design and publication status. We separate predictive associations from causal findings and model what-if scenarios from observed interventions. Preprints are labeled as preprints. Synthetic datasets are never presented as employer evidence.

Four questions worth asking

  • Does the model generalize? A high score inside one company, role or benchmark says little about another workforce without external validation.
  • Was class imbalance handled? Turnover is often the minority outcome, so accuracy can look reassuring while the model misses the people it is meant to identify.
  • Is the recommendation causal? A feature, SHAP value or counterfactual can explain a prediction without proving that changing the feature will change retention.
  • Can employees see and challenge the use? Consent, benefit, access control, audit trails and recourse are part of the system, not compliance details added at the end.

Annotated bibliography

Ribes, Touahri and Perthame (2017)

What it adds: a rare end-to-end case study linking prediction to policy design. Best finding: imbalance handling and targeted action mattered in the model simulation. Boundary: one role in one company, with policy results simulated rather than observed. Read the preprint.

Derrazi and Pourmostafa Roshan Sharami (2025)

What it adds: a clean comparison of tree ensembles, a tabular transformer and hybrid models. Best finding: the simpler tree models won on discrimination, speed and feature-level interpretation in this experiment. Boundary: one public, unusually balanced dataset. Read on arXiv or open the version of record.

Mohiuddin et al. (2023)

What it adds: an accessible SHAP and what-if workflow. Best finding: it demonstrates how an attractive accuracy number can coexist with a weak attrition-class F1 score. Boundary: IBM's synthetic 1,470-row dataset, with scenarios that are not causal interventions. Read on arXiv or open the version of record.

AlKetbi et al. (2025)

What it adds: professional network history from a large public financial register. Best finding: recent peer departures carried predictive information beyond individual and firm variables. Boundary: association in Hong Kong financial services, published as a preprint. Read the preprint.

Balthrop and Jung (2024)

What it adds: communications linked to job spells in a high-turnover industry. Best finding: a reported problem can reinforce fit when employee and employer incentives align around a remedy. Boundary: observational trucking data with self-reporting and unobserved-driver concerns. Read the preprint.

Artelt and Gregoriades (2023)

What it adds: diverse counterfactual explanations for groups of employees. Best finding: counterfactuals can make a model's decision boundary easier to inspect. Boundary: synthetic data, and suggested changes are not causal retention prescriptions. Read the paper.

Zander and Zieglmeier (2023)

What it adds: employee benefit as a first-class design requirement for people analytics. Best finding: benefits were preferred but did not significantly increase consent in the study. Boundary: a small hypothetical user study. Read on arXiv or open the version of record.

Zieglmeier and Pretschner (2023)

What it adds: inverse transparency, where employees can audit how their data is used. Best finding: a working design was technically feasible and well received in a controlled study. Boundary: only 15 participants in a student workplace simulation. Read on arXiv or open the version of record.

Zieglmeier, Gierlich-Joas and Pretschner (2023)

What it adds: a taxonomy of values, benefits and incentives that may shape employee data sharing. Best finding: it broadens the design space beyond one-time consent. Boundary: qualitative expert evidence, not a test of which strategy works. Read on arXiv or open the version of record.