Employee Turnover Prediction and Interpretability Analysis Based on SMOTE and SHAP

Authors

  • Ruoqin Liu School of Economics and Management, Southwest Petroleum University, Chengdu, China Author

DOI:

https://doi.org/10.71222/15snp991

Keywords:

employee turnover, machine learning, data resampling, model interpretability, ensemble learning, human resource management

Abstract

Employee turnover imposes substantial financial and operational burdens on organizations by increasing recruitment and training expenditures while undermining institutional stability and knowledge continuity. In the era of data-driven decision-making, accurately identifying employees at risk of departure has emerged as a critical challenge in human resource management. This study leverages the IBM HR Analytics Employee Attrition dataset to develop and evaluate predictive models for employee turnover. Guided by prior literature and the imperative of model interpretability, ten key features are systematically selected, encompassing demographic, financial, and organizational variables such as age, monthly income, stock option level, years with the current manager, and overtime status. To mitigate the inherent class imbalance in the dataset, the Synthetic Minority Over-sampling Technique (SMOTE) is employed to generate balanced training samples. Three predictive models—Logistic Regression, Random Forest, and XGBoost—are constructed and rigorously compared. Model outputs are further interpreted using SHapley Additive exPlanations (SHAP) to elucidate feature-level contributions. Empirical results demonstrate that XGBoost achieves the highest overall predictive accuracy at 82.31%, followed by Random Forest at 80.73% and Logistic Regression at 74.38%. Regarding the area under the receiver operating characteristic curve (AUC), Logistic Regression attains 0.770, marginally surpassing XGBoost (0.759) and Random Forest (0.756). Notably, Logistic Regression yields the highest recall of 0.732 in identifying departing employees, substantially outperforming Random Forest (0.296) and XGBoost (0.282). SHAP-based interpretability analysis reveals that overtime significantly elevates turnover risk, whereas higher monthly income and stock option levels exert a protective effect. By constructing a parsimonious yet effective feature system, this study successfully balances predictive performance with interpretability, offering actionable insights for evidence-based human resource management.

References

1. İ. T. Baydili and B. Tasci, "Predicting employee attrition: XAI-powered models for managerial decision-making," Systems, vol. 13, no. 7, p. 583, 2025.

2. P. Ajit, "Prediction of employee turnover in organizations using machine learning algorithms," Algorithms, vol. 4, no. 5, p. C5, 2016.

3. D. Avrahami, D. Pessach, G. Singer, and H. Chalutz Ben-Gal, "A human resources analytics and machine-learning examination of turnover: implications for theory and practice," Int. J. Manpower, vol. 43, no. 6, pp. 1405–1424, 2022.

4. L. Breiman, "Random forests," Mach. Learn., vol. 45, no. 1, pp. 5–32, 2001.

5. N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, "SMOTE: synthetic minority over-sampling technique," J. Artif. Intell. Res., vol. 16, pp. 321–357, 2002.

6. C. H. Chang, H. W. Lin, W. H. Tsai, W. L. Wang, and C. T. Huang, "Employee satisfaction, corporate social responsibility and financial performance," Sustainability, vol. 13, no. 18, p. 9996, 2021.

7. R. C. Chen et al., "An end to end of scalable tree boosting system," Sylwan, vol. 165, no. 1, pp. 1–11, 2020.

8. S. P. Das and J. Samal, "People analytics for employee turnover prediction using Extreme Learning Machine," J. Bus. Analytics, vol. 9, no. 2, pp. 108–123, 2026.

9. Y. Fang and Z. Zhang, "Employee Turnover Prediction Model Based on Feature Selection and Imbalanced Data Handling," IEEE Access, 2025.

10. A. Bujold, I. Roberge-Maltais, X. Parent-Rocheleau, J. Boasen, S. Sénécal, and P. M. Léger, "Responsible artificial intelligence in human resources management: a review of the empirical literature," AI and Ethics, vol. 4, no. 4, pp. 1185–1200, 2024.

11. L. Kohoutová et al., "Toward a unified framework for interpreting machine-learning models in neuroimaging," Nat. Protoc., vol. 15, no. 4, pp. 1399–1435, 2020.

12. D. Yin, X. Lu, and T. Mei, "Explainable deep learning for healthcare workforce attrition: a methodological study on the Watson healthcare synthetic benchmark," Front. Public Health, vol. 14, p. 1861450, 2026.

13. A. Dimri, P. Kumar, and D. Garg, "Exploring job embeddedness in the era of artificial intelligence (AI) and machine learning: a bibliometric analysis and future trends," Benchmarking: An Int. J., vol. 33, no. 7, pp. 1991–2018, 2026.

14. S. Chowdhury, S. Joel-Edgar, P. K. Dey, S. Bhattacharya, and A. Kharlamov, "Embedding transparency in artificial intelligence machine learning models: managerial implications on predicting and explaining employee turnover," Int. J. Hum. Resour. Manag., vol. 34, no. 14, pp. 2732–2764, 2023.

15. W. Wu and S. Fukui, "Using human resources data to predict turnover of community mental health employees: Prediction and interpretation of machine learning methods," Int. J. Ment. Health Nurs., vol. 33, no. 6, pp. 2180–2192, 2024.

Downloads

Published

12 July 2026

Issue

Section

Article

How to Cite

Liu, R. (2026). Employee Turnover Prediction and Interpretability Analysis Based on SMOTE and SHAP. Economics and Management Innovation, 3(3), 8-18. https://doi.org/10.71222/15snp991