Analysis of Life Insurance Underwriting Risk Classification Using Ordinal Logistic Regression and XGboost
DOI:
https://doi.org/10.32493/sm.v8i1.59268Keywords:
Life Insurance, Risk Classification, Ordinal Logistic Regression, XGBoost, Quadratic Weighted KappaAbstract
The underwriting process in life insurance is a critical step in determining the risk classification of prospective policyholders, which impacts premium setting and the company’s sustainability. This study aims to analyze underwriting risk classification using the Ordinal Logistic Regression and XGBoost methods. The data used is the Prudential Life Insurance Assessment dataset, consisting of 59,381 training data points and 19,765 test data points with over 120 variables. The research methodology includes data preprocessing, variable selection using XGBoost, and modeling using Ordinal Logistic Regression and XGBoost. Model evaluation was conducted using the accuracy metric and Quadratic Weighted Kappa (QWK). The results indicate that variables related to health conditions and medical history, such as Medical_History, Medical_Keyword, and BMI, have a significant influence on risk classification. The Ordinal Logistic Regression model offers an advantage in interpreting relationships between variables, while XGBoost demonstrates fairly good classification performance with an accuracy of 0.568 and a QWK of 0.540. Overall, this study demonstrates that a combination of statistical and machine learning approaches can support a more effective underwriting process in the life insurance industry.
References
Abidin, S., Anugrawati, S. D., & Nurwahidah., I. (2023). Penentuan Premi Asuransi Kseshatan Dengan Faktor Underwriting Menggunakan Generalized Linear. 6(1), 17–22.
Aulia, D., & Hendri, R. (2020). XGBoost in handling missing values for life insurance risk prediction. SN Applied Sciences, 2(8), 1–10. https://doi.org/10.1007/s42452-020-3128-y
Bronsema, J., Brouwer, S., Boer, M. R. De, & Groothoff, J. W. (2025). The Added Value of Medical Testing in Underwriting Life Insurance. 1–11. https://doi.org/10.1371/journal.pone.0145891
Dong, P. (2024). Automated Machine Learning in Insurance. 1–45.
German, G. G., & Kristiyanti, D. A. (2024). INTELLIGENT SYSTEMS AND APPLICATIONS IN ENGINEERING Machine Learning Algorithm Comparison using Sampling Techniques for Car Insurance Claim Classification. 0–1.
Gunawan, K., & Purnaba, I. G. P. (2022). Penerapan Analisis Regresi Logistik Ordinal pada Asuransi Kredit Perdagangan Domestik. 6(2), 366–380.
Hukum, P., Kejadian, T., Underwriting, R., & Kesehatan, D. A. (2023). Jurnal Surya Kencana Satu : Dinamika Masalah Hukum dan Keadilan. 14(1), 49–60.
Hutagaol, B. J., & Mauritsius, T. (2020). Risk Level Prediction of Life Insurance Applicant using Machine Learning. 9(2).
Kalyanasundaram, S., & Selvakumar, S. (2025). A Hybrid Machine Learning Framework for Personalized Risk Prediction in Health Insurance Underwriting. 13(4), 1–6.
Karuniasari, W., & Prathivi, R. (2025). Komparasi Algoritma Random Forest dan XGBoost dalam Prediksi Premi Asuransi Kesehatan. 23(1), 118–125.
Malali, N. (2025). Artificial Intelligence in Life Insurance Underwriting : A Risk Assessment and Ethical Implications. 12(1), 36–49.
Peddamukkula, P. K. (2024). The Impact of AI-Driven Automated Underwriting on the Life Insurance Industry. 13(9). https://doi.org/10.15680/IJIRSET.2024.1309082
Sahai, R., Al-ataby, A., B, S. A., Jayabalan, M., & Liatsis, P. (n.d.). Insurance Risk Prediction Using Machine. 3, 419–433.
Silva, H. S., Antonio, P., Saraiva, R., De, E. F., Godoy, R. V, Ambrosio, L. A., & Becker, M. (2025). The Impact of Feature Scaling In Machine Learning : Effects on Regression and Classification Tasks. https://doi.org/10.1109/ACCESS.2025.3635541
Sumantiawan, D. I., & Karangturi, U. N. (2024). Metode Analisis Menggunakan Algoritma Random Forest . 1(1), 1–8.
Wang, Y. P. (2021). Predictive machine learning for underwriting life and health insurance. October, 19–22.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Wenny Susanty, Dewi Fortuna Silaban

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
As an Author, you have the right to a variety of uses for your article, including institutions or companies. The author's rights might do without the need for special permission.
Authors who publish in the Jurnal Jurnal Statistika dan Matematika (Statmat) have broad rights to use their works for education and scientific purposes without permission, including:
Used to discuss in a class by the author or the author's body and presentations at meetings or conferences and participant approval;
Used for internal training by the author's company;
Distribution to colleagues for the use of their research;
Used in preparation for further author's works;
Included in a thesis or dissertation;
Partial or extra reuse of articles in other works (with full acknowledgment of the last item);
Prepare derivatives (other than for commercial purposes);
Post voluntarily on a website opened by the author or approve the author for scientific purposes (follow CC with a SA License).
