Branch-Level Mortgage Credit Risk Classification Using Loan Origination System Data for Managerial Decision Support

Authors

  • Harmansyah Nasution Universitas Ekuitas
  • Patah Herwanto Universitas Ekuitas
  • Agridan Nasywan Universitas EKuitas
  • Mochamad Dexan Oktan Ramdhan Universitas Ekuitas

Keywords:

Credit Risk Prediction, Mortgage Loans (KPR), Loan Origination System (LOS), Machine Learning, Random Forest, Logistic Regression, Decision Support System, Banking Risk Management

Abstract

This study addresses the challenge of early and accurate mortgage credit risk classification using operational data from the Loan Origination System (LOS) of Bank BJB. In this study, LOS consistently refers to the Loan Origination System used to record and manage mortgage loan operational data. The core research problem lies in the limited availability of granular borrower-level data in many banking environments, where credit risk information is commonly managed in aggregated branch-level form. To address this gap, this study develops a branch-level classification framework adapted to panel data by applying ratio-based, growth-based, volatility-based, and lag-based feature engineering techniques to enrich a limited set of original LOS variables. Two supervised classification models, Logistic Regression and Random Forest, were implemented and evaluated using mortgage loan data from 10 branch offices covering the period 2019–2024. The dataset consists of periodic branch-level observations constructed from three original LOS variables: outstanding loan amount, number of active accounts, and non-performing loan value. The binary risk label was defined using the 5% NPL ratio threshold, resulting in an imbalanced but operationally meaningful classification problem between high-risk and low-risk branches. Both models demonstrated strong discriminative performance, with AUC values of 0.91 for Logistic Regression and 0.90 for Random Forest. Logistic Regression produced better high-risk recall and a lower false negative rate, with recall of 0.79 and a false negative rate of 20.59%, while Random Forest achieved higher overall accuracy and high-risk precision but produced a higher false negative rate for high-risk cases. Feature importance analysis confirmed that lagged NPL ratio and lagged NPL value are the most influential predictors of future branch-level credit risk classification, followed by NPL volatility and average loan size. The selected classification model was integrated into a Decision Support System (DSS) dashboard that converts risk probabilities into ranked risk categories and management recommendations. Although the study is limited by its use of aggregated branch-level data from a single bank and one credit product, the findings show that meaningful branch-level risk classification signals can be extracted from routinely available LOS data without requiring individual borrower profiles.

References

[1] X. Zhang and L. Yu, “Consumer credit risk assessment: A review from the state-of-the-art classification algorithms, data traits, and learning methods,” Expert Systems with Applications, vol. 237, p. 121484, Mar. 2024, doi: 10.1016/j.eswa.2023.121484.

[2] E. Dumitrescu, S. Hué, C. Hurlin, and S. Tokpavi, “Machine learning for credit scoring: Improving logistic regression with non-linear decision-tree effects,” European Journal of Operational Research, vol. 297, no. 3, pp. 1178–1192, Mar. 2022, doi: 10.1016/j.ejor.2021.06.053.

[3] S. Hossain et al., “Comparative Analysis of Machine Learning Models for Credit Risk Prediction in Banking Systems.,” tajet, vol. 07, no. 04, pp. 22–33, Apr. 2025, doi: 10.37547/tajet/Volume07Issue04-04.

[4] N. Bussmann, P. Giudici, D. Marinelli, and J. Papenbrock, “Explainable Machine Learning in Credit Risk Management,” Comput Econ, vol. 57, no. 1, pp. 203–216, Jan. 2021, doi: 10.1007/s10614-020-10042-0.

[5] R. B. Tareaf, M. AbuJarour, and F. Zinn, “Revolutionizing Credit Risk: A Deep Dive into Gradient-Boosting Techniques in AI-Driven Finance,” in 2024 International Conference on Information Networking (ICOIN), Ho Chi Minh City, Vietnam: IEEE, Jan. 2024, pp. 322–327. doi: 10.1109/ICOIN59985.2024.10572140.

[6] M. Abdullah, M. A. F. Chowdhury, A. Uddin, and S. Moudud‐Ul‐Huq, “Forecasting nonperforming loans using machine learning,” Journal of Forecasting, vol. 42, no. 7, pp. 1664–1689, Nov. 2023, doi: 10.1002/for.2977.

[7] J. Wang, “Bridging the Gap between Business Practice and Data Science Approaches,” TMLAI, vol. 13, no. 02, pp. 01–09, Apr. 2025, doi: 10.14738/tmlai.1302.18289.

[8] D. Rashad, M. ِAbdel-Salam, and E. Yehia, “The Prediction of Non-Performing Loans Using Artificial Intelligence - A literature Review,” المجلة العلمية للبحوث والدراسات التجارية, vol. 39, no. 3, pp. 1467–1497, Sep. 2025, doi: 10.21608/sjrbs.2025.348316.1855.

[9] M. S. Uddin, G. Chi, M. A. M. Al Janabi, and T. Habib, “Leveraging random forest in micro‐enterprises credit risk modelling for accuracy and interpretability,” Int. J Fin Econ, vol. 27, no. 3, pp. 3713–3729, Jul. 2022, doi: 10.1002/ijfe.2346.

[10] S. Yan, L. Zhu, B. Wang, L. Wang, S. Liu, and H. Chen, “Research on Machine Learning-based Credit Risk Prediction Models and Algorithms,” in 2025 International Conference on Advanced Machine Learning and Data Science (AMLDS), Tokyo, Japan: IEEE, Jul. 2025, pp. 570–575. doi: 10.1109/AMLDS63918.2025.11159424.

[11] M. R. Machado and S. Karray, “Assessing credit risk of commercial customers using hybrid machine learning algorithms,” Expert Systems with Applications, vol. 200, p. 116889, Aug. 2022, doi: 10.1016/j.eswa.2022.116889.

[12] X. Zhu, Q. Chu, X. Song, P. Hu, and L. Peng, “Explainable prediction of loan default based on machine learning models,” Data Science and Management, vol. 6, no. 3, pp. 123–133, Sep. 2023, doi: 10.1016/j.dsm.2023.04.003.

[13] Mb. Salas, P. Lamothe, E. Delgado, A. L. Fernández-Miguélez, and L. Valcarce, “Determinants of Nonperforming Loans: A Global Data Analysis,” Comput Econ, vol. 64, no. 5, pp. 2695–2716, Nov. 2024, doi: 10.1007/s10614-023-10543-8.

[14] “Peraturan BI No. 7/2/PBI/2005,” Database Peraturan | JDIH BPK. Accessed: Jun. 03, 2026. [Online]. Available: http://peraturan.bpk.go.id/Details/137781/peraturan-bi-no-72pbi2005

[15] T. Hua, “Machine Learning Approaches to Creditworthiness Classification,” in Proceedings of the 2025 International Conference on Big Data, Artificial Intelligence and Digital Economy, Kunming China: ACM, Jul. 2025, pp. 127–134. doi: 10.1145/3767052.3767072.

[16] H. Wu, “From Scores to Decisions: Comparing Logistic Regression, Random Forest, and XGBoost for Calibrated, Cost‑Sensitive Credit Default Prediction,” AEMPS, vol. 240, no. 1, pp. 70–79, Nov. 2025, doi: 10.54254/2754-1169/2025.BL29296.

[17] T. Xu, “Comparative Analysis of Machine Learning Algorithms for Consumer Credit Risk Assessment,” TCSISR, vol. 4, pp. 60–67, Jun. 2024, doi: 10.62051/r1m3pg16.

[18] C. Li and J. Zhang, “Research on Credit Risk Prediction Models Based on Machine Learning,” in 2024 6th International Conference on Machine Learning, Big Data and Business Intelligence (MLBDBI), Hangzhou, China: IEEE, Nov. 2024, pp. 1–4. doi: 10.1109/MLBDBI63974.2024.10824021.

[19] S. Lin, “Research on Credit Card Default Prediction Based on Random Forest and Logistic Regression Models,” AEMPS, vol. 185, no. 1, pp. 61–69, Jun. 2025, doi: 10.54254/2754-1169/2025.LH23945.

Published

2026-09-29