Machine Learning versus Multiple Linear Regression for Predicting Concrete Compressive Strength: A Group-Aware Validation Study on a Pooled 431-Record Database

Authors

  • Ali Hasan Faculty of Engineering and Petroleum, University of Benghazi, Benghazi, Libya
  • Wasfi Albadry Faculty of Engineering, University of Benghazi, Benghazi, Libya
  • Mohammed Almisteeri Faculty of Engineering and Petroleum, University of Benghazi, Benghazi, Libya
  • Mohammed Zobi College of Science and Technology, Qaminis, Libya

Keywords:

concrete compressive strength, machine learning, random forest, gradient boosting, neural network, multiple linear regression, data leakage, group-aware cross-validation

Abstract

Multiple linear regression (MLR) is widely used for early prediction of concrete compressive strength because it is transparent and accessible to practising engineers, but recent literature suggests that machine-learning (ML) methods achieve higher accuracy on large, heterogeneous datasets. This study re-examines that question on a pooled, multi-source database of 431 concrete mixture records, comparing three MLR specifications against Random Forest, Gradient Boosting, and a neural network (multilayer perceptron). A key methodological issue is addressed explicitly: 96 unique mix designs in the database are each tested at multiple curing ages, so a naive random train/validation split places near-duplicate records of the same mix design on both sides of the split, inflating apparent ML accuracy through data leakage. A group-aware split that holds out entire mix designs was therefore used, yielding 379 calibration records (73 mix designs) and 52 independent validation records (23 previously unseen mix designs). Under this rigorous validation scheme, the linear models generalised poorly to unseen mix designs (validation R² between -0.30 and 0.18, mean absolute percentage error 32-49%), while Gradient Boosting achieved the best performance (validation R² = 0.894, MAPE = 12.4%), followed by Random Forest (R² = 0.848, MAPE = 16.5%); the neural network was intermediate (R² = 0.672, MAPE = 21.5%). The water-cement ratio and curing age were the dominant predictors in both ensemble models, consistent with established concrete technology. The results indicate that, once evaluated under a leakage-free, mix-design-independent validation protocol, ensemble machine-learning methods provide a substantially more reliable prediction of concrete strength than classical multiple linear regression on this type of pooled dataset, while also demonstrating that validation methodology itself materially affects the conclusions drawn from such comparisons.

Dimensions

Published

2026-09-24

How to Cite

Ali Hasan, Wasfi Albadry, Mohammed Almisteeri, & Mohammed Zobi. (2026). Machine Learning versus Multiple Linear Regression for Predicting Concrete Compressive Strength: A Group-Aware Validation Study on a Pooled 431-Record Database. African Journal of Advanced Pure and Applied Sciences, 5(3), 458–462. Retrieved from https://aaasjournals.com/index.php/ajapas/article/view/2190

Issue

Section

Articles