Machine learning predicts restaurant food waste from daily data
A Machine Learning Framework for Predicting Restaurant Food Waste to Support Sustainable Food Management
Machine Learning
Summary
Food waste in restaurants is a big problem for the environment and money management. The paper shows how machine learning models can estimate how much food gets wasted each day using info like customer demand, weather, and special events. Since exact waste measurements aren’t available, the researchers carefully created an estimated waste target from operational details. They found that smarter models, like Random Forest, predicted waste better than simple methods. Important factors for prediction included how varied the menu is, the restaurant’s size, and daily activity patterns.
machine learningfood wastesupervised regressionRandom Forestfeature importancetime-series cross-validationmean absolute errorroot mean squared erroroperational dataenvironmental sustainability
Authors
Md Mehedi Hasan Naeem, Md Ashraful Islam, Moumita Barua, Ishtiyak Ahmmad Araf, Md. Arefin Haque Mahir
Abstract
Food waste in the restaurant sector poses a substantial challenge to environmental sustainability and economic efficiency. This paper presents an exploratory machine learning framework for estimating daily restaurant food waste quantities from operational and contextual features. A structured dataset was constructed by integrating restaurant demand records, meteorological data and temporal event indicators, yielding 77,980 records across 27 features. Because large-scale ground-truth food waste measurements are not publicly available, the target variable was derived from operationally justified assumptions, with the complete construction formula and controlled stochastic variability disclosed for full reproducibility. Four supervised regression models, namely Linear Regression, Decision Tree, Random Forest and Gradient Boosting, were evaluated under a chronological 70-30 train-test split that respects the temporal ordering of restaurant operations, augmented by 5-fold time-series cross-validation. All reported metrics are explicitly scoped to performance against the constructed target and do not imply validation against measured food waste. Ensemble methods consistently outperformed linear baselines. Random Forest attained an MAE of 6.19 kg, RMSE of 8.36 kg and $R^2$ of 0.817 on the realistic feature subset following systematic exclusion of algebraically leakage-prone variables. Feature importance analysis identified menu diversity, operational area and temporal activity patterns as the primary predictive drivers. The full dataset, target construction formula, codebase and experimental configurations are publicly released to support reproducibility and future extension to empirically measured waste data.