Machine Learning-Based Runtime Prediction and Energy Optimization for HPC Job Scheduling Using the NREL Eagle Supercomputer Dataset
Abstract
Accurate runtime prediction is essential for efficient HPC job scheduling, yet users chronically overesti- mate their jobs’ requirements. We analyze 7.3 million completed jobs from the NREL Eagle supercomputer and find that the problem is far worse than previously reported: median time-limit utilization is just 6.7%, with users consuming a median of 10.6 minutes against 4-hour requests. We train ensemble models (Ran- dom Forest, Gradient Boosting, HistGradientBoosting) and an MLP neural network enriched with user behavioral features—historical runtimes, utilization habits, submission frequency—and temporal context. Our central finding concerns evaluation methodology: under the random train/test splits common in prior work, Random Forest reaches R2 = 0.602 (MAE = 0.99 h), but under a realistic temporal split (train <2022, test ≥2022) the same models achieve only R2 = 0.172–0.233—a 71% loss attributable to concept drift, unreported in prior Eagle studies. This gap indicates that random-split results, including our own, overstate deployable accuracy, and that temporal evaluation should be the standard for credible claims. Ablation under random split confirms that user history (∆R2 = −0.078) outweighs the time-limit signal (∆R2 = −0.042), and HistGradientBoosting proves most drift-resilient (R2 = 0.233). Crucially, energy optimization remains viable despite degraded predictions: because overprovisioning is so extreme, even the conservative 200% safety margin required under realistic conditions yields 64.8% weighted reduction in over-reserved processor-hours. Our results reframe HPC runtime prediction as a non-stationary user- behavior calibration problem.References
[1] D. Duplyakin and K. Menear, “NREL Eagle supercomputer jobs dataset,” Open Energy
Data Initiative (OEDI), 2023. [Online]. Available:
https://data.openei.org/submissions/5860
[2] K. Menear, A. Nag, J. Perr-Sauer, M. Lunacek,
K. Potter, and D. Duplyakin, “Mastering HPC
runtime prediction: From observing patterns to a
methodological approach,” in PEARC ’23, ACM,
2023, pp. 75–85. DOI: 10.1145/3569951.3593598
[3] E. Strubell, A. Ganesh, and A. McCallum, “Energy and policy considerations for deep learning
in NLP,” in ACL 2019, pp. 3645–3650.
[4] G. P. Rodrigo, E. Elmroth, P.-O. Ostberg, ¨
and L. Ramakrishnan, “Enabling workflow-aware
scheduling on HPC systems,” in HPDC ’17,
ACM, 2017, pp. 3–14.
[5] A. Klimentov et al., “HPC job runtime prediction
using machine learning,” J. Phys.: Conf. Ser.,
vol. 1525, no. 1, p. 012014, 2020.
6] S. Ramachandran et al., “Combining machine
learning techniques and genetic algorithm for predicting run times of HPC jobs,” Appl. Soft Comput., vol. 170, p. 112053, 2024.
[7] K. Menear, K. Konate, K. Potter, and D. Duplyakin, “Tandem predictions for HPC jobs,”
in PEARC ’24, ACM, 2024, pp. 1–9. DOI:
10.1145/3626203.3670547
[8] K. Menear, K. Konate, and D. Duplyakin, “Predictive modeling of HPC job queue times: Improving user decision-making and resource utilization,” in PEARC ’25, ACM, 2025. DOI:
10.1145/3708035.3736067
[9] B. Kocot, P. Czarnul, and J. Proficz, “Energyaware scheduling for high-performance computing
systems: A survey,” Energies, vol. 16, no. 2, p.
890, 2023.
[10] A. A. Springborg, M. Albano, and S. X. de
Souza, “Automatic energy-efficient job scheduling in HPC: A novel SLURM plugin approach,”
in SC ’23 Workshops, ACM, 2023, pp. 1961–1969.
[11] A. Hossain, K. Menear, K. Konate, and
D. Duplyakin, “Power-aware scheduling for
multi-center HPC electricity cost optimization,”
arXiv:2503.11011, 2025.
[12] J. Boyle, A. Buluc, and L. Oliker, “Classification
of HPC job power consumption using a rich feature set,” in PEARC ’25, ACM, 2025.
[13] C. Hao et al., “ORA: Job runtime prediction for
high-performance computing platforms,” in ICS
’25, ACM, 2025. DOI: 10.1145/3721145.3721165
[14] Y. Gao, Z. Zhuang, and X.-H. Sun, “Evaluating
HPC job run time predictions using application
input data,” in ICS ’23, ACM, 2023, pp. 252–264.
DOI: 10.1145/3577193.3593722
[15] G. E. Gorbet and B. Demeler, “Optimizing UltraScan job scheduling with deep learning-based
runtime prediction,” in PEARC ’25, ACM, 2025.
DOI: 10.1145/3708035.3736016
[16] S. Wang, S. Chen, and Y. Shi, “Utilizationprediction-aware energy optimization approach
for heterogeneous GPU clusters,” J. Supercomput., vol. 80, no. 7, pp. 9554–9578, 2024.
[17] D. Tsafrir, Y. Etsion, and D. G. Feitelson, “Backfilling using system-generated predictions rather
than user runtime estimates,” IEEE Trans. Parallel Distrib. Syst., vol. 18, no. 6, pp. 789–803,
2007.
[18] M. Tanash, B. Dunn, D. Andresen, W.
Hsu, S. Yang, and A. Okanlawon, “Improving HPC system performance by predicting
job resources via supervised machine learning,”
in PEARC ’19, ACM, 2019, pp. 1–8. DOI:
10.1145/3332186.3333041
[19] M. Tanash, D. Andresen, and W. Hsu, “Ensemble prediction of job resources to improve
system performance for Slurm-based HPC systems,” in PEARC ’21, ACM, 2021. DOI:
10.1145/3437359.3465574
[20] X. Chen, H. Zhang, H. Bai, C. Yang, X. Zhao, and
B. Li, “Runtime prediction of high-performance
computing jobs based on ensemble learning,” in
Proc. 4th Int. Conf. High Performance Compilation, Computing and Communications (HP3C),
ACM, 2020, pp. 56–62.
[21] E. Gaussier, D. Glesser, V. Reis, and D. Trystram, “Improving backfilling by using machine
learning to predict running times,” in SC ’15:
Proc. Int. Conf. High Performance Computing,
Networking, Storage and Analysis, ACM, 2015,
pp. 1–10.
22] S. Madireddy, P. Balaprakash, P. Carns, R.
Latham, G. K. Lockwood, R. Ross, S. Snyder,
and S. M. Wild, “Adaptive learning for concept
drift in application performance modeling,” in
Proc. 48th Int. Conf. Parallel Processing (ICPP),
ACM, 2019, pp. 1–11.
[23] S. Zrigui, R. Y. de Camargo, A. Legrand, and D.
Trystram, “Improving the performance of batchschedulers using online job runtime classification,” IEEE Trans. Parallel Distrib. Syst., vol.
33, no. 8, pp. 1907–1920, 2022.
[24] Y. Fan, P. Rich, W. E. Allcock, M. E. Papka,
and Z. Lan, “Trade-off between prediction accuracy and underestimation rate in job runtime estimates,” in IEEE Int. Conf. Cluster Computing
(CLUSTER), IEEE, 2017, pp. 1–10.
[25] C. B. Lee, Y. Schwartzman, J. Hardy, and A.
Snavely, “Are user runtime estimates inherently
inaccurate?,” in Job Scheduling Strategies for
Parallel Processing (JSSPP), Springer, 2004, pp.
253–263.
[26] M. Naghshnejad and M. Singhal, “A hybrid
scheduling platform: a runtime prediction reliability aware scheduling platform to improve HPC
scheduling performance,” The Journal of Supercomputing, vol. 76, pp. 8551–8570, 2020.
[27] F. Chen, “Job runtime prediction of HPC cluster
based on PC-Transformer,” The Journal of Supercomputing, vol. 79, no. 17, pp. 20208–20234,
2023.
[28] D. G. Feitelson, D. Tsafrir, and D. Krakov,
“Experience with using the Parallel Workloads
Archive,” J. Parallel Distrib. Comput., vol. 74,
no. 10, pp. 2967–2982, 2014.
[29] W. Tang, N. Desai, D. Buettner, and Z. Lan, “Analyzing and adjusting user runtime estimates to
improve job scheduling on the Blue Gene/P,” in
IEEE Int. Symp. Parallel & Distributed Processing (IPDPS), IEEE, 2010, pp. 1–11.
[30] A. B. Yoo, M. A. Jette, and M. Grondona,
“SLURM: Simple Linux Utility for Resource
Management,” in Job Scheduling Strategies for
Parallel Processing (JSSPP), Springer, 2003, pp.
44–60.
DOI:
https://doi.org/10.31449/inf.v50i15.15344Keywords:
High-Performance Computing, Machine Learning, Job Scheduling, Runtime Prediction, Energy Optimization, Concept Drift, User Behavior Modeling, Temporal Generalization, Eagle SupercomputerDownloads
Published
Issue
Section
License
Authors retain copyright in their work. By submitting to and publishing with Informatica, authors grant the publisher (Slovene Society Informatika) the non-exclusive right to publish, reproduce, and distribute the article and to identify itself as the original publisher.
All articles are published under the Creative Commons Attribution license CC BY 3.0. Under this license, others may share and adapt the work for any purpose, provided appropriate credit is given and changes (if any) are indicated.
Authors may deposit and share the submitted version, accepted manuscript, and published version, provided the original publication in Informatica is properly cited.







