Machine Learning-Based Runtime Prediction and Energy Optimization for HPC Job Scheduling Using the NREL Eagle Supercomputer Dataset

Abstract

Accurate runtime prediction is essential for efficient HPC job scheduling, yet users chronically overesti- mate their jobs’ requirements. We analyze 7.3 million completed jobs from the NREL Eagle supercomputer and find that the problem is far worse than previously reported: median time-limit utilization is just 6.7%, with users consuming a median of 10.6 minutes against 4-hour requests. We train ensemble models (Ran- dom Forest, Gradient Boosting, HistGradientBoosting) and an MLP neural network enriched with user behavioral features—historical runtimes, utilization habits, submission frequency—and temporal context. Our central finding concerns evaluation methodology: under the random train/test splits common in prior work, Random Forest reaches R2 = 0.602 (MAE = 0.99 h), but under a realistic temporal split (train <2022, test ≥2022) the same models achieve only R2 = 0.172–0.233—a 71% loss attributable to concept drift, unreported in prior Eagle studies. This gap indicates that random-split results, including our own, overstate deployable accuracy, and that temporal evaluation should be the standard for credible claims. Ablation under random split confirms that user history (∆R2 = −0.078) outweighs the time-limit signal (∆R2 = −0.042), and HistGradientBoosting proves most drift-resilient (R2 = 0.233). Crucially, energy optimization remains viable despite degraded predictions: because overprovisioning is so extreme, even the conservative 200% safety margin required under realistic conditions yields 64.8% weighted reduction in over-reserved processor-hours. Our results reframe HPC runtime prediction as a non-stationary user- behavior calibration problem.

References

[1] D. Duplyakin and K. Menear, “NREL Eagle supercomputer jobs dataset,” Open Energy

Data Initiative (OEDI), 2023. [Online]. Available:

https://data.openei.org/submissions/5860

[2] K. Menear, A. Nag, J. Perr-Sauer, M. Lunacek,

K. Potter, and D. Duplyakin, “Mastering HPC

runtime prediction: From observing patterns to a

methodological approach,” in PEARC ’23, ACM,

2023, pp. 75–85. DOI: 10.1145/3569951.3593598

[3] E. Strubell, A. Ganesh, and A. McCallum, “Energy and policy considerations for deep learning

in NLP,” in ACL 2019, pp. 3645–3650.

[4] G. P. Rodrigo, E. Elmroth, P.-O. Ostberg, ¨

and L. Ramakrishnan, “Enabling workflow-aware

scheduling on HPC systems,” in HPDC ’17,

ACM, 2017, pp. 3–14.

[5] A. Klimentov et al., “HPC job runtime prediction

using machine learning,” J. Phys.: Conf. Ser.,

vol. 1525, no. 1, p. 012014, 2020.

6] S. Ramachandran et al., “Combining machine

learning techniques and genetic algorithm for predicting run times of HPC jobs,” Appl. Soft Comput., vol. 170, p. 112053, 2024.

[7] K. Menear, K. Konate, K. Potter, and D. Duplyakin, “Tandem predictions for HPC jobs,”

in PEARC ’24, ACM, 2024, pp. 1–9. DOI:

10.1145/3626203.3670547

[8] K. Menear, K. Konate, and D. Duplyakin, “Predictive modeling of HPC job queue times: Improving user decision-making and resource utilization,” in PEARC ’25, ACM, 2025. DOI:

10.1145/3708035.3736067

[9] B. Kocot, P. Czarnul, and J. Proficz, “Energyaware scheduling for high-performance computing

systems: A survey,” Energies, vol. 16, no. 2, p.

890, 2023.

[10] A. A. Springborg, M. Albano, and S. X. de

Souza, “Automatic energy-efficient job scheduling in HPC: A novel SLURM plugin approach,”

in SC ’23 Workshops, ACM, 2023, pp. 1961–1969.

[11] A. Hossain, K. Menear, K. Konate, and

D. Duplyakin, “Power-aware scheduling for

multi-center HPC electricity cost optimization,”

arXiv:2503.11011, 2025.

[12] J. Boyle, A. Buluc, and L. Oliker, “Classification

of HPC job power consumption using a rich feature set,” in PEARC ’25, ACM, 2025.

[13] C. Hao et al., “ORA: Job runtime prediction for

high-performance computing platforms,” in ICS

’25, ACM, 2025. DOI: 10.1145/3721145.3721165

[14] Y. Gao, Z. Zhuang, and X.-H. Sun, “Evaluating

HPC job run time predictions using application

input data,” in ICS ’23, ACM, 2023, pp. 252–264.

DOI: 10.1145/3577193.3593722

[15] G. E. Gorbet and B. Demeler, “Optimizing UltraScan job scheduling with deep learning-based

runtime prediction,” in PEARC ’25, ACM, 2025.

DOI: 10.1145/3708035.3736016

[16] S. Wang, S. Chen, and Y. Shi, “Utilizationprediction-aware energy optimization approach

for heterogeneous GPU clusters,” J. Supercomput., vol. 80, no. 7, pp. 9554–9578, 2024.

[17] D. Tsafrir, Y. Etsion, and D. G. Feitelson, “Backfilling using system-generated predictions rather

than user runtime estimates,” IEEE Trans. Parallel Distrib. Syst., vol. 18, no. 6, pp. 789–803,

2007.

[18] M. Tanash, B. Dunn, D. Andresen, W.

Hsu, S. Yang, and A. Okanlawon, “Improving HPC system performance by predicting

job resources via supervised machine learning,”

in PEARC ’19, ACM, 2019, pp. 1–8. DOI:

10.1145/3332186.3333041

[19] M. Tanash, D. Andresen, and W. Hsu, “Ensemble prediction of job resources to improve

system performance for Slurm-based HPC systems,” in PEARC ’21, ACM, 2021. DOI:

10.1145/3437359.3465574

[20] X. Chen, H. Zhang, H. Bai, C. Yang, X. Zhao, and

B. Li, “Runtime prediction of high-performance

computing jobs based on ensemble learning,” in

Proc. 4th Int. Conf. High Performance Compilation, Computing and Communications (HP3C),

ACM, 2020, pp. 56–62.

[21] E. Gaussier, D. Glesser, V. Reis, and D. Trystram, “Improving backfilling by using machine

learning to predict running times,” in SC ’15:

Proc. Int. Conf. High Performance Computing,

Networking, Storage and Analysis, ACM, 2015,

pp. 1–10.

22] S. Madireddy, P. Balaprakash, P. Carns, R.

Latham, G. K. Lockwood, R. Ross, S. Snyder,

and S. M. Wild, “Adaptive learning for concept

drift in application performance modeling,” in

Proc. 48th Int. Conf. Parallel Processing (ICPP),

ACM, 2019, pp. 1–11.

[23] S. Zrigui, R. Y. de Camargo, A. Legrand, and D.

Trystram, “Improving the performance of batchschedulers using online job runtime classification,” IEEE Trans. Parallel Distrib. Syst., vol.

33, no. 8, pp. 1907–1920, 2022.

[24] Y. Fan, P. Rich, W. E. Allcock, M. E. Papka,

and Z. Lan, “Trade-off between prediction accuracy and underestimation rate in job runtime estimates,” in IEEE Int. Conf. Cluster Computing

(CLUSTER), IEEE, 2017, pp. 1–10.

[25] C. B. Lee, Y. Schwartzman, J. Hardy, and A.

Snavely, “Are user runtime estimates inherently

inaccurate?,” in Job Scheduling Strategies for

Parallel Processing (JSSPP), Springer, 2004, pp.

253–263.

[26] M. Naghshnejad and M. Singhal, “A hybrid

scheduling platform: a runtime prediction reliability aware scheduling platform to improve HPC

scheduling performance,” The Journal of Supercomputing, vol. 76, pp. 8551–8570, 2020.

[27] F. Chen, “Job runtime prediction of HPC cluster

based on PC-Transformer,” The Journal of Supercomputing, vol. 79, no. 17, pp. 20208–20234,

2023.

[28] D. G. Feitelson, D. Tsafrir, and D. Krakov,

“Experience with using the Parallel Workloads

Archive,” J. Parallel Distrib. Comput., vol. 74,

no. 10, pp. 2967–2982, 2014.

[29] W. Tang, N. Desai, D. Buettner, and Z. Lan, “Analyzing and adjusting user runtime estimates to

improve job scheduling on the Blue Gene/P,” in

IEEE Int. Symp. Parallel & Distributed Processing (IPDPS), IEEE, 2010, pp. 1–11.

[30] A. B. Yoo, M. A. Jette, and M. Grondona,

“SLURM: Simple Linux Utility for Resource

Management,” in Job Scheduling Strategies for

Parallel Processing (JSSPP), Springer, 2003, pp.

44–60.

Authors

  • Renato Quispe-Vargas Universidad Nacional del Altiplano image/svg+xml
  • Dina Maribel Yana-Yucra
  • Richar Andre Vilca-Solorzano
  • Vladimiro Ibañez-Quispe
  • Fred Torres-Cruz

DOI:

https://doi.org/10.31449/inf.v50i15.15344

Keywords:

High-Performance Computing, Machine Learning, Job Scheduling, Runtime Prediction, Energy Optimization, Concept Drift, User Behavior Modeling, Temporal Generalization, Eagle Supercomputer

Downloads

Published

09/09/2026

How to Cite

Quispe-Vargas, R., Yana-Yucra, D. M., Vilca-Solorzano, R. A., Ibañez-Quispe, V., & Torres-Cruz, F. (2026). Machine Learning-Based Runtime Prediction and Energy Optimization for HPC Job Scheduling Using the NREL Eagle Supercomputer Dataset. Informatica, 50(15). https://doi.org/10.31449/inf.v50i15.15344