Cost-Bounded, Accuracy-Aware Hyperparameter Optimization Integrated with Rule-Based Book Recommendation for Developer Profiling on Stack Overflow
Abstract
This paper introduces HyperR, an integrated framework that combines hyperparameter optimization with a rule-based book recommender for developer profiling on Stack Overflow. The proposed optimizationmethod employs a cost-bounded, accuracy-aware strategy with early stopping and adaptive search to balance computational efficiency and search effectiveness. The framework evaluates both model-free (Grid Search, Random Search, Nelder–Mead, DEoptim) and model-based (Bayesian Optimization, AutoFT, mlrMBO,TuRBO) optimizers across multiple classifiers on a dataset of 20,000 Stack Overflow questions, using stratified 70/30 train-test splits with 20 repeated runs per configuration. Experimental results show that the proposed method achieves the highest macro-averaged precision (0.671) and significantly outperformsclassical baselines (adjusted p < 0.008), while remaining competitive with state-of-the-art modelbased approaches (adjusted p > 0.05). In addition, it reduces peak memory usage by up to 25% compared with exhaustive search methods. The recommendation module achieves strong performance (Precision@3= 0.82 and Recall@3 = 0.77), demonstrating the effectiveness of the proposed framework in delivering personalized learning resources.References
[1] Ndukwe I, Licorish S, and MacDonell S (2023) Perceptions on the utility of community question and answer websites like Stack Overflow to software developers, IEEE Transactions on Software Engineering, IEEE, vol. 49, pp. 2413–2425. https://doi.org/10.1109/TSE.2022.3220236
[2] Meldrum S, Licorish S, Owen C, and Savarimuthu B (2020) Understanding Stack Overflow code quality: A recommendation of caution, Science of Computer Programming, Elsevier, vol. 199, p. 102516. https://doi.org/10.1016/j.scico.2020.102516
[3] Siegmund J, Kästner C, Liebig J, Apel S, and Hanenberg S (2013) Measuring and modeling programming experience, Empirical Software Engineering, Springer, vol. 19, no. 5, pp. 1299–1334. https://doi.org/10.1007/s10664-013-9286-4
[4] Wu Y, Wang S, Bezemer C, and Inoue K (2018) How do developers utilize source code from Stack Overflow?, Empirical Software Engineering, Springer, vol. 24, pp. 637–673. https://doi.org/10.1007/s10664-018-9634-5
[5] Sela A, Hadar L, Morgan S, and Maimaran M (2019) Variety-seeking and perceived expertise, Journal of Consumer Psychology, Wiley. https://doi.org/10.1002/jcpy.1110
[6] Li C, Ishak I, Ibrahim H, Zolkepli M, Sidi F, and Li C (2023) Deep learning-based recommendation system: Systematic review and classification, IEEE Access, IEEE, vol. 11, pp. 113790–113835. https://doi.org/10.1109/ACCESS.2023.3323353
[7] Wei H, Jia H, Li Y, and Xu Y (2020) Verify and measure the quality of rule-based machine learning, Knowledge-Based Systems, Elsevier, vol. 205, p. 106300. https://doi.org/10.1016/j.knosys.2020.106300
[8] Wang M, Fu W, He X, Hao S, and Wu X (2020) A survey on large-scale machine learning, IEEE Transactions on Knowledge and Data Engineering, IEEE, vol. 34, pp. 2574–2594. https://doi.org/10.1109/TKDE.2020.3015777
[9] Bischl B, Binder M, Lang M, et al. (2023) Hyperparameter optimization: Foundations, algorithms, best practices, and open challenges, Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, Wiley, vol. 13, no. 2, p. e1484. https://doi.org/10.1002/widm.1484
[10] Majidi F, Openja M, Khomh F, and Li H (2022) An empirical study on the usage of automated machine learning tools, Proceedings of the 38th IEEE International Conference on Software Maintenance and Evolution (ICSME), IEEE, pp. 190–200. https://doi.org/10.1109/ICSME55016.2022.00014
[11] Bergstra J and Bengio Y (2012) Random search for hyper-parameter optimization, Journal of Machine Learning Research, MIT Press, vol. 13, no. 2. [Online]. Available: https://www.jmlr.org/papers/v13/bergstra12a.html
[12] Snoek J, Larochelle H, and Adams RP (2012) Practical Bayesian optimization of machine learning algorithms, Advances in Neural Information Processing Systems (NeurIPS), vol. 25, pp. 2951–2959
[13] Vasilescu B, Filkov V, and Serebrenik A (2013) Stack Overflow and GitHub: Associations between software development and crowdsourced knowledge, Proceedings of the 2013 International Conference on Social Computing, IEEE, pp. 188–195. https://doi.org/10.1109/SocialCom.2013.35
[14] Chen X, Xu F, Huang Y, Zhou X, and Zheng Z (2024) An empirical study of code reuse between GitHub and Stack Overflow during software development, Journal of Systems and Software, Elsevier, p. 111964. https://doi.org/10.1016/j.jss.2024.111964
[15] Nasehi SM, Sillito J, Maurer F, and Burns C (2012) What makes a good code example? A study of programming Q&A in Stack Overflow, Proceedings of the 28th IEEE International Conference on Software Maintenance (ICSM), IEEE, pp. 25–34. https://doi.org/10.1109/ICSM.2012.6405249
[16] Storey M-A, Zagalsky A, Figueira Filho F, Singer L, and German DM (2016) How social and communication channels shape and challenge a participatory culture in software development, IEEE Transactions on Software Engineering, IEEE, vol. 43, no. 2, pp. 185–204. https://doi.org/10.1109/TSE.2016.2584053
[17] He J, Xu B, Yang Z, Han D, Yang C, and Lo D (2022) PTM4Tag: Sharpening tag recommendation of Stack Overflow posts with pretrained models, Proceedings of the 30th IEEE/ACM International Conference on Program Comprehension, ACM/IEEE, pp. 1–11. https://doi.org/10.1145/3524610.3527897
[18] Joorabchi A, English M, and Mahdi AE (2016) Text mining Stack Overflow: An insight into challenges and subject-related difficulties faced by computer science learners, Journal of Enterprise Information Management, Emerald, vol. 29, no. 2, pp. 255–275. https://doi.org/10.1108/JEIM-11-2014-0109
[19] Harrag F and Khamliche M (2020) Mining Stack Overflow: A recommender systems-based model, Preprints. https://doi.org/10.20944/preprints202008.0265.v2
[20] Silva CC, Galster M, and Gilson F (2021) Topic modeling in software engineering research, Empirical Software Engineering, Springer, vol. 26, no. 6, p. 120. https://doi.org/10.1007/s10664-021-10026-0
[21] Nelder JA and Mead R (1965) A simplex method for function minimization, The Computer Journal, Oxford University Press, vol. 7, no. 4, pp. 308–313. https://doi.org/10.1093/comjnl/7.4.308
[22] Storn R and Price K (1997) Differential evolution: A simple and efficient heuristic for global optimization over continuous spaces, Journal of Global Optimization, Springer, vol. 11, no. 4, pp. 341–359. https://doi.org/10.1023/A:1008202821328
[23] Liashchynskyi P and Liashchynskyi P (2019) Grid search, random search, genetic algorithm: A big comparison for NAS, arXiv preprint arXiv:1912.06059. https://doi.org/10.48550/arXiv.1912.06059
[24] Bergstra J, Yamins D, and Cox D (2013) Making a science of model search: Hyperparameter optimization in hundreds of dimensions for vision architectures, Proceedings of the International Conference on Machine Learning, PMLR, pp. 115–123. [Online]. Available: https://proceedings.mlr.press/v28/bergstra13.html
[25] Lagarias JC, Reeds JA, Wright MH, and Wright PE (1998) Convergence properties of the Nelder–Mead simplex method in low dimensions, SIAM Journal on Optimization, SIAM, vol. 9, no. 1, pp. 112–147. https://doi.org/10.1137/S1052623496303470
[26] Eriksson D, Pearce M, Gardner JR, Turner RD, and Poloczek M (2019) Scalable global optimization via local Bayesian optimization, Advances in Neural Information Processing Systems (NeurIPS), vol. 32, pp. 5497–5508
[27] Bischl B, Richter J, Bossek J, Horn D, Thomas J, and Lang M (2017) mlrMBO: A modular framework for model-based optimization of expensive black-box functions, arXiv preprint, arXiv:1703.03373. https://doi.org/10.48550/arXiv.1703.03373
[28] Choi C, Lee Y, Chen A, Zhou A, Raghunathan A, and Finn C (2024) AutoFT: Learning an objective for robust fine-tuning, arXiv preprint, arXiv:2401.10220. https://doi.org/10.48550/arXiv.2401.10220
[29] Li L, Jamieson K, DeSalvo G, Rostamizadeh A, and Talwalkar A (2017) Hyperband: A novel bandit-based approach to hyperparameter optimization, Journal of Machine Learning Research, MIT Press, vol. 18, no. 1, pp. 6765–6816. https://doi.org/10.48550/arXiv.1603.06560
[30] Mallik N, Bergman E, Hvarfner C, Stoll D, Janowski M, Lindauer M, and Hutter F (2023) PriorBand: Practical hyperparameter optimization in the age of deep learning, Advances in Neural Information Processing Systems (NeurIPS), vol. 36, pp. 7377–7391
[31] Liu T, Astorga N, Seedat N, and van der Schaar M (2024) Large language models to enhance Bayesian optimization, arXiv preprint, arXiv:2402.03921. https://doi.org/10.48550/arXiv.2402.03921
[32] Montaner M, López B, and De La Rosa JL (2003) A taxonomy of recommender agents on the internet, Artificial Intelligence Review, Springer, vol. 19, pp. 285–330. https://doi.org/10.1023/A:1022850703159
[33] He X, Liao L, Zhang H, Nie L, Hu X, and Chua TS (2017) Neural collaborative filtering, Proceedings of the 26th International Conference on World Wide Web (WWW), ACM, pp. 173–182. https://doi.org/10.1145/3038912.3052569
[34] Mantovani RG, Rossi AL, Alcobaça E, Vanschoren J, and de Carvalho AC (2019) A meta-learning recommender system for hyperparameter tuning: Predicting when tuning improves SVM classifiers, Information Sciences, Elsevier, vol. 501, pp. 193–221. https://doi.org/10.1016/j.ins.2019.06.005
[35] Yang C, Akimoto Y, Kim DW, and Udell M (2019) OBOE: Collaborative filtering for AutoML model selection, Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, ACM, New York, USA, pp. 1173–1183. https://doi.org/10.1145/3292500.3330909
[36] Feurer M, Klein A, Eggensperger K, Springenberg J, Blum M, and Hutter F (2019) Auto-sklearn: Efficient and robust automated machine learning, in Automated Machine Learning, Springer, Cham, pp. 113–134. https://doi.org/10.1007/978-3-030-05318-5_6
[37] Erickson N, Mueller J, Shirkov A, Zhang H, Larroy P, Li M, and Smola A (2020) AutoGluon-Tabular: Robust and accurate AutoML for structured data, arXiv preprint, arXiv:2003.06505. https://doi.org/10.48550/arXiv.2003.06505
[38] Wang C, Wu Q, Weimer M, and Zhu E (2021) FLAML: A fast and lightweight AutoML library, arXiv preprint, arXiv:1911.04706. https://doi.org/10.48550/arXiv.1911.04706
[39] Komer B, Bergstra J, and Eliasmith C (2014) Hyperopt-sklearn: Automatic hyperparameter configuration for Scikit-learn, Proceedings of the 13th Python in Science Conference, pp. 32–37. https://doi.org/10.25080/Majora-14bd3278-006
[40] Akiba T, Sano S, Yanase T, Ohta T, and Koyama M (2019) Optuna: A next-generation hyperparameter optimization framework, Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, ACM, pp. 2623–2631. https://doi.org/10.1145/3292500.3330701
[41] Yuan B, Gui S, Zhang Q, Wang Z, Wen J, Mao B, Liu J, and Yao X (2024) FairerML: An extensible platform for analysing, visualising, and mitigating biases in machine learning, IEEE Computational Intelligence Magazine, IEEE, vol. 19, no. 2, pp. 129–141. https://doi.org/10.1109/MCI.2024.3364430
[42] Mansion T, Braud R, Amrani A, Chaouche S, Adjed F, and Cantat L (2024) Debiai: Open-source toolkit for data analysis, visualisation and evaluation in machine learning, Proceedings of the 20th International Conference on Autonomic and Autonomous Systems, IARIA, Athens, Greece
[43] Jia Y, Wang H, and Chen D (2023) LAMDA-SSL: Label augmentation and manifold-based data alignment for semi-supervised learning, Pattern Recognition, Elsevier, vol. 144, p. 109690. https://doi.org/10.1016/j.patcog.2023.109690
[44] Altınsoy F and Öztürk MM (2026) Hyperparameter optimization web tool: HyperOpt, Sigma Journal of Engineering and Natural Sciences, vol. 44, no. 2, pp. 1219–1239. https://doi.org/10.14744/sigma.2026.2033
[45] Zhu Y, Li W, and Li T (2023) A hybrid artificial immune optimization for high-dimensional feature selection, Knowledge-Based Systems, Elsevier, vol. 260, p. 110111. https://doi.org/10.1016/j.knosys.2022.110111
[46] Pan H, Chen S, and Xiong H (2023) A high-dimensional feature selection method based on modified gray wolf optimization, Applied Soft Computing, Elsevier, vol. 135, p. 110031. https://doi.org/10.1016/j.asoc.2023.110031
[47] Hsu CW, Chang CC, and Lin CJ (2003) A practical guide to support vector classification, Technical Report, Department of Computer Science and Information Engineering, National Taiwan University, Taipei, pp. 1–12. Available: https://www.csie.ntu.edu.tw/~cjlin/papers/guide/guide.pdf
[48] Montañés E, Quevedo JR, and del Coz JJ (2014) Dependent binary relevance models for multi-label classification, Pattern Recognition, Elsevier, vol. 47, no. 3, pp. 1494–1508. https://doi.org/10.1016/j.patcog.2013.09.029
[49] Bergstra J, Bardenet R, Bengio Y, and Kégl B (2011) Algorithms for hyper-parameter optimization, Advances in Neural Information Processing Systems (NeurIPS), vol. 24, pp. 2546–2554
[50] Kaggle (2023) Kaggle dataset: Stack Overflow - Stacksample. Available: https://www.kaggle.com/datasets/stackoverflow/stacksample, accessed: 2023-10-17
[51] Manning CD, Raghavan P, and Schütze H (2008) Introduction to information retrieval, Cambridge University Press. https://doi.org/10.1017/CBO9780511809071
[52] Aggarwal CC and Zhai C (2012) Mining text data, Springer. https://doi.org/10.1007/978-1-4614-3223-4
[53] Porter MF (1980) An algorithm for suffix stripping, Program, Emerald, vol. 14, no. 3, pp. 130–137. https://doi.org/10.1108/eb046814
[54] Salton G and Buckley C (1988) Term-weighting approaches in automatic text retrieval, Information Processing & Management, Elsevier, vol. 24, no. 5, pp. 513–523. https://doi.org/10.1016/0306-4573(88)90021-0
[55] Jurafsky D and Martin JH (2025) Speech and language processing: An introduction to natural language processing, computational linguistics, and speech recognition (3rd ed., draft), Available: https://web.stanford.edu/~jurafsky/slp3/
[56] Mikolov T, Sutskever I, Chen K, Corrado GS, and Dean J (2013) Distributed representations of words and phrases and their compositionality, Advances in Neural Information Processing Systems (NeurIPS), vol. 26. https://doi.org/10.48550/arXiv.1310.4546
[57] Deerwester S, Dumais ST, Furnas GW, Landauer TK, and Harshman R (1990) Indexing by latent semantic analysis, Journal of the American Society for Information Science, Wiley, vol. 41, no. 6, pp. 391–407. https://doi.org/10.1002/(SICI)1097-4571(199009)41:6<391::AID-ASI1>3.0.CO;2-9
[58] Peruma A, Simmons S, AlOmar EA, Newman CD, Mkaouer MW, and Ouni A (2022) How do I refactor this? An empirical study on refactoring trends and topics in Stack Overflow, Empirical Software Engineering, Springer, vol. 27, no. 1, p. 11. https://doi.org/10.1007/s10664-021-10045-x
[59] West E (2022) Buy now: How Amazon branded convenience and normalized monopoly, MIT Press
[60] Kuhn M (2008) Building predictive models in R using the caret package, Journal of Statistical Software, vol. 28, no. 5, pp. 1–26. https://doi.org/10.18637/jss.v028.i05
[61] Dimitriadou E, Hornik K, Leisch F, Meyer D, and Weingessel A (2009) E1071: Misc functions of the Department of Statistics (E1071), TU Wien
[62] Ridgeway G (2007) Generalized boosted models: A guide to the gbm package. Available: https://cran.r-project.org/web/packages/gbm/gbm.pdf, accessed: 2021
[63] de Carvalho AC and Freitas AA (2009) A tutorial on multi-label classification techniques, in Foundations of Computational Intelligence Volume 5: Function Approximation and Classification, Springer, vol. 5, pp. 177–195. https://doi.org/10.1007/978-3-642-01536-6_8
[64] Read J, Pfahringer B, and Holmes G (2008) Multi-label classification using ensembles of pruned sets, Proceedings of the Eighth IEEE International Conference on Data Mining (ICDM), IEEE, pp. 995–1000. https://doi.org/10.1109/ICDM.2008.74
[65] Liaw A and Wiener M (2002) Classification and regression by randomForest, R News, vol. 2, pp. 18–22
[66] Yan Y (2016) rBayesOptimization: Bayes optimization of hyperparameters, CRAN: Contributed Packages
[67] Probst P, Wright MN, and Boulesteix AL (2019) Hyperparameters and tuning strategies for random forest, Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, Wiley, vol. 9, no. 3, p. e1301. https://doi.org/10.1002/widm.1301
DOI:
https://doi.org/10.31449/inf.v50i2.11275Downloads
Published
Issue
Section
License
Authors retain copyright in their work. By submitting to and publishing with Informatica, authors grant the publisher (Slovene Society Informatika) the non-exclusive right to publish, reproduce, and distribute the article and to identify itself as the original publisher.
All articles are published under the Creative Commons Attribution license CC BY 3.0. Under this license, others may share and adapt the work for any purpose, provided appropriate credit is given and changes (if any) are indicated.
Authors may deposit and share the submitted version, accepted manuscript, and published version, provided the original publication in Informatica is properly cited.







