RMAM: A Rule-based Filter Method with Modified Huffman Ranking for Variable Selection in Supervised Learning
Abstract
Most existing filter-based variable selection methods compute variable scores using static mathematical formulations over fixed contingency tables and do not provide principled cut-off mechanisms to distinguish influential from redundant variables thus leaving final subset selection to end users. To address this limitation, this paper proposes a new filter-based variable selection method called Rule Mining Analysis Method (RMAM). For classification tasks involving supervised learning, this method dynamically generates rules from the training data by iteratively selecting variable values that maximize information gain with respect to the target class. Each generated rule is assigned a weight based on rule coverage and uncovered training data instances, and variable scores are subsequently computed by aggregating the weights of all rules in which they appear. To automate variable subset determination, RMAM integrates a new Modified Huffman Ranker (MHR) that computes an adaptive cut-off threshold to distinguish essential from redundant variables. Experimental evaluation across 27 UCI benchmark datasets demonstrates that RMAM reduced variable search space by an average of 75.2% relative to the original datasets thereby outperforming Information Gain and Chi-Square methods which achieved average reductions of 44.7% and 50.0%, respectively. When evaluated using Naïve Bayes classification algorithm, RMAM achieved comparable or superior predictive performance on the majority of datasets while offering substantially smaller and more interpretable variable subsets. These results demonstrate RMAM’s effectiveness as a practical variable selection approach for supervised learning applications.References
[1] Al-Dhaheri, S. (2021). A New Variable Selection Method Based on Class Association Rule (Doctoral dissertation, City University of New York).
[2] Azar, A. T., Boubellouta, A., Bouzeriba, A., & Zouari, F. (2025). Novel fixed-time fuzzy sliding-mode synchronization schemes of two uncertain dissimilar chaotic systems. Soft Computing. Advance online publication
[3] Begum, A. M., Mondal, M. R. H., Podder, P., & Kamruzzaman, J. (2024). Weighted Rank Difference Ensemble: A New Form of Ensemble Variable Selection Method for Medical Datasets. BioMedInformatics, 4(1), 477-488.
[4] Boulkroune, A., Boubellouta, A., Bouzeriba, A., & Zouari, F. (2025). Practical finite-time fuzzy synchronization of chaotic systems with non-integer orders: Two chattering-free approaches. Journal of Systems Science and Systems Engineering, 34, 334–359.
[5] Dua, D., & Graff, C. (2019). UCI Machine Learning Repository. University of California, Irvine, School of Information and Computer Sciences. https://doi.org/10.24432/C5J30R
[6] Hammoud, M. (2009). Introduction to key-value stores. Proceedings of the IEEE Computer Society Annual Symposium on VLSI, 198-203. https://doi.org/10.1109/ISVLSI.2009.44
[7] Huffman, D. A. (1952). A method for the construction of minimum-redundancy codes. Proceedings of the IRE, 40(9), 1098-1101. https://doi.org/10.1109/JRPROC.1952.273898
[8] John, G. H., & Langley, P. (1995). Estimating continuous distributions in Bayesian classifiers. Proceedings of the Eleventh Conference on Uncertainty in Artificial Intelligence (UAI '95), 338-345. Morgan Kaufmann Publishers Inc. https://doi.org/10.5555/2074158.2074196
[9] Karimi, K., Ghodratnama, A., & Tavakkoli-Moghaddam, R. (2023). Two new variable selection methods based on learn-heuristic techniques for breast cancer prediction: a comprehensive analysis. Annals of Operations Research, 328(1), 665-700.
[10] Khan, M. A., & Afzal, H. (2019). Efficient variable selection for sentiment analysis in social media data using hybrid method. Journal of Ambient Intelligence and Humanized Computing, 10(9), 3527-3545. https://doi.org/10.1007/s12652-019-01367-6
[11] Kumar, M., & Kaur, H. (2020). Variable selection for high-dimensional data using hybrid model based on information gain and chi-square. International Journal of Advanced Computer Science and
Applications (IJACSA), 11(4), 435-440. https://doi.org/10.14569/IJACSA.2020.0110455
[12] Li, Chao, Xiao Luo, Yanpeng Qi, Zhenbo Gao, and Xiaohui Lin (2020) A new variable selection algorithm based on relevance, redundancy and complementarity. Computers in biology and medicine 119 (2020): 103667.
[13] Li, J., Cheng, K., Wang, S., Morstatter, F., Trevino, R. P., Tang, J., & Liu, H. (2018). Variable selection: A data perspective. ACM Computing Surveys (CSUR), 50(6), 1-45.
[14] McHugh, M. L. (2013). The Chi-square test of independence. Biochemia Medica, 23(2), 143-149. https://doi.org/10.11613/BM.2013.018
[15] Quinlan, J. R. (1986). Induction of decision trees. Machine Learning, 1(1), 81-106. https://doi.org/10.1007/BF00116251
[16] Rajab, M. D., Taketa, T., Wharton, S. B., Wang, D., & Cognitive Function and Ageing Neuropathology Study, and for the Alzheimer's Disease Neuroimaging Initiative. (2024). Ranking and filtering of neuropathology variables in the machine learning evaluation of dementia studies. Brain Pathology, e13247.
[17] Rigatos, G., Busawon, K., Abbaszadeh, M., & Zouari, F. (2020). Flatness-based adaptive fuzzy control for the Uzawa–Lucas endogenous growth model. AIP Conference Proceedings, 2293(1), 310003.
[18] Rigatos, G., Abbaszadeh, M., & Zouari, F. (2026). Flatness-based control in successive loops for dual-arm robotic manipulators. Journal of Vibration and Control, 32(3–4), 488–507.
[19] Rigatos, G., Busawon, K., Abbaszadeh, M., Pomares, J., Gao, Z., & Zouari, F. (2024). Flatness-based control in successive loops for autonomous quadrotors. Journal of Dynamic Systems, Measurement, and Control, 146(2), Article 011106
[20] Salehi, H., Rashidi, H., & Abadeh, M. S. (2021). Hybrid variable selection methods: A survey and critique of filter, wrapper, and embedded strategies. Artificial Intelligence Review, 54, 4775-4827. https://doi.org/10.1007/s10462-021-09984-x
[21] Thabtah F., Kamalov F., Hammoud S., & Shahamiri S. R. (2020). Least Loss: A simplified filter method for variable selection. Information Science Journal, Volume 534, September 2020, Pages 1-15
[22] Tran, H. N., & Huynh, H. T. (2022). Variable selection methods for high-dimensional data: A comparative study and a new method based on metaheuristics. Knowledge-Based Systems, 243, 108424. https://doi.org/10.1016/j.knosys.2022.108424
[23] Yu, T., Liu, Z., Liu, Y., Wang, H., & Adilov, N. (2020, November). A New Variable Selection Method for Intrusion Detection System Dataset–TSDR method. In 2020 16th international conference on computational intelligence and security (CIS) (pp. 362-365). IEEE.
[24] Zhang, Y., Liu, X., & Wang, J. (2023). A comprehensive review on variable selection methods in machine learning. IEEE Access, 11, 10574-10588. https://doi.org/10.1109/ACCESS.2023.3245678
[25] Zhao, Y., Huang, Y., Wang, Z., & Liu, X. (2024). A new variable selection method based on importance measures for crude oil return forecasting. Neurocomputing, 581, 127470.
[26] Zhao, Z., & Liu, H. (2020). On the challenges in variable selection with big data. IEEE Intelligent Systems, 35(2), 40-45. https://doi.org/10.1109/MIS.2020.2972094
[27] Zubair, I. M., Lee, Y. S., & Kim, B. (2024). A New Permutation-Based Method for Ranking and Selecting Group Variables in Multiclass Classification. Applied Sciences, 14(8), 3156.
DOI:
https://doi.org/10.31449/inf.v50i14.14292Keywords:
Dimensionality Reduction, Feature selection, Machine Learning, Rules, Supervised LearningDownloads
Published
Issue
Section
License
Authors retain copyright in their work. By submitting to and publishing with Informatica, authors grant the publisher (Slovene Society Informatika) the non-exclusive right to publish, reproduce, and distribute the article and to identify itself as the original publisher.
All articles are published under the Creative Commons Attribution license CC BY 3.0. Under this license, others may share and adapt the work for any purpose, provided appropriate credit is given and changes (if any) are indicated.
Authors may deposit and share the submitted version, accepted manuscript, and published version, provided the original publication in Informatica is properly cited.







