KMHC: A Parallel Hybrid Clustering Algorithm Integrating K-Means and Hierarchical Agglomerative Clustering with Voting-Based Consensus
Abstract
This paper proposes KMHC, a parallel hybrid clustering algorithm that combines K-means and Hierarchi- cal Agglomerative Clustering (HAC) through an overlap-based label alignment and deterministic voting consensus mechanism. The objective of KMHC is to exploit the complementarity between centroid-based compactness and hierarchical structural information in order to improve clustering consistency and reduce assignment errors. Unlike conventional ensemble clustering methods that rely on repeated resampling, large pools of base partitions, or co-association matrices, KMHC uses a lightweight two-clusterer architec- ture with explicit cluster correspondence estimation. The proposed method is formally defined through an overlap matrix, a cluster mapping function, and a deterministic tie-breaking rule. Its computational com- plexity is also analyzed, showing that the main bottleneck is the HAC component, while the alignment and voting stages introduce limited additional cost. KMHC is evaluated on benchmark and synthetic datasets with different dimensionalities, noise levels, and cluster distributions. Its performance is compared with K-means, HAC, DBSCAN, Spectral Clustering, and Gaussian Mixture Models using precision, normal- ized mutual information, adjusted Rand index, error rate, runtime, and stability over multiple independent runs. Statistical significance tests and an ablation study are also conducted to assess the contribution of K-means, HAC, label alignment, and voting. The results show that KMHC can improve clustering perfor- mance when the two base clusterers provide complementary partitions, particularly in datasets where both local compactness and hierarchical structure are informative. However, the method remains limited by the quadratic cost of HAC and by its current batch-mode design. Future work will investigate approximate hierarchical clustering, incremental extensions, and streaming-data scenarios.References
DOI:
https://doi.org/10.31449/inf.v50i15.11331Downloads
Published
Issue
Section
License
Authors retain copyright in their work. By submitting to and publishing with Informatica, authors grant the publisher (Slovene Society Informatika) the non-exclusive right to publish, reproduce, and distribute the article and to identify itself as the original publisher.
All articles are published under the Creative Commons Attribution license CC BY 3.0. Under this license, others may share and adapt the work for any purpose, provided appropriate credit is given and changes (if any) are indicated.
Authors may deposit and share the submitted version, accepted manuscript, and published version, provided the original publication in Informatica is properly cited.







