A Novel Method for Automatic Parameter Tuning in Density-Based Clustering using Local Density Variations
Abstract
Density based clustering is clustering technique which has the ability to find the arbitrary shaped clusters. One of the most important reason for the success of density-based clustering is its ability to obtain the clusters of arbitrary shapes. DBSCAN is pioneer and one of the most studied algorithms in this category of clustering. The success of DBSCAN attracted the attention of researchers. Due to its ability of forming the arbitrary shape clusters and noise detection, DBSCAN is used as a benchmark algorithm in density-based clustering technique. DBSCAN also have few limitations, one of the major limitations it suffers is the manual input parameters. DBSCAN used two manual values as an input namely Epsilon and MinPts. Epsilon define the searching area and hence plays a very important and decisive role in cluster formation. With very small value of epsilon, one may obtain too many clusters. It will be difficult to detect the noise with small value of epsilon. When the value of epsilon is larger, the size of cluster will be larger. The large size of cluster will affect the ability of algorithm to detect noise. As the epsilon plays a very important role in quality of clusters. The proposed approach addressed this issue by automatic determination of epsilon. The proposed algorithm provides an optimal value of epsilon.References
Yin, L., Hu, H., Li, K., Zheng, G., Qu, Y., & Chen, H. (2023). Improvement of DBSCAN Algorithm Based on K-Dist Graph for Adaptive Determining Parameters. Electronics, 12(15), 3213. https://doi.org/10.3390/electronics12153213
Delgado, L., & Morales, E. F. (2021). DBSCAN Parameter Selection Based on K-NN. In Advances in Computational Intelligence (pp. 187–198). Springer. https://doi.org/10.1007/978-3-030-89817-5_14
Sharma, A., & Sharma, A. (2017). KNN-DBSCAN: Using k-nearest neighbor information for parameter-free density based clustering. In Proceedings of the 2017 International Conference on Intelligent Computing, Instrumentation and Control Technologies (ICICICT). IEEE. https://doi.org/10.1109/ICICICT1.2017.8342664
Wang, Y., Ye, Z., Du, Y., Mao, Y., Liu, Y., Wu, Z., & Wang, J. (2022). AMD-DBSCAN: An Adaptive MultiDensity DBSCAN for Datasets of Extremely Variable Density. arXiv preprint arXiv:2210.08162.
Zhang, R., Peng, H., Dou, Y., Wu, J., Sun, Q., Zhang, J., & Yu, P. S. (2022). Automating DBSCAN via Deep Reinforcement Learning. arXiv preprint arXiv:2208.04537. https://arxiv.org/abs/2208.04537
Chen, Y., Ruys, W., & Biros, G. (2020). KNN-DBSCAN: A DBSCAN in high dimensions. arXiv preprint arXiv:2009.04552. https://arxiv.org/abs/2009.04552
Chowdhury, S., & Amorim, R. C. de. (2018). An efficient density-based clustering algorithm using reverse nearest neighbour. arXiv preprint arXiv:1811.07615. https://arxiv.org/abs/1811.07615.
Zhang, L., & Xu, Z. (2021). A Novel Approach to Determining the Radius of the Neighborhood Required for the DBSCAN Algorithm. Journal of Applied Intelligence, 51(5), 2789–2803. https://doi.org/10.1007/s10462-020-09891-4
Liu, X., & Wang, Y. (2020). A fast DBSCAN algorithm using a bi-directional HNSW index structure for big data. Neurocomputing, 414, 1–10. https://doi.org/10.1016/j.neucom.2020.07.008
Zhang, Y., & Li, X. (2021). A Distributed Neighbourhood DBSCAN Algorithm for Effective Data Clustering in Wireless Sensor Networks. Sensors, 21(16), 5351 https://doi.org/10.3390/s21165351
Zhou, J., & Li, J. (2021). A Novel Density-Based Clustering Algorithm Using Reverse Nearest Neighbour. Pattern Recognition Letters, 141, 1–8. https://doi.org/10.1016/j.patrec.2020.11.013
Zhang, R., & Liu, T. (2022). DBSCAN Parameter Selection Using KNN Distance Graphs for Varying Density Data. IEEE Transactions on Cybernetics, 52(7), 6453-6462.
Kang, J., & Lee, K. (2023). Improving DBSCAN Clustering with Adaptive K-Nearest Neighbor Search for Parameter Tuning. Applied Intelligence, 53(6), 7345-7356.
Wang, H., & Jiang, Y. (2024). An Adaptive Density-Based DBSCAN Algorithm Using K-Nearest Neighbors for High-Dimensional Data. Expert Systems with Applications, 185, 115636.
Xu, J., & Zhang, W. (2023). Automating Radius Parameter Selection for DBSCAN via K-Nearest Neighbor Distance Distribution. Neural Computing and Applications, 35(15), 11371-11383.
Wang, Y., & Liu, X. (2021) A Fast DBSCAN Algorithm Using a Bi-Directional HNSW Index Structure for Big Data. Neurocomputing, 414, 1–10. https://doi.org/10.1016/j.neucom.2020.07.008
Gao, H., & Sun, Y. (2024). DBSCAN Parameter Tuning via KNN-Distances for Unsupervised Clustering in Image Processing. Signal Processing, 179, 107843
Li, H., & Zhao, Y. (2023). KNN-DBSCAN: A Density-Based Clustering Algorithm for Handling Large, Noisy Data. Journal of Data Mining and Knowledge Discovery, 37(1), 123-136.
Yang, J., & Zhao, X. (2022). Optimizing DBSCAN Parameters Automatically Using KNN and Curvature Analysis. Pattern Recognition Letters, 156, 79-87.
Yang, T., & Zhang, Z. (2023). KNN-Based Method for Parameter Selection in DBSCAN for Large-Scale Environmental Data Clustering. Environmental Modelling & Software, 166, 105015.
Ma, Z., & Zhang, T. (2023). Automatic Parameter Estimation for DBSCAN Clustering Using KNN and Density Distribution Analysis. Data & Knowledge Engineering, 147, 101912.
Zhou, H., & Li, X. (2021). An Adaptive Eps Parameter of DBSCAN Algorithm for Identifying Clusters in Spatial Data. Computers, Environment and Urban Systems, 85, 101568. https://doi.org/10.1016/j.compenvurbsys.2020.101568
Wu, Y., & Shi, X. (2024). Optimizing DBSCAN for Multi-Dimensional Data Using KNN Distance Metrics. Data Science and Engineering, 9(2), 96-104.
Huang, Z., & Li, H. (2024). A KNN-Based Density Estimation for DBSCAN Parameter Selection in High-Noise Data. Journal of Computational Intelligence, 32(3), 459-470.
Tan, L., & Wang, C. (2023). Enhancing DBSCAN with KNN-Generated Epsilon for Real-Time Data Stream Clustering. IEEE Transactions on Neural Networks and Learning Systems, 34(5), 1968-1979.
Wang, Q., & Liu, X. (2021). A Fast DBSCAN Algorithm Using a Bi-Directional HNSW Index Structure for Big Data. Neurocomputing, 414, 1–10. https://doi.org/10.1016/j.neucom.2020.07.008
Zhang, S., & Liu, Y. (2021). An Efficient Density-Based Clustering Algorithm Using Reverse Nearest Neighbour. Pattern Recognition Letters, 141, 1–8. https://doi.org/10.1016/j.patrec.2020.11.013
Wang, X., & Zhang, L. (2021). A Novel Density-Based Clustering Algorithm Using Nearest Neighbor Information. Journal of Computer Science and Technology, 36(3), 561–573. https://doi.org/10.1007/s11390-021-1354-3
Zhang, J., & Zhang, Y. (2021). A Distributed Neighbourhood DBSCAN Algorithm for Effective Data Clustering in Wireless Sensor Networks. Sensors, 21(16), 5351. https://doi.org/10.3390/s21165351
Kulkarni, O., & Burhanpurwala, A. (2024, February). A survey of advancements in DBSCAN clustering algorithms for big data. In 2024 3rd International conference on Power Electronics and IoT Applications in Renewable Energy and its Control (PARC) (pp. 106-111). IEEE.
DOI:
https://doi.org/10.31449/inf.v50i2.11055Downloads
Published
Issue
Section
License
Authors retain copyright in their work. By submitting to and publishing with Informatica, authors grant the publisher (Slovene Society Informatika) the non-exclusive right to publish, reproduce, and distribute the article and to identify itself as the original publisher.
All articles are published under the Creative Commons Attribution license CC BY 3.0. Under this license, others may share and adapt the work for any purpose, provided appropriate credit is given and changes (if any) are indicated.
Authors may deposit and share the submitted version, accepted manuscript, and published version, provided the original publication in Informatica is properly cited.







