A dynamic core evolutionary clustering algorithm based on saturated memory

Autonomous Intelligent Systems - Tập 3 - Trang 1-16 - 2023
Haibin Xie1, Peng Li1, Zhiyong Ding1
1College of Intelligence Science and Technology, National University of Defense Technology, Changsha, China

Tóm tắt

Because the number of clustering cores needs to be set before implementing the K-means algorithm, this type of algorithm often fails in applications with increasing data and changing distribution characteristics. This paper proposes an evolutionary algorithm DCC, which can dynamically adjust the number of clustering cores with data change. DCC algorithm uses the Gaussian function as the activation function of each core. Each clustering core can adjust its center vector and coverage based on the response to the input data and its memory state to better fit the sample clusters in the space. The DCC algorithm model can evolve from 0. After each new sample is added, the winning dynamic core can be adjusted or split by competitive learning, so that the number of clustering cores of the algorithm always maintains a better adaptation relationship with the existing data. Furthermore, because its clustering core can split, it can subdivide the densely distributed data clusters. Finally, detailed experimental results show that the evolutionary clustering algorithm DCC based on the dynamic core method has excellent clustering performance and strong robustness.

Tài liệu tham khảo

A. Saxena, M. Prasad, A. Gupta, N. Bharill, O.P. Patel, A. Tiwari, M.J. Er, W. Ding, C.T. Lin, A review of clustering techniques and developments. Neurocomputing 267, 664–681 (2017). https://doi.org/10.1016/j.neucom.2017.06.053 F. Li, H. Qiao, B. Zhang, Discriminatively boosted image clustering with fully convolutional auto-encoders. Pattern Recognit. 83, 161–173 (2017). https://doi.org/10.48550/arXiv.1703.07980 A.M. Bagirov, J. Ugon, D. Webb, Fast modified global k-means algorithm for incremental cluster construction. Pattern Recognit. 44(4), 866–876 (2011). https://doi.org/10.1016/j.patcog.2010.10.018 X. Yi, Y. Zhang, Equally contributory privacy-preserving k-means clustering over vertically partitioned data. Inf. Sci. 38(1), 97–107 (2013). https://doi.org/10.1016/j.is.2012.06.001 P. Fränti, S. Sieranoja, K-means properties on six clustering benchmark datasets. Appl. Intell. 48(12), 4743–4759 (2018). https://doi.org/10.1007/s10489-018-1238-7 N. Tsapanos, A. Tefas, N. Nikolaidis, I. Pitas, A distributed framework for trimmed kernel k-means clustering. Pattern Recognit. 48(8), 2685–2698 (2015). https://doi.org/10.1016/j.patcog.2015.02.020 G. Tzortzis, A. Likas, The minmax k-means clustering algorithm. Pattern Recognit. 47(7), 2505–2516 (2014). https://doi.org/10.1016/j.patcog.2014.01.015 K.-P. Lin, A novel wvolutionary kernel intuitionistic fuzzy c-means clustering algorithm. IEEE Trans. Fuzzy Syst. 22(5), 1074–1087 (2014). https://doi.org/10.1109/TFUZZ.2013.2280141 M.E. Celebi, H.A. Kingravi, P.A. Vela, A comparative study of efficient initialization methods for the k-means clustering algorithm. Expert Syst. Appl. 40(1), 200–210 (2013). https://doi.org/10.1016/j.eswa.2012.07.021 J. Wu, H. Liu, H. Xiong, J. Cao, J. Chen, K-means based consensus clustering: a unified view. IEEE Trans. Knowl. Data Eng. 27(1), 155–169 (2015). https://doi.org/10.1109/TKDE.2014.2316512 J. Saha, J. Mukherjee, Cnak: cluster number assisted k-means. Pattern Recognit. 110, 107625 (2021). https://doi.org/10.1016/j.patcog.2020.107625 Y. Zhang, K. Tangwongsan, S. Tirthapura, Fast streaming k-means clustering with coreset caching. IEEE Trans. Knowl. Data Eng. 34, 2740–2754 (2022). https://doi.org/10.1109/TKDE.2020.3018744 F.D. Bortoloti, E. de Oliveira, P.M. Ciarelli, Supervised kernel density estimation k-means. Expert Syst. Appl. 168, 114350 (2021). https://doi.org/10.1016/j.eswa.2020.114350 R. Mehmood, G. Zhang, R. Bie, H. Dawood, H. Ahmad, Clustering by fast search and find of density peaks via heat diffusion. Neurocomputing 208, 210–217 (2016). https://doi.org/10.1016/j.neucom.2016.01.102 A. Rodriguez, A. Laio, Clustering by fast search and find of density peaks. Science 344(6191), 1492–1496 (2014). https://doi.org/10.1126/science.1242072 X. Xu, S. Ding, Z. Shi, An improved density peaks clustering algorithm with fast finding cluster centers. Knowl.-Based Syst. 158, 65–74 (2018). https://doi.org/10.1016/j.knosys.2018.05.034 Z. Li, Y. Tang, Comparative density peaks clustering. Expert Syst. Appl. 95, 236–247 (2018). https://doi.org/10.1016/j.eswa.2017.11.020 M.-S. Yang, C.-Y. Lai, C.-Y. Lin, A robust EM clustering algorithm for Gaussian mixture models. Pattern Recognit. 45(11), 3950–3961 (2012). https://doi.org/10.1016/j.patcog.2012.04.031 M. Ester, H.-P. Kriegel, J. Sander, X. Xu, A density-based algorithm for discovering clusters in large spatial databases with noise, in Proceedings of the Second International Conference on Knowledge Discovery and Data Mining. KDD’96 (AAAI Press, Menlo Park, 1996), pp. 226–231 E. Schubert, J. Sander, M. Ester, H.P. Kriegel, X. Xu, Dbscan revisited, revisited: why and how you should (still) use dbscan. ACM Trans. Database Syst. 42(3), 19 (2017). https://doi.org/10.1145/3068335 K. Mahesh Kumar, A. Rama Mohan Reddy, A fast dbscan clustering algorithm by accelerating neighbor searching using groups method. Pattern Recognit. 58, 39–48 (2016). https://doi.org/10.1016/j.patcog.2016.03.008 D. Luchi, A. Loureiros Rodrigues, F. Miguel Varejão, Sampling approaches for applying dbscan to large datasets. Pattern Recognit. Lett. 117, 90–96 (2019). https://doi.org/10.1016/j.patrec.2018.12.010 T. Kohonen, The self-organizing map. Neurocomputing 21(1), 1–6 (1998). https://doi.org/10.1016/S0925-2312(98)00030-7 A. Kobren, N. Monath, A. Krishnamurthy, A. McCallum, A hierarchical algorithm for extreme clustering, in Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. KDD’17 (Association for Computing Machinery, New York, 2017), pp. 255–264. https://doi.org/10.1145/3097983.3098079 T.T. Nguyen, M.T. Dang, A.V. Luong, A.W.-C. Liew, T. Liang, J. McCall, Multi-label classification via incremental clustering on an evolving data stream. Pattern Recognit. 95, 96–113 (2019). https://doi.org/10.1016/j.patcog.2019.06.001 A. Shafeeq, Dynamic clustering of data with modified k-means algorithm, in International Conference on Information and Computer Networks, vol. 27 (2012). https://doi.org/10.13140/2.1.4972.3840 E. Lughofer, A dynamic split-and-merge approach for evolving cluster models. Evolv. Syst. 3, 135–151 (2012). https://doi.org/10.1007/s12530-012-9046-5 M.M. Black, R.J. Hickey, The use of time stamps in handling latency and concept drift in online learning. Evolv. Syst. 3, 203–220 (2012). https://doi.org/10.1007/s12530-012-9055-4 L. Zheng, Improved K-means clustering algorithm based on dynamic clustering. Int. J. Adv. Res. Big Data Manag. Syst. 4, 17–26 (2019). https://doi.org/10.21742/IJARBMS.2020.4.1.02 H.-J. Li, Z. Bu, Z. Wang, J. Cao, Dynamical clustering in electronic commerce systems via optimization and leadership expansion. IEEE Trans. Ind. Inform. 16(8), 5327–5334 (2020). https://doi.org/10.1109/TII.2019.2960835 F. Bernstein, S. Modaresi, D. Sauré, A dynamic clustering approach to data-driven assortment personalization. Manag. Sci. 65(5), 2095–2115 (2019). https://doi.org/10.1287/mnsc.2018.3031 I. Khan, Z. Luo, J.Z. Huang, W. Shahzad, Variable weighting in fuzzy k-means clustering to determine the number of clusters. IEEE Trans. Knowl. Data Eng. 32(9), 1838–1853 (2020). https://doi.org/10.1109/TKDE.2019.2911582 P. Guo, C.L.P. Chen, M.R. Lyu, Cluster number selection for a small set of samples using the Bayesian Ying-Yang model. IEEE Trans. Neural Netw. 13(3), 757–763 (2002). https://doi.org/10.1109/TNN.2002.1000144 Y. Yao, Y. Li, B. Jiang, H. Chen, Multiple kernel k-means clustering by selecting representative kernels. IEEE Trans. Neural Netw. Learn. Syst. 32, 4983–4996 (2021). https://doi.org/10.1109/TNNLS.2020.3026532 X.-F. Wang, D.-S. Huang, A novel density-based clustering framework by using level set method. IEEE Trans. Knowl. Data Eng. 21(11), 1515–1531 (2009). https://doi.org/10.1109/TKDE.2009.21 D. Huang, C.-D. Wang, J.-H. Lai, Locally weighted ensemble clustering. IEEE Trans. Cybern. 48(5), 1460–1473 (2018). https://doi.org/10.1109/TCYB.2017.2702343 E. Min, X. Guo, Q. Liu, G. Zhang, J. Cui, J. Long, A survey of clustering with deep learning: from the perspective of network architecture. IEEE Access 6, 39501–39514 (2018). https://doi.org/10.1109/ACCESS.2018.2855437 L. Yang, W. Fan, N. Bouguila, Clustering analysis via deep generative models with mixture models. IEEE Trans. Neural Netw. Learn. Syst. 33, 340–350 (2022). https://doi.org/10.1109/TNNLS.2020.3027761 N. Monath, A. Kobren, A. Krishnamurthy, M.R. Glass, A. McCallum, Scalable hierarchical clustering with tree grafting, in Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. KDD’19 (Association for Computing Machinery, New York, 2019), pp. 1438–1448. https://doi.org/10.1145/3292500.3330929 H. Xie, P. Li, A density-based evolutionary clustering algorithm for intelligent development. Eng. Appl. Artif. Intell. 104, 104396 (2021). https://doi.org/10.1016/j.engappai.2021.104396 Z. Yu, P. Luo, J. You, H.-S. Wong, H. Leung, S. Wu, J. Zhang, G. Han, Incremental semi-supervised clustering ensemble for high dimensional data clustering. IEEE Trans. Knowl. Data Eng. 28(3), 701–714 (2016). https://doi.org/10.1109/TKDE.2015.2499200 H. Yu, J. Lu, G. Zhang, Online topology learning by a Gaussian membership-based self-organizing incremental neural network. IEEE Trans. Neural Netw. Learn. Syst. 31(10), 3947–3961 (2020). https://doi.org/10.1109/TNNLS.2019.2947658 D.A. Berg, Y. Su, D. Jimenez-Cyrus, A. Patel, N. Huang, D. Morizet, S. Lee, R. Shah, F.R. Ringeling, R. Jain, J.A. Epstein, Q.-F. Wu, S. Canzar, G.-L. Ming, H. Song, A.M. Bond, A common embryonic origin of stem cells drives developmental and adult neurogenesis. Cell 177(3), 654–66815 (2019). https://doi.org/10.1016/j.cell.2019.02.010 A. Fahad, N. Alshatri, Z. Tari, A. Alamri, I. Khalil, A.Y. Zomaya, S. Foufou, A. Bouras, A survey of clustering algorithms for big data: taxonomy and empirical analysis. IEEE Trans. Emerg. Top. Comput. 2(3), 267–279 (2014). https://doi.org/10.1109/TETC.2014.2330519 S. Łukasik, P.A. Kowalski, M. Charytanowicz, P. Kulczycki, Clustering using flower pollination algorithm and Calinski-Harabasz index, in 2016 IEEE Congress on Evolutionary Computation (CEC) (2016), pp. 2724–2728. https://doi.org/10.1109/CEC.2016.7744132 A. Strehl, J. Ghosh, Cluster esembles—a knowledge reuse framework for combining multiple partitions. J. Mach. Learn. Res. 3, 583–617 (2003). https://doi.org/10.1162/153244303321897735 S. Chakraborty, N.K. Nagwani, Analysis and study of incremental k-means clustering algorithm. Commun. Comput. Inf. Sci. 169, 338–341 (2011). https://doi.org/10.1007/978-3-642-22577-2_46 L. Dey, S. Chakraborty, N.K. Nagwani, Performance comparison of incremental k-means and incremental DBSCAN algorithms. Comput. Sci. 27(11), 14–18 (2013). http://doi.org/10.5120/3346-4611 B. Fritzke, A growing neural gas network learns topologies, in Proceedings of the 7th International Conference on Neural Information Processing Systems (1994), pp. 625–632 S. Marsland, J. Shapiro, U. Nehmzow, A self-organising network that grows when required. Neural Netw. 15, 1041–1058 (2002). https://doi.org/10.1016/S0893-6080(02)00078-3 S. Furao, O. Hasegawa, An incremental network for on-line unsupervised classification and topology learning. Neural Netw. 19(1), 90–106 (2016). https://doi.org/10.1016/j.neunet.2005.04.006