Clustering applications of IFDBSCAN algorithm with comparative analysis


Unver M., Erginel N.

Journal of Intelligent and Fuzzy Systems, cilt.39, sa.5, ss.6099-6108, 2020 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 39 Sayı: 5
  • Basım Tarihi: 2020
  • Doi Numarası: 10.3233/jifs-189082
  • Dergi Adı: Journal of Intelligent and Fuzzy Systems
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Academic Search Premier, Aerospace Database, Applied Science & Technology Source, Business Source Elite, Business Source Premier, Communication Abstracts, Compendex, Computer & Applied Sciences, INSPEC, Metadex, zbMATH, Civil Engineering Abstracts
  • Sayfa Sayıları: ss.6099-6108
  • Anahtar Kelimeler: clustering, clustering validation indices, DBSCAN, IFDBSCAN, intuitionistic fuzzy sets, Unsupervised machine learning
  • Maltepe Üniversitesi Adresli: Hayır

Özet

Density Based Spatial Clustering of Application with Noise (DBSCAN) is one of the mostly preferred algorithm among density based clustering approaches in unsupervised machine learning, which uses epsilon neighborhood construction strategy in order to discover arbitrary shaped clusters. DBSCAN separates dense regions from low density regions and simultaneously assigns points that lie alone as outliers to unearth the hidden cluster patterns in the datasets. DBSCAN identifies dense regions by means of core point definition, detection of which are strictly dependent on input parameter definitions: ϵ is distance of the neighborhood or radius of hypersphere and MinPts is minimum density constraint inside ϵ radius hypersphere. Contrarily to classical DBSCAN's crisp core point definition, intuitionistic fuzzy core point definition is proposed in our preliminary work to make DBSCAN algorithm capable of detecting different patterns of density by two different combinations of input parameters, particularly is a necessity for the density varying large datasets in multidimensional feature space. In this study, preliminarily proposed DBSCAN extension is studied: IFDBSCAN. The proposed extension is tested by computational experiments on several machine learning repository real-time datasets. Results show that, IFDBSCAN is superior to classical DBSCAN with respect to external internal performance indices such as purity index, adjusted rand index, Fowlkes-Mallows score, silhouette coefficient, Calinski-Harabasz index and with respect to clustering structure results without increasing computational time so much, along with the possibility of trying two different density patterns on the same run and trying intermediary density values for the users by manipulating α margin.