A novel grid-based clustering algorithm
Authors:
- Artur Starczewski,
- Magdalena M. Scherer,
- Wojciech Książek,
- Maciej Dębski,
- Lipo Wang
Abstract
Data clustering is an important method used to discover naturally occurring structures in datasets. One of the most popular approaches is the grid-based concept of clustering algorithms. This kind of method is characterized by a fast processing time and it can also discover clusters of arbitrary shapes in datasets. These properties allow these methods to be used in many different applications. Researchers have created many versions of the clustering method using the grid-based approach. However, the key issue is the right choice of the number of grid cells. This paper proposes a novel grid-based algorithm which uses a method for an automatic determining of the number of grid cells. This method is based on the kdist function which computes the distance between each element of a dataset and its kth nearest neighbor. Experimental results have been obtained for several different datasets and they confirm a very good performance of the newly proposed method.
- Record ID
- CUT47902c131be24160a9e1289cb7efd3f7
- Publication categories
- ;
- Author
- Journal series
- Journal of Artificial Intelligence and Soft Computing Research, ISSN 2083-2567, e-ISSN 2449-6499
- Issue year
- 2021
- Vol
- 11
- No
- 4
- Pages
- 319-330
- Other elements of collation
- rys.; tab.; wykr.; Bibliografia (na s.) - 328-329; Bibliografia (liczba pozycji) - 28; Oznaczenie streszczenia - Abstr.; Numeracja w czasopiśmie - Vol. 11, No. 4
- Keywords in English
- data mining, grid-based clustering, grid structure
- DOI
- DOI:10.2478/jaiscr-2021-0019 Opening in a new tab
- URL
- https://www.sciendo.com/article/10.2478/jaiscr-2021-0019 Opening in a new tab
- Language
- eng (en) English
- License
- Score (nominal)
- 140
- Uniform Resource Identifier
- https://cris.pk.edu.pl/info/article/CUT47902c131be24160a9e1289cb7efd3f7/
- URN
urn:pkr-prod:CUT47902c131be24160a9e1289cb7efd3f7
* presented citation count is obtained through Internet information analysis, and it is close to the number calculated by the Publish or PerishOpening in a new tab system.