Data Clustering in Data Lakes Using the BRKGA Metaheuristic

Vol 57, 2025 - 340328
Complete Articles (CA)
Favorite this paper
How to cite this paper?
Abstract

Data Lakes have been gaining prominence in the database field due to their ability to store large amounts of data in a flexible and scalable manner. Typically, clustering methods are used to explore data in Data Lakes, given their unsupervised nature. This work proposes the application of the Biased Random-Key Genetic Algorithm (BRKGA) to the automatic clustering problem in Data Lakes. The approach was evaluated on 17 Data Lakes and compared to the reference algorithm RÓMULO. Both clustering quality and computational efficiency were analyzed. For the quality analysis, the Silhouette, Davies-Bouldin, and Calinski-Harabasz metrics were used to verify the compactness and separation of the obtained clusters. BRKGA outperformed RÓMULO in both analyses, yielding consistent values across the evaluation metrics and a shorter execution time.

Share your ideas or questions with the authors!

Did you know that the greatest stimulus in scientific and cultural development is curiosity? Leave your questions or suggestions to the author!

Sign in to interact

Have a question or suggestion? Share your feedback with the authors!

Institutions
  • 1 Universidade Federal de Alagoas
  • 2 Universidade Federal de Alagoas | (Universidade Federal de Alagoas)
  • 3 Federal Institute of Education, Science and Technology Alagoas
Track
  • AS&DS – Data Analysis and Science
Keywords
Data Lakes
Automatic Clustering
Genetic Algorithms