To cite this paper use one of the standards below:
Data Lakes have been gaining prominence in the database field due to their ability to store large amounts of data in a flexible and scalable manner. Typically, clustering methods are used to explore data in Data Lakes, given their unsupervised nature. This work proposes the application of the Biased Random-Key Genetic Algorithm (BRKGA) to the automatic clustering problem in Data Lakes. The approach was evaluated on 17 Data Lakes and compared to the reference algorithm RÓMULO. Both clustering quality and computational efficiency were analyzed. For the quality analysis, the Silhouette, Davies-Bouldin, and Calinski-Harabasz metrics were used to verify the compactness and separation of the obtained clusters. BRKGA outperformed RÓMULO in both analyses, yielding consistent values across the evaluation metrics and a shorter execution time.
With nearly 200,000 papers published, Galoá empowers scholars to share and discover cutting-edge research through our streamlined and accessible academic publishing platform.
Learn more about our products:
This proceedings is identified by a DOI , for use in citations or bibliographic references. Attention: this is not a DOI for the paper and as such cannot be used in Lattes to identify a particular work.
Check the link "How to cite" in the paper's page, to see how to properly cite the paper