To cite this paper use one of the standards below:
This paper proposes to combine Information Theory metrics and learning models to identify languages in texts. The metrics used are Shannon Entropy, Statistical Complexity, and Fisher Information, extracted from the symbolic representation of Bandt–Pompe over the time series of UTF-8 codes. We used 29 thousand texts from Wikipedia, distributed in 29 languages (one thousand texts per language). Initially, we reproduced the method of Hassanpour et al. [2021], converting the texts into time series of UTF-8 codepoints, grouped into six clusters based on the average of the codes. Next, we extracted 32 features of energy per signal and trained a multilayer neural network per cluster, in 10 cycles, with 1, 5, 7, and 12 artificial spaces after two consecutive words. We extended the pipeline with informational features, adding to the input vectors the centroids per language in the Entropy–Complexity and Shannon–Fisher planes. obtaining an overall average of 76.01% and accuracy of 78.18% in 12 spaces.
With nearly 200,000 papers published, Galoá empowers scholars to share and discover cutting-edge research through our streamlined and accessible academic publishing platform.
Learn more about our products:
This proceedings is identified by a DOI , for use in citations or bibliographic references. Attention: this is not a DOI for the paper and as such cannot be used in Lattes to identify a particular work.
Check the link "How to cite" in the paper's page, to see how to properly cite the paper