Characterization and identification of languages in texts based on Information Theory

Vol 57, 2025 - 340273
Complete Articles (CA)
Favorite this paper
How to cite this paper?
Abstract

This paper proposes to combine Information Theory metrics and learning models to identify languages in texts. The metrics used are Shannon Entropy, Statistical Complexity, and Fisher Information, extracted from the symbolic representation of Bandt–Pompe over the time series of UTF-8 codes. We used 29 thousand texts from Wikipedia, distributed in 29 languages (one thousand texts per language). Initially, we reproduced the method of Hassanpour et al. [2021], converting the texts into time series of UTF-8 codepoints, grouped into six clusters based on the average of the codes. Next, we extracted 32 features of energy per signal and trained a multilayer neural network per cluster, in 10 cycles, with 1, 5, 7, and 12 artificial spaces after two consecutive words. We extended the pipeline with informational features, adding to the input vectors the centroids per language in the Entropy–Complexity and Shannon–Fisher planes. obtaining an overall average of 76.01% and accuracy of 78.18% in 12 spaces.

Share your ideas or questions with the authors!

Did you know that the greatest stimulus in scientific and cultural development is curiosity? Leave your questions or suggestions to the author!

Sign in to interact

Have a question or suggestion? Share your feedback with the authors!

Institutions
  • 1 Universidade Federal de Alagoas
Track
  • AS&DS – Data Analysis and Science
Keywords
Language identification
Information theory
Signal processing