To cite this paper use one of the standards below:
An Efficient Coding Technique for Stochastic Processes
Jesus Enrique Garcia
University of Campinas
Now you could share with me your questions, observations and congratulations
Create a topicIn the framework of coding theory, under the assumption of a Markov process (X_t) on a finite alphabet A, the compressed representation of the data will be composed of a description of the model used to code the data and the encoded data. Given the model, the Huffman algorithm is optimal for the number of bits needed to encode the data, see [1]. On the other hand, modeling (X_t) through a Partition Markov Model (PMM) - see [2] - promotes a reduction in the number of transition probabilities needed to define the model. This paper shows how the use of Huffman code with a PMM reduces the number of bits needed in this process. We prove the estimation of a PMM allows estimating the entropy of (X_t), providing an estimator of the minimum expected codeword length per symbol. We show the efficiency of the new methodology on a simulation study and, through a real problem of compression of DNA sequences of SARS-CoV-2, obtaining in the real data at least a reduction of 10.4%.
[1] Cover TM. Elements of Information Theory. Wiley Series in Telecommunications and Signal Processing. Wiley-Interscience; 2006.
[2] García JE, González-López VA (2017) Consistent Estimation of Partition Markov Models. Entropy, 19(4): 160. https://doi.org/10.3390/e19040160
With nearly 200,000 papers published, Galoá empowers scholars to share and discover cutting-edge research through our streamlined and accessible academic publishing platform.
Learn more about our products:
This proceedings is identified by a DOI , for use in citations or bibliographic references. Attention: this is not a DOI for the paper and as such cannot be used in Lattes to identify a particular work.
Check the link "How to cite" in the paper's page, to see how to properly cite the paper