Related Experiment Videos
Entropy and complexity of finite sequences as fluctuating quantities
Miguel A Jiménez-Montaño1, Werner Ebeling, Thomas Pohl
1Departamento de Física y Matemáticas, Universidad de las Américas/Puebla, Sta. Catarina Mártir, 72820, Puebla, Mexico. jimm@mail.udlap.mx
Bio Systems
|January 5, 2002
Summary
This study analyzes digitized sequences using entropy and complexity measures, revealing consistent results across diverse biological and artificial data. It highlights the fluctuating nature of these measures and their application in understanding sequence randomness.
Area of Science:
- Information theory
- Computational biology
- Statistical analysis
Background:
- Digitized sequences (real numbers, discrete strings) are analyzed using entropy and complexity.
- The random characteristics and fluctuation spectrum of these quantities are examined.
- Applications include neural spike-trains and DNA sequences.
Purpose of the Study:
- To analyze digitized sequences by concepts of entropy and complexity.
- To investigate the random character and fluctuation spectrum of these quantities.
- To demonstrate consistent results from diverse complexity measures across different sequence types.
Main Methods:
- Analysis of digitized sequences using n-gram entropies and context-free grammatical complexity.
- Consideration of sequences as finite realizations of random processes with surrogate sequences and processes.
- Study of fluctuation distributions for entropy and complexity measures.
Main Results:
- N-gram entropies and context-free grammatical complexity are identified as fluctuating quantities.
- Different complexity measures reveal distinct sequence characteristics.
- Despite differing scales, entropy and context-free grammatical complexity provide consistent rankings for biological and artificial sequences.
Conclusions:
- Entropy and context-free grammatical complexity are valuable, albeit different, measures for sequence analysis.
- These measures offer consistent insights into sequence randomness and characteristics.
- The findings have implications for analyzing complex biological and artificial data sequences.