Related Experiment Video
Updated: Jan 19, 2026

Single Cell Multiplex Reverse Transcription Polymerase Chain Reaction After Patch-clamp
Published on: June 20, 2018
Identification of Words in Biological Sequences Under the Semi-Markov Hypothesis
Brenda Ivette Garcia-Maya1, Nikolaos Limnios1
1Laboratory of Applied Mathematics (LMAC), Université de Technologie de Compiègne, Sorbonne University, Compiègne, France.
Abstract:
Identifying a word (pattern) in a long sequence of letters is not an easy task. To achieve this objective, several models have been proposed under the assumption that the sequence of letters is described by a Markov chain. The Markovian hypothesis imposes restrictions on the distribution of the sojourn time in a state, which has geometric distribution in a discrete process. This is the main drawback when applying Markov chains to real problems. By contrast, semi-Markov processes are generalized. In semi-Markov processes, the sojourn time in a state can be governed by any distribution function. The goal of this article is to compute the first hitting time (position) of a word (pattern) in a semi-Markov sequence. To achieve this objective, we use the auxiliary prefix and backward chain. To give an example of the applications of the proposed model, the model is tested in a bacteriophage DNA sequence that is lacking the enzyme SmaI. We compute the probability that a word occurs for the first time after n nucleotides in a DNA sequence. The corresponding probability distribution, the mean waiting position, the variance, and rate of the occurrence of the word are obtained.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Modern Molecular Taxonomy
Maxam-Gilbert Sequencing
Challenges of the Maxam-Gilbert Method
The...
Signal Sequences and Sorting Receptors
Cis-regulatory Sequences

