Related Experiment Video
Updated: Aug 13, 2025

10:44
Single Cell Multiplex Reverse Transcription Polymerase Chain Reaction After Patch-clamp
Published on: June 20, 2018
9.9K
End-to-End Transcript Alignment of 17th Century Manuscripts: The Case of Moccia Code
Giuseppe De Gregorio1, Giuliana Capriolo2, Angelo Marcelli1
1Department of Information and Electrical Engineering and Applied Mathematics, University of Salerno, Via Giovanni Paolo II, 132, 84084 Fisciano, Italy.
Journal of Imaging
|January 20, 2023
Summary
We developed a new learning-free method to automatically align word transcriptions with historical document images. This technique accurately links text to images, aiding scholars and digital processing of historical documents.
Area of Science:
- Digital Humanities
- Computer Vision
- Document Image Analysis
Background:
- Digital libraries increasingly house historical handwritten documents.
- Accurate word-level alignment between transcriptions and document images is crucial for scholarly study and automated processing.
- Existing methods may struggle with the complexities of historical scripts and document layouts.
Purpose of the Study:
- To propose a novel, learning-free method for automatic word-level text-to-image alignment in historical documents.
- To address challenges in segmenting text lines with curved baselines.
- To handle word-level under- and over-segmentation errors during alignment.
Main Methods:
- A learning-free approach for transcription-to-image alignment.
- A line-level segmentation algorithm designed for curved baselines.
- A text-to-image alignment algorithm robust to segmentation errors.
Main Results:
- The line segmentation algorithm achieved 92% accuracy on a 17th-century Italian manuscript.
- The text-to-image alignment achieved an accuracy greater than 68%.
- Performance favorably compares with state-of-the-art methods on widely used datasets.
Conclusions:
- The proposed method offers an effective solution for aligning transcriptions to historical document images.
- The approach is valuable for digital humanities research and automated document analysis.
- The learning-free nature makes it broadly applicable without extensive training data.

