Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Survival Tree01:19

Survival Tree

Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Ā Building a Survival Tree
Constructing a survival tree begins...
Contingency Table01:29

Contingency Table

A contingency table provides a way of portraying data that can facilitate calculating probabilities. It is a method of displaying a frequency distribution as a table with rows and columns to show how two variables may be dependent (contingent) upon each other; The table helps determine conditional probabilities quite quickly and can help systematically organize, analyze and quantify data. The table displays sample values concerning two variables that may be dependent or contingent on one...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Responsible AI in mental healthcare: policy directions and stakeholder insights.

Frontiers in public healthĀ·2026
Same author

Adversarial Defense without <i>Adversarial Defense</i>: Enhancing Language Model Robustness via Instance-level Principal Component Removal.

Transactions of the Association for Computational LinguisticsĀ·2025
Same authorSame journal

Tabular context-aware optical character recognition and tabular data reconstruction for historical records.

International journal on document analysis and recognition (Online)Ā·2025
Same author

Identifying key challenges and needs in digital mental health moderation practices supporting users exhibiting risk behaviours to develop responsible AI tools: the case study of Kooth.

SN social sciencesĀ·2022

Related Experiment Video

Updated: Jun 24, 2026

Generation of Comprehensive Thoracic Oncology Database - Tool for Translational Research
11:18

Generation of Comprehensive Thoracic Oncology Database - Tool for Translational Research

Published on: January 22, 2011

Data rescue of historical tables through semi-supervised table structure recognition.

Loitongbam Gyanendro Singh1, Stuart E Middleton1

  • 1School of Electronics and Computer Science, University of Southampton, Southampton, UK.

International Journal on Document Analysis and Recognition (Online)
|June 23, 2026
PubMed
Summary

Semi-supervised learning enhances tabular structure recognition for historical documents, reducing annotation needs and improving accuracy. This approach aids in digitizing archives and preserving historical data.

Keywords:
Data Annotation TechniquesHistorical Document DigitizationSemi-Supervised LearningTabular Structure Recognition

More Related Videos

Databases to Efficiently Manage Medium Sized, Low Velocity, Multidimensional Data in Tissue Engineering
09:43

Databases to Efficiently Manage Medium Sized, Low Velocity, Multidimensional Data in Tissue Engineering

Published on: November 22, 2019

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
12:27

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations

Published on: February 15, 2017

Related Experiment Videos

Last Updated: Jun 24, 2026

Generation of Comprehensive Thoracic Oncology Database - Tool for Translational Research
11:18

Generation of Comprehensive Thoracic Oncology Database - Tool for Translational Research

Published on: January 22, 2011

Databases to Efficiently Manage Medium Sized, Low Velocity, Multidimensional Data in Tissue Engineering
09:43

Databases to Efficiently Manage Medium Sized, Low Velocity, Multidimensional Data in Tissue Engineering

Published on: November 22, 2019

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
12:27

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations

Published on: February 15, 2017

Area of Science:

  • Computer Science
  • Digital Humanities
  • Information Retrieval

Background:

  • Tabular Structure Recognition (TSR) is vital for digitizing historical documents.
  • Challenges include document degradation and inconsistent handwriting, complicating annotation for model training.
  • Existing methods often require extensive labeled data, which is costly and time-consuming to produce for historical archives.

Purpose of the Study:

  • To investigate if semi-supervised learning can reduce the need for expensive data annotations in TSR.
  • To determine if semi-supervised training enhances model robustness for historical document analysis.
  • To improve the accessibility and analysis of historical tabular data through efficient digitization.

Main Methods:

  • A novel semi-supervised learning framework was developed and applied.
  • The CascadeTabNet model was employed for Tabular Structure Recognition.
  • Methodology was tested on historical (GloSAT, ICDAR-2019) and modern (PubTabNet) document datasets.

Main Results:

  • Semi-supervised learning significantly increased TSR accuracy across datasets.
  • The approach reduced dependency on large amounts of labeled data.
  • Improved model robustness was observed, particularly for historical documents.

Conclusions:

  • Semi-supervised learning offers a robust and efficient solution for large-scale digitization of historical documents.
  • This framework enhances the preservation and accessibility of valuable historical data.
  • The study demonstrates the effectiveness of semi-supervised approaches in overcoming annotation challenges in archival research.