Related Experiment Video
Updated: Oct 5, 2026

Robust DNA Isolation and High-throughput Sequencing Library Construction for Herbarium Specimens
Published on: March 8, 2018
Integrating scanned literature and online sources to build a data text collection for plant biodiversity
Nicolas Turenne1,2, Eric Chenin3, Youcef Sklab1
1IRD, Sorbonne Université, UMMISCO, Paris, France IRD, Sorbonne Université, UMMISCO Paris France https://ror.org/02en5vm52.
Background:
Plant taxonomic descriptions are essential for botany, but remain scattered in unstructured texts. Automatic trait extraction is difficult due to linguistic variation, specialised terminology and the need for large, structured datasets.
New Information:
Methods: We developed WordGen, an algorithm for species name recognition tolerant to OCR and typographical errors, using curated botanical dictionaries to reliably identify species sections in noisy and web-scraped texts.Results: We created a 345,000 species file dataset with natural language descriptions of global plant species from monographs, Wikipedia and databases. It achieved over 99% coverage and high recall on New Caledonia and Cameroon Floras datasets, enabling advanced AI-based plant taxonomy analysis.
Related Concept Videos
Plant Breeding and Biotechnology
Non-vascular Seedless Plants
Light Acquisition
Biodiversity and Human Values
Pollination and Flower Structure
Responses to Drought and Flooding

