Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Clinical Trials: Overview01:11

Clinical Trials: Overview

4.7K
Clinical development focuses on how the drug will interact with the human body and encompasses four key phases of clinical trials, each serving a specific purpose in assessing the safety and effectiveness of new drugs. These phases overlap and build upon one another. Phase I involves a small group of healthy volunteers (typically 20-80 individuals) or, in cases where significant toxicity is expected, patients with the targeted disease, such as cancer or AIDS. The volunteers are tested for...
4.7K
Clinical Trials01:16

Clinical Trials

8.5K
Clinical trials are prospective experimental studies conducted on humans to determine the safety and efficacy of treatments, drugs, diet methods, and medical devices. Using statistics in clinical trials enables researchers to derive reasonable and accurate conclusions from the collected data, allowing them to make wise decisions in uncertain situations. In medical research, statistical methods are crucial for preventing errors and bias.
There are four phases in a clinical trial. A phase one...
8.5K
Statistical Software for Data Analysis and Clinical Trials01:12

Statistical Software for Data Analysis and Clinical Trials

1.7K
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
1.7K
Investigation of Disease Outbreaks01:23

Investigation of Disease Outbreaks

74
Multistate foodborne outbreaks pose significant public health risks and require meticulous investigation to identify sources and implement control measures. The Centers for Disease Control and Prevention (CDC) utilizes a dynamic seven-step process for these investigations, integrating data from laboratories, interviews, and environmental assessments to protect public health.Outbreak Detection: The detection of multistate outbreaks typically begins with PulseNet, the CDC's national laboratory...
74
Data Collection by Observations01:08

Data Collection by Observations

11.0K
Data collection refers to a systematic way of obtaining, observing, measuring, and analyzing accurate information. Observational studies are one of the most widely used methods of data collection. It involves collecting data by observing the behavior and physical characteristics of a sample without making any modifications to the sample.
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
11.0K
Statistical Methods for Analyzing Epidemiological Data01:25

Statistical Methods for Analyzing Epidemiological Data

1.3K
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
1.3K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Therapeutic Response to Anti-Vascular Endothelial Growth Factor Retreatment Among Eyes With Stabilized Vision After Being Lost to Follow-Up in Diabetic Macular Edema: A Nationwide, Registry-Based Cohort Study.

American journal of ophthalmology·2026
Same author

Erratum: Efficacy and Safety of Camrelizumab Plus Apatinib in Patients With Refractory Chordoma: A Phase II Clinical Trial.

Journal of clinical oncology : official journal of the American Society of Clinical Oncology·2026
Same author

Patients' response to doctors with different consultation lengths: the moderating role of power distance belief.

Psychology & health·2026
Same author

Anti-GLDN antibody-associated CIDP (nodopathy): transient IVIg response and B-cell depletion remission.

BMC neurology·2026
Same author

Efficacy and Safety of Camrelizumab Plus Apatinib in Patients With Refractory Chordoma: A Phase II Clinical Trial.

Journal of clinical oncology : official journal of the American Society of Clinical Oncology·2026
Same author

Cytochrome P450-catalyzed allylic oxidation of pentalenene to 1-deoxypentalenic acid in pentalenolactone biosynthesis.

Engineering microbiology·2026

Related Experiment Video

Updated: May 1, 2026

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
07:50

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts

Published on: September 20, 2018

15.8K

Using large clinical corpora for query expansion in text-based cohort identification.

Dongqing Zhu1, Stephen Wu2, Ben Carterette1

  • 1Department of Computer and Information Sciences, University of Delaware, 440 Smith Hall, Newark, DE 19716, USA.

Journal of Biomedical Informatics
|April 1, 2014
PubMed
Summary

Using a large clinical corpus for query expansion significantly improves patient cohort identification. Adding the Mayo Clinic corpus enhanced retrieval performance, demonstrating the value of curated clinical data over simply using all available data.

Keywords:
Clinical textCohort identificationElectronic medical recordsInformation retrievalQuery expansion

More Related Videos

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.3K
Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry
05:53

Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry

Published on: June 21, 2018

9.2K

Related Experiment Videos

Last Updated: May 1, 2026

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
07:50

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts

Published on: September 20, 2018

15.8K
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.3K
Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry
05:53

Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry

Published on: June 21, 2018

9.2K

Area of Science:

  • Natural Language Processing
  • Medical Informatics
  • Information Retrieval

Background:

  • Clinical text presents challenges like polysemy, synonymy, and hyponymy, complicating patient cohort identification.
  • Existing information retrieval (IR) methods struggle with the nuances of clinical language.

Purpose of the Study:

  • To investigate if using a large, in-domain clinical corpus for query expansion can improve patient cohort identification.
  • To evaluate the effectiveness of auxiliary collections in IR-based cohort retrieval.

Main Methods:

  • Utilized four auxiliary collections for query expansion in the Text REtrieval Conference (TREC) task.
  • Employed a mixture of relevance models for cohort retrieval from the Pittsburgh NLP Repository.
  • Assessed the impact of collection size, query difficulty, and collection interactions on retrieval performance, measured by mean average precision (MAP).

Main Results:

  • Performance improved over the baseline query likelihood model (MAP=0.373) with any auxiliary resource (MAP=0.386+).
  • Adding the Mayo Clinic collection significantly improved performance (MAP=0.4223 with all four collections).
  • Retrieval gains plateaued after including 2.5 billion term instances from the Mayo Clinic collection, indicating diminishing returns and the importance of data curation.

Conclusions:

  • Query expansion with a large clinical corpus, specifically the Mayo Clinic corpus, consistently and significantly improves IR-based cohort identification.
  • The study highlights that "more data is not necessarily better," emphasizing the value of curated clinical datasets for effective query expansion.
  • Access to large clinical corpora is beneficial for IR query expansion tasks in healthcare settings.