Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Archival Research01:40

Archival Research

16.4K
Some researchers gain access to large amounts of data without interacting with a single research participant. Instead, they use existing records to answer various research questions. This type of research approach is known as archival research. Archival research relies on looking at past records or data sets to look for interesting patterns or relationships. For example, a researcher might access the academic records of all individuals who enrolled in college within the past ten years and...
16.4K
Proteomics01:33

Proteomics

7.9K
A proteome is the entire set of proteins that a cell type produces. We can study proteomes using the knowledge of genomes because genes code for mRNAs, and the mRNAs encode proteins. Although mRNA analysis is a step in the right direction, not all mRNAs are translated into proteins.
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
7.9K
Surveys02:16

Surveys

15.3K
Often, psychologists develop surveys as a means of gathering data. Surveys are lists of questions to be answered by research participants, and can be delivered as paper-and-pencil questionnaires, administered electronically, or conducted verbally. Generally, the survey itself can be completed in a short time, and the ease of administering a survey makes it easy to collect data from a large number of people.
15.3K
NMR Spectroscopy of Aromatic Compounds01:14

NMR Spectroscopy of Aromatic Compounds

5.0K
Aromatic compounds can be identified or analyzed using proton NMR and carbon‐13 NMR. Typically, aromatic hydrogens or hydrogens directly bonded to the aromatic rings are strongly deshielded by the aromatic ring current. Therefore, they absorb in the range of 6.5–8.0 ppm in proton NMR spectra. For instance, aromatic hydrogens directly bonded to the benzene ring absorb at 7.3 ppm. However, aromatic hydrogens of larger rings absorb farther upfield or downfield than the ideal range.
5.0K
Data Collection by Observations01:08

Data Collection by Observations

12.8K
Data collection refers to a systematic way of obtaining, observing, measuring, and analyzing accurate information. Observational studies are one of the most widely used methods of data collection. It involves collecting data by observing the behavior and physical characteristics of a sample without making any modifications to the sample.
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
12.8K
Data Collection I01:30

Data Collection I

6.6K
Data collection gathers information needed to make accurate judgments about a patient's present condition. During a health history interview, subjective data is collected from the patient, their caregivers, or family members, and objective data is collected through observations and physical assessment. Patients are the primary source of subjective data. Thus information gathered from patients through interviews, observations, and physical examination is primary data. Secondary sources of...
6.6K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Benzimidazole derived hydrazone Schiff bases as potent cholinesterase inhibitors: synthesis, <i>in vitro</i> and <i>in silico</i> approaches.

Future medicinal chemistry·2026
Same author

Design, Synthesis, Computational Studies, and Antidiabetic Evaluation of Hydrazide Derivative: In Vitro, In Vivo and In Silico Investigation.

Chemistry & biodiversity·2026
Same author

RETRACTED: Sabeela et al. Reactive Mesoporous pH-Sensitive Amino-Functionalized Silica Nanoparticles for Efficient Removal of Coomassie Blue Dye. <i>Nanomaterials</i> 2019, <i>9</i>, 1721.

Nanomaterials (Basel, Switzerland)·2025
Same author

Synthesis, spectroscopic properties, structural characterization, and computational studies of a new non-centrosymmetric hybrid compound: bis-cyanato-<i>N</i> chromium(iii) <i>meso</i>-arylporphyrin complex.

RSC advances·2025
Same author

Interpretation of plasma pharmacokinetics and urinary excretion of phenolic metabolites and catabolites derived from (poly)phenols following ingestion of mango by ileostomists and subjects with a full gastrointestinal tract: complications associated with endogenous and ingested phenylalanine and tyrosine.

Food & function·2025
Same author

Retraction of "Surface Activity of Smart Hybrid Polysiloxane-<i>co</i>-<i>N</i>‑isopropylacrylamide Microgels".

ACS omega·2025

Related Experiment Video

Updated: Sep 10, 2025

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
04:04

Asthma Detection Research Based on Voice Signal Processing and Machine Learning

Published on: July 22, 2025

405

Open source Arabic research paper dataset for natural language processing.

Tahani M Almutairi1, Shireen R Saifuddin2, Reem M Alotaibi2

  • 1Department of Information Technology, King Abdulaziz University, Jeddah, Saudi Arabia. tjuhaidelalmutairi@stu.kau.edu.sa.

Scientific Reports
|August 27, 2025
PubMed
Summary

A new Arabic Research Papers Dataset (ARPD) was created to support academic research and natural language processing (NLP) tasks. The ARPD dataset shows promising results for classification and clustering, outperforming traditional methods.

Keywords:
Academic corpusAcademic datasetArabic corpusArabic databaseArabic datasetArabic retrievalResearch paper

More Related Videos

Comparing Bibliometric Analysis Using PubMed, Scopus, and Web of Science Databases
05:02

Comparing Bibliometric Analysis Using PubMed, Scopus, and Web of Science Databases

Published on: October 24, 2019

32.0K
A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
07:50

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts

Published on: September 20, 2018

16.0K

Related Experiment Videos

Last Updated: Sep 10, 2025

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
04:04

Asthma Detection Research Based on Voice Signal Processing and Machine Learning

Published on: July 22, 2025

405
Comparing Bibliometric Analysis Using PubMed, Scopus, and Web of Science Databases
05:02

Comparing Bibliometric Analysis Using PubMed, Scopus, and Web of Science Databases

Published on: October 24, 2019

32.0K
A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
07:50

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts

Published on: September 20, 2018

16.0K

Area of Science:

  • Natural Language Processing (NLP)
  • Applied Linguistics
  • Data Mining
  • Information Retrieval
  • Machine Translation

Background:

  • Existing Arabic corpora are primarily sourced from social media or news, creating a gap for academic research.
  • There is a growing need for specialized datasets to advance applications like NLP and machine translation.
  • The Arabic Research Papers Dataset (ARPD) is introduced to fill this void.

Purpose of the Study:

  • To introduce and describe the methodology behind the creation of the ARPD.
  • To provide a specialized, publicly available dataset for Arabic academic research.
  • To evaluate the dataset's utility in classification and clustering tasks.

Main Methods:

  • The ARPD was constructed with seven distinct classes.
  • The dataset is available in multiple formats to maximize accessibility for researchers.
  • Experiments involved evaluating classical clustering algorithms, bio-inspired algorithms (PSO, GWO), and classification algorithms (SVM).

Main Results:

  • Bio-inspired algorithms (PSO, GWO) outperformed classical clustering algorithms on the ARPD, as measured by the Davies-Bouldin index.
  • The Support Vector Machine (SVM) algorithm achieved the highest accuracy in classification tasks (89%-99%).
  • Other classification algorithms also demonstrated high performance on the ARPD.

Conclusions:

  • The ARPD is a valuable resource for advancing Arabic academic research.
  • The dataset facilitates the development and evaluation of advanced NLP models.
  • The findings underscore the potential of specialized corpora for improving AI applications in Arabic.