Related Experiment Video
Updated: Sep 15, 2025

Comparing Bibliometric Analysis Using PubMed, Scopus, and Web of Science Databases
Published on: October 24, 2019
Addressing the Problem of Hard-to-Reach Unpublished Data from Theses in University Repositories
Héctor L Venegas-Quiñones1, Pablo A Garcia-Chevesich1, Madeleine Guillen2
1Department of Civil and Environmental Engineering, Colorado School of Mines, Golden, CO.
Abstract:
Researchers frequently encounter challenges in accessing valuable data encapsulated within university theses, which are predominantly archived in PDF format and remain unpublished in repositories. These documents often encompass original research, including vital environmental and hydrological data, yet they pose difficulties for searching or analysis due to inconsistent formatting and inefficient repository search tools such as keyword searches, which lead to an overwhelming list of documents. Our research team, engaged in developing a groundwater database for the Arequipa region of Peru, encountered this issue directly, with numerous relevant theses dispersed across local university repositories. The manual review process proved excessively time-consuming, necessitating the development of an innovative, automated solution. Our multi-step methodology commenced with optical character recognition (OCR) and Python scripts for keyword scoring, followed by the employment of Large Language Models (LLMs), notably Google's Gemini and the locally hosted Ollama, to semantically analyze content. This facilitated the identification and extraction of pertinent data (e.g., water quality parameters, well locations) and its organization into usable formats such as Excel spreadsheets; subsequent manual checks confirmed a high level of accuracy. The final system enables users to query an extensive number of documents swiftly and contextually, effectively overcoming traditional keyword search limitations. The tool is presently being disseminated among local researchers and institutions, offering a robust solution for accessing and managing regional groundwater data. This methodology possesses the potential for global scaling and adaptation, thereby enhancing access to gray literature and expediting scientific discovery across various disciplines.
Related Concept Videos
Archival Research
Types of Records II: Educational and Administrative Records
Methods of Documentation I: Source-Oriented Records
In an SOR, each discipline involved in patient care maintains a separate medical record section. This record-keeping method enables easy tracking of patient progress and ensures healthcare staff have access to up-to-date information.
Key Attributes include the following:
X-ray Diffraction of Biological Samples
According to Bragg's law, when X-rays strike the sample positioned on a stage, the rays are scattered by the electron clouds around the sample atoms. The X-ray diffraction or scattering is caused by constructive interference of the X-ray waves that reflect off the internal...
Data Collection II
Systematic Error: Methodological and Sampling Errors
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...

