Related Experiment Video
Updated: Jun 16, 2025

Global and Current Research Trends of Single-Cell Sequencing in Cancer: A Bibliometric and Visualization Study
Published on: April 18, 2025
Patent data-driven analysis of literature associations with changing innovation trends
Adrian Sven Geissler1, Jan Gorodkin1, Stefan Ernst Seemann1
1Center for non-coding RNA in Technology and Health, Department of Veterinary and Animal Sciences, University of Copenhagen, Frederiksberg, Denmark.
Abstract:
Patents are essential for transferring scientific discoveries to meaningful products that benefit societies. While the academic community focuses on the number of citations to rank scholarly works according to their "scientific merit," the number of citations is unrelated to the relevance for patentable innovation. To explore associations between patents and scholarly works in publicly available patent data, we propose to utilize statistical methods that are commonly used in biology to determine gene-disease associations. We illustrate their usage on patents related to biotechnological trends of high relevance for food safety and ecology, namely the CRISPR-based gene editing technology (>60,000 patents) and cyanobacterial biotechnology (>33,000 patents). Innovation trends are found through their unexpected large changes of patent numbers in a time-series analysis. From the total set of scholarly works referenced by all investigated patents (~254,000 publications), we identified ~1,000 scholarly works that are statistical significantly over-represented in the references of patents from changing innovation trends that concern immunology, agricultural plant genomics, and biotechnological engineering methods. The detected associations are consistent with the technical requirements of the respective innovations. In summary, the presented data-driven analysis workflow can identify scholarly works that were required for changes in innovation trends, and, therefore, is of interest for researches that would like to evaluate the relevance of publications beyond the number of citations.
Related Concept Videos
Outliers and Influential Points
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Correlation and Regression
Trends in Lattice Energy: Ion Size and Charge
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.

