Related Experiment Video
Updated: Jul 17, 2025

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Capturing drug use patterns at a glance: An n-ary word sufficient statistic for repeated univariate categorical
Gabriel J Odom1, Laura Brandt2, Clinton Castro3
1Department of Biostatistics, Florida International University, Miami, FL, United States of America.
Introduction:
The efficacy of treatments for substance use disorders (SUD) is tested in clinical trials in which participants typically provide urine samples to detect whether the person has used certain substances via urine drug screenings (UDS). UDS data form the foundation of treatment outcome assessment in the vast majority of SUD clinical trials. However, existing methods to calculate treatment outcomes are not standardized, impeding comparability between studies and prohibiting reproducibility of results.
Methods:
We extended the concept of a binary UDS variable to multiple categories: "+" [positive for substance(s) of interest], "-" [negative for substance(s)], "o" [patient failed to provide sample], "*" [inconclusive or mixed results], and "_" [no specimens required per study design]. This construct can be used to create a standardized and sufficient representation of UDS datastreams and sufficiently collapses longitudinal records into a single, compact "word", which preserves all information contained in the original data.
Results:
We developed the R software package CTNote (available on CRAN) as a tool to enable computers to parse these "words". The software package contains five groups of routines: detect a substance use pattern, account for a specific trial protocol, handle missing UDS data, measure the longest period of consecutive behavior, and count substance use events. Executing permutations of these routines result in algorithms which can define SUD clinical trial endpoints. As examples, we provide three algorithms to define primary endpoints from seminal SUD clinical trials.
Discussion:
Representing substance use patterns as a "word" allows researchers and clinicians an "at a glance" assessment of participants' responses to treatment over time. Further, machine readable use pattern summaries are a standardized method to calculate treatment outcomes and are therefore useful to all future SUD clinical trials. We discuss some caveats when applying this data summarization technique in practice and areas of future study.
Related Concept Videos
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Drug Concentration Versus Time Correlation
Two pivotal parameters are the minimum effective concentration (MEC) and the minimum toxic concentration (MTC). The MEC is the...
Drug Dependence
Contingency Table
Introduction to Test of Independence
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
Statistical Methods for Analyzing Epidemiological Data

