Related Experiment Video
Updated: Jan 9, 2026

Author Spotlight: UAV Remote Sensing for Efficient Invasive Plant Biomass Estimation
Published on: February 9, 2024
A New Streaming K-Nearest Neighbor Algorithm for Status Prediction in Block-Sparse, Autocorrelated, Irregular
Xin Zhao1, Xiaokai Nie2,3,4, Yu Zhao5
1School of Mathematics, Southeast University, Nanjing, People's Republic of China.
None:
In streaming longitudinal data, status prediction becomes challenging when input variables are block-sparse, autocorrelated, and irregular in both dimension and distribution. General methods cannot model such data directly, especially when the classes are extremely imbalanced. This research proposes a K-Nearest Neighbor (KNN) algorithm where distance is measured by Kullback-Leibler (KL) divergence. The algorithm uses features extracted from metric conditional density, both with and without first-order lag. The developed streaming KNN algorithm is further applied to simulation data. Results show that when differences originate from the location hyperparameters of the Gaussian distribution or both the shape and scale hyperparameters of the inverse gamma distribution, the method performs quite well, as expected, with an AUC close to 1. Additionally, a numerical method is proposed for general distributions that lack an analytical expression in real data. This method is applied to a big medical streaming dataset with similar properties. Results indicate that the AUC value gradually increases to 0.913, with a sensitivity of 0.851 and a specificity of 0.816.
Related Concept Videos
End Point Prediction: Gran Plot
For potentiometric titration, the Gran plot is created by plotting...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...