Related Experiment Video
Updated: Nov 27, 2025

Measuring Attention and Visual Processing Speed by Model-based Analysis of Temporal-order Judgments
Published on: January 23, 2017
Estimating Predictive Rate-Distortion Curves via Neural Variational Inference.
Michael Hahn1, Richard Futrell2
1Department of Linguistics, Stanford University, Stanford, CA 94305, USA.
We introduce Neural Predictive Rate-Distortion (NPRD), a new method to estimate the Predictive Rate-Distortion curve for complex processes like natural language. NPRD scales effectively, offering improved bounds over existing techniques.
Area of Science:
- Information Theory
- Machine Learning
- Computational Linguistics
Background:
- The Predictive Rate-Distortion (PRD) curve measures the trade-off between past information compression and future prediction accuracy.
- Current PRD estimation methods struggle with complex processes like natural language due to large alphabets and unknown causal states.
- Existing methods based on clustering or known causal states lack scalability.
Purpose of the Study:
- To develop a scalable method for estimating the PRD curve for complex stochastic processes.
- To leverage neural networks for estimating PRD curves where analytical solutions are intractable.
- To provide improved PRD bounds for natural language processing.
Main Methods:
- Introduced Neural Predictive Rate-Distortion (NPRD), a novel estimation technique.
- Utilized the universal approximation capabilities of neural networks.
- Computed a variational bound on the PRD curve using only time series data.
Main Results:
- NPRD demonstrates scalability for processes with large alphabets and long dependencies.
- The method was validated on processes with analytically known PRD curves.
- NPRD provided improved PRD bounds for natural language compared to sequence clustering.
Conclusions:
- NPRD offers a scalable approach to estimating the PRD curve for complex systems.
- The PRD curve is a more effective characterization tool than statistical complexity for highly complex processes like natural language.
- This work advances the application of information theory to natural language and other complex data.
Related Concept Videos
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
Linear Approximation in Time Domain
For a simple pendulum with a mass evenly distributed along its length and the center of mass located at half the pendulum's length,...
Neural Regulation
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Calibration Curves: Linear Least Squares
For data that follow a straight line, the standard method for fitting is the linear...
Variability: Analysis
The range is a simple measure of variability, indicating the difference between the highest and...
