Dissociating model architectures from inference computations
Noor Sajid1,2, Johan Medrano3,4,5
1Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University, Cambridge, USA.
None:
Parr et al., 2025 examines how auto-regressive and deep temporal models differ in their treatment of non-Markovian sequence modelling. Building on this, we highlight the need for dissociating model architectures-i.e., how the predictive distribution factorises-from the computations invoked at inference. We demonstrate that deep temporal computations are mimicked by autoregressive models by structuring context access during iterative inference. Using a transformer trained on next-token prediction, we show that inducing hierarchical temporal factorisation during iterative inference maintains predictive capacity while instantiating fewer computations. This emphasises that processes for constructing and refining predictions are not necessarily bound to their underlying model architectures.
Related Concept Videos
Theories of Dissolution: Diffusion Layer Model
This process starts with a thin layer, saturated with the drug, forming at the interface between the solid and liquid. The solute then diffuses from this layer into the main solution. The Noyes-Whitney equation suggests that the rate of dissolution relies on the diffusion...
Mechanistic Models: Overview of Compartment Models
Multicompartment Models: Overview
These models offer a more comprehensive representation of drug behavior in the body than one-compartment models. They accommodate the complexity of drug distribution,...
Theories of Dissolution: The Danckwerts' Model and Interfacial Barrier Model
Storage
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...


