Related Experiment Video
Updated: Jan 22, 2026

Involving Individuals with Developmental Language Disorder and Their Parents/Carers in Research Priority Setting
Published on: June 6, 2020
Scaling laws and representation learning in simple hierarchical languages: Transformers versus convolutional
Francesco Cagnetta1, Alessandro Favero2, Antonio Sclocchi3
1Scuola Internazionale Superiore di Studi Avanzati, (SISSA), Via Bonomea 265, 34136 Trieste, Italy.
Neural language models learn language structure through next-token prediction. Convolutional networks outperform transformers due to architectural alignment with data structure, clarifying neural scaling laws.
Area of Science:
- Computational linguistics
- Machine learning theory
- Deep learning architectures
Background:
- Neural language models (NLMs) are trained on next-token prediction tasks.
- Understanding how NLMs acquire linguistic structure is a key research question.
- Previous work established a theory of representation learning based on data correlations.
Purpose of the Study:
- To derive theoretical scaling laws for NLMs trained on synthetic data.
- To investigate the impact of model architecture on learning linguistic structure.
- To extend existing representation learning theory to account for architectural differences.
Main Methods:
- Developed theoretical scaling laws for neural network performance.
- Utilized synthetic datasets from the random hierarchy model (RHM).
- Extended representation learning theory to incorporate architectural variations.
Main Results:
- Convolutional networks show faster performance scaling than transformer models.
- Architectural biases in neural scaling laws are identified.
- Empirically validated predictions regarding network performance.
Conclusions:
- Model architecture significantly influences how NLMs learn linguistic structure.
- The alignment between network architecture and data's statistical properties is crucial.
- Findings clarify the interplay between architecture, data, and representation learning in NLMs.
Related Concept Videos
Convolution Properties II
The width property indicates that if the durations of input signals are T1 and T2, then the width of the output response equals the sum of both durations, irrespective of the shapes of the two functions. For instance, convolving two rectangular pulses with durations of 2 seconds and 1 second results in a function with a width of 3 seconds.
The area property asserts that the area under the...
Second Law of Thermodynamics
First Law of Thermodynamics
First Law of Thermodynamics
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Gas Laws: Boyle's, Gay-Lussac, Charles', Avogadro's, and Ideal Gas Law

