A Comprehensive Survey of Dataset Distillation
IEEE Transactions on Pattern Analysis and Machine Intelligence
|October 6, 2023
Summary
Dataset distillation synthesizes small datasets from large ones to improve deep learning efficiency. This review explores frameworks, challenges, and future directions for dataset distillation techniques.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Computer Science
Background:
- Deep learning's rapid advancement relies on massive data and computing power.
- Unlimited data growth challenges limited computational resources.
- Dataset distillation offers a solution by creating smaller, representative datasets.
Purpose of the Study:
- To provide a comprehensive overview of dataset distillation.
- To analyze existing distillation frameworks and algorithms.
- To identify limitations and future research avenues.
Main Methods:
- Categorizing dataset distillation into meta-learning and data matching frameworks.
- Exploring factorized dataset distillation approaches.
- Reviewing performance comparisons and applications.
Main Results:
- Dataset distillation effectively compresses large datasets.
- Current methods face limitations with high-resolution data and complex label spaces.
- The paper offers a holistic understanding of the field.
Conclusions:
- Dataset distillation is a promising technique for efficient data processing in deep learning.
- Further research is needed to address current limitations.
- The review highlights key challenges and future directions for promoting the field.
Related Concept Videos
Distillation: Vapor–Liquid Equilibria
2.8K
Distillation is a separation technique that takes advantage of the boiling point properties of disparate elements in a mixture. To perform distillation, we begin by heating a miscible mixture of two liquids with a significant difference in boiling points (at least 20°C). As the solution heats up and reaches the bubble point of the more volatile component, some molecules of the more volatile component transition into the gas phase and travel upward into the condenser, which is a glass tube...
2.8K
Distribution Reliability and Automation
113
Distribution reliability in electrical power systems is critical for ensuring an uninterrupted power supply to consumers at minimal cost. According to IEEE Standard Terms, reliability is the probability that a device will function without failure over a specified time period or amount of usage. For electric power distribution, this translates to maintaining continuous power supply and addressing customer concerns over power outages. Several indices, as defined by IEEE Standard 1366-2012, are...
113
Sampling Distribution
12.9K
Given simple random samples of size n from a given population with a measured characteristic such as mean, proportion, or standard deviation for each sample, the probability distribution of all the measured characteristics is called a sampling distribution. How much the statistic varies from one sample to another is known as the sampling variability of a statistic. You typically measure the sampling variability of a statistic by its standard error. The standard error of the mean is an example...
12.9K
Cluster Sampling Method
12.0K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
12.0K
Extraction: Partition and Distribution Coefficients
2.5K
The distribution law or Nernst's distribution law is the law that governs the distribution of a solute between two immiscible solvents. This law, also known as the partition law, states that if a solute is added to the mixture of two immiscible solvents at a constant temperature, the solute is distributed between the two solvents in such a way that the ratio of solute concentrations in the solvents remains constant at equilibrium.
For extracting a solute from an aqueous phase into an...
For extracting a solute from an aqueous phase into an...
2.5K
Distributions to Estimate Population Parameter
4.1K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.1K

