Related Experiment Videos
Optimal Transport-based Difficulty-aware Contribution Allocation for Dataset Distillation
Abstract:
The increasing reliance on large-scale datasets imposes significant storage and computational burdens on training deep learning models. Dataset distillation methods, particularly those based on sample generation, aim to condense large original datasets into compact synthetic sets while preserving essential information. Existing subset synthesis approaches typically minimize a homogeneous distance, assigning uniform contributions from all real instances to the construction of each synthetic sample. We show that such equal allocation neglects instance-level relationships between real-synthetic pairs, leading to inadequate modeling of the geometric structural discrepancies between the distilled and original datasets. In this work, we reformulate homogeneous distance minimization as a bi-level optimization problem via a matching-and-approximating paradigm. In the matching stage, we employ an optimal transport matrix to dynamically allocate contributions from real instances. Building upon this transport-based allocation, we further introduce a difficulty-aware marginal reweighting mechanism to emphasize informative instances while preserving global geometric consistency. In the subsequent approximation stage, synthetic samples are refined according to the established allocation scheme to better approximate the real data distribution. This strategy enables a more faithful characterization of intricate geometric structures and improved handling of intra-class variations, thereby enhancing distillation fidelity. Extensive experiments across diverse architectures, modalities, and learning paradigms demonstrate that the proposed framework consistently improves performance, with gains observed in standard supervised, federated, continual, and multimodal learning settings.
Related Concept Videos
Short-distance Transport of Resources
Distributed Loads: Problem Solving
Maxwell-Boltzmann Distribution: Problem Solving
This distribution function f(v) is defined by saying that the expected number N (v1,v2) of particles with speeds between v1 and v2 is given by
Choosing Between z and t Distribution
Analyte Adsorption and Distribution