Related Experiment Video
Updated: Aug 15, 2025

05:39
Generating Strictly Controlled Stimuli for Figure Recognition Experiments
Published on: March 18, 2019
5.3K
Image Generation from Text Using StackGAN with Improved Conditional Consistency Regularization.
Rihito Tominaga1, Masataka Seo1
1Osaka Institute of Technology, Graduate School of Robotics and Design Engineering, 1-45 Chayamachi, Kita-ku, Osaka 530-0013, Japan.
Sensors (Basel, Switzerland)
|January 8, 2023
Summary
Researchers improved text-to-image generation by introducing Improved Conditional Consistency Regularization (ICCR) to StackGAN. This method significantly reduces mode collapse and enhances image quality, outperforming existing models like AttnGAN.
Area of Science:
- Multimodal learning
- Deep learning for image generation
- Generative Adversarial Networks (GANs)
Background:
- Image generation from natural language is a rapidly advancing field.
- Stacked Generative Adversarial Networks (StackGAN) can produce high-resolution images but suffer from unintelligible outputs and mode collapse.
- Existing models like AttnGAN also show limitations in text-to-image synthesis.
Purpose of the Study:
- To address the limitations of StackGAN, specifically unintelligible image generation and mode collapse.
- To enhance the fidelity and diversity of images generated from text descriptions.
- To improve the performance of text-to-image models in conditional generation tasks.
Main Methods:
- Incorporation of Improved Consistency Regularization (ICCR), a novel technique, into the StackGAN architecture.
- ICCR stabilizes adversarial learning by matching semantic information before and after data augmentation, suppressing mode collapse.
- Modification of ICR to ICCR to eliminate generator-induced negative impacts and enhance conditional generation.
Main Results:
- StackGAN with ICCR achieved a 16% higher Inception Score than original StackGAN on the CUB dataset.
- The proposed model outperformed AttnGAN by 4% on the Inception Score.
- StackGAN with ICCR completely eliminated mode collapse (0% probability compared to 20% in original StackGAN) and received higher user ratings.
Conclusions:
- The proposed ICCR method significantly enhances StackGAN's ability to generate high-quality, relevant images from text.
- ICCR is more effective than ICR for conditional generation tasks, improving both image quality and stability.
- The developed model represents a substantial advancement in text-to-image generation, offering improved performance and reduced mode collapse.
More Related Videos
Related Concept Videos
Improving Translational Accuracy
11.8K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.8K
Genetic Lingo
103.7K
Overview
103.7K
Aliasing
189
Accurate signal sampling and reconstruction are crucial in various signal-processing applications. A time-domain signal's spectrum can be revealed using its Fourier transform. When this signal is sampled at a specific frequency, it results in multiple scaled replicas of the original spectrum in the frequency domain. The spacing of these replicas is determined by the sampling frequency.
If the sampling frequency is below the Nyquist rate, these replicas overlap, preventing the original...
If the sampling frequency is below the Nyquist rate, these replicas overlap, preventing the original...
189
Stereotype Content Model
14.8K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.8K
Reducing Line Loss
184
In a three-phase circuit, line loss is an indicator of energy dissipated as heat due to the resistance of transmission lines. To address this, incorporating transformers into the system—a step-up transformer at the source and a step-down transformer at the load—is a strategic solution. Two three-phase transformers are introduced to improve this.
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss...
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss...
184
Survival Tree
131
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
131

