Deep Convolutional Generative Adversarial Networks

Comparing activation functions and learning rates on CIFAR-10 and CelebA.

University of California, San Diego

Abstract

I trained DC-GANs on CIFAR-10 and CelebA to compare ReLU and ELU activations across several learning-rate settings.

The experiments focused on two questions: which setups stayed stable, and which produced sharper samples. ReLU performed better on both datasets, while CIFAR-10 benefited more from an asymmetric learning rate.

This page documents the architecture, training runs, results, and limits of the evaluation.

1 Introduction

A GAN trains two networks together. The generator produces samples from noise, while the discriminator learns to tell generated samples from real ones.

DC-GANs replace fully connected layers with convolutional blocks, which makes them a useful baseline for image generation.

They are still sensitive to activation functions, learning rates, and the balance between the two networks.

I kept the architecture fixed where possible so the activation and learning-rate changes were easier to compare.

Research Objectives

Train the same DC-GAN baseline on both datasets
Compare ReLU and ELU activations
Test balanced and asymmetric learning rates
Track stability, loss behavior, and sample quality

2 Methodology

2.1 Architecture Design

My DC-GAN implementation loosely follows the architectural guidelines established by Radford et al., with systematic variations to explore the impact of different design choices. The architecture consists of two competing networks working in an adversarial framework.

Generator Network

Employs a series of transposed-convolution layers to progressively up-sample random-noise vectors into full-resolution images, beginning with a dense layer that reshapes the noise into a small spatial feature map.

Transposed convolutions for upsampling
Batch normalization for stability
ReLU and ELU activations (varied across experiments)
Tanh output activation

Discriminator Network

Progressively down-samples input images to a binary classification, using convolutions and LeakyReLU activations with spectral normalization for CIFAR-10.

Convolutional layers for feature extraction
LeakyReLU (α = 0.2) activations
Spectral normalization (CIFAR-10)
Binary classification output

2.2 Training Strategy

The training loop uses the standard GAN minimax objective, updating the discriminator and generator in alternating steps. I added several stabilization techniques and tracked how each run behaved.

Stabilization Techniques

Spectral normalization (CIFAR-10)
Exponential-moving-average (EMA) weight tracking
Instance-noise decay
Label smoothing
Careful weight initialization
Mixed-precision with gradient scaling
Adam optimizer (β₁ = 0.5, β₂ = 0.999)

Learning Rate Schedules

Balanced:
G_LR = D_LR = 1 × 10⁻⁴
Asymmetric (CIFAR-10):
G_LR = 5 × 10⁻⁵, D_LR = 1 × 10⁻⁴
Face-tuned (CelebA):
G_LR = 2 × 10⁻⁴, D_LR = 3 × 10⁻⁴

2.3 Experimental Design

For each activation × learning-rate setting I train three seeds, log losses, checkpoint every five epochs, and compute Inception Score (50k samples, 10 splits). CIFAR-10 runs for 100 epochs; CelebA converges by epoch 25.

Datasets

Two fundamentally different image datasets were selected to evaluate DC-GAN performance across varied domains: CIFAR-10 for diverse object categories and CelebA for high-fidelity human faces.

CIFAR-10

60,000 images • 32×32 pixels • 10 classes

Contains 60,000 32×32 colour images across ten classes: airplanes, automobiles, birds, cats, deer, dogs, frogs, horses, ships, and trucks.

Low resolution ideal for initial GAN testing
Challenging due to object diversity
Normalized pixels to [-1, 1]
Random flips and small rotations applied
Challenge: Diverse textures and object categories make the dataset harder to model

CelebA

~50,000 images • 64×64 pixels • Celebrity faces

Contains more than 200,000 aligned celebrity faces with 40 binary attributes. I used about 50,000 images, center-cropped and resized to 64×64.

Higher resolution for detailed features
Structured domain (human faces)
Centre-cropped and aligned faces
Corrupted images removed during preprocessing
Challenge: Fine-grained detail in skin texture, symmetry, and facial features

Preprocessing Pipeline

CIFAR-10:
• Pixel normalization [-1, 1]
• Random horizontal flips
• Small rotation augmentation
CelebA:
• Quality filtering
• Centre crop faces
• Resize to 64×64
• Pixel normalization [-1, 1]

4 CIFAR-10 Results & Discussion

4.1 Training Dynamics Analysis

The CIFAR-10 experiments revealed significant differences in training stability and output quality between activation functions and learning rate configurations. Three key scenarios emerged from the systematic evaluation.

ReLU Activation

Best Performance: IS 5.49 ± 1.8

ReLU's sparse activations preserve high-frequency detail essential for diverse object generation. Performs best with asymmetric learning rates.

Sharp, colorful object generation
Requires asymmetric LR (D = 2 × G)
Healthy adversarial dynamics
Nearly doubled ELU score

ELU Activation

Inception Score: 2.87 ± 0.98

ELU's smooth negative region led to oversmoothing on CIFAR-10's diverse textures, resulting in mode collapse and poor sample quality.

Desaturated, blurry outputs
Dominant discriminator dynamics
Mode collapse evident
Consistent underperformance

CIFAR-10 Training Results

Key Training Scenarios

Best: ReLU + Asymmetric LR

G: 1e-4, D: 2e-4 → Healthy dynamics, vivid objects, IS 5.49

Worst: ELU + Balanced LR

Diverging losses, mode collapse, blurry blobs, IS 2.87

Problematic: ReLU + Balanced LR

Flat losses, generator complacency, uniform grey patches

CIFAR-10 Performance

IS: 5.49 ± 1.8

Optimal Configuration:
ReLU + Asymmetric Learning Rate

Key Insights

Activation choice crucial for object diversity
Asymmetric LR prevents discriminator dominance
Early loss divergence indicates mode collapse
ReLU preserves high-frequency textures

4.2 Architecture Impact Assessment

ELU's smooth negative region oversmooths outputs on CIFAR-10, while ReLU's sparse activations preserve high-frequency detail. Activation alone was not enough: ReLU needed a higher discriminator learning rate to work best on the more varied object images.

5 CelebA Results & Discussion

5.1 Architecture Impact Assessment

CelebA's facial geometry stabilizes training for both activations, but distinct differences emerge in output quality and fine-grained detail preservation. The structured nature of faces allows both ReLU and ELU to achieve reasonable stability, highlighting the importance of activation choice for detail rendering.

CelebA Training Results Comparison

ReLU: Photorealistic Detail

Inception Score: 6.82 ± 1.4

ReLU better preserves hair strands and skin pores, generating photorealistic faces with sharp detail and accurate anatomy across diverse demographics.

Sharp skin texture and pores
Detailed hair strand rendering
Accurate facial anatomy
Varied lighting and demographics
Consistent reproducibility across seeds

ELU: Airbrushed Softness

Best Score: 4.91 ± 1.2

ELU yields softer, airbrushed faces with less texture detail. While aesthetically pleasing, lacks the fine-grained realism achieved by ReLU.

Smoother facial features
Airbrushed skin appearance
Less hair detail
Good overall structure
Lower inception scores consistently

5.2 Hyperparameter Optimization

CelebA was less sensitive to learning-rate changes than CIFAR-10. The repeated facial structure produced more balanced training, though ReLU still made sharper samples in these runs.

CelebA Performance

IS: 6.82 ± 1.4

Optimal Configuration:
ReLU + Face-tuned LR
(G: 2e-4, D: 3e-4)

CelebA Insights

Activation choice outweighs LR tuning
Facial structure stabilizes training
ReLU preserves fine details better
Higher resolution shows clear differences
Consistent quality across demographics

5.3 Sample Quality Evaluation

Quality Comparison Summary

ReLU Characteristics:
• Photorealistic skin texture
• Sharp hair definition
• Detailed facial features
• Natural lighting effects
ELU Characteristics:
• Smooth, airbrushed skin
• Softer hair rendering
• Less textural detail
• Pleasant but less realistic

ReLU kept more skin and hair detail, while ELU produced softer features and lower Inception Scores. The same pattern appeared across the seeds I tested.

6 Analysis & Conclusion

6.1 Comparative Analysis

CIFAR-10 worked best with ReLU and a discriminator learning rate twice as high as the generator's. CelebA was less sensitive to that balance, but still favored ReLU for sharp detail.

CIFAR-10 Requirements

Asymmetric learning rates essential
ReLU critical for texture preservation
High sensitivity to hyperparameters
Diverse object categories challenging

CelebA Characteristics

Learning rate robust
ReLU still superior for detail
Structured domain stabilizes training
Fine-grained texture differences

6.2 Training Insights

Early Warning Signs

Discriminator loss > 1.5 and generator loss pinned at ≈ 0.7 within 20 epochs signal collapse. These indicators proved consistent across all failed experiments.

Activation choice was the main stability factor. Learning-rate asymmetry mattered chiefly for CIFAR-10, and most failed runs showed warning signs within the first 20 epochs.

6.3 Practical Implementation Guidelines

Recommended Best Practices

General Guidelines
Use ReLU activations in DC-GAN generators
Monitor loss dynamics in first 20 epochs
Implement multiple stabilization techniques
Train with multiple random seeds
Diverse Datasets (CIFAR-10-like)
Set D ≈ 2 × G learning rate
Monitor loss divergence carefully
Use spectral normalization
Expect longer convergence times
Structured Datasets (Faces)
Balanced learning rates from 1×10⁻⁴ to 5×10⁻⁴ were enough
Focus on activation function choice
Higher resolution reveals differences
Quality assessment via fine details

Key Findings

Activation-function choice significantly impacts DC-GAN performance: ReLU consistently outperforms ELU across object and face domains. Optimal configurations achieved IS 5.49 ± 1.8 (CIFAR-10) and IS 6.82 ± 1.4 (CelebA).

ReLU was the more consistent activation across both datasets. Learning-rate changes mattered most on CIFAR-10.

Conclusion & Future Work

The best setup depended on the dataset. ReLU worked better in these runs, but CIFAR-10 needed more careful learning-rate balance than CelebA.

The useful lesson was practical: keep the architecture stable, change one variable at a time, and watch the early loss curves before committing to a long run.

Future Research Directions

Extension to Progressive-GAN and StyleGAN architectures
Exploration of transfer learning strategies
Development of more generalizable generative models
Investigation of attention mechanisms in GANs
Cross-domain style transfer applications

A follow-up could repeat the comparison with multiple random seeds and stronger image-quality metrics before moving to newer architectures.

Full Research Paper

First page of GAN research paper

Research Paper Preview

Preview of research paper
Download Full Paper (PDF)

"Deep Convolutional Generative Adversarial Networks on CIFAR-10 and CelebA"

Explore the Code

View on GitHub