Abstract
I trained DC-GANs on CIFAR-10 and CelebA to compare ReLU and ELU activations across several learning-rate settings.
The experiments focused on two questions: which setups stayed stable, and which produced sharper samples. ReLU performed better on both datasets, while CIFAR-10 benefited more from an asymmetric learning rate.
This page documents the architecture, training runs, results, and limits of the evaluation.
1 Introduction
A GAN trains two networks together. The generator produces samples from noise, while the discriminator learns to tell generated samples from real ones.
DC-GANs replace fully connected layers with convolutional blocks, which makes them a useful baseline for image generation.
They are still sensitive to activation functions, learning rates, and the balance between the two networks.
I kept the architecture fixed where possible so the activation and learning-rate changes were easier to compare.
Research Objectives
2 Methodology
2.1 Architecture Design
My DC-GAN implementation loosely follows the architectural guidelines established by Radford et al., with systematic variations to explore the impact of different design choices. The architecture consists of two competing networks working in an adversarial framework.
Generator Network
Employs a series of transposed-convolution layers to progressively up-sample random-noise vectors into full-resolution images, beginning with a dense layer that reshapes the noise into a small spatial feature map.
Discriminator Network
Progressively down-samples input images to a binary classification, using convolutions and LeakyReLU activations with spectral normalization for CIFAR-10.
2.2 Training Strategy
The training loop uses the standard GAN minimax objective, updating the discriminator and generator in alternating steps. I added several stabilization techniques and tracked how each run behaved.
Stabilization Techniques
Learning Rate Schedules
G_LR = D_LR = 1 × 10⁻⁴
G_LR = 5 × 10⁻⁵, D_LR = 1 × 10⁻⁴
G_LR = 2 × 10⁻⁴, D_LR = 3 × 10⁻⁴
2.3 Experimental Design
For each activation × learning-rate setting I train three seeds, log losses, checkpoint every five epochs, and compute Inception Score (50k samples, 10 splits). CIFAR-10 runs for 100 epochs; CelebA converges by epoch 25.
Datasets
Two fundamentally different image datasets were selected to evaluate DC-GAN performance across varied domains: CIFAR-10 for diverse object categories and CelebA for high-fidelity human faces.
CIFAR-10
Contains 60,000 32×32 colour images across ten classes: airplanes, automobiles, birds, cats, deer, dogs, frogs, horses, ships, and trucks.
CelebA
Contains more than 200,000 aligned celebrity faces with 40 binary attributes. I used about 50,000 images, center-cropped and resized to 64×64.
Preprocessing Pipeline
• Pixel normalization [-1, 1]
• Random horizontal flips
• Small rotation augmentation
• Quality filtering
• Centre crop faces
• Resize to 64×64
• Pixel normalization [-1, 1]
4 CIFAR-10 Results & Discussion
4.1 Training Dynamics Analysis
The CIFAR-10 experiments revealed significant differences in training stability and output quality between activation functions and learning rate configurations. Three key scenarios emerged from the systematic evaluation.
ReLU Activation
ReLU's sparse activations preserve high-frequency detail essential for diverse object generation. Performs best with asymmetric learning rates.
ELU Activation
ELU's smooth negative region led to oversmoothing on CIFAR-10's diverse textures, resulting in mode collapse and poor sample quality.
CIFAR-10 Training Results
Key Training Scenarios
Best: ReLU + Asymmetric LR
G: 1e-4, D: 2e-4 → Healthy dynamics, vivid objects, IS 5.49
Worst: ELU + Balanced LR
Diverging losses, mode collapse, blurry blobs, IS 2.87
Problematic: ReLU + Balanced LR
Flat losses, generator complacency, uniform grey patches
CIFAR-10 Performance
Optimal Configuration:
ReLU + Asymmetric Learning Rate
Key Insights
4.2 Architecture Impact Assessment
ELU's smooth negative region oversmooths outputs on CIFAR-10, while ReLU's sparse activations preserve high-frequency detail. Activation alone was not enough: ReLU needed a higher discriminator learning rate to work best on the more varied object images.
5 CelebA Results & Discussion
5.1 Architecture Impact Assessment
CelebA's facial geometry stabilizes training for both activations, but distinct differences emerge in output quality and fine-grained detail preservation. The structured nature of faces allows both ReLU and ELU to achieve reasonable stability, highlighting the importance of activation choice for detail rendering.
CelebA Training Results Comparison
ReLU: Photorealistic Detail
ReLU better preserves hair strands and skin pores, generating photorealistic faces with sharp detail and accurate anatomy across diverse demographics.
ELU: Airbrushed Softness
ELU yields softer, airbrushed faces with less texture detail. While aesthetically pleasing, lacks the fine-grained realism achieved by ReLU.
5.2 Hyperparameter Optimization
CelebA was less sensitive to learning-rate changes than CIFAR-10. The repeated facial structure produced more balanced training, though ReLU still made sharper samples in these runs.
CelebA Performance
Optimal Configuration:
ReLU + Face-tuned LR
(G: 2e-4, D: 3e-4)
CelebA Insights
5.3 Sample Quality Evaluation
Quality Comparison Summary
• Photorealistic skin texture
• Sharp hair definition
• Detailed facial features
• Natural lighting effects
• Smooth, airbrushed skin
• Softer hair rendering
• Less textural detail
• Pleasant but less realistic
ReLU kept more skin and hair detail, while ELU produced softer features and lower Inception Scores. The same pattern appeared across the seeds I tested.
6 Analysis & Conclusion
6.1 Comparative Analysis
CIFAR-10 worked best with ReLU and a discriminator learning rate twice as high as the generator's. CelebA was less sensitive to that balance, but still favored ReLU for sharp detail.
CIFAR-10 Requirements
CelebA Characteristics
6.2 Training Insights
Early Warning Signs
Discriminator loss > 1.5 and generator loss pinned at ≈ 0.7 within 20 epochs signal collapse. These indicators proved consistent across all failed experiments.
Activation choice was the main stability factor. Learning-rate asymmetry mattered chiefly for CIFAR-10, and most failed runs showed warning signs within the first 20 epochs.
6.3 Practical Implementation Guidelines
Recommended Best Practices
General Guidelines
Diverse Datasets (CIFAR-10-like)
Structured Datasets (Faces)
Key Findings
Activation-function choice significantly impacts DC-GAN performance: ReLU consistently outperforms ELU across object and face domains. Optimal configurations achieved IS 5.49 ± 1.8 (CIFAR-10) and IS 6.82 ± 1.4 (CelebA).
ReLU was the more consistent activation across both datasets. Learning-rate changes mattered most on CIFAR-10.
Conclusion & Future Work
The best setup depended on the dataset. ReLU worked better in these runs, but CIFAR-10 needed more careful learning-rate balance than CelebA.
The useful lesson was practical: keep the architecture stable, change one variable at a time, and watch the early loss curves before committing to a long run.
Future Research Directions
A follow-up could repeat the comparison with multiple random seeds and stronger image-quality metrics before moving to newer architectures.