A cinematic, first-principles mathematical orchestration explaining optimization in Machine Learning.
This repository hosts a production-grade visualization project designed using Manim (Community Edition) to build a geometric, first-principles understanding of Gradient Descent. The animations avoid static slides and instead rely on fluid, dynamic coordinate systems to build a robust mental model of machine learning optimization.
The complete orchestrated render. Controls, high-definition playbacks, and automated looping supported.
Traditional educational resources often present Gradient Descent as an abstract, dry algebraic update rule:
While mathematically complete, this notation obscures the rich underlying geometry. Learners routinely struggle to grasp:
- How the gradient vector behaves as a directional compass.
- How the learning rate (
$\alpha$ ) physically scales parameters across the loss landscape. - Why convergence step sizes naturally contract as the loss function flattens.
- How optimization dynamics scale when moving from 2D parameter curves to 3D loss terrains.
By rendering these mathematical vectors, tangents, and coordinate trajectories in high-fidelity motion, this project transforms dry calculus into a highly readable, intuitive, and memorable visual narrative.
- Dynamic Tangent & Derivative Updaters: Real-time evaluation of derivative slopes, adjusting local tangent lines dynamically as values slide along the loss curve.
-
Unified Translucent Dashboard: A custom layout system displaying current parameter value
$x$ , numerical gradient value$dJ/dx$ , and step-direction state without overlapping. - Parallel Hyperparameter Compilers: Side-by-side comparative panels simulating Too Small, Optimal, and Too Large learning rates simultaneously to demonstrate stability constraints.
- 3D Terrain Projections: Smooth orbital camera choreography tracing convergence trajectories down a 3D bowl-shaped surface representing multiple weights and biases.
- Modular Animation Pipeline: Clean, object-oriented scene architecture separating core mathematical utilities from camera and rendering rigs.
The complete production is divided into 7 carefully sequenced chapters that build mathematical intuition from the ground up:
| Chapter | Title | Mathematical Purpose | Visual Mechanics | Engineering Highlight |
|---|---|---|---|---|
| 1 | The Intro | Core update formula breakdown | Sleek dark-mode overlay tracking parameter elements | Mathematical typesetting and smooth group animations |
| 2 | The Loss Landscape | Visualizing loss terrains geometrically | 2D parabolic tracking highlighting High vs. Lower Loss states | Area fill generators and custom vector annotations |
| 3 | Understanding the Gradient | Defining derivative vectors geometrically | Slidable coordinate dot with dynamic tangent lines | Re-computed tangent sloper linked to ValueTrackers |
| 4 | Gradient Descent in 2D | Demonstrating step iteration mechanics | Converging coordinate steps with trailing ghost markers | Step path history tracing and convergence banners |
| 5 | Learning Rate Comparison | Visualizing hyperparameter boundaries | Triple parallel horizontal comparative axes panels | Multi-axis value clamping and comparative loop grids |
| 6 | Gradient Descent in 3D | Scaling parameters to multi-variable terrains | 3D bowl-shaped surface optimization with camera orbits | Orbital camera rotation and trajectory path generation |
| 7 | The Outro | Summarizing core takeaway concepts | Four interconnected conceptual nodes fading into code | Node layout trees and localized coordinate transitions |
An elegant presentation breaking down the mathematical terms of the gradient descent formula.
θ_{t+1} = θ_t - α · ∇J(θ_t)
│ │ │ │
[New Parameter] [Current] [Learning] [Gradient/Slope]
[Value] [Rate] [Direction]
A real-time geometric representation of derivatives, recalculating slope tangent lines as parameters slide.
Comparing convergence speeds and numerical behaviors across distinct step coefficients:
-
$\alpha = 0.05$ (Too Small): Extremely slow steps; fails to reach the minimum in the allotted timeframe. -
$\alpha = 0.20$ (Optimal): Steady, exponential deceleration direct to the global minimum. -
$\alpha = 1.10$ (Too Large): Oversteps and diverges, oscillating up the curves.
The codebase leverages a modular, inheritance-based class structure designed for rendering scalability and visual consistency.
graph TD
classDef base fill:#0B0F19,stroke:#00D4AA,stroke-width:2.5px,color:#fff;
classDef scene fill:#151922,stroke:#FFD93D,stroke-width:2px,color:#fff;
classDef video fill:#151922,stroke:#FF6B6B,stroke-width:2px,color:#fff;
classDef method fill:#0F1219,stroke:#6BCB77,stroke-width:1px,color:#E8E8E8;
Base["GradientDescentBase Base Class<br>(Global configurations, logo layers, and panel templates)"]:::base
Base --> IntroScene["IntroScene<br>(Standalone Title & Formulas Chapter)"]:::scene
Base --> CompleteVideo["CompleteVideo Orchestrator<br>(Sequentially executes all presentation chapters)"]:::video
CompleteVideo --> play_intro["play_intro()"]:::method
CompleteVideo --> play_loss_landscape["play_loss_landscape()"]:::method
CompleteVideo --> play_gradient_understanding["play_gradient_understanding()"]:::method
CompleteVideo --> play_descent_2d["play_descent_2d()"]:::method
CompleteVideo --> play_learning_rates["play_learning_rates()"]:::method
CompleteVideo --> play_descent_3d["play_descent_3d()"]:::method
CompleteVideo --> play_outro["play_outro()"]:::method
To avoid screen clutter and overlapping text, the legend dashboard inside the Gradient Understanding scene is packed into nested, auto-aligning virtual groups:
legend = always_redraw(lambda: self.make_panel(
VGroup(
VGroup(
Text("x = ", font_size=20, color=C_DOT),
DecimalNumber(x_tracker.get_value(), num_decimal_places=2, font_size=22, color=C_DOT)
).arrange(RIGHT, buff=0.1),
VGroup(
Text("dJ/dx = ", font_size=20, color=C_TANGENT),
DecimalNumber(grad_1d(x_tracker.get_value()), num_decimal_places=2, font_size=22, color=C_TANGENT)
).arrange(RIGHT, buff=0.1),
get_sign_text()
).arrange(DOWN, buff=0.18, aligned_edge=LEFT),
pos=[4.2, 2.2, 0]
))This design completely removes hardcoded absolute coordinates, ensuring the bounding box and borders dynamically stretch and adapt to shifting string dimensions.
Calculations are decoupled from visual pixel coordinates. Mobject transformations utilize specialized coordinate transformation wrappers:
dot = always_redraw(lambda: Dot(
axes.c2p(x_tracker.get_value(), loss_1d(x_tracker.get_value())),
color=C_DOT, radius=0.12
))This guarantees mathematical accuracy by directly linking the numerical state (x_tracker.get_value()) to visual coordinate placements (axes.c2p).
- The Problem: During rendering, the dynamic labels for the
xvalue anddJ/dxderivative were displaying as extremely dark, illegible gray text, while the indicator text underneath was bright. - The Cause: The semi-transparent panel background (
panel_bg) was assigned az_indexof40to float above the coordinate grid. Since the text labels defaulted to az_indexof0, they were being masked behind the 95% opacity background rectangle, obscuring them. - The Solution: Promoted all text, numerical labels, and indicator mobjects to an explicit
z_index = 50. This forced the renderer to layer them cleanly on top of the translucent panel, establishing excellent readability.
- The Problem: Real-time updates caused components to overlap or shift awkwardly as numerical widths contracted/expanded (e.g., transitions from
POSITIVE → LefttoZERO). - The Solution: Refactored the hardcoded component layout into a unified virtual layout. Wrapping the row
VGroupsinsidealways_redraw(lambda: self.make_panel(...))ensures the panel dynamically resizes, aligns to the left, and adjusts boundaries on every frame.
- The Problem: When generating parallel panels for the learning rate comparison, all animated dots mapped to the same coordinate tracking vector, causing them to move in unison rather than showing distinct speeds.
- The Cause: Python closures capture loop variables by reference rather than by value inside standard lambda scopes.
- The Solution: Implemented default parameter bindings (
lambda t=tracker, a=ax: ...) to capture unique coordinate tracking instances at each loop step, creating distinct comparative timelines.
- The Problem: Dynamic frame updates occasionally triggered thread errors or memory leaks during scene transitions.
- The Cause: Active updaters on faded-out objects continued to execute background processes, attempting to poll coordinates from trackers that no longer existed.
- The Solution: Implemented strict lifecycle cleanup by calling
.clear_updaters()on all active objects (likelegend,sphere, andcurrent_dot) immediately before fading them out.
-
Loss Terminology: Formulating cost/loss curves
$J(\theta)$ to measure prediction errors. - Gradient Directionality: Proving geometrically why the path of steepest descent is opposite to the gradient vector.
- Tangent Vectors: Visualizing derivatives as local, instantaneous slopes.
- Hyperparameter Sensitivity: Illustrating how learning rates scale updating vectors.
- Decelerative Convergence: Proving why steps naturally contract as loss gradients approach zero.
- Multivariate Terrains: Scaling optimization models to 3D surfaces representing multiple weights and biases.
gradient-descent-manim/
├── .git/ # Git version control metadata
├── manim-env/ # Isolated Python virtual environment (Manim, NumPy, FFmpeg)
├── media/ # Cached and compiled mathematical visual assets
│ ├── Tex/ # Cached LaTeX equation vector assets (SVG paths)
│ ├── texts/ # Cached mathematical text components (vector layouts)
│ └── videos/
│ └── gradient_descent_visualization/
│ ├── 480p15/ # Low-fidelity draft renders (480p, 15 FPS)
│ │ ├── partial_movie_files/ # Incremental render caches for rapid iterations
│ │ └── CompleteVideo.mp4 # Compiled draft movie file
│ └── 2160p60/ # Cinematic production master renders (4K, 60 FPS)
│ ├── partial_movie_files/ # High-resolution incremental caches
│ └── CompleteVideo.mp4 # Final high-fidelity orchestrator output file
├── .gitignore # Excludes python build caches, raw logs, and env configs
├── README.md # Comprehensive developer, engineer, and recruiter documentation
├── chapter3_preview.png # Preview asset representing chapter 3 tangent calibration
├── script.txt # Expressive voiceover narration script optimized for TTS systems
├── clue.txt # Frame-by-frame visual markers and audio-sync timeline cues
└── gradient_descent_visualization.py # Consolidated rendering codebase holding all mathematical scenes
This repository comes with an isolated virtual environment containing all dependencies including Manim, NumPy, and FFmpeg codecs.
git clone https://github.com/himanshu-jadhav108/gradient-descent-manim.git
cd gradient-descent-manimActivate the environment to load correct libraries and packages:
- Windows (PowerShell):
.\manim-env\Scripts\Activate.ps1 - Windows (Command Prompt):
.\manim-env\Scripts\activate.bat
- macOS / Linux:
source manim-env/bin/activate
Check that the active installation of Manim is accessible and correct:
manim --versionTo render the animations, run the following commands with the active virtual environment.
Compile all chapters sequentially into a single cinematic presentation:
# Low Quality (Fast render for testing - 480p, 15fps)
manim -pql gradient_descent_visualization.py CompleteVideo
# High Quality (Standard production grade - 1080p, 60fps)
manim -pqh gradient_descent_visualization.py CompleteVideo
# Ultra Quality (Mastering grade - 2160p 4K, 60fps)
manim -pqk gradient_descent_visualization.py CompleteVideoIf you want to view, modify, or render individual visual scenes:
# Chapter 1: Introduction Scene
manim -pql gradient_descent_visualization.py IntroScene
# Chapter 2: Loss Landscape
manim -pql gradient_descent_visualization.py LossFunctionScene
# Chapter 3: Dynamic Tangent Calculations
manim -pql gradient_descent_visualization.py GradientVisualizationScene
# Chapter 4: Step-by-Step Descent
manim -pql gradient_descent_visualization.py GradientDescent2DScene
# Chapter 5: Comparative Learning Rates
manim -pql gradient_descent_visualization.py LearningRateComparisonScene
# Chapter 6: 3D Surface Optimization
manim -pqh gradient_descent_visualization.py ThreeDGradientDescentScene
# Chapter 7: Takeaway Outro Map
manim -pql gradient_descent_visualization.py OutroScene-p: Automatically preview the video file once compilation completes.-q: Specifies the output quality parameter.l(Low): 480p at 15 FPS. Ideal for rapid code iteration and design verification.h(High): 1080p at 60 FPS. Perfect for standard, clean web distribution.k(4K): 2160p at 60 FPS. Extreme presentation grade.
- Render Latencies: Low-quality (480p/15fps) scenes compile in <5 seconds; Complete 4K master files compile in ~120 seconds (on a 24-thread configuration).
- Target Container: MPEG-4 (
.mp4) wrapped using H.264 video encoding. - Audio Multiplexing: Formatted with default empty silent audio channels, fully compatible with external voiceover dubbing pipelines.
- Code Optimization: Dynamic vectors and matrices are evaluated using vectorized NumPy operations, reducing computational latency.
Our visualization relies on a carefully selected Cyberpunk Obsidian color palette designed for high contrast and modern developer-oriented presentation:
| Token Name | Hex Code | Visual Application | Conceptual Association |
|---|---|---|---|
| 🌌 C_BG | #0B0F19 |
Deep Obsidian Blue | Outer space backdrop |
| 🧪 C_CURVE | #00D4AA |
Luminous Teal | The Loss Function |
| 🎯 C_DOT | #FF6B6B |
Coral Red | Current parameter marker / dot |
| 📐 C_TANGENT | #FFD93D |
Vibrant Gold | Derivative vector / tangent slope |
| 🌿 C_ACCENT | #6BCB77 |
Mint Green | Optimization success / Global Minimum |
#FF4757 |
Neon Red | High loss states / parameter divergence | |
| 🎛️ C_PANEL_BG | #151922 |
Translucent Charcoal | Translucent panel dashboards |
- Stochastic and Mini-Batch Optimizers: Adding dynamic plots comparing standard GD with SGD and mini-batch trajectories.
- Adaptive Optimizers: Visualizing modern adaptive learning rate strategies including Adam, RMSProp, and Adagrad.
- Multivariate Dimensional Projections: Visualizing neural network loss curves by applying PCA dimensional reductions to complex manifolds.
- Web-Based Interactive Controls: Exporting scenes into JavaScript-controlled canvas wrappers, supporting real-time drag-and-drop parameter optimization on the web.
Contributions to scale and refine these animations are highly welcome!
- Fork this repository.
- Create a feature branch:
git checkout -b feature/amazing-animation. - Commit your modifications:
git commit -m 'feat: add Adam optimizer scene'. - Push changes:
git push origin feature/amazing-animation. - Submit a pull request.
This project is licensed under the terms of the MIT License. For full licensing parameters, see LICENSE.
Loved this visualization? Support mathematical open-source code by giving this repository a star! ⭐
