# From First Principles

*A self-study path from rusty pre-university physics to training reinforcement-learning policies for drones and quadruped robots.*

**7 phases · 25–44 weeks · 8–12 focused hours a week** · last revised 2026-09-29

Understand the physics and maths that force each engineering decision, instead of memorising the decision. A quadcopter needs four motors for a reason you can derive. A GPU is fast at AI for a reason you can measure. Every phase starts with that reason, then gives you the few resources that teach it and a build that proves you have it.

## Contents

- [How to work through it](#how-to-work-through-it)
- [The path at a glance](#the-path-at-a-glance)
- [Phase 0 · Maths & Physics Refresher](#phase-0)
- [Phase 1 · Machine Learning & Deep Learning](#phase-1)
- [Phase 2 · Reinforcement Learning](#phase-2)
- [Phase 3 · NVIDIA GPU Architecture](#phase-3)
- [Phase 4 · Drones: Dynamics, Estimation & Control](#phase-4)
- [Phase 5 · Quadruped Robots](#phase-5)
- [Phase 6 · Capstone: Project Pegasus](#phase-6)

## How to work through it

- **Finish the core, skim the rest.** Each phase has at most four core resources. Finish them. Optional depth is there for when something does not click, not as a reading list to clear.
- **Build every phase.** The build is the real exit exam. Reading and watching without building will not get you to understanding why.
- **Explain it back.** Before moving on, answer every exit question out loud or in writing without notes. If you cannot, loop back to the part you are shaky on.
- **Keep a learning log.** Write a short weekly note: what you built, what broke, what you still do not trust. It becomes your Project Pegasus design notebook.

**A 10-hour week**

| Block | Hours | What |
|---|---|---|
| Study | 4 | Core lectures and reading, with pen and paper derivations. |
| Build | 4 | The phase build. Commit to a public or private repo every session. |
| Explain | 1 | Write the learning-log entry and answer one exit question from memory. |
| Review | 1 | Re-derive one earlier idea (Bellman equation, PID loop, chain rule) from scratch. |

## The path at a glance

| Phase | You come out able to… | Time | Starts |
|---|---|---|---|
| 0 · Maths & physics | Rotations, gradients and torque feel like tools, not symbols. | 2–4 wk | Start here |
| 1 · Deep learning | Backpropagation is the chain rule on a graph, and you can prove it in code. | 4–8 wk | After Phase 0 |
| 2 · Reinforcement learning | You have trained a PPO agent you wrote yourself, and can defend every line of its loss. | 6–10 wk | After Phase 1 |
| 3 · GPU architecture | You can say why a kernel is slow by looking at its memory access pattern. | 3–5 wk (alongside 2) | After Phase 1 |
| 4 · Drones | You can trace a position command all the way to four motor signals. | 4–6 wk | After Phase 0 |
| 5 · Quadrupeds | A policy you trained walks over rough simulated terrain. | 5–8 wk | After Phase 2 + Phase 3 |
| 6 · Capstone | The four tracks work as one system you built. | 4–8 wk | After Phase 4 + Phase 5 |

Totals assume the GPU phase runs alongside reinforcement learning. Run it on its own and add 3–5 weeks.

```
0 Maths & physics ─┬─► 1 Deep learning ─┬─► 2 Reinforcement learning ─┬─► 5 Quadrupeds ─┐
                   │                    └─► 3 GPU architecture ───────┘                 ├─► 6 Capstone
                   └─► 4 Drones ─────────────────────────────────────────────────────────┘
```

---

<a id="phase-0"></a>
## Phase 0 · Maths & Physics Refresher

**2–4 weeks · Start here**

> Rotations, gradients and torque feel like tools, not symbols.

### Why this matters

Every later phase is linear algebra, calculus, probability and Newtonian mechanics in different clothes. A neural network is matrix multiplications trained by following a gradient. A GPU is fast because it is built for the parallel structure of matrix multiplication. A drone flies by Newton's third law and stays level through feedback on torque and angular momentum. A quadruped balances by solving rigid-body dynamics in real time.

### Explain these without notes

- Why the gradient points in the direction of steepest ascent.
- What an eigenvector is, geometrically, and why it matters for stability.
- Why torque depends on the lever arm (τ = r × F), not just force magnitude.
- How a rotation matrix and a quaternion describe the same orientation.
- Why a propeller produces thrust: it accelerates air downward (momentum change), not mainly "Bernoulli".

### Core: finish these

- [ ] **[Essence of Linear Algebra, then MIT 18.06 problem sets](https://ocw.mit.edu/courses/18-06-linear-algebra-spring-2010/)** — 3Blue1Brown · Gilbert Strang *(Course, free)*. Geometric intuition first (about 3 hours of video), then rigour.
- [ ] **[Essence of Calculus, then MIT 18.02 Multivariable Calculus](https://ocw.mit.edu/courses/18-02-multivariable-calculus-fall-2007/)** — 3Blue1Brown · MIT OCW *(Course, free)*. Gradients, Jacobians and the chain rule. The chain rule is backpropagation.
- [ ] **[Stat 110: Probability](https://projects.iq.harvard.edu/stat110)** — Joe Blitzstein, Harvard *(Course, free)*. Expectation, conditional probability and Gaussians. Needed for RL and Kalman filters.
- [ ] **[Classical Mechanics (8.01)](https://ocw.mit.edu/courses/8-01sc-classical-mechanics-fall-2016/)** — MIT OCW, or Morin's Introduction to Classical Mechanics *(Course, free)*. Focus on Newton's laws, torque, angular momentum, moment of inertia and oscillation.

### Build: Rotations and a falling body in NumPy

1. Represent 3D rotations as rotation matrices and as quaternions, and convert between them. Check round-trips numerically.
2. Simulate a point mass under gravity plus one external force using Euler integration, and plot the trajectory.
3. Swap Euler for semi-implicit Euler or RK4 and explain why the energy drift changes.

### Ready to move on when…

- [ ] I can explain gradients, eigenvectors and torque without notes.
- [ ] My rotation code round-trips matrix ↔ quaternion correctly.

### Optional depth

- **[Linear Algebra and Physics tracks](https://www.khanacademy.org)** — Khan Academy *(Course, free)*. If MIT OCW moves too fast.
- **Div, Grad, Curl, and All That** — H. M. Schey *(Book, paid)*. Vector calculus intuition before robot dynamics.
- **[Python, NumPy, Git and the Linux shell](https://numpy.org/learn/)** — Any short tutorial *(Tool, free)*. Every build from here on assumes these.

---

<a id="phase-1"></a>
## Phase 1 · Machine Learning & Deep Learning

**4–8 weeks · After Phase 0**

> Backpropagation is the chain rule on a graph, and you can prove it in code.

### Why this matters

Reinforcement learning, the software on NVIDIA GPUs and the controllers of modern legged robots and drones are all deep neural networks trained by gradient descent. You need to see how a network learns as repeated application of the chain rule from Phase 0. Without that, RL looks like magic.

### Explain these without notes

- Backpropagation as the chain rule applied to a computation graph.
- Why we use mini-batches, and what the learning rate trades off.
- Overfitting, the bias–variance trade-off and what regularisation does.
- Why convolutions share weights, and what that buys for images.

### Core: finish these

- [ ] **[Neural Networks: Zero to Hero](https://karpathy.ai/zero-to-hero.html)** — Andrej Karpathy *(Course, free)*. Build micrograd, then a language model, then a GPT, with nothing hidden. Type every notebook yourself.
- [ ] **[Dive into Deep Learning](https://d2l.ai)** — Zhang, Lipton, Li & Smola *(Book, free)*. Code-first reference text to read alongside everything else.
- [ ] **[CS231n: Deep Learning for Computer Vision](https://cs231n.stanford.edu)** — Stanford *(Course, free)*. Best-taught "why does this architecture work" course. Vision matters for obstacle avoidance and terrain perception.
- [ ] **[PyTorch tutorials: Learn the Basics](https://pytorch.org/tutorials/)** — PyTorch *(Tool, free)*. The framework everything downstream uses.

### Build: micrograd from memory, then a CNN

1. Rebuild a tiny autograd engine (micrograd) without looking at Karpathy's code.
2. Use it to train a small classifier and check its gradients against finite differences.
3. Train a small CNN on CIFAR-10 in PyTorch and beat a naive baseline. Log loss curves.

### Ready to move on when…

- [ ] I can write a feed-forward network's forward and backward pass without a framework.
- [ ] I can read and debug PyTorch training code comfortably.

### Optional depth

- **[Machine Learning Specialization](https://www.coursera.org/specializations/machine-learning-introduction)** — Andrew Ng, DeepLearning.AI *(Course, paid)*. Gentler classical-ML on-ramp: regression, gradient descent, bias–variance.
- **[Deep Learning](https://www.deeplearningbook.org)** — Goodfellow, Bengio & Courville *(Book, free)*. The mathematical reference when d2l is not deep enough.
- **[Practical Deep Learning for Coders](https://course.fast.ai)** — fast.ai *(Course, free)*. Top-down alternative if bottom-up is not working for you.

---

<a id="phase-2"></a>
## Phase 2 · Reinforcement Learning

**6–10 weeks · After Phase 1**

> You have trained a PPO agent you wrote yourself, and can defend every line of its loss.

### Why this matters

RL is the maths of an agent that improves through trial, error and reward. It is how a quadruped learns to walk on rough ground and how a drone learns an agile manoeuvre. Understanding it from first principles means understanding the Markov Decision Process (the agent acts, the world returns a new state and a reward) and why value-based, policy-gradient and actor-critic methods are justified solutions to it, not arbitrary tricks.

### Explain these without notes

- The Bellman equation, derived from the definition of return.
- On-policy vs off-policy learning, and why it matters for sample efficiency.
- Why policy gradients have high variance, and how baselines and advantages reduce it.
- What PPO's clipped objective prevents, and why GAE trades bias for variance.
- Why continuous control (motor torques) needs different algorithms from Atari-style discrete actions.

### Core: finish these

- [ ] **[Reinforcement Learning: An Introduction (2nd ed.)](http://incompleteideas.net/book/the-book-2nd.html)** — Sutton & Barto *(Book, free)*. The field's textbook. Read parts I and II in order; do not skip the bandit chapters.
- [ ] **[UCL Course on RL](https://www.davidsilver.uk/teaching/)** — David Silver *(Course, free)*. Tracks Sutton & Barto lecture for lecture. Watch alongside the book.
- [ ] **[Spinning Up in Deep RL](https://spinningup.openai.com)** — OpenAI *(Course, free)*. The bridge from theory to the deep RL used on robots: policy gradients, PPO, SAC.
- [ ] **[CS285: Deep Reinforcement Learning](https://rail.eecs.berkeley.edu/deeprlcourse/)** — Sergey Levine, UC Berkeley *(Course, free)*. The research-level version, with assignments.

### Papers, in reading order

1. **[Playing Atari with Deep Reinforcement Learning (DQN)](https://arxiv.org/abs/1312.5602)** — Mnih et al., 2013. Deep networks + Q-learning, stabilised by replay and target networks.
2. **[Trust Region Policy Optimization](https://arxiv.org/abs/1502.05477)** — Schulman et al., 2015. Why policy updates must be kept small.
3. **[High-Dimensional Continuous Control Using GAE](https://arxiv.org/abs/1506.02438)** — Schulman et al., 2015. The advantage estimator PPO uses.
4. **[Proximal Policy Optimization Algorithms](https://arxiv.org/abs/1707.06347)** — Schulman et al., 2017. The default algorithm for legged-robot training today.
5. **[Continuous Control with Deep RL (DDPG)](https://arxiv.org/abs/1509.02971)** — Lillicrap et al., 2015. Off-policy actor-critic for continuous actions.
6. **[Soft Actor-Critic](https://arxiv.org/abs/1801.01290)** — Haarnoja et al., 2018. Maximum-entropy RL; sample-efficient off-policy control.

### Build: Three agents, from scratch

1. Tabular Q-learning on a gridworld. Plot the learned value function.
2. DQN on CartPole (Gymnasium) with a replay buffer and target network.
3. PPO on Pendulum-v1, then HalfCheetah in MuJoCo. No RL library.

### Ready to move on when…

- [ ] I can derive the Bellman equation on paper.
- [ ] My own PPO visibly improves on a continuous-control task, and I can explain its clipping term and advantage estimate.

### Optional depth

- **[CS234: Reinforcement Learning](https://web.stanford.edu/class/cs234/)** — Emma Brunskill, Stanford *(Course, free)*. Problem sets for rigour.
- **[The 37 Implementation Details of PPO](https://iclr-blog-track.github.io/2022/03/25/ppo-implementation-details/)** — Huang et al., ICLR Blog Track 2022 *(Article, free)*. Read when your PPO does not learn and the paper says it should.
- **[CleanRL](https://github.com/vwxyzjn/cleanrl)** — Costa Huang et al. *(Tool, free)*. Single-file reference implementations to diff your own code against.

---

<a id="phase-3"></a>
## Phase 3 · NVIDIA GPU Architecture

**3–5 weeks · After Phase 1 · runs alongside Phase 2**

> You can say why a kernel is slow by looking at its memory access pattern.

### Why this matters

Every RL policy you train, and every large simulated-robot run in Phase 5, is limited by how fast you can do matrix multiplications and how fast you can feed them data. A CPU has a few powerful cores built for sequential, branchy work. A GPU has thousands of simple lanes executing the same instruction on different data (SIMT). That structural difference is why GPUs, not CPUs, made modern deep learning and massively parallel robot simulation possible.

### Explain these without notes

- Why a "CUDA core" is not comparable to a CPU core: a warp of 32 threads shares one instruction stream.
- The memory hierarchy (registers → shared memory/L1 → L2 → HBM) and why coalesced access matters.
- What a Tensor Core does: a whole small matrix multiply-accumulate per instruction, instead of one scalar operation.
- The memory wall: why bandwidth, not FLOPs, is often the real limit (arithmetic intensity, roofline).
- Why thousands of simulated robots fit on one GPU.

### Core: finish these

- [ ] **Programming Massively Parallel Processors (4th ed.)** — Hwu, Kirk & El Hajj *(Book, paid)*. The standard first-principles CUDA text. Chapters on SMs, warps, memory and tiling are essential.
- [ ] **[GPU MODE lectures](https://github.com/gpu-mode/lectures)** — GPU MODE community *(Course, free)*. Practitioners walking through real kernel optimisation. Pairs chapter-for-chapter with PMPP.
- [ ] **[One NVIDIA architecture whitepaper, read closely](https://www.nvidia.com/en-us/data-center/resources/)** — NVIDIA (Ampere, Hopper or Blackwell) *(Paper, free)*. Ground the book in a real, current chip: SM layout, Tensor Core generations, memory bandwidth.
- [ ] **[llm.c](https://github.com/karpathy/llm.c)** — Andrej Karpathy *(Tool, free)*. GPT-2 training in plain CUDA. Connects Phase 1 maths directly to Phase 3 hardware.

> **Hardware:** You need an NVIDIA GPU for this phase and Phase 5. A cloud GPU rented by the hour is fine; a consumer RTX card works for everything here.

### Build: Matmul, three ways

1. Write a naive CUDA matrix-multiply kernel and benchmark it.
2. Write a tiled, shared-memory version and benchmark again.
3. Compare both against cuBLAS and explain each gap with the memory hierarchy.

### Ready to move on when…

- [ ] I can explain the naive → tiled → cuBLAS speed gaps from first principles.
- [ ] I can explain why robot RL trains in GPU-parallel simulators rather than one robot at a time.

### Optional depth

- **[CUDA C++ Programming Guide](https://docs.nvidia.com/cuda/cuda-c-programming-guide/)** — NVIDIA *(Article, free)*. The reference for syntax and the execution model.
- **[Making Deep Learning Go Brrrr From First Principles](https://horace.io/brrr_intro.html)** — Horace He *(Article, free)*. Compute-, memory- or overhead-bound: where training time actually goes.
- **[How to Optimize a CUDA Matmul Kernel](https://siboehm.com/articles/22/CUDA-MMM)** — Simon Boehm *(Article, free)*. A step-by-step companion for this phase's build.

---

<a id="phase-4"></a>
## Phase 4 · Drones: Dynamics, Estimation & Control

**4–6 weeks · After Phase 0**

> You can trace a position command all the way to four motor signals.

### Why this matters

A multirotor is an underactuated rigid body: six degrees of freedom (three translational, three rotational) controlled by only four motor thrusts. Why a quadcopter needs four motors, and how differential motor speeds create roll, pitch and yaw torques, is the core of flight. Flight controllers, PID tuning and sensor fusion all exist to solve that underactuated control problem hundreds of times a second.

*Only needs Phase 0, so it can start earlier if you want a break from ML. It sits here so the flight-control ideas are fresh for quadrupeds.*

### Size classes are different physics problems

| Class | Mass | What changes |
|---|---|---|
| Nano | < 250 g | Crazyflie-class. Drag and battery energy density dominate, flights last minutes and payload is tiny. Cheap and safe to crash, which is why it is the classic RL research platform. |
| Micro | 250 g – 2 kg | FPV and hobby builds. Room for a flight controller, camera and a Raspberry Pi-class computer. |
| Full-size | 2 – 25 kg | Room for Jetson-class onboard GPUs and lidar. The class used in serious autonomy research and delivery or inspection work. |

### Explain these without notes

- Why propellers spin in alternating directions (cancelling reactive yaw torque) and how differential thrust produces roll, pitch and yaw.
- Why attitude estimation needs sensor fusion: gyros drift over time, accelerometers are noisy moment to moment.
- The cascaded control structure: slow position/attitude loop feeding a fast angular-rate loop.
- What PID gains do to a step response, and when you need more than PID (LQR, MPC).
- Why sim-to-real transfer is essential and imperfect.

### Core: finish these

- [ ] **Control fundamentals: PID, state space, LQR** — Brian Douglas (Control Systems Lectures) or Steve Brunton (Control Bootcamp), YouTube *(Course, free)*. The control theory the rest of this phase assumes. Do this first.
- [ ] **[Small Unmanned Aircraft: Theory and Practice](https://github.com/randybeard/uavbook)** — Randal Beard & Timothy McLain *(Book, free)*. Derives 6-DOF rigid-body dynamics, then sensors, Kalman filtering and control, with a companion simulator.
- [ ] **[Aerial Robotics](https://www.coursera.org/learn/robotics-flight)** — Vijay Kumar, Penn (Coursera) *(Course, paid)*. First-principles quadrotor dynamics and control.
- [ ] **[Kalman and Bayesian Filters in Python](https://github.com/rlabbe/Kalman-and-Bayesian-Filters-in-Python)** — Roger Labbe *(Book, free)*. Interactive notebooks. Makes sensor fusion concrete.

### Papers, in reading order

1. **[Champion-level drone racing using deep reinforcement learning](https://www.nature.com/articles/s41586-023-06419-4)** — Kaufmann et al., Nature 2023. An RL-trained quadrotor beating human world champions. Read after your PID build.

### Build: Cascaded PID in simulation, then on hardware

1. Write a complementary filter (then a Kalman filter) that estimates attitude from simulated IMU data.
2. Implement a cascaded PID controller that holds a stable hover in simulation.
3. Track a circular trajectory and plot tracking error.
4. If you have a Crazyflie, fly the same controller and write down every sim-to-real difference.

### Ready to move on when…

- [ ] I can explain the full path from desired position to individual motor commands.
- [ ] My simulated quad hovers and tracks a circle with bounded error.

### Optional depth

- **[Crazyflie 2.x platform and docs](https://www.bitcraze.io)** — Bitcraze *(Hardware, paid)*. Open hardware and firmware nano quad, widely used in RL-for-drones research.
- **[PX4 / ArduPilot source and developer docs](https://docs.px4.io)** — Dronecode · ArduPilot *(Tool, free)*. Read the attitude-control and estimator code: Phase 0 quaternions as real C++.
- **[gym-pybullet-drones](https://github.com/utiasDSL/gym-pybullet-drones)** — UTIAS Dynamic Systems Lab *(Tool, free)*. Crazyflie-style simulator with a Gymnasium interface, ready for RL.

---

<a id="phase-5"></a>
## Phase 5 · Quadruped Robots

**5–8 weeks · After Phase 2 + Phase 3**

> A policy you trained walks over rough simulated terrain.

### Why this matters

Legged locomotion is harder than flight. Which feet touch the ground changes the equations of motion from moment to moment, and those contacts depend on unpredictable friction and terrain. Hand-derived controllers handle flat, predictable ground but struggle elsewhere. That is why the field moved to RL policies trained over millions of trials in GPU-parallel simulation. Phases 2 and 3 meet here.

### Explain these without notes

- Why contact makes legged locomotion a hybrid (continuous + switching) dynamical system.
- Why domain randomisation (friction, mass, terrain, motor strength) closes the sim-to-real gap.
- Model-based control (MIT Cheetah: derive the model, solve MPC) vs learned control (ETH: train a policy), and why state of the art blends them.
- Why actuator models and observation noise matter as much as the reward function.

### Core: finish these

- [ ] **[Modern Robotics: Mechanics, Planning, and Control](https://hades.mech.northwestern.edu/index.php/Modern_Robotics)** — Kevin Lynch & Frank Park *(Book, free)*. Rigid-body kinematics and dynamics with free lecture videos. Focus on chapters 2–5 and 8.
- [ ] **Legged Robots That Balance** — Marc Raibert *(Book, paid)*. The classical foundation: hopping, balance and the three-part control decomposition.
- [ ] **[Isaac Lab](https://isaac-sim.github.io/IsaacLab/)** — NVIDIA *(Tool, free)*. GPU-parallel robot learning: thousands of robots on one GPU. Phase 3 pays off here.
- [ ] **[MuJoCo + MuJoCo Menagerie](https://mujoco.org)** — Google DeepMind *(Tool, free)*. Accurate contact physics with ready quadruped models (Unitree Go1/Go2, ANYmal).

### Papers, in reading order

1. **[Learning Agile and Dynamic Motor Skills for Legged Robots](https://arxiv.org/abs/1901.08652)** — Hwangbo et al., Science Robotics 2019. RL policy on real ANYmal hardware, with a learned actuator model.
2. **[Learning Quadrupedal Locomotion over Challenging Terrain](https://arxiv.org/abs/2010.11251)** — Lee et al., Science Robotics 2020. Robust RL locomotion on rocks, mud and snow.
3. **[Learning to Walk in Minutes Using Massively Parallel Deep RL](https://arxiv.org/abs/2109.11978)** — Rudin et al., CoRL 2021. The paper that joins Phase 3 and Phase 5: thousands of robots, minutes of training.
4. **[RMA: Rapid Motor Adaptation for Legged Robots](https://arxiv.org/abs/2107.04034)** — Kumar et al., RSS 2021. Adapting online to terrain the robot was never trained on.
5. **Dynamic Locomotion in the MIT Cheetah 3 Through Convex MPC** — Di Carlo et al., IROS 2018. The model-based counterpoint to the RL papers above.

### Build: Teach a simulated quadruped to walk

1. Using your own PPO from Phase 2, train a Go1/Go2 or ANYmal model to walk forward on flat ground.
2. Add domain randomisation (friction, mass, motor strength) and observation noise.
3. Add rough terrain with a curriculum and measure robustness before and after.
4. Write a one-page sim-to-real plan for your setup: what would break on hardware and why.

### Ready to move on when…

- [ ] My policy visibly walks on uneven simulated terrain.
- [ ] I can explain what would need to change to deploy it on a physical robot.

### Optional depth

- **[legged_gym and rsl_rl](https://github.com/leggedrobotics/legged_gym)** — ETH Zurich RSL *(Tool, free)*. Reference code for the Rudin et al. training setup.
- **Unitree Go2 models and docs** — Unitree *(Hardware, free)*. A widely supported reference quadruped to target in simulation.

---

<a id="phase-6"></a>
## Phase 6 · Capstone: Project Pegasus

**4–8 weeks · After Phase 4 + Phase 5**

> The four tracks work as one system you built.

### Why this matters

By now the tracks should feel like one connected picture. The capstone proves it by building something that uses all of them, and it becomes the first real Project Pegasus prototype.

### Pick one or more

- **RL vs PID on a nano drone.** Train an RL policy on GPU to fly a simulated Crazyflie-class quad through hover and trajectory tracking. Compare it head to head with your Phase 4 PID controller on the same dynamics model.
- **Quadruped with a job.** Extend the Phase 5 walker with a harder objective: carrying a variable payload, or avoiding obstacles using a depth camera.
- **Reproduce a recent paper.** Pick a drone or legged-robot RL paper from the last two years and reproduce its core result within your compute budget.

### Ready to move on when…

- [ ] I shipped at least one capstone with code, results and a write-up.
- [ ] I can name the next three things Project Pegasus needs, and why.

---

*Generated from `roadmap-data.js` by `tools/build-roadmap.js`. Edit the data file, not this one.*
