As a fellow neanderthal and after going through several tutorials and math videos on how PPO works, i still fail to grasp its intuition. I wanted to believe that I am not a dumb person, so instead, I blame PPO and RL for lacking a cohesive story that ties parts of the math to its mechanisms. Here is a fabricated story that strives to make an attempt. It is mainly for myself but if it helps please let me know.

References: https://claude.ai/public/artifacts/de1b44ef-e3a0-4bf1-99ed-f0b1edb9338e

Preface: An introduction to RL

The PPO story

Good to know

Tuning experiences

Legacy stuff but still useful

Actual implementation

Metrics

ML utils

Pytorch related

General Python stuff

Version control issues


Linear algebra insights