Introduction
John Schulman is an American research scientist who co-founded OpenAI and was a primary designer of ChatGPT's reinforcement learning systems before joining Anthropic in 2024.
PPO Algorithm
Schulman completed his PhD at UC Berkeley, where he worked on reinforcement learning. In 2017, he authored the paper describing Proximal Policy Optimization (PPO), which became the industry standard algorithm for training robotic agents and aligning language models.
ChatGPT Alignment
Schulman served as the leader of OpenAI's reinforcement learning from human feedback (RLHF) alignment team, helping turn base language models into the conversational interface known as ChatGPT.