Skip to main content
Skip to main content
Back to Blog
ResearchAI SafetyCommunity

Rich Sutton: The Complete Guide to the Father of Reinforcement Learning and 2024 Turing Award Winner

From inventing temporal difference learning to winning the 2024 Turing Award, Rich Sutton built Edmonton into the world capital of reinforcement learning—and charts a contrarian path toward AGI that bets everything on learning by doing.

A

AGI House Canada

Community Team

January 11, 202615 min read
Rich Sutton: The Complete Guide to the Father of Reinforcement Learning and 2024 Turing Award Winner

Rich Sutton didn't just contribute to artificial intelligence—he invented how machines learn to act. The 2024 Turing Award recognized what AI researchers have known for decades: modern reinforcement learning exists because of Sutton's foundational work. From AlphaGo defeating world champions to ChatGPT learning from human feedback, the algorithms trace back to his lab in Edmonton.

Who Is Rich Sutton?

Richard S. Sutton is a Canadian-American computer scientist who co-invented reinforcement learning as we know it today. His work on temporal difference learning, policy gradient methods, and the textbook that trained a generation of researchers earned him the 2024 ACM A.M. Turing Award—the "Nobel Prize of Computing"—shared with his longtime collaborator Andrew Barto.

DetailInformation
Born1957/1958, United States
EducationBA Psychology, Stanford (1978); MA Psychology, UMass (1980); PhD Computer Science, UMass (1984)
Dissertation AdvisorAndrew Barto
Current RolesProfessor, University of Alberta; Chief Scientific Advisor, Amii; Co-founder, Keen Technologies
PreviousDistinguished Research Scientist, DeepMind (2017-2023)
Textbook Citations75,000+ (Reinforcement Learning: An Introduction)
Awards2024 Turing Award, IJCAI Research Excellence Award, IEEE Neural Network Pioneer Award

Sutton moved to Edmonton in 2003, transforming the University of Alberta into the global epicenter of reinforcement learning research. His students went on to create AlphaGo, lead DeepMind's research, and build the algorithms powering modern AI systems.

"Reinforcement learning is about understanding your world, whereas large language models are about mimicking people." — Rich Sutton, Interview with New Scientist, 2024

What Is Reinforcement Learning?

Reinforcement learning (RL) is the branch of AI where agents learn by trial and error—taking actions, receiving rewards or penalties, and improving over time. Unlike supervised learning (which requires labeled examples) or unsupervised learning (which finds patterns), RL teaches machines to act in the world.

Learning TypeData RequiredOutputExample
SupervisedLabeled examplesPredictionsImage classification
UnsupervisedUnlabeled dataPatternsCustomer clustering
ReinforcementEnvironment + rewardsActionsGame playing, robotics

Every system that learns from interaction—robots, game AI, recommendation engines, RLHF (Reinforcement Learning from Human Feedback)—builds on foundations Sutton established.

What Did Rich Sutton Invent?

Temporal Difference Learning (1988)

Sutton's most fundamental contribution: temporal difference (TD) learning combines ideas from dynamic programming and Monte Carlo methods to let agents learn from incomplete episodes—without waiting for final outcomes.

ConceptPre-TD LearningTD Learning
When learning happensOnly after episode endsAfter each step
Information neededComplete trajectoryCurrent transition
Real-world applicabilityLimitedImmediate

TD learning powers everything from game AI to trading algorithms. The breakthrough: machines could learn from predictions about predictions, not just final results.

Policy Gradient Methods (1999)

Sutton's policy gradient theorem provided the mathematical foundation for directly optimizing policies—the rules that determine what action to take. This enabled:

  • Continuous action spaces: Robots controlling motors, not just choosing from menus
  • Stochastic policies: Probabilistic decisions rather than deterministic rules
  • Actor-critic architectures: Combining value estimation with direct policy optimization

Modern systems like PPO (Proximal Policy Optimization)—the algorithm behind ChatGPT's RLHF—descend directly from Sutton's policy gradient work.

The Textbook (1998, 2018)

Reinforcement Learning: An Introduction, co-authored with Andrew Barto, is the field's defining text. With over 75,000 citations, it has trained virtually every modern RL researcher.

EditionYearPagesKey Additions
First1998322Foundational framework
Second2018526Deep RL, function approximation

The book's influence extends beyond academia—it's the reference engineers at Google, DeepMind, and OpenAI keep on their desks.

"For 40 years, Sutton and Barto have pursued a bold vision: that a machine, like a person, can develop skills only through trial and error. Reinforcement learning, as pioneered by Barto and Sutton, directly answers Turing's challenge to develop machines that learn by exploration." — Jeff Dean, Chief Scientist, Google DeepMind

Why Did Rich Sutton Win the 2024 Turing Award?

The ACM A.M. Turing Award—computer science's highest honor, carrying a $1 million prize—recognized Sutton and Barto for "foundational contributions to the field of reinforcement learning."

The ACM Citation

"Barto and Sutton's work demonstrates the immense potential of applying a multidisciplinary approach to develop powerful algorithms to help computers autonomously make decisions in complex environments." — Yannis Ioannidis, ACM President, 2024

The Impact Chain

Sutton's work traces directly to systems that changed the world:

YearSystemConnection to Sutton
2013DQN (Atari)TD learning + deep learning
2016AlphaGoBuilt by Sutton's student David Silver
2017AlphaZeroPolicy gradients + MCTS
2020GPT-3 RLHFReinforcement learning from human feedback
2023ChatGPTPPO (policy gradient descendant)

Every major AI breakthrough in the last decade used reinforcement learning algorithms that trace to Sutton's foundational work.

How Did Rich Sutton Build the Alberta RL Ecosystem?

When Sutton moved to the University of Alberta in 2003, Edmonton was not on the AI map. Two decades later, it's the undisputed world capital of reinforcement learning.

The University of Alberta Program

Sutton didn't just research RL—he built the infrastructure to train the next generation:

ComponentImpact
Reinforcement Learning and Artificial Intelligence Lab (RLAI)100+ graduates now leading AI worldwide
Canada CIFAR AI ChairAttracted federal investment
PhD studentsDavid Silver (AlphaGo), Doina Precup (DeepMind), Adam White (Keen)

DeepMind's Choice of Edmonton

In 2017, Google DeepMind established its only North American research lab in Edmonton—not San Francisco, not Toronto, not Montreal. Why? Because Sutton had built the world's best RL talent pipeline.

"I first met with Rich—our first ever adviser—seven years ago when DeepMind was just a handful of people with a big idea. He saw our potential and encouraged us from day one." — Demis Hassabis, CEO, Google DeepMind

Amii's Growth

As Chief Scientific Advisor to Amii (Alberta Machine Intelligence Institute), Sutton helped transform a regional initiative into one of Canada's three national AI institutes:

MetricValue
Amii Fellows36
Canada CIFAR AI Chairs26
Startup investments29
2025 federal funding$29M

For more on Amii's programs and ecosystem, see our Amii Complete Guide.

Who Are Rich Sutton's Most Notable Students?

Sutton's students didn't just learn RL—they deployed it to change the world.

David Silver: The Creator of AlphaGo

David Silver completed his PhD under Sutton at the University of Alberta, then led the team at DeepMind that created AlphaGo—the first AI to defeat a world champion at Go.

"I've got this inner Rich Sutton who sits on my shoulder and says 'well, it's experience that really matters'... AlphaGo was the culmination of 12 years of research that I began back in Alberta." — David Silver, Principal Research Scientist, Google DeepMind

Silver's AlphaGo moment in 2016—defeating Lee Sedol 4-1—marked AI's coming of age. The system used TD learning and policy gradients directly descended from Sutton's work.

Doina Precup: Leading DeepMind Montreal

Doina Precup co-created the options framework with Sutton (hierarchical RL), then went on to lead DeepMind Montreal while maintaining her McGill professorship.

Adam White: Sutton's Co-founder

Adam White, another Sutton PhD graduate, left DeepMind to co-found Keen Technologies with Sutton—their bet on building "full intelligence" through RL.

The Student Network

GraduateCurrent RoleContribution
David SilverDeepMindAlphaGo, AlphaZero
Doina PrecupDeepMind Montreal, McGillHierarchical RL, options framework
Adam WhiteKeen TechnologiesApplied RL, robotics
Michael BowlingDeepMind Alberta, UAlbertaPoker AI (Cepheus, DeepStack)
Patrick PilarskiUAlbertaMedical RL, prosthetics

"Rich has been an incredible mentor to both his students and his community." — Cam Linke, CEO, Amii

What Are Rich Sutton's Views on AI Safety?

Sutton holds contrarian views on AI safety that put him at odds with other AI pioneers like Geoffrey Hinton and Yoshua Bengio.

Sutton's Position

TopicSutton's ViewMainstream Concern
Existential risk"The doomers are out of line and the concerns are overblown"AI could pose existential threat
LLMs and AGI"I don't think that's the direction that's going to lead to full intelligence"LLMs may be path to AGI
Regulation urgencyOpposes hasty regulationCalls for immediate governance

The Fundamental Disagreement

While Hinton left Google to warn about AI dangers and Bengio pivoted Mila toward safety research, Sutton argues that:

  1. Current AI isn't close to dangerous — The systems that concern safety researchers (LLMs) aren't the path to AGI anyway
  2. RL systems are more controllable — Agents that learn from rewards can be shaped by what we reward them for
  3. Fear is counterproductive — Overregulation could slow beneficial AI development

"It's really too bad we have this name 'artificial intelligence.' Let's just call it intelligence." — Rich Sutton, Interview, 2024

The Canadian AI Debate

This creates productive intellectual tension within Canadian AI. Sutton at Amii debates Bengio at Mila on approaches, timelines, and risks—a disagreement that shapes Canadian AI policy discussions.

What Is Rich Sutton's AGI Timeline?

Sutton is more bullish on AGI timelines than many expect—but through reinforcement learning, not language models.

Sutton's Predictions

ProbabilityTimeframe
25% chanceBy 2030
50% chanceBy 2040
Almost certainThis century

Why Faster Than Expected?

Sutton believes AGI will arrive quickly once the field pivots from LLMs to genuine learning systems:

"I expect the transition to come sooner than others do, but I think that in part because I think large language models and things like them are going to crap out sooner than others do." — Rich Sutton, Dwarkesh Podcast, 2024

What Is The Alberta Plan?

Sutton outlined his vision for achieving "full intelligence" in a 12-stage framework called The Alberta Plan.

The 12 Stages

StageFocusGoal
1PredictionAccurate world models
2ControlActing effectively
3PlanningLookahead and simulation
4OptionsHierarchical abstraction
5KnowledgeFactual representation
6LanguageCommunication
7ExplorationCuriosity and discovery
8GeneralizationTransfer learning
9Multiple timescalesLong-term planning
10Continual learningNever stop improving
11ScalabilityEfficient with compute
12IntegrationAll components working together

Why This Matters

The Alberta Plan represents Sutton's counterproposal to the LLM scaling hypothesis. Rather than making language models bigger, he argues we need systems that actually learn about the world through interaction.

What Is Keen Technologies?

In 2024, Sutton co-founded Keen Technologies with former students to pursue his vision of "full intelligence" through reinforcement learning.

DetailInformation
Founded2024
Co-foundersRich Sutton, Adam White, others
FocusReinforcement learning systems for "full intelligence"
LocationEdmonton, Alberta
ApproachLearning from interaction, not just data

Keen represents Sutton's bet that the path to AGI runs through RL, not LLMs—and that Edmonton can be where it happens.

What Is The Bitter Lesson?

Sutton's most famous essay, "The Bitter Lesson" (2019), argues that AI progress comes from scaling compute, not clever human-designed features.

The Core Argument

"The biggest lesson that can be read from 70 years of AI research is that general methods that leverage computation are ultimately the most effective, and by a large margin." — Rich Sutton, "The Bitter Lesson," 2019

The Lesson's Impact

Human-Designed FeatureCompute-Based ApproachWinner
Chess knowledgeSearch + learningCompute
Go intuitionMonte Carlo Tree SearchCompute
Linguistic rulesStatistical learningCompute
Vision featuresDeep learningCompute

The essay became required reading in AI labs, influencing the shift toward scaling that produced GPT-3, GPT-4, and beyond.

The Irony

While Sutton's Bitter Lesson influenced the LLM scaling paradigm, he himself doesn't believe LLMs are the path to intelligence—they're just very good at pattern matching, not genuine understanding.

How Can You Learn from Rich Sutton?

Free Resources

ResourceAccessWhat You'll Learn
RL BookOnline freeComplete RL foundations
"The Bitter Lesson"Blog postSutton's philosophy
Course lecturesYouTube, UAlbertaAcademic-level instruction
PapersGoogle ScholarLatest research

Academic Pathway

To work directly in Sutton's ecosystem:

  1. Apply to University of Alberta — PhD in Computing Science
  2. Join RLAI Lab — Sutton's research group
  3. Connect through Amii — Fellowship and research programs

Industry Connection

  • Keen Technologies — Sutton's startup
  • DeepMind Alberta — Connected ecosystem
  • Amii-affiliated companies — 29+ invested startups

Frequently Asked Questions

Why is Rich Sutton called the "father of reinforcement learning"?

While others contributed to RL's foundations, Sutton created the mathematical framework (temporal difference learning, policy gradients) and the textbook that defined the field. His work transformed RL from a theoretical curiosity into the practical foundation of modern AI systems.

What's the difference between Rich Sutton and Geoffrey Hinton?

Both are Canadian AI pioneers, but they work on different problems. Hinton pioneered deep learning (neural network training), while Sutton pioneered reinforcement learning (learning from interaction). Hinton is at Vector Institute in Toronto; Sutton is at Amii in Edmonton. They also disagree on AI safety—Hinton warns of existential risks, Sutton considers those concerns overblown.

Did Rich Sutton create AlphaGo?

No, but his student David Silver led the AlphaGo team at DeepMind. The system used reinforcement learning algorithms that descend directly from Sutton's foundational work. Silver credits his PhD training under Sutton as formative for AlphaGo's development.

What does Rich Sutton think about ChatGPT and LLMs?

Sutton is skeptical that LLMs are the path to general intelligence. He argues that language models "mimic people" rather than "understand the world," and predicts they'll "crap out" before achieving AGI. He believes reinforcement learning—learning by doing—is the correct path.

Is Rich Sutton worried about AI safety?

No. Unlike Hinton and Bengio, Sutton considers AI safety concerns "overblown" and the "doomers out of line." He believes current AI systems aren't close to dangerous and that RL systems can be shaped through their reward functions.

Why did Rich Sutton move to Edmonton?

Sutton joined the University of Alberta in 2003, attracted by the computing science department's strength and the opportunity to build a research program from the ground up. His decision to stay—rather than move to a major tech hub—transformed Edmonton into the global center of reinforcement learning research.

What is the Alberta Plan?

Sutton's 12-stage outline for achieving "full intelligence" through reinforcement learning rather than language models. It emphasizes prediction, control, planning, hierarchical abstraction, and continual learning—capabilities he believes LLMs lack.

How can I study under Rich Sutton?

Apply to the PhD program in Computing Science at the University of Alberta and indicate interest in the RLAI (Reinforcement Learning and Artificial Intelligence) Lab. Competition is intense—Sutton's lab attracts top candidates globally.

How Can You Connect with the RL Community?

Edmonton

  • Amii events — Regular seminars and workshops
  • Upper Bound Conference — Western Canada's premier AI conference
  • RLAI Lab talks — Public research presentations

Online

  • r/reinforcementlearning — Active community
  • RL Discord servers — Real-time discussion
  • Papers With Code — Implementation resources

AGI House Canada

Join our community to connect with RL researchers and practitioners:


Related Reading

The Other Godfathers of AI

Sutton's Legacy: Western Canada AI

The Broader Canadian AI Ecosystem


Join AGI House Canada to connect with Canada's reinforcement learning community. Scan the QR code to join our WhatsApp group or subscribe to our newsletter.

Data current as of January 2026. Statistics from ACM announcements, Amii public reports, and published interviews. Quote attributions verified from public sources.

Share this post
A

AGI House Canada

Community Team