跳到主要内容

Reinforcement Learning: An Introduction — 书籍拆解

读到哪:未读。 readState 不是 read/partial 的书不能当锚

作者Richard S. Sutton、Andrew G. Barto
版次2nd edition (2018, MIT Press;作者站点免费 PDF)
格式pdf | 文本源 pdftotext -layout
许可作者官网免费发布(MIT Press 版权)
来源作者 Richard Sutton 个人站点 incompleteideas.net 教材页(the-book-2nd.html)挂出的免费 PDF,2026-08-23 取
清洗删页眉页脚 329 行、页码 255 行、断词接回 208 处、按书规则修复 2231 处

我们重写的拆解(0 章)

(还没写。拆解是这本书对我们的真正产出——底下的元数据只是索引。)

为什么收它

强化学习的标准教材。以后碰到「agent 怎么从反馈里学」这类问题,它是最硬的原理出处。

合法性

作者官网免费发布(MIT Press 版权,作者获准放出电子版)。原始文件不入库,转码文本入库(私有库)。

转码注记

这本 PDF 的字体编码有坏字( 实为 ff、 实为 *),已在 library/reinforcement-learning-sutton-barto/clean-rules.json 配了修复规则, 本轮修了 2231 处;另删掉页眉页脚 329 行、页码 255 行。

它大概覆盖什么、不覆盖什么

覆盖:RL 的完整理论体系(赌博机、MDP、动态规划、蒙特卡洛、TD、函数近似、策略梯度)。 不覆盖:深度 RL 的工程实践、LLM 时代的 RLHF——出版年代早于那些。

它覆盖什么、不覆盖什么

(还没读到能下判断的程度。claims / notCovered 空着就是空着,不猜。)

怎么引用它

(依据: book=reinforcement-learning-sutton-barto §Introduction)

章节名对不上会被 lab:validate 拦下;页码锚(§p.123)同样可用。

结构(22 段,共 1876k 字符)

章节规模
01Preface to the Second Editionp.13–1614.8k
02Preface to the First Editionp.17–186.3k
03Summary of Notationp.19–228.9k
04Introductionp.23–4677.6k
05Multi-armed Banditsp.47–6864.3k
06Finite Markov Decision Processesp.69–9488.8k
07Dynamic Programmingp.95–11258.9k
08Monte Carlo Methodsp.113–14086.8k
09Temporal-Difference Learningp.141–16277.4k
10n-step Bootstrappingp.163–18047.3k
11Planning and Learning with Tabular Methodsp.181–218116.9k
12On-policy Prediction with Approximationp.219–264145.1k
13On-policy Control with Approximationp.265–27842.4k
14*Off-policy Methods with Approximationp.279–308206.3k
15Eligibility Tracesp.309–342111.8k
16Policy Gradient Methodsp.343–36255.2k
17Psychologyp.363–398120.3k
18Neurosciencep.399–442149.8k
19Applications and Case Studiesp.443–480164.8k
20Frontiersp.481–50276.3k
21Referencesp.503–540131.5k
22Indexp.541–54924.8k

我们自己的读书笔记(0 篇)

(还没有。读完某章后写进 docs/reinforcement-learning-sutton-barto/notes/,那才是这本书对我们的产出。)


本页由 node scripts/book-build.mjs 生成:表格来自转码结果,散文来自书卡正文。不要手改本页。