<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>InstructGPT on AI 实战派 · 从技术到赚钱</title><link>https://guijiagi.com/tags/instructgpt/</link><description>Recent content in InstructGPT on AI 实战派 · 从技术到赚钱</description><generator>Hugo</generator><language>zh-cn</language><copyright>本站内容采用 CC BY-NC-SA 4.0 国际许可协议授权</copyright><lastBuildDate>Wed, 07 Oct 2026 18:06:00 +0800</lastBuildDate><atom:link href="https://guijiagi.com/tags/instructgpt/index.xml" rel="self" type="application/rss+xml"/><item><title>论文解读：从 InstructGPT 到 DPO，对齐方法经历了什么</title><link>https://guijiagi.com/posts/2026-10-07-paper-instructgpt-rlhf-dpo/</link><pubDate>Wed, 07 Oct 2026 18:06:00 +0800</pubDate><guid>https://guijiagi.com/posts/2026-10-07-paper-instructgpt-rlhf-dpo/</guid><description>RLHF 让模型学会人类偏好，DPO 又绕开了强化学习。拆解这套「对齐三部曲」的问题、方法与演进逻辑。</description></item></channel></rss>