<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>大模型 / 模型评测 on AI 实战派 · 从技术到赚钱</title><link>https://guijiagi.com/categories/%E5%A4%A7%E6%A8%A1%E5%9E%8B-/-%E6%A8%A1%E5%9E%8B%E8%AF%84%E6%B5%8B/</link><description>Recent content in 大模型 / 模型评测 on AI 实战派 · 从技术到赚钱</description><generator>Hugo</generator><language>zh-cn</language><copyright>本站内容采用 CC BY-NC-SA 4.0 国际许可协议授权</copyright><lastBuildDate>Wed, 07 Oct 2026 18:24:00 +0800</lastBuildDate><atom:link href="https://guijiagi.com/categories/%E5%A4%A7%E6%A8%A1%E5%9E%8B-/-%E6%A8%A1%E5%9E%8B%E8%AF%84%E6%B5%8B/index.xml" rel="self" type="application/rss+xml"/><item><title>长上下文评测：大海捞针之外，RULER 测的是真本事</title><link>https://guijiagi.com/posts/2026-10-07-long-context-needle-ruler-eval/</link><pubDate>Wed, 07 Oct 2026 18:24:00 +0800</pubDate><guid>https://guijiagi.com/posts/2026-10-07-long-context-needle-ruler-eval/</guid><description>宣称支持百万上下文的模型，真能找到针吗？从 Needle-in-a-Haystack 到 RULER，拆解长上下文能力的硬核测法。</description></item><item><title>MMLU 与 MMLU-Pro：知识基准为什么需要不断加难</title><link>https://guijiagi.com/posts/2026-10-07-mmlu-pro-knowledge-benchmark/</link><pubDate>Wed, 07 Oct 2026 18:18:00 +0800</pubDate><guid>https://guijiagi.com/posts/2026-10-07-mmlu-pro-knowledge-benchmark/</guid><description>MMLU 用 57 科多选题测通识，却被高分刷到饱和；MMLU-Pro 加干扰项、上推理题。拆解知识基准的演进与局限。</description></item><item><title>SWE-bench 透视：从 2% 到 97%，编码智能体的一年狂奔</title><link>https://guijiagi.com/posts/2026-10-07-swe-bench-coding-agent-eval/</link><pubDate>Wed, 07 Oct 2026 18:12:00 +0800</pubDate><guid>https://guijiagi.com/posts/2026-10-07-swe-bench-coding-agent-eval/</guid><description>SWE-bench Verified 用真实 GitHub issue 考 Agent。从最初的个位数到 2026 年逼近人类水平，拆解这个基准测的到底是什么。</description></item><item><title>论文解读：从 InstructGPT 到 DPO，对齐方法经历了什么</title><link>https://guijiagi.com/posts/2026-10-07-paper-instructgpt-rlhf-dpo/</link><pubDate>Wed, 07 Oct 2026 18:06:00 +0800</pubDate><guid>https://guijiagi.com/posts/2026-10-07-paper-instructgpt-rlhf-dpo/</guid><description>RLHF 让模型学会人类偏好，DPO 又绕开了强化学习。拆解这套「对齐三部曲」的问题、方法与演进逻辑。</description></item><item><title>论文解读：检索增强生成 RAG，让模型学会开卷考试</title><link>https://guijiagi.com/posts/2026-10-07-paper-retrieval-augmented-generation/</link><pubDate>Wed, 07 Oct 2026 18:00:00 +0800</pubDate><guid>https://guijiagi.com/posts/2026-10-07-paper-retrieval-augmented-generation/</guid><description>Lewis 2020 提出的 RAG 把外部检索塞进生成过程。拆解它如何让闭卷模型开卷答题，以及幻觉与时效性怎么治。</description></item><item><title>论文解读：混合专家 MoE，如何用稀疏激活换来参数规模暴涨</title><link>https://guijiagi.com/posts/2026-10-07-paper-moe-mixture-of-experts/</link><pubDate>Wed, 07 Oct 2026 17:54:00 +0800</pubDate><guid>https://guijiagi.com/posts/2026-10-07-paper-moe-mixture-of-experts/</guid><description>从 Outrageously Large Neural Networks 到 Mixtral、DeepSeekMoE。拆解稀疏 MoE 的路由机制、负载均衡与推理成本账。</description></item><item><title>论文重读：Attention Is All You Need，Transformer 为什么改写了一切</title><link>https://guijiagi.com/posts/2026-10-07-paper-attention-is-all-you-need/</link><pubDate>Wed, 07 Oct 2026 17:48:00 +0800</pubDate><guid>https://guijiagi.com/posts/2026-10-07-paper-attention-is-all-you-need/</guid><description>2017 年那篇只用注意力、不用循环与卷积的论文。拆解自注意力、多头机制与位置编码，看它为何成为大模型地基。</description></item></channel></rss>