<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>技术 / 工具 on AI 实战派 · 从技术到赚钱</title><link>https://guijiagi.com/categories/%E6%8A%80%E6%9C%AF-/-%E5%B7%A5%E5%85%B7/</link><description>Recent content in 技术 / 工具 on AI 实战派 · 从技术到赚钱</description><generator>Hugo</generator><language>zh-cn</language><copyright>本站内容采用 CC BY-NC-SA 4.0 国际许可协议授权</copyright><lastBuildDate>Wed, 07 Oct 2026 18:54:00 +0800</lastBuildDate><atom:link href="https://guijiagi.com/categories/%E6%8A%80%E6%9C%AF-/-%E5%B7%A5%E5%85%B7/index.xml" rel="self" type="application/rss+xml"/><item><title>GPTQ 与 AWQ：权重量化两派，到底该选哪个</title><link>https://guijiagi.com/posts/2026-10-07-gptq-awq-quantization-deep-dive/</link><pubDate>Wed, 07 Oct 2026 18:54:00 +0800</pubDate><guid>https://guijiagi.com/posts/2026-10-07-gptq-awq-quantization-deep-dive/</guid><description>同样是把权重压到 4bit，GPTQ 用二阶误差补偿，AWQ 保护显著权重。拆解两派原理、质量差异与选型建议。</description></item><item><title>FlashAttention 与 KV Cache：长上下文为什么这么吃显存</title><link>https://guijiagi.com/posts/2026-10-07-flashattention-kv-cache-deep-dive/</link><pubDate>Wed, 07 Oct 2026 18:48:00 +0800</pubDate><guid>https://guijiagi.com/posts/2026-10-07-flashattention-kv-cache-deep-dive/</guid><description>注意力是 O(N²) 的显存怪兽，KV Cache 又随并发与长度膨胀。拆解 FlashAttention 怎么省显存，以及长上下文的成本账。</description></item><item><title>让模型稳定吐 JSON：结构化输出的三层防线</title><link>https://guijiagi.com/posts/2026-10-07-prompt-structured-output-json/</link><pubDate>Wed, 07 Oct 2026 18:42:00 +0800</pubDate><guid>https://guijiagi.com/posts/2026-10-07-prompt-structured-output-json/</guid><description>模型偶尔吐出的 JSON 缺括号、带多余文字。从 prompt 约束到 JSON mode 再到语法约束，三层手段把格式锁死。</description></item><item><title>Ollama 与 LM Studio 上手避坑：本地模型怎么调才顺</title><link>https://guijiagi.com/posts/2026-10-07-ollama-lmstudio-local-tips/</link><pubDate>Wed, 07 Oct 2026 18:36:00 +0800</pubDate><guid>https://guijiagi.com/posts/2026-10-07-ollama-lmstudio-local-tips/</guid><description>Ollama 偏命令行、LM Studio 偏图形界面。从模型选型、上下文长度到 API 调用，讲清两个本地工具的分工与坑。</description></item><item><title>投机解码实操：本地推理翻倍提速的正确姿势与坑</title><link>https://guijiagi.com/posts/2026-10-07-speculative-decoding-speedup/</link><pubDate>Wed, 07 Oct 2026 18:30:00 +0800</pubDate><guid>https://guijiagi.com/posts/2026-10-07-speculative-decoding-speedup/</guid><description>draft 模型先猜、主模型批量验证，吞吐能提 1.5 到 3 倍。但词表不匹配会反噬。手把手讲清怎么开、怎么调。</description></item></channel></rss>