<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>开源 on AI 实战派 · 从技术到赚钱</title><link>https://guijiagi.com/categories/%E5%BC%80%E6%BA%90/</link><description>Recent content in 开源 on AI 实战派 · 从技术到赚钱</description><generator>Hugo</generator><language>zh-cn</language><copyright>本站内容采用 CC BY-NC-SA 4.0 国际许可协议授权</copyright><lastBuildDate>Wed, 07 Oct 2026 17:18:00 +0800</lastBuildDate><atom:link href="https://guijiagi.com/categories/%E5%BC%80%E6%BA%90/index.xml" rel="self" type="application/rss+xml"/><item><title>LLaMA-Factory 实战：零代码微调如何跑通一份指令数据集</title><link>https://guijiagi.com/posts/2026-10-07-llama-factory-finetuning/</link><pubDate>Wed, 07 Oct 2026 17:18:00 +0800</pubDate><guid>https://guijiagi.com/posts/2026-10-07-llama-factory-finetuning/</guid><description>LLaMA-Factory 把 LoRA、QLoRA、全参微调压进一个 Web UI。从数据集格式到训练参数，讲清一次微调怎么落地。</description></item><item><title>vLLM 服务端解构：PagedAttention 如何把吞吐推上数量级</title><link>https://guijiagi.com/posts/2026-10-07-vllm-paged-attention-serving/</link><pubDate>Wed, 07 Oct 2026 17:06:00 +0800</pubDate><guid>https://guijiagi.com/posts/2026-10-07-vllm-paged-attention-serving/</guid><description>从 PagedAttention 到连续批处理，再到 2026 年的 Model Runner V2，vLLM 是开源推理服务的吞吐标杆。拆解它的调度内核。</description></item><item><title>llama.cpp 深度拆解：消费级显卡凭什么跑起 320B 大模型</title><link>https://guijiagi.com/posts/2026-10-07-llama-cpp-inference-engine/</link><pubDate>Wed, 07 Oct 2026 17:00:00 +0800</pubDate><guid>https://guijiagi.com/posts/2026-10-07-llama-cpp-inference-engine/</guid><description>从 GGUF 量化到 MTP 投机解码，llama.cpp v0.6.0 把 320B MoE 拉进了消费级显存。拆解它的工程内核与适用边界。</description></item></channel></rss>