<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Speculative-Decoding on Musings of an AI Wrangler</title><link>https://mazurov.dev/tags/speculative-decoding/</link><description>Recent content in Speculative-Decoding on Musings of an AI Wrangler</description><generator>Hugo -- 0.148.0</generator><language>en-us</language><lastBuildDate>Mon, 07 Sep 2026 15:33:53 -0600</lastBuildDate><atom:link href="https://mazurov.dev/tags/speculative-decoding/index.xml" rel="self" type="application/rss+xml"/><item><title>Speculative Decoding on GLM-5.3-Flash: DFlash vs Native MTP</title><link>https://mazurov.dev/posts/glm-speculation-benchmark/</link><pubDate>Mon, 07 Sep 2026 11:00:00 -0700</pubDate><guid>https://mazurov.dev/posts/glm-speculation-benchmark/</guid><description>A four-arm benchmark of speculative decoding on GLM-5.3-Flash (NVFP4, TP2 over two GB10 nodes): DFlash BF16, DFlash FP8 draft KV, native MTP, and no speculation. Native MTP&amp;#39;s bundled draft head costs about 20% of short-prompt throughput and leaves an 89% larger KV pool - a good trade for long-context serving.</description></item></channel></rss>