I keep seeing AI inference described as if the only job is to make the next token arrive faster.
But a large language model normally advances one token at a time, revisiting its weights and working with its growing KV state at every target-model step.
Speculative decoding asks a clever question: what if a smaller model writes several likely next tokens first, and the larger model checks them together?
If several guesses survive, one expensive target-model pass has done the work of several serial steps. If they do not, the extra draft work and memory can become part of the bill. Which side wins?
Previous Tech(EN) post: MoE Uses Fewer Experts Per Token. Why Does It Still Need So Much Memory?
Key Takeaways
- Speculative decoding lets a smaller drafter propose several output tokens before a larger target model verifies them in parallel.
- The potential gain comes from accepting more useful tokens per target-model iteration, not from making target weights or KV state unnecessary.
- Acceptance length, draft overhead, context length, batch size, and resident-memory pressure determine whether the method helps.
- Exact speculative sampling preserves the target distribution only under the required acceptance, rejection, and correction rules; it does not mean two random runs will always display the same text.
- Interactive chat latency and fleet-level throughput are different optimization goals. A good result for one does not automatically establish a good result for the other.
AI-generated original conceptual illustration for this article series; no third-party images, paper figures, logos or trademarks were reproduced.
The Serial Work That Speculation Tries to Rearrange
Autoregressive decoding is naturally serial. The target model produces one token, adds that token to the context, and then runs again to produce the next one. Each step uses the model’s weights and the active KV state built from the prior context.
At a low effective batch size, this repeated target execution can be constrained by memory movement and launch overhead as much as by raw arithmetic. The model may have plenty of calculations to do, yet it still has to repeatedly bring the right weights into use and access the state needed for attention.
Speculative decoding does not remove those target-model resources. It tries to make each target pass accomplish more. A smaller draft model proposes D possible next tokens. The target then evaluates that candidate sequence together, accepting proposals in order until it finds a mismatch.
NVIDIA calls the resulting number of output tokens per target iteration the acceptance length, or AL. With D proposed tokens, AL can range from 1 to D+1. The higher the useful acceptance length, the more serial target steps the system may avoid.
The key word is “may.” The drafter must generate candidates, the target must verify them, and the system must hold the weights and state needed for both sides. Speculation changes the shape of the work; it does not make inference free.
One Simple Verification Flow
A simple example makes the mechanism easier to see. Suppose the draft model proposes four tokens after a prompt: A, B, C, and D.
- The draft model produces A, B, C, and D sequentially.
- The target model verifies that proposed sequence in a larger pass.
- If the target agrees with A, B, and C but disagrees at D, the system accepts the first three candidates and uses the target’s result at the mismatch.
- The next round begins from that accepted prefix.
In that example, one verification pass produced several usable output tokens. But if the target disagrees immediately, the draft effort may deliver little benefit. This is why acceptance length matters more than the marketing-friendly idea that the system “guesses ahead.”
NVIDIA’s September 2, 2026 technical explainer on speculative decoding describes this trade-off directly: draft length should be chosen together with acceptance length and draft overhead. Increasing the number of candidates is not automatically faster, especially after verification becomes compute-bound or when longer verification adds attention and memory pressure.
“Same Output” Needs a Decoding-Mode Footnote
One of the most common shortcuts in this topic is saying that speculative decoding gives “the same output, only faster.” That can be true in a carefully defined sense, but the definition matters.
For stochastic generation, the original speculative decoding paper by Leviathan, Kalman, and Matias presents an exact sampling method. When its proposal distribution, acceptance and rejection rules, and correction sampling are applied correctly, it preserves the target model’s output distribution.
Distribution-preserving does not mean two random generations will always print the identical sentence. Random sampling can produce different valid draws from the same target distribution, whether speculation is used or not.
Greedy decoding is a different case. A deterministic verifier can keep only the tokens confirmed by the target and fall back to the target at the first mismatch. Under the same tokenizer, logit processing, tie-breaking, and acceptance rule, that can preserve the standard greedy sequence.
Approximate variants require another level of care. If a method relaxes acceptance, changes the model, or changes the routing behavior, it should be measured for quality and output behavior rather than described as exact distribution-preserving decoding.
Why the Draft Model Is Not Free
The draft model has to come from somewhere. An external drafter brings separate weights and a full KV cache of its own. Target-attached approaches, such as draft heads, can change that footprint, but they still add weights, hidden-state handling, or a smaller cache requirement.
That makes memory capacity part of the choice. A serving stack needs room for target weights, target KV cache, draft resources, candidate verification, and concurrent requests. Long contexts can make the KV portion more prominent; a high number of simultaneous users can turn modest per-request overhead into a system-level constraint.
There is also a workload question. A short, interactive request with a small batch may benefit from fewer serial target iterations. A high-concurrency fleet may already use batching to keep the target busy. In that setting, extra verification, draft computation, communication, and attention work can move the optimal point.
In other words, a faster chat response and a higher fleet token throughput are not the same claim. Both should be measured at a stated model, precision, context length, batch size, hardware topology, and latency objective.
Memory Bandwidth Still Matters—Just Differently
Speculative decoding is sometimes described as a way around the memory wall. I think of it more as a way to get more useful output from a costly target-model pass when conditions are favorable.
Where serial decode is memory-bound, verifying multiple candidates can raise the amount of useful work done per target iteration. That may improve arithmetic intensity. But a longer draft can also add target verification work, drafter work, candidate-related state, and pressure on KV capacity.
The result depends on acceptance. If proposed tokens are often accepted, the target pass is spread across more delivered tokens. If rejections arrive early, the draft’s sequential work becomes overhead. If context is long or available memory is tight, the extra state can change the balance again.
That is why it is too simple to conclude that speculation lowers HBM demand. It may improve utilization at one serving point, while total memory need still depends on the target, the drafter, concurrency, context, and the overall amount of AI work customers choose to run.
Related Companies: System Exposure Is Not a Supply Claim
NVIDIA (Nasdaq: NVDA) has direct software and accelerator exposure to this conversation. Its TensorRT-LLM materials list speculative decoding, quantization, paged KV caching, and other inference optimizations for NVIDIA GPUs. That makes runtime efficiency and accelerator utilization relevant platform variables. It does not show that a named customer uses a particular speculative method, or that the feature determines GPU revenue in one direction.
AMD (Nasdaq: AMD) is relevant through the accelerator-memory system. Its Instinct MI350 Series page, checked September 28, 2026, lists up to 288 GB of HBM3E and up to 8 TB/s of peak theoretical memory bandwidth per GPU. Those specifications are inputs to a target-and-draft memory design; they are not application-throughput evidence or proof that the cited research has run on MI350.
Micron (Nasdaq: MU) participates across memory and storage tiers. Its June 24, 2026 fiscal Q3 release discusses HBM4 shipments, PCIe Gen6 SSD production, and its DRAM, NAND, and storage portfolio. HBM, server DRAM, and enterprise SSDs can each have different roles in an AI system. That product exposure does not establish that an SSD replaces HBM for target decoding, that Micron supplied a research system, or that speculative decoding has a known revenue contribution.
The longer business path is worth preserving: technical feature, benchmark, customer qualification, deployment, shipment, product mix, and margin are separate events. A single algorithm article cannot collapse them into one conclusion.
Investment Watchpoints
For readers following AI infrastructure, I would put these questions next to any speculative-decoding performance claim.
- Acceptance-length distribution: What happens across real prompts and workloads, not just the headline average?
- Draft overhead and resident memory: How much capacity goes to target weights, draft weights, target KV, draft KV, and concurrent requests together?
- Interactive latency versus fleet throughput: Are p50, p95, and p99 token latency reported separately from aggregate tokens per second and cost per token?
- Achieved rather than peak hardware performance: What do bandwidth, interconnect utilization, attention time, and queueing look like in the actual system?
- Quality and re-tuning: Does the target update require the drafter to be adapted again, and does quality remain acceptable under the intended decoding rule?
- Commercial evidence: Are qualification, deployment, shipment, product-mix, or filing disclosures available beyond a technical feature page or benchmark?
The important caveat is that efficiency can pull demand in both directions. A service may need less infrastructure per request, then attract more usage or support a larger workload. The net effect on accelerators or memory cannot be inferred from a speedup claim alone.
This article is for industry context and does not recommend buying or selling any security.
Appendix: Where Speculative Decoding Fits Best
Speculation is easiest to understand in an interactive assistant where each new token matters and the target model would otherwise advance one serial step at a time. A useful drafter repeatedly proposes tokens the target would accept, allowing the target to validate more than one useful token in a pass.
It is less a magic speed button than a co-design choice. The right draft method and length depend on the target model, output style, context size, batch level, hardware, and the service’s definition of success. The real puzzle is not whether the system can speculate; it is whether its guesses are accepted often enough to justify everything required to make them.
Sources and Date Checked
- Fast Inference from Transformers via Speculative Decoding, submitted November 30, 2022; revised May 18, 2023; checked September 28, 2026.
- NVIDIA: Co-Designing AI Models Using Speculative Decoding for Faster LLM Inference, September 2, 2026; checked September 28, 2026.
- NVIDIA TensorRT-LLM, product documentation, checked September 28, 2026.
- NVIDIA FY2026 fourth-quarter results, February 25, 2026; checked September 28, 2026.
- AMD Instinct MI350 Series, product page, checked September 28, 2026.
- AMD second-quarter 2026 results, August 4, 2026; checked September 28, 2026.
- Micron fiscal Q3 2026 results, June 24, 2026; checked September 28, 2026.