Speculative Decoding: Faster AI Tokens, or a New Memory Trade-Off?

I keep seeing AI inference described as if the only job is to make the next token arrive faster.

But a large language model normally advances one token at a time, revisiting its weights and working with its growing KV state at every target-model step.

Speculative decoding asks a clever question: what if a smaller model writes several likely next tokens first, and the larger model checks them together?

If several guesses survive, one expensive target-model pass has done the work of several serial steps. If they do not, the extra draft work and memory can become part of the bill. Which side wins?

Previous Tech(EN) post: MoE Uses Fewer Experts Per Token. Why Does It Still Need So Much Memory?



Key Takeaways

  • Speculative decoding lets a smaller drafter propose several output tokens before a larger target model verifies them in parallel.
  • The potential gain comes from accepting more useful tokens per target-model iteration, not from making target weights or KV state unnecessary.
  • Acceptance length, draft overhead, context length, batch size, and resident-memory pressure determine whether the method helps.
  • Exact speculative sampling preserves the target distribution only under the required acceptance, rejection, and correction rules; it does not mean two random runs will always display the same text.
  • Interactive chat latency and fleet-level throughput are different optimization goals. A good result for one does not automatically establish a good result for the other.
Conceptual diagram contrasting one-token-at-a-time generation with a draft model proposing several tokens for a larger model to verify, accept or reject.
A draft model proposes candidate tokens, and a larger model verifies them together. This conceptual illustration does not show measured speed, memory use or cost savings.

AI-generated original conceptual illustration for this article series; no third-party images, paper figures, logos or trademarks were reproduced.



The Serial Work That Speculation Tries to Rearrange

Autoregressive decoding is naturally serial. The target model produces one token, adds that token to the context, and then runs again to produce the next one. Each step uses the model’s weights and the active KV state built from the prior context.

At a low effective batch size, this repeated target execution can be constrained by memory movement and launch overhead as much as by raw arithmetic. The model may have plenty of calculations to do, yet it still has to repeatedly bring the right weights into use and access the state needed for attention.

Speculative decoding does not remove those target-model resources. It tries to make each target pass accomplish more. A smaller draft model proposes D possible next tokens. The target then evaluates that candidate sequence together, accepting proposals in order until it finds a mismatch.

NVIDIA calls the resulting number of output tokens per target iteration the acceptance length, or AL. With D proposed tokens, AL can range from 1 to D+1. The higher the useful acceptance length, the more serial target steps the system may avoid.

The key word is “may.” The drafter must generate candidates, the target must verify them, and the system must hold the weights and state needed for both sides. Speculation changes the shape of the work; it does not make inference free.



One Simple Verification Flow

A simple example makes the mechanism easier to see. Suppose the draft model proposes four tokens after a prompt: A, B, C, and D.

  1. The draft model produces A, B, C, and D sequentially.
  2. The target model verifies that proposed sequence in a larger pass.
  3. If the target agrees with A, B, and C but disagrees at D, the system accepts the first three candidates and uses the target’s result at the mismatch.
  4. The next round begins from that accepted prefix.

In that example, one verification pass produced several usable output tokens. But if the target disagrees immediately, the draft effort may deliver little benefit. This is why acceptance length matters more than the marketing-friendly idea that the system “guesses ahead.”

NVIDIA’s September 2, 2026 technical explainer on speculative decoding describes this trade-off directly: draft length should be chosen together with acceptance length and draft overhead. Increasing the number of candidates is not automatically faster, especially after verification becomes compute-bound or when longer verification adds attention and memory pressure.



“Same Output” Needs a Decoding-Mode Footnote

One of the most common shortcuts in this topic is saying that speculative decoding gives “the same output, only faster.” That can be true in a carefully defined sense, but the definition matters.

For stochastic generation, the original speculative decoding paper by Leviathan, Kalman, and Matias presents an exact sampling method. When its proposal distribution, acceptance and rejection rules, and correction sampling are applied correctly, it preserves the target model’s output distribution.

Distribution-preserving does not mean two random generations will always print the identical sentence. Random sampling can produce different valid draws from the same target distribution, whether speculation is used or not.

Greedy decoding is a different case. A deterministic verifier can keep only the tokens confirmed by the target and fall back to the target at the first mismatch. Under the same tokenizer, logit processing, tie-breaking, and acceptance rule, that can preserve the standard greedy sequence.

Approximate variants require another level of care. If a method relaxes acceptance, changes the model, or changes the routing behavior, it should be measured for quality and output behavior rather than described as exact distribution-preserving decoding.



Why the Draft Model Is Not Free

The draft model has to come from somewhere. An external drafter brings separate weights and a full KV cache of its own. Target-attached approaches, such as draft heads, can change that footprint, but they still add weights, hidden-state handling, or a smaller cache requirement.

That makes memory capacity part of the choice. A serving stack needs room for target weights, target KV cache, draft resources, candidate verification, and concurrent requests. Long contexts can make the KV portion more prominent; a high number of simultaneous users can turn modest per-request overhead into a system-level constraint.

There is also a workload question. A short, interactive request with a small batch may benefit from fewer serial target iterations. A high-concurrency fleet may already use batching to keep the target busy. In that setting, extra verification, draft computation, communication, and attention work can move the optimal point.

In other words, a faster chat response and a higher fleet token throughput are not the same claim. Both should be measured at a stated model, precision, context length, batch size, hardware topology, and latency objective.



Memory Bandwidth Still Matters—Just Differently

Speculative decoding is sometimes described as a way around the memory wall. I think of it more as a way to get more useful output from a costly target-model pass when conditions are favorable.

Where serial decode is memory-bound, verifying multiple candidates can raise the amount of useful work done per target iteration. That may improve arithmetic intensity. But a longer draft can also add target verification work, drafter work, candidate-related state, and pressure on KV capacity.

The result depends on acceptance. If proposed tokens are often accepted, the target pass is spread across more delivered tokens. If rejections arrive early, the draft’s sequential work becomes overhead. If context is long or available memory is tight, the extra state can change the balance again.

That is why it is too simple to conclude that speculation lowers HBM demand. It may improve utilization at one serving point, while total memory need still depends on the target, the drafter, concurrency, context, and the overall amount of AI work customers choose to run.



Related Companies: System Exposure Is Not a Supply Claim

NVIDIA (Nasdaq: NVDA) has direct software and accelerator exposure to this conversation. Its TensorRT-LLM materials list speculative decoding, quantization, paged KV caching, and other inference optimizations for NVIDIA GPUs. That makes runtime efficiency and accelerator utilization relevant platform variables. It does not show that a named customer uses a particular speculative method, or that the feature determines GPU revenue in one direction.

AMD (Nasdaq: AMD) is relevant through the accelerator-memory system. Its Instinct MI350 Series page, checked September 28, 2026, lists up to 288 GB of HBM3E and up to 8 TB/s of peak theoretical memory bandwidth per GPU. Those specifications are inputs to a target-and-draft memory design; they are not application-throughput evidence or proof that the cited research has run on MI350.

Micron (Nasdaq: MU) participates across memory and storage tiers. Its June 24, 2026 fiscal Q3 release discusses HBM4 shipments, PCIe Gen6 SSD production, and its DRAM, NAND, and storage portfolio. HBM, server DRAM, and enterprise SSDs can each have different roles in an AI system. That product exposure does not establish that an SSD replaces HBM for target decoding, that Micron supplied a research system, or that speculative decoding has a known revenue contribution.

The longer business path is worth preserving: technical feature, benchmark, customer qualification, deployment, shipment, product mix, and margin are separate events. A single algorithm article cannot collapse them into one conclusion.



Investment Watchpoints

For readers following AI infrastructure, I would put these questions next to any speculative-decoding performance claim.

  • Acceptance-length distribution: What happens across real prompts and workloads, not just the headline average?
  • Draft overhead and resident memory: How much capacity goes to target weights, draft weights, target KV, draft KV, and concurrent requests together?
  • Interactive latency versus fleet throughput: Are p50, p95, and p99 token latency reported separately from aggregate tokens per second and cost per token?
  • Achieved rather than peak hardware performance: What do bandwidth, interconnect utilization, attention time, and queueing look like in the actual system?
  • Quality and re-tuning: Does the target update require the drafter to be adapted again, and does quality remain acceptable under the intended decoding rule?
  • Commercial evidence: Are qualification, deployment, shipment, product-mix, or filing disclosures available beyond a technical feature page or benchmark?

The important caveat is that efficiency can pull demand in both directions. A service may need less infrastructure per request, then attract more usage or support a larger workload. The net effect on accelerators or memory cannot be inferred from a speedup claim alone.

This article is for industry context and does not recommend buying or selling any security.



Appendix: Where Speculative Decoding Fits Best

Speculation is easiest to understand in an interactive assistant where each new token matters and the target model would otherwise advance one serial step at a time. A useful drafter repeatedly proposes tokens the target would accept, allowing the target to validate more than one useful token in a pass.

It is less a magic speed button than a co-design choice. The right draft method and length depend on the target model, output style, context size, batch level, hardware, and the service’s definition of success. The real puzzle is not whether the system can speculate; it is whether its guesses are accepted often enough to justify everything required to make them.



Sources and Date Checked

추측 디코딩, 토큰을 미리 쓰면 메모리 병목이 풀릴까

요즘 AI 추론 속도 이야기를 보면 ‘토큰을 미리 만든다’는 표현이 자주 보인다.

처음에는 작은 모델이 정답을 대신 맞히는 기술처럼 들렸는데, 조금 더 들여다보니 큰 모델이 무거운 일을 몇 번 덜 하게 만드는 쪽에 가깝다.

다만 후보를 미리 만든다고 HBM 병목이 사라지는 것은 아니다. 초안을 만드는 모델과 검증, 캐시도 결국 자리를 차지한다.

그렇다면 추측 디코딩은 AI 메모리의 가치를 낮추는 기술일까, 아니면 같은 메모리를 더 바쁘게 쓰게 만드는 기술일까?

이전 Tech(KR) 글: MoE가 덜 계산해도, 전문가 가중치는 어디에 둘까



핵심 요약

  • 추측 디코딩은 작은 draft 모델이 여러 후보 토큰을 먼저 제안하고, 큰 target 모델이 이를 한 번에 검증하는 추론 방식이다.
  • 실제로 받아들여진 후보가 많을 때에만 target 모델의 한 번의 실행을 여러 유효 토큰에 나눠 쓸 수 있다.
  • draft의 가중치·KV 캐시, 검증 연산, 후보 길이, batch와 요청 유형이 추가 비용을 만들기 때문에 항상 빨라지는 것은 아니다.
  • 정확한 확률 샘플링 방식은 target의 출력 분포를 보존할 수 있지만, 매번 눈으로 같은 문장이 나온다는 뜻은 아니다.
초안 모델이 여러 토큰 후보를 만들고, 큰 모델이 이를 검증한 뒤 일부를 수락하거나 다시 계산하는 speculative decoding 흐름 개념도.
Speculative decoding은 작은 초안 모델의 후보를 큰 모델이 한 번에 검증해, 수락된 구간만 다음 생성으로 이어 가는 방식이다.

이 글을 위해 AI로 생성한 자체 제작 오리지널 개념도이며, 제3자 이미지·논문 도식·로고·상표를 사용하지 않았습니다.



큰 모델이 여는 문을 한 번에 여러 후보 앞에서 열어 본다

일반적인 LLM 출력 과정은 토큰 하나를 만든 뒤, 그 결과를 다시 입력에 붙여 다음 토큰을 만드는 순서로 이어진다. 이때 큰 target 모델은 토큰마다 다시 실행되고, 가중치와 KV 캐시에 반복적으로 접근한다. 특히 작은 batch의 대화형 추론에서는 이 반복이 지연 시간과 메모리 대역폭의 영향을 크게 받을 수 있다.

추측 디코딩에서는 더 가벼운 draft 모델이 다음 후보를 여러 개 먼저 제안한다. 이후 target 모델이 후보들을 한 번의 검증 과정에서 순서대로 확인한다. 후보 가운데 실제로 받아들여진 토큰이 여러 개라면, 원래 여러 차례였을 target 실행을 더 많은 유효 출력 토큰에 나눠 쓴 셈이 된다.

여기서 중요한 숫자는 후보를 몇 개 만들었는지가 아니라 acceptance length, 즉 target 검증 뒤 실제로 살아남은 토큰 수다. NVIDIA의 2026년 9월 기술 설명도 전체 시간에는 target 검증뿐 아니라 draft 생성 시간이 함께 들어가며, draft 길이·acceptance·메모리 부담을 같이 봐야 한다고 설명한다.



‘출력이 같다’는 말도 두 가지로 나눠 봐야 한다

추측 디코딩을 설명할 때 결과가 같다고 말하는 경우가 있다. 확률 샘플링에서는 정확한 acceptance·rejection과 보정 샘플링을 올바르게 적용할 때 target 모델의 샘플링 분포를 보존한다는 뜻이다. 무작위성이 있는 생성에서 매번 똑같은 문장이 나온다는 뜻은 아니다.

greedy decoding처럼 결정적으로 다음 토큰을 고르는 경우에는 조건이 또 다르다. 같은 tokenizer, logit 처리, tie-breaking, acceptance 규칙 아래에서 target의 토큰과 일치하는 제안만 유지하고 처음 어긋난 지점부터 target 결과로 돌아가면, 표준 greedy decode와 같은 결정 경로를 만들 수 있다.

반대로 후보를 더 쉽게 받아들이거나 draft·router·target 자체를 바꾸는 방법은 별도의 품질 검증이 필요하다. 더 빠르다는 이유만으로 ‘동일 출력’이나 ‘정확한 분포 보존’까지 따라오는 것은 아니다.



HBM 병목이 사라지는 것이 아니라, 유효 토큰당 부담이 달라진다

target 모델이 한 번 실행될 때 받아들여지는 토큰이 늘어나면, 반복적인 target 가중치 접근을 유효 출력 토큰에 나눠 볼 여지가 생긴다. NVIDIA는 이 과정이 일부 조건에서 target의 arithmetic intensity를 높여 decode의 성격을 memory-bound에서 compute-bound 쪽으로 옮길 수 있다고 설명한다. 그렇다고 HBM 대역폭의 가치가 없어지는 것은 아니다.

draft도 공짜가 아니다. 외부 draft 모델은 자체 가중치와 KV 캐시가 필요하고, target에 보조 구조를 붙이는 방식도 추가 레이어·상태 또는 hidden state를 사용한다. 후보 길이를 길게 잡으면 target 검증과 attention, 통신, KV 관련 부담도 함께 커질 수 있다.

결국 이득은 조건부다. acceptance length가 낮으면 draft가 만든 후보와 검증 비용을 회수하지 못할 수 있고, 이미 compute-bound인 환경에서는 후보를 더 늘리는 것이 오히려 부담이 될 수 있다. 긴 문맥, 높은 동시성, 모델 업데이트 뒤 달라진 요청 분포도 모두 다시 측정해야 하는 변수다.



관련 기업

NVIDIA(NASDAQ: NVDA)는 TensorRT-LLM에서 speculative decoding, quantization, KV cache 등을 포함한 GPU용 LLM 추론 최적화 기능을 제공한다. 이 글의 target·draft runtime과 GPU 메모리 활용을 읽는 데는 관련이 있지만, 특정 최적화가 특정 고객의 배포나 GPU 추가 구매로 이어졌다고 볼 근거는 별도로 필요하다.

AMD(NASDAQ: AMD)의 Instinct MI350 Series 공식 제품 페이지는 GPU당 최대 288GB HBM3E와 최대 8TB/s의 peak theoretical memory bandwidth를 제시한다. 이런 용량과 대역폭은 target·draft 구조가 고려하는 하드웨어 변수지만, 제품 사양만으로 실제 애플리케이션의 지연 시간이나 특정 추측 디코딩 구현의 성능을 추론할 수는 없다.

국내에서는 SK hynix(KRX 000660)의 2026년 2분기 자료에서 HBM, AI 서버 DRAM, eSSD 제품군을, Samsung Electronics(KRX 005930)의 2026년 2분기 자료와 PM1763 발표에서 HBM·서버 DRAM·enterprise SSD 사업 노출을 확인할 수 있다. HBM은 target 가중치와 연산 가까이의 고대역폭 메모리, 서버 DRAM과 eSSD는 다른 계층의 용량·저장 역할로 나눠 봐야 한다. 이 제품군이 특정 draft 모델이나 논문의 장비에 채택됐다는 뜻은 아니며, 효율 개선이 곧바로 매출이나 주가로 이어진다는 주장도 아니다.



투자 체크포인트

  • acceptance length의 평균뿐 아니라 요청 유형별 분포가 실제로 유지되는지
  • draft 생성 시간과 target 검증 시간이 합쳐진 뒤에도 사용자당 토큰 지연이 개선되는지
  • p50 평균값과 함께 p95·p99 지연, GPU당 동시성, fleet 전체 처리량을 따로 보는지
  • target·draft의 상주 가중치와 KV 캐시가 GPU 메모리를 얼마나 차지하는지
  • 기업의 경우 고객 인증, 실제 출하, HBM·DRAM·NAND 제품 mix와 재고·설비투자가 공식 자료로 확인되는지

추론 최적화 기사 하나로 메모리 반도체 수요나 개별 종목의 방향을 확정하기는 어렵다. 이 글은 매수·매도 추천이 아니라, AI 추론 구조와 메모리 공급망을 읽기 위한 정보다.



빠른 추론이 메모리 수요를 한 방향으로 정하지 않는 이유

추측 디코딩이 잘 작동하면 같은 장비에서 더 많은 요청을 처리하거나, 사용자가 느끼는 토큰 간 지연을 줄일 가능성이 생긴다. 하지만 서비스 사업자는 확보한 HBM을 절약하는 대신 더 높은 동시성에 쓸 수도 있고, draft와 KV 캐시를 위해 다른 메모리 자원을 더 쓸 수도 있다. 알고리즘 효율이 메모리 수요를 늘리는지 줄이는지는 서비스 규모와 서버 설계까지 봐야 한다.

그래서 HBM, 서버 DRAM, eSSD를 서로 대체재처럼 놓기보다는 각 계층이 어떤 상태와 가중치를 맡는지 보는 편이 낫다. 추측 디코딩은 HBM을 덜 쓰는 마법이라기보다, 큰 모델의 비싼 한 회전을 더 많은 유효 토큰에 배분할 수 있는지 시험하는 방식에 가깝다.



Appendix. draft 모델과 target 모델은 무엇이 다를까

target 모델은 최종 검증을 맡는 큰 모델이고, draft 모델은 다음 토큰 후보를 빠르게 제안하는 가벼운 모델 또는 보조 구조다. target이 후보를 모두 받아들이는 경우도 있고, 중간에서 거절한 뒤 target의 결과로 이어지는 경우도 있다.

따라서 추측 디코딩의 성패는 draft가 얼마나 빨리 후보를 만들었는지와, target이 그 후보를 얼마나 자주 받아들였는지의 균형에 달려 있다. 후보가 자주 맞을수록 이득을 기대할 수 있지만, 후보를 만드는 비용과 메모리 상태까지 함께 계산해야 한다.



출처 및 확인일

MoE Uses Fewer Experts Per Token. Why Does It Still Need So Much Memory?

Mixture-of-experts models promise a neat idea: use only the specialists needed for each token, rather than running the entire model every time.

That can lower active compute. It does not make the model’s full collection of expert weights disappear.

Those weights still need to wait somewhere, and the moment they leave fast accelerator memory, the problem becomes less about storage capacity and more about timing.

Can the right expert arrive before the GPU needs it? That small question may decide whether offloading is a cost-saving design or a new source of delay.

Previous Tech(EN) post: Prefix Caching Saves AI Compute—But It Does Not Make the Memory Problem Disappear



Key Takeaways

  • MoE activates only some experts for a token, but the full expert-weight pool must still be stored somewhere in the system.
  • Keeping more weights resident in fast memory can make latency more predictable; offloading can increase usable capacity but adds a transfer deadline.
  • A prefetch is useful only if the needed weights arrive before the relevant layer runs. A nominal cache hit that arrives late can still stall inference.
  • Research on expert offloading is promising but conditional. Benchmark results, model changes, hardware configuration, and quality trade-offs need to be read together.
  • NVIDIA, AMD, and Micron have different infrastructure exposure to the issue. That does not establish that they supply the cited research systems or that any one technique determines revenue or share performance.
Abstract diagram of a large MoE expert-weight pool, a small set of experts resident in GPU HBM, and backing storage tiers, with timely prefetch and delayed-transfer paths.
Keeping only some MoE experts in GPU memory makes the timing of a weight transfer as important as where the weights are stored.

Original illustration generated for this publication team; no third-party artwork, logos, or copied diagrams used.



A Sparse Model Still Has a Full Closet of Weights

An MoE model has multiple expert networks and a router that chooses a small subset for each token. The appeal is obvious: the model does not need to run every expert on every step.

But sparse activation is not the same thing as a small total model. A hotel may have hundreds of rooms even when only a few guests check in at once. Likewise, an MoE system may activate only a few experts while still needing access to a large pool of weights.

The simplest way to serve that model is to keep all the relevant weights resident close to compute. That favors predictable response time, but it consumes accelerator memory. At a larger model size or higher concurrency, holding everything nearby can become expensive or impossible within a chosen hardware configuration.

Offloading changes the arrangement. A serving system can keep a working set of likely experts in fast memory and store less active experts in a lower tier. The full model remains available, but not every part is immediately available at GPU-memory speed.

That is why the active-parameter headline needs a second question beside it: how large is the full weight pool, and how reliably can the next required expert be made ready?



Offloading Is Really a Deadline Problem

It is tempting to describe offloading as a simple capacity fix: put the weights that do not fit in HBM somewhere else. The harder part is bringing them back before the model reaches the layer that needs them.

Imagine a production line that needs a particular tool at the next station. The tool can be kept in a cheaper warehouse, but that helps only if it reaches the station before the work arrives. If it shows up late, the line waits.

MoE inference has a comparable sequence. The router selects experts as the model processes tokens and layers. A runtime may try to predict which expert will be needed next, start the transfer early, and evict a lower-value resident expert to make room. The system succeeds when the selected weights arrive before their compute deadline.

The September 11, 2026 SeqMoE preprint proposes this kind of predictive memory management, combining sequence prediction, deadline-aware prefetching, cache eviction, and a graph-compatible runtime. In its evaluation, the paper reports 45% expert residency, a 96.97% average hit rate, and 80.22% of full-load performance. Those are paper-specific results under the authors’ test conditions, not a general service-level guarantee or a hardware benchmark for every MoE workload.

This is the useful distinction: a high cache hit rate tells us that a requested expert was found in a planned location. An on-time hit tells us that it was ready early enough to avoid a stall. For an interactive AI service, the second measure is usually the one a user feels.



Why an SSD Is Not a Drop-In Substitute for HBM

HBM, server DRAM, and enterprise SSDs all store data, but they are not interchangeable places to run an MoE model. They differ in capacity, bandwidth, latency, connection path, and the software work required to use them well.

HBM is close to the accelerator and designed for high-bandwidth compute workloads. Server DRAM can offer a larger memory tier with different access characteristics. SSDs can provide much more economical capacity for colder data, but an expert weight stored there has to travel through a longer path before it can be used by the GPU.

The September 16, 2026 Edge0 preprint makes the timing challenge very clear. Its authors argue that merely placing experts on an SSD is insufficient because the next-layer routing decision normally arrives too late to hide the read. Their proposed answer is a prerouter that makes a one-token-ahead routing choice, plus a recovery LoRA trained to address quality loss associated with routing replacement and INT4 quantization.

That detail is important. Edge0 is not simply a lossless cache-management recipe. It changes part of the routing behavior and adds a quality-recovery component. The paper reports a 35B-class model serving at 20 tokens per second with 3 GiB of peak active memory on a single 24 GB system. That reported outcome belongs to its particular model, hardware, precision, and evaluation setup; it should not be read as proof that SSD-backed MoE serving is already a universal production solution.

Transfer speed alone also does not settle the question. Queueing, interconnect topology, host overhead, prefetch accuracy, and concurrency can change whether loading a weight beats leaving it resident or recomputing a different design choice. Peak specification numbers do not tell the whole serving story.



Some of the Memory Problem Can Move Into Model Design

There is more than one response to a growing expert pool. A runtime can move weights around more cleverly, but model designers can also reduce how much distinct expert material must be available at once.

The MoRE preprint, revised September 17, 2026, explores one such direction: sharing an expert pool across adjacent layers while keeping layer-specific routers. This is an architectural approach to parameter growth, not an expert-offloading runtime.

That separation matters for readers following AI infrastructure. A software cache, a storage tier, an interconnect strategy, and a redesigned model can all improve a capacity problem in different ways. Treating them as the same solution would hide the trade-offs.

MoRE is marked as accepted to COLM 2026, but its reported experiments cover models from 114 million to 1.15 billion parameters. That is meaningful research evidence, not enough to generalize its results to every frontier-scale or commercial serving environment.



Expert Parallelism and Expert Offloading Are Different Moves

There is another way to avoid putting every expert on one accelerator: distribute experts across more GPUs. That is expert parallelism, not offloading to a lower memory tier.

NVIDIA documents wide expert parallelism in TensorRT-LLM and describes a multi-GPU implementation in its Wide Expert Parallelism technical post. The idea is to reduce the number of expert weights each GPU must hold by spreading the model across a larger system.

That may relieve per-GPU memory pressure, but it replaces one constraint with others: inter-GPU communication, topology, system scale, and scheduling. It should not be confused with the SSD or host-memory offloading explored in the research papers above.

The same distinction applies to NVIDIA Dynamo. NVIDIA’s March 16, 2026 announcement describes Dynamo as production-grade inference software, and its public materials discuss KV and state-data management. That is relevant context for inference systems, but it is not evidence that Dynamo offloads MoE expert weights or that it implements SeqMoE or Edge0.



Related Companies: Different Exposure, Different Evidence

NVIDIA (Nasdaq: NVDA) is exposed through accelerators, networking, and inference software. Its public TensorRT-LLM and Wide Expert Parallelism materials show that large-scale MoE serving is a platform consideration. A broader MoE deployment cycle could make memory capacity, interconnect design, runtime support, and system utilization more important. It does not follow that NVIDIA supplies the cited papers, that a named customer uses their methods, or that efficiency changes GPU demand in one automatic direction.

AMD (Nasdaq: AMD) has relevant accelerator-memory attributes through its Instinct portfolio. AMD’s MI350 Series page, checked September 28, 2026, lists up to 288 GB of HBM3E and up to 8 TB/s of peak theoretical memory bandwidth per GPU. Those are hardware specifications, not an MoE-offload benchmark or proof that the system runs the research methods discussed here.

Micron (Nasdaq: MU) participates across memory and storage categories. Its June 24, 2026 fiscal Q3 release discusses HBM4 shipments, PCIe Gen6 SSD production, and a broader DRAM, NAND, and storage portfolio. Those products can occupy different infrastructure tiers, but the release does not show that Micron supplies the research systems or quantify a revenue contribution from MoE offloading.

The business path remains longer than a technical diagram: workload need, system design, qualification, customer deployment, shipment, product mix, and margin. Each step needs its own evidence.



Investment Watchpoints

For readers trying to follow the infrastructure implications, the more useful questions are often operational rather than promotional.

  • Resident working set: How much expert weight must remain close to compute at the target model size, precision, and concurrency?
  • On-time prefetch rate: What share of required experts reaches the accelerator before the relevant layer needs it, rather than merely appearing somewhere in the cache hierarchy?
  • Measured transfer path: What do real bandwidth, latency, queueing, and topology measurements show—not just peak component specifications?
  • Quality trade-off: Does the approach alter routing, quantization, or training objectives, and what happens to task quality as well as tokens per second?
  • Commercial evidence: Are there named qualifications, deployments, shipments, product-mix disclosures, or filing commentary beyond a paper or technical demonstration?

There is a genuine two-sided demand question here. Better serving efficiency can reduce infrastructure required for one request, yet it can also enable lower-cost services, higher concurrency, larger models, or more total requests. The net effect on accelerator, HBM, DRAM, or SSD demand cannot be inferred from an efficiency claim alone.

This article is for industry context, not a recommendation to buy or sell any security.



Appendix: Where This Matters in Practice

MoE memory design becomes easier to see in an AI service that has a large model, many simultaneous users, and uneven expert reuse. One user’s request may repeatedly favor a small set of experts while another triggers a different group. The serving system has to decide which weights are worth keeping nearby and which can wait in another tier.

That is why “sparse” does not mean “memory no longer matters.” It means the system has a new choice: reduce active compute by selecting fewer experts, then manage the full set of expert weights so the selected few arrive on time.



Sources and Date Checked

MoE가 덜 계산해도, 전문가 가중치는 어디에 둘까

요즘 MoE 연구를 따라가다 보면 ‘활성 파라미터가 작다’는 말이 먼저 눈에 들어온다.

MoE는 토큰마다 일부 전문가만 골라 계산하니, 처음엔 메모리 부담도 함께 줄어드는 구조처럼 보인다.

그런데 선택받지 않은 전문가의 가중치까지 사라지는 것은 아니다. 필요해지는 순간을 맞추려면, 결국 어딘가에 보관해 두고 제시간에 가져와야 한다.

AI 추론에서 다음 병목은 연산량보다 ‘전문가 가중치를 언제 옮길 수 있나’가 될 수도 있겠다는 생각이 든다.

이전 Tech(KR) 글: KV 캐시 재사용, HBM은 덜 필요할까



핵심 요약

  • MoE는 한 토큰에서 일부 전문가만 활성화해 계산량을 낮출 수 있지만, 모델 전체의 전문가 가중치를 보관해야 하는 문제까지 자동으로 없애지는 않는다.
  • HBM에 자주 쓰는 전문가만 남기고 나머지를 DRAM이나 SSD 같은 다른 계층에 둘 수는 있지만, 그때 병목은 용량에서 전송 지연과 대역폭으로 옮겨간다.
  • 중요한 것은 단순 캐시 적중률보다 필요한 전문가 가중치가 다음 연산 시점 전에 도착했는지다. 이를 위해 예측·프리페치·퇴출 정책이 함께 필요해진다.
  • 최근 SeqMoE, Edge0, MoRE 같은 연구는 서로 다른 답을 제시하지만, 각 실험 결과를 모든 모델·모든 데이터센터에 그대로 적용할 수는 없다.


전체 전문가 가중치 풀과 GPU HBM에 상주하는 일부 전문가, 서버 저장 계층에서 제시간에 프리페치되는 경로와 늦게 도착해 GPU가 기다릴 수 있는 경로를 표현한 개념도.
전문가 가중치를 모두 GPU에 올리지 않을 때는 필요한 데이터를 연산 전에 가져오는 시간이 중요해진다.

이 글을 위해 자체 제작한 개념도이며, 제3자 이미지나 로고를 사용하지 않았습니다.



적게 계산하는 모델과 크게 보관해야 하는 모델은 다르다

MoE(Mixture of Experts)는 모든 파라미터를 매번 동시에 계산하는 대신, 라우터가 토큰별로 일부 전문가를 선택해 처리하는 구조다. 그래서 활성 파라미터만 보면 계산 효율이 좋아 보인다. 하지만 서비스가 다음 토큰에서 어떤 전문가를 부를지 미리 확신할 수 없다면, 선택되지 않은 전문가 가중치도 완전히 잊어버릴 수는 없다.

여기서 HBM은 GPU 가까이에 두고 바로 꺼내 쓸 수 있는 작업 공간에 가깝다. 모든 전문가를 HBM에 상주시킬 수 있다면 단순하겠지만, 모델의 전체 용량과 여러 요청의 동시 처리까지 생각하면 공간은 제한된다. 반대로 서버 DRAM이나 SSD로 일부를 내리면 더 많은 가중치를 보관할 여지는 생기지만, 다음 연산에 맞춰 GPU로 되돌리는 시간이 새 변수로 들어온다.

그래서 MoE의 메모리 문제는 “가중치를 어디에 넣을까”에서 끝나지 않는다. 어떤 전문가는 계속 쓰고, 어떤 전문가는 오랫동안 안 쓰다가 갑자기 필요해진다. 이 불규칙한 재사용 패턴 때문에 용량, 링크 대역폭, 전송 지연, GPU의 대기 시간이 한 덩어리로 묶인다.



프리페치는 적중보다 도착 시간이 중요하다

전문가 오프로딩은 HBM에서 비활성 가중치를 비워 공간을 만드는 방법이다. 대신 GPU가 다음 전문가를 요청한 뒤에야 가져오기 시작하면 늦을 수 있다. 이 경우 GPU는 연산할 데이터를 기다리고, 사용자 입장에서는 토큰이 이어지는 속도가 느려질 수 있다.

그래서 프리페치는 ‘다음에 필요할 전문가’를 조금 먼저 예측해 옮겨두는 방식이다. 잘 맞히면 HBM을 모두 채우지 않고도 필요한 가중치를 제시간에 준비할 수 있다. 반대로 예측이 빗나가면 링크를 쓴 만큼의 전송이 낭비되고, 정작 필요한 전문가는 늦게 도착할 수 있다.

이 차이는 생각보다 크다. 캐시가 맞았다는 기록만으로는 충분하지 않고, 추론의 마감 시점 안에 도착했는지를 봐야 한다. 서버 DRAM이나 SSD의 표기 대역폭만 보고 실제 토큰 지연을 계산하기 어려운 이유도 큐 대기, 경로 구성, 다른 요청과의 경합, 실행 환경이 모두 끼어들기 때문이다.



SeqMoE·Edge0·MoRE는 같은 해법이 아니다

9월 11일 arXiv에 등록된 SeqMoE 사전공개 논문은 전문가 선택의 순서를 예측하고, 필요한 가중치를 마감 시간에 맞춰 미리 가져오는 관리 방식을 다룬다. 핵심은 전체 전문가를 항상 HBM에 올려두기보다, 다음 연산 그래프에 맞춰 이동을 겹치게 하는 데 있다. 저자들은 45% 상주 조건에서 평균 96.97% 적중률과 전체 전문가를 GPU 메모리에 상주시킨 경우 대비 80.22% 성능을 보고했지만, 이는 논문의 특정 설정에서 나온 값이지 특정 서비스의 성능 보장은 아니다.

9월 17일 수정본이 올라온 Edge0 사전공개 논문은 조금 다른 방향이다. SSD 계층을 쓰는 전문가 서빙을 위해 다음 레이어의 라우팅을 한 토큰 앞서 prerouter로 예측하고, 이를 실제 route로 쓰며 회복용 LoRA까지 학습한다. 저자들이 제시한 것은 35B 4-bit 모델을 단일 24GB 소비자용 장비에서 다룬 실험으로, 초당 20 token·GPU 메모리 3GiB라는 결과도 그 조건에 묶여 있다. 이것은 단순히 캐시를 잘 비우고 채우는 접근보다 모델의 라우팅과 품질 조건을 함께 바꾸는 설계다. 따라서 일반적인 무손실 오프로딩과 같은 선에서 비교하기보다, 품질 회귀와 운영 복잡도까지 함께 봐야 한다.

MoRE 사전공개 논문은 레이어마다 분리해 두던 전문가 풀을 공유하는 방식으로 접근한다. 다만 이 연구의 실험 범위는 1억1400만~11억5000만 파라미터 규모로 제시돼 있다. 훨씬 큰 최전선 모델에서도 같은 효율이 난다고 단정하기보다, 메모리 제약이 모델 구조 자체를 바꾸게 할 수 있다는 사례로 보는 편이 맞다.



관련 기업

국내 투자자 관점에서는 SK hynix(KRX 000660)의 HBM·AI 서버용 DRAM·eSSD 제품군을 먼저 따로 볼 수 있다. 회사의 2026년 2분기 실적 자료는 이 제품군이 회사 사업에서 다뤄진다는 근거다. 다만 이것이 SeqMoE나 Edge0 같은 특정 연구를 도입했다는 뜻도, 그 연구가 곧바로 공급 계약이나 매출 증가로 이어진다는 뜻도 아니다.

Samsung Electronics(KRX 005930)도 공식 2026년 2분기 실적 발표에서 HBM·서버 DRAM·eSSD를 다루고 있으며, 차세대 AI 인프라용 PM1763 SSD 양산 발표도 확인할 수 있다. MoE에서 가중치를 다른 계층에 두는 방식이 실제로 확산된다면 서버 메모리와 eSSD를 함께 읽어야 하는 이유는 생긴다. 그래도 제품군의 존재와 특정 MoE 서빙 구조의 고객 채택, 그리고 실적 기여는 서로 다른 확인 단계다.

국내 메모리 공급망 전체를 볼 때도 ‘MoE 오프로딩이 늘면 HBM이 줄어든다’ 또는 ‘SSD가 곧바로 늘어난다’고 단순하게 연결하기 어렵다. 오프로딩은 HBM의 한정된 공간을 더 효율적으로 쓰게 할 수 있지만, 서비스 사업자가 확보한 HBM을 더 많은 동시 요청에 배분하는 방향으로도 활용할 수 있다. 멀리 있는 계층에 가중치를 보관하는 구조가 실제로 채택되더라도, 필요한 것은 단순 용량만이 아니라 전송 성능, 내구성, 소프트웨어 호환성, 고객사의 운영 방식이다.

가속기 쪽에서는 NVIDIA TensorRT MoE 추론 문서와 GPU 플랫폼, AMD의 Instinct·ROCm 생태계를 함께 확인할 수 있다. 다만 이들 플랫폼 문서나 제품 사양이 SeqMoE·Edge0·MoRE의 상용 도입 또는 공동 공급을 뜻하는 것은 아니다. AMD의 MI350 제품 페이지는 288GB HBM3E와 최대 8TB/s라는 메모리 사양을 제시하지만, 이 사양은 특정 MoE 오프로딩 논문의 성능을 증명하는 수치가 아니다. 메모리 업체 역시 DRAM·NAND·HBM의 제품 구성이 실제 고객 채택과 어떤 방식으로 연결되는지를 실적 자료에서 따로 확인해야 한다.



투자 체크포인트

  • 실제 서비스에서 HBM에 상주하는 expert working set이 어느 정도인지
  • 예측한 전문가를 제시간에 준비한 비율과, 불필요하게 옮긴 전송의 비율이 어떻게 달라지는지
  • 링크 사용률, 토큰 지연의 p95·p99, GPU당 동시 처리 요청이 함께 개선되는지
  • 라우팅이나 공유 expert pool 도입이 정확도·안전성·운영 복잡도에 어떤 비용을 남기는지
  • 가속기·메모리·스토리지 업체가 실제 고객 도입, 제품 인증, 출하와 관련한 근거를 공식 자료로 제시하는지

특히 이 분야에서는 ‘적중률이 높다’는 한 숫자보다 실패했을 때의 지연이 더 중요할 수 있다. 드문 miss가 길게 이어지면 평균 처리량이 좋아도 체감 품질은 나빠질 수 있기 때문이다. 이 글은 매수·매도 추천이 아니라, AI 추론과 메모리 공급망을 읽기 위한 정보다.



효율 개선이 메모리 수요를 한 방향으로 정하지 않는 이유

MoE 오프로딩과 프리페치가 좋아질수록 한 GPU에서 더 큰 모델을 다룰 수 있거나, 같은 장비에서 더 많은 요청을 처리할 가능성은 생긴다. 하지만 그 결과가 HBM, DRAM, SSD 중 어느 한쪽의 수요를 일괄적으로 키우거나 줄인다고 말할 수는 없다. 모델 구조, 요청 길이, 동시성, 서버 구성, 서비스 사업자의 투자 결정이 모두 다르기 때문이다.

오히려 지금 단계에서는 메모리 한계가 소프트웨어를 바꾸고, 바뀐 소프트웨어가 다시 어떤 메모리 계층을 얼마나 쓰게 만드는지의 순환을 보는 편이 흥미롭다. MoE가 덜 계산하는 모델이라는 설명 뒤에는, ‘모든 전문가를 어느 순간까지 어디에 둘 것인가’라는 더 현실적인 서버 문제가 남아 있다.



Appendix. MoE expert offloading은 무엇을 옮기는 기술일까

MoE expert offloading은 현재 GPU HBM에 꼭 필요하지 않은 전문가 가중치를 서버 DRAM이나 SSD 등 다른 저장 계층에 두었다가, 필요할 때 GPU 쪽으로 가져오는 운영 방식이다. 입력 문맥을 저장하는 KV 캐시와는 대상이 다르다. 이 글의 초점은 대화 상태가 아니라, 모델 안의 전문가 가중치와 그 이동 시간에 있다.

좋은 오프로딩은 단지 많은 가중치를 밖으로 내보내는 것이 아니다. 다음에 쓰일 가능성이 높은 전문가를 예측하고, 필요한 시점보다 앞서 옮기며, 잘못 옮긴 비용과 늦게 도착한 비용을 함께 줄이는 일에 가깝다.



출처 및 확인일

Prefix Caching Saves AI Compute—But It Does Not Make the Memory Problem Disappear

AI inference keeps getting framed as a race to do less work. That is true, but it leaves out an awkward detail: the work we skip often has to be remembered somewhere.

That is why prefix caching matters. It can make repeated AI requests feel much faster, especially when many users begin with the same long instructions, policy document, or tool description.

But the saved state does not vanish. It becomes a cache, and that cache competes for valuable memory capacity.

The next question is more interesting than “does caching work?” Where should that state live, when should it be moved, and when is it cheaper to calculate it again?

Previous Tech(EN) post: Beyond HBM: Where AI Data Center Spending Goes Next



Key Takeaways

  • Prefix caching reuses work from the input-processing stage, known as prefill. It does not eliminate the step-by-step decode work required to generate a new answer.
  • A retained KV cache uses capacity even when no one is asking a question. Keeping more state can save future compute, but it can also reduce room for active requests.
  • Moving KV state out of HBM can free fast memory, yet it introduces a restore path whose transfer time can hurt first-token latency.
  • A cache-hit percentage is not enough on its own. Reusing a short prefix and reusing a long, expensive prefix can produce the same hit count but very different compute savings.
  • NVIDIA, SK hynix, and Samsung have different points of exposure to this system design. Product announcements and technical demonstrations are not proof of shared contracts, revenue, or a directional stock outcome.


Abstract diagram showing repeated prompt prefixes flowing into a KV cache, with paths for local reuse, storage restore, and recomputation.
An original illustration of three ways a repeated prompt prefix can be served: reuse, restore, or recompute.

Original illustration generated for this post; no third-party artwork or logos used.



What Prefix Caching Actually Saves

Think of an AI assistant that reads a 100-page employee handbook before answering each question. If every request begins with the exact same token sequence—the same handbook text in the same order—making the model process that shared prefix from scratch every time is wasteful. Prefix caching is about exact reusable state, not merely prompts that mean something similar.

In a transformer model, processing that input produces key-value, or KV, states. Prefix caching keeps the KV states for the shared beginning of the request. The next user with the same beginning can reuse those states instead of repeating that part of the input calculation.

That is the practical win: less repeated prefill work and often a faster time to first token. The vLLM Automatic Prefix Caching documentation, checked September 28, 2026, makes an important limitation explicit: the feature does not reduce decode time itself.

Decode is the part where the model produces a new answer token by token. Even with a perfectly reusable handbook, the model still has to generate a fresh response to a fresh question. A long answer can therefore remain slow even when the first token arrives sooner.

That distinction matters because “AI got more efficient” can mean several different things. It may mean less input computation, lower first-token delay, higher concurrency, lower cost per successful request, or some mix of them. Those are related outcomes, not interchangeable ones.



The Compute-Memory Trade-Off Hiding Behind a Cache Hit

A cached prefix is like a prepared ingredient in a restaurant kitchen. Preparing it ahead of time saves work when the next order arrives. But the ingredient still occupies shelf space, and the kitchen cannot reserve every shelf forever for dishes that might be ordered again.

KV state creates the same tension in an inference server. Active conversations need memory now. Recently used prefixes may be worth retaining for later. Long-lived agent sessions, shared tool instructions, and recurring enterprise documents can all compete for the same pool of fast memory.

This is why hit rate can be a misleading headline number. Suppose one cache hit avoids processing ten tokens, while another avoids processing ten thousand. Counting both as one hit hides the difference in avoided compute, bytes retained, and user-facing latency.

A September 24, 2026 preprint studying prefix-cache eviction makes a related point: cache policy should consider more than simple reuse. The paper analyzes traces from two companies and 14 eviction policies, so it is useful evidence rather than a universal rule. Its scope is not a reason to assume every workload will behave the same way.

The useful operating question is not merely, “Was this prefix reused?” It is, “How much expensive prefill did it avoid, how much capacity did we reserve for it, and what happened to the next user while we kept it?”



Why Moving KV State Off HBM Is Not a Free Upgrade

HBM is close to the GPU and designed for very high-bandwidth work, but its capacity is finite and costly. That makes it tempting to keep only the hottest KV state nearby and move colder state to host memory, local storage, or a remote tier.

The trade-off is straightforward in principle. A system can free HBM capacity by offloading KV state, then restore it when a matching request returns. NVIDIA’s Dynamo KV-cache offloading documentation, checked September 28, 2026, describes host-memory and disk-based offload paths.

In practice, the restore is the part that decides whether the idea helps. If fetching the state takes longer than recomputing the prefix, the cache may save storage but lose time. Transfer bandwidth matters, but so do topology, queueing, software overhead, and whether the system can overlap movement with other work.

This is why memory tiering should not be described as HBM replacement. It is a placement decision. Fast memory can be reserved for the working set that needs it immediately, while slower tiers may hold state that is useful but not urgent.

The result can change how an AI service is designed. Prompt templates may put stable shared instructions first. Routers may favor workers that already hold the right cache. Eviction policies may prioritize prefixes that are both likely to return and expensive to rebuild.



Cache-Aware Routing Turns Memory Location Into a Scheduling Decision

There is another wrinkle. Even a valuable cache is less valuable if it sits on a busy worker.

A request router can send a new prompt to the GPU that already holds the matching KV state, avoiding a rebuild. But that same GPU might have a long decode queue. Sending the request elsewhere may create more compute work but deliver the answer sooner.

That is the balancing act behind cache-aware routing: reuse location, queue pressure, and latency targets all matter at once. NVIDIA’s April 17, 2026 Dynamo technical overview discusses routing based on KV-cache overlap alongside decode load and tiered cache management.

For global cloud and enterprise AI operators, this is a serving-economics issue. A system that delivers a better first-token experience with the same installed hardware may improve utilization. A system that protects cache hits but worsens tail latency may not. The winner is not the one with the prettiest cache statistic; it is the one that meets a service target at an acceptable cost.



Related Companies: Three Different Paths Into the Same Problem

NVIDIA (Nasdaq: NVDA) sits at the accelerator and inference-software layer. Its Dynamo materials show that cache-aware routing and memory tiering are active software design areas. That establishes technology exposure, not a simple revenue equation. Better software efficiency could support more workload on an installed GPU fleet, while it could also reduce hardware needed for a given workload. Total demand, pricing, and customer deployment choices decide the business result.

SK hynix (KRX: 000660) has exposure across HBM, DRAM, and NAND. At its September 17, 2026 AI Infra Summit, the company demonstrated SALT-KV, a concept that places KV state across HBM, DRAM, and SSD according to reuse value and storage cost. It is a useful example of the tiering problem becoming a memory-system conversation. It is not evidence that a particular customer has adopted the design or that a specific product shipment will follow.

Samsung Electronics (KRX: 005930) offers CXL-based memory expansion products, including the CMM-D MD220. CXL expansion addresses server-memory capacity and pooling, a different role from GPU-attached HBM. A larger memory tier could be relevant to some inference architectures, but this product page does not establish deployment in a particular prefix-cache system, nor does it make CXL DRAM a direct substitute for HBM.

The lesson is that one AI workload can touch several memory categories without making them interchangeable. HBM serves a fast working set. DRAM can expand capacity closer to the server. SSD can hold colder data more cheaply. The system architecture determines what is purchased, in what configuration, and on what timetable.



What to Watch Instead of a Single Cache-Hit Number

When companies talk about inference efficiency, I would put a few questions next to the headline metric.

  • How much prefill compute was avoided? Reused tokens and estimated compute savings are more revealing than raw request-level hits.
  • What happened to time to first token? Look at both typical and tail latency, not just an average.
  • How much state remained resident? Cache retention time, eviction behavior, and memory pressure reveal the capacity cost of the gain.
  • How expensive was restore? A lower HBM footprint is not automatically a faster or cheaper service if offloaded state is slow to retrieve.
  • Did the implementation reach production? Qualification, deployment scale, product mix, shipment disclosures, and margin commentary are separate questions from a technical demo.

There are also clear cases where prefix caching may matter less. Requests with little shared context produce fewer useful hits. Long generated answers shift attention back toward decode. Strict tenant isolation or rapidly changing instructions can limit what is safely shareable. And a poorly chosen offload tier can turn a cache into a delay.

So the clean conclusion is not that caching reduces memory demand, or that caching creates endless memory demand. Prefix caching changes the mix of compute, capacity, and data movement. It asks infrastructure teams to decide what is worth remembering—and how quickly they can bring it back.

This article is for industry and company context, not a recommendation to buy or sell any security.



Appendix: Where Prefix Caching Shows Up

The easiest examples are enterprise AI assistants that repeatedly receive the same policy manual, customer-support systems that continue a long conversation, and agent workflows that send the same tool definitions before each task. In all three cases, the reusable item is not the final answer. It is the intermediate state created while processing the shared beginning of the request.

That small distinction is why this topic belongs in the broader AI infrastructure story. Less repeated computation can be valuable, but the retained state still has to occupy a real place in a real memory hierarchy.



Sources and Date Checked

KV 캐시 재사용, HBM은 덜 필요할까

처음엔 KV 캐시를 재사용한다는 말이 메모리를 덜 쓴다는 뜻처럼 들렸다.

그런데 SK하이닉스가 9월 17일 소개한 SALT-KV를 보니, 질문이 조금 달라진다. 계산한 결과를 HBM·DRAM·SSD 중 어디에 남겨둘지가 핵심이다.

국내 메모리 기업을 볼 때도 HBM 하나의 수요만 따라가면 놓치는 부분이 생기겠다 싶다.

한 번 계산한 것을 다시 쓰면, HBM은 정말 덜 필요할까?

이전 Tech(KR) 글: AI 데이터센터, HBM 다음은 전력·냉각·기판일까





핵심 요약

  • 같은 입력 앞부분을 재사용하면 입력 처리인 prefill을 줄일 수 있다. 새 답을 만드는 decode까지 없어지는 것은 아니다.
  • 동일한 KV 블록을 공유하면 중복 저장을 줄일 수 있지만, 다음 요청에 쓰려고 남긴 캐시도 공간을 차지한다.
  • 메모리가 부족해지면 프롬프트 구성, 요청 배분, 캐시 삭제 순서까지 함께 바뀐다.
  • SK하이닉스의 계층화, 삼성전자의 CXL 메모리 확장, NVIDIA의 요청 배분은 서로 다른 사업 경로다. 기술 소개와 실제 매출 기여는 구분해야 한다.
같은 입력 앞부분의 KV 캐시를 재사용하면 prefill 계산이 줄지만 상태 보관 공간이 필요하다. HBM의 빠른 재사용과 외부 메모리의 복원 전송 비용을 비교하는 도식.
vLLM·NVIDIA 공식 문서를 바탕으로 자체 제작. 4 GiB는 가정에 따른 설명용 수치이며 실제 모델 측정치가 아니다.

자체 도식 · 자료 확인 2026-09-28





계산을 건너뛰는데 왜 HBM을 계속 읽을까?

KV 캐시 재사용은 이미 계산한 입력의 결과를 꺼내 쓰는 방식이다. 새 답을 생성할 때 그 결과를 참조하는 일은 남는다.

예를 들어 회사 규정집을 읽고 답하는 AI를 생각해보자. 여러 질문 앞에 같은 규정집이 붙는다면, 매번 처음부터 입력 처리를 반복할 이유가 없다. 공통 앞부분, 즉 prefix의 KV를 재사용하면 첫 답이 나오기까지의 일을 줄일 수 있다.

다만 답은 질문마다 새로 만들어야 한다. 전체 문맥에 주의를 기울이는 full attention 모델은 새 토큰을 만들 때 과거 문맥의 KV를 참조한다. 저장된 규정집 부분도 여기에 들어간다. 다시 계산하지 않는 것과 다시 읽지 않는 것은 다르다.

vLLM 공식 문서도 prefix caching이 줄이는 것은 prefill이며 decode 시간 자체를 줄이는 기능은 아니라고 설명한다. 답변 생성이 긴 작업이라면 체감 효과가 작을 수 있다.

그래서 나는 캐시 적중률 옆에 생성 토큰당 지연도 같이 놓고 싶다. 입력 처리만 빨라졌는지, 사용자가 끝까지 기다리는 시간도 좋아졌는지 구분하기 위해서다. 실제 HBM 트래픽은 attention 커널, 배치, 하드웨어 캐시에도 영향을 받으므로 재사용률만으로 계산할 수 없다.





같은 4GiB도 열 번 복사하면 이야기가 달라진다

규모를 보기 위해 가상의 모델을 하나 놓아보자. 32개 층, KV head 8개, head 차원 128, 원소당 2바이트, 문맥 32,768토큰인 full-attention GQA 모델이다. 요청 하나의 KV 데이터는 다음과 같다.

2(K와 V) × 32 × 8 × 128 × 2바이트 × 32,768 = 4GiB

이는 설명용 자체 계산이다. 모델 가중치, 관리 정보, 메모리 단편화는 제외했다. GQA는 여러 query head가 KV head를 공유하는 구조이며, 여기서는 KV head 수를 8로 가정했다. 모든 모델의 32K 문맥이 4GiB라는 뜻은 아니다.

이 전체 문맥이 열 요청의 완전히 동일한 prefix이고, PagedAttention 논문에서 설명한 것처럼 하나의 저장본을 공유하는 구현이라면 공통 부분은 논리적으로 4GiB면 된다. 요청마다 따로 복제하면 같은 부분만 40GiB다. 각 요청의 개별 질문과 출력 부분은 별도다.

그렇다고 실제 서버에서 항상 36GiB가 줄어드는 것은 아니다. 서로 다른 GPU에 복사본이 필요하거나 블록 경계가 맞지 않을 수 있다. 한 서버에서 아낀 공간을 더 많은 동시 요청에 쓰면 전체 점유량은 다시 늘 수 있다. 이 계산은 중복 제거의 가능성을 보여주며, HBM 구매량 전망은 아니다.





메모리 부족은 프롬프트와 요청 배분도 바꾼다

메모리가 빠듯할수록 같은 내용을 얼마나 오래, 어디에 보관할지 선택해야 한다. 그 선택이 소프트웨어 설계로 되돌아온다.

먼저 입력의 순서다. vLLM의 설계 문서에서 캐시 블록의 식별에는 해당 토큰뿐 아니라 앞선 prefix도 들어간다. 내용이 비슷하다는 이유만으로 공유되는 구조가 아니다. 이를 응용하면 고정 규칙과 도구 설명은 앞에, 매번 바뀌는 질문은 뒤에 두는 구성이 유리할 수 있다. 다만 순서를 바꿨을 때 답의 품질도 함께 검증해야 한다.

다음은 보관과 삭제다. 다시 오지 않을 요청의 캐시가 자리를 오래 차지하면 새로운 요청이 들어갈 여지가 줄어든다. 반대로 곧 이어질 대화의 캐시를 버리면 입력 처리를 다시 해야 한다. 큰 캐시를 남기는 비용과 버렸다가 다시 만드는 비용 사이의 선택이다.

9월 24일 공개된 prefix cache 교체 정책 연구는 두 회사의 운영 기록과 14개 정책을 분석했다. 해당 조건에서는 복잡한 전통적 정책의 개선 폭이 LRU보다 크지 않았다고 보고한다. 최근에 사용한 캐시를 중시하는 기본 판단부터 튼튼해야 한다는 얘기다. 아직 프리프린트이며 여기서 실험을 재현하지 않았으므로 모든 서비스에 일반화할 수는 없다.

마지막은 요청 배분이다. 캐시가 있는 GPU로 보내면 재계산을 피할 수 있지만, 그 GPU의 대기열이 길면 오히려 늦어진다. NVIDIA는 Dynamo 기술 설명에서 캐시 겹침과 decode 부하를 함께 고려하는 배분을 소개한다. 메모리의 위치가 트래픽을 보내는 기준이 된 셈이다.





SSD와 CXL은 HBM의 어떤 일을 나눠 맡을까?

다시 쓸 가능성은 있지만 당장 계산에 쓰이지 않는 캐시를 보관하는 일이 먼저다. SSD나 확장 메모리를 늘렸다고 GPU 가까이의 HBM과 같은 속도로 읽을 수 있는 것은 아니다.

계층화의 성패를 가르는 질문은 간단하다. 보관한 KV를 찾아 다시 가져오는 시간이, 그것을 재계산하는 시간보다 유리한가? 전송 크기뿐 아니라 연결 대역폭, 대기열, 복원 경로까지 영향을 준다. 복원 때문에 첫 응답이 늦어지면 높은 적중률도 만족스러운 서비스로 이어지지 않는다.

반대로 필요한 데이터를 미리 올려놓거나, 잠깐 멈춘 대화의 캐시를 넓은 보관 공간에 남길 수 있다면 HBM을 더 활발한 요청에 배정할 여지가 생긴다. 이것은 설계상의 조건부 효과다. 저장 장치 용량 증가가 곧 추론 성능 증가라는 주장은 아니다.





국내 메모리 기업과 NVIDIA를 잇는 실제 경로

SK하이닉스(KRX: 000660)는 9월 17일 공식 행사 소개에서 자사 eSSD를 장착한 SALT-KV 시연을 공개했다. 문맥별 KV 조각의 재사용 가치와 저장 비용을 따져 HBM·DRAM·SSD에 배치한다고 설명한다. 국내 투자자가 확인할 경로는 캐시 보관 계층의 도입이 실제 eSSD·DRAM 구성과 출하로 이어지는지다. 시연만으로 고객 채택량이나 매출 기여를 알 수는 없다.

삼성전자(KRX: 005930)는 CMM-D MD220 제품 문서에서 CXL 기반 DRAM 메모리 확장을 제시한다. 이는 서버가 확보할 수 있는 메모리 용량과 관련된 직접적인 제품 노출이다. 다만 이 문서는 특정 KV 서비스 채택이나 SALT-KV 연동을 확인해주지 않는다. KV 보관용으로 쓰일 가능성은 분석상의 연결이며, 실제 경로는 서버 지원, 소프트웨어 통합, 고객 채택으로 검증해야 한다.

NVIDIA(NASDAQ: NVDA)의 Dynamo는 GPU 추론 시스템에서 캐시를 재사용하도록 요청을 배분하는 소프트웨어 경로다. 같은 GPU 구성으로 목표 지연을 지키며 더 많은 요청을 처리할 수 있는지가 중요한 검증 대상이다. 소프트웨어 효율 개선을 곧바로 GPU 판매 증가로 환산할 수는 없다. 필요한 장비 수가 줄어드는 효과와 서비스 사용량이 늘어나는 효과를 모두 봐야 한다.





적중률 다음에 확인할 숫자들

기업 자료를 읽을 때 나는 다음 항목을 함께 확인하려 한다.

  • 얼마나 비싼 계산을 아꼈나: 요청 수 기준 적중률만 보지 말고, 재사용한 토큰과 절약한 prefill 계산량을 확인한다.
  • 복원이 얼마나 걸렸나: HBM 밖에서 가져오는 시간과 첫 토큰 지연의 상위 구간을 함께 본다.
  • 동시에 얼마나 처리했나: 같은 모델·문맥 길이·응답 길이·지연 목표에서 처리량을 비교한다.
  • 제품 매출로 연결됐나: 시연 이후 고객 검증, 실제 탑재, 출하·매출 공시가 이어지는지 확인한다.

반대 사례도 분명하다. 입력이 제각각이면 공유할 prefix가 적고, 답이 길면 decode 비용이 더 중요해진다. 모델이나 보안 격리 조건이 다르면 같은 문장도 공유하지 못할 수 있다. 확장 메모리에 많이 남겨놨는데 전송 경로가 막히는 경우도 있다.

결국 제목의 질문에는 조건부로 답해야 한다. KV 공유는 요청당 중복 저장을 줄일 수 있지만, 그것만으로 HBM 수요 감소를 결론낼 수 없다. 국내 기업을 볼 때는 HBM을 얼마나 대체하는가와 함께 DRAM·eSSD·CXL 제품이 어떤 역할을 추가로 맡는지 살펴볼 만하다. 이 글은 매수·매도 추천이 아니라 산업·기업을 읽기 위한 정보다.





Appendix. KV 캐시 재사용은 어디에 쓰일까

동일한 규정집을 두고 여러 질문을 받는 업무, 긴 대화를 이어가는 상담, 같은 도구 설명을 반복해서 보내는 에이전트가 이해하기 쉬운 사례다. 재사용되는 것은 최종 답변이 아니라 입력 처리 과정의 중간 결과다. 따라서 캐시에 적중해도 사용자의 새 질문에 대한 답은 다시 생성한다.

여기서 사용한 4GiB는 이진 단위로 4,294,967,296바이트다. 모델 구조와 KV 정밀도가 바뀌면 값도 바뀐다. 용량 계산과 실제 메모리 대역폭 측정은 별도의 작업이다.





출처와 확인일

자료 확인일은 2026년 9월 28일이다. 아래 발표들을 최근 72시간의 신규 발표로 취급하지 않았다.

AI 데이터센터, HBM 다음은 전력·냉각·기판일까

요즘 AI 반도체 이야기가 나오면 가장 먼저 HBM부터 떠올리게 된다.

그런데 AI 데이터센터를 조금 더 길게 들여다보면, 메모리 다음에 필요한 것은 결국 전기와 열을 다루는 설비, 그리고 칩을 시스템으로 연결하는 기판일 수 있다.

서버 한 대가 늘어나는 문제가 아니다. 랙 하나에 더 많은 연산을 넣으려 할수록 전력은 더 안정적으로 들어와야 하고, 열은 더 빨리 밖으로 빠져나와야 한다.

HBM 다음을 찾는다는 말은 새로운 유행어를 찾자는 뜻보다는, AI 데이터센터 안에서 병목이 어디로 이동하는지 보자는 이야기다.

영국 국립문서보관소의 서버실 내부
사진: 영국 국립문서보관소 서버실 · EduVolunteer · CC BY 3.0 · 원문: https://commons.wikimedia.org/wiki/File:A_view_of_the_server_room_at_The_National_Archives.jpg


핵심 요약

  • AI 데이터센터 투자는 GPU나 HBM 한 가지 부품으로 끝나지 않고 전력·냉각·기판까지 이어지는 긴 공급망을 만든다.
  • AI 서버 집적도가 높아질수록 데이터센터 운영에서는 전력 공급과 발열 관리의 중요성이 함께 커질 수 있다.
  • 고성능 반도체를 안정적으로 연결하려면 패키지 기판과 고사양 PCB·소재의 역할도 계속 확인할 필요가 있다.
  • 산업의 관심이 곧바로 특정 기업의 실적이나 주가로 이어지는 것은 아니다. 실제 수주, 고객 인증, 생산능력과 매출 비중을 따로 봐야 한다.


AI 데이터센터의 전력 밀도

AI 데이터센터를 이야기할 때 서버 대수만 보면 그림이 조금 단순해진다. 같은 공간 안에 더 많은 연산 장비를 넣고, 더 높은 성능의 가속기를 가동하려 하면 랙 단위의 전력 사용량과 발열도 함께 올라간다. 이때부터 데이터센터는 단순히 서버를 많이 꽂아두는 공간이 아니라, 전기를 얼마나 안정적으로 받아 나누고 열을 얼마나 효율적으로 빼내느냐가 중요한 설비 산업에 가까워진다.

AI 서버 랙을 중심으로 전력 설비, 액체 냉각 장치, 반도체 기판이 연결된 데이터센터 인프라 일러스트
AI 데이터센터의 전력·냉각·기판 연결 구조를 표현한 자체 생성 이미지

IEA는 데이터센터 전력소비가 2025년 485TWh에서 2030년 약 950TWh로 늘 수 있다고 봤다. 이 수치 하나만으로 어느 기업의 실적을 예측할 수는 없지만, AI 투자가 커질수록 전력망 연결, 변압기, 스위치기어, UPS, 전력 분배 장치가 함께 논의되는 이유는 분명해 보인다.

전력은 눈에 잘 보이지 않지만 데이터센터에서는 서버가 멈추지 않게 만드는 가장 기본적인 조건이다. GPU와 HBM 주문이 잡혔다고 해도 전력 인입과 배전, 백업 전원 설계가 따라가지 못하면 계획한 만큼의 장비를 가동하기 어렵다. AI 투자 뉴스가 나올 때 전력망과 전력기기 이야기가 함께 나오는 이유도 여기에 있다.



냉각은 부가 장비가 아니라 운영 변수

연산량이 높아질수록 열은 더 큰 문제가 된다. 공기를 이용한 냉각은 여전히 넓게 쓰이지만, 고집적 AI 서버 환경에서는 액체 냉각을 포함한 여러 방식이 검토된다. 어떤 방식이 표준이 될지는 서버 구조와 데이터센터 설계에 따라 달라지겠지만, 열을 어떻게 처리할 것인가가 투자비와 운영비를 함께 좌우할 수 있다는 점은 분명하다.

여기서 볼 것은 냉각 장비 한 종류가 아니다. 칠러, 열교환기, 펌프, 냉각 분배 장치, 배관, 제어 시스템처럼 여러 구성요소가 묶여 들어간다. 데이터센터 운영사가 어떤 냉각 방식을 채택하는지에 따라 실제 매출이 발생하는 위치도 달라진다.

국내 투자자 입장에서는 냉각이라는 단어만 보고 기업을 연결하기보다, 해당 회사가 데이터센터용 제품을 공식적으로 공급하고 있는지, 시험·인증 단계인지, 반복 수주가 가능한 구조인지를 구분해 보는 편이 좋다. 기대감은 빠르지만 매출은 보통 고객사의 설계와 증설 일정 뒤에 따라오기 때문이다.



기판은 칩과 시스템 사이의 조용한 연결고리

HBM과 AI 가속기의 성능이 높아질수록 칩을 어떻게 패키징하고 연결할지도 더 중요해질 수 있다. 여기서 기판은 단순히 부품을 올려두는 판이 아니다. 고속 신호와 전력을 안정적으로 전달하고, 복잡한 칩 구성을 시스템으로 묶는 역할을 한다.

패키지 기판, 고다층 PCB, 동박적층판 같은 분야가 AI 서버 공급망에서 자주 거론되는 이유도 이 때문이다. 다만 같은 기판 산업 안에서도 모바일, 네트워크, 자동차, 일반 서버, AI 서버의 비중은 서로 다르다. AI 관련 수요가 커진다고 해서 모든 기판 업체가 같은 속도로 움직인다고 보기는 어렵다.

그래서 기판은 “AI 수혜”라는 표현보다 제품 믹스와 기술 난이도를 보는 쪽이 더 흥미롭다. 고사양 제품 비중이 실제로 높아지는지, 고객사 인증이 끝났는지, 생산능력을 늘릴 만큼 주문 가시성이 생겼는지가 더 중요한 질문이 된다.



관련 기업

AI 데이터센터 공급망을 국내 증시 관점에서 볼 때는 기업 이름을 먼저 고르기보다 역할을 먼저 나누는 편이 정리가 쉽다. 메모리와 후공정은 AI 서버의 연산 성능을 끌어올리는 쪽이고, 전력기기와 전력 인프라는 서버가 안정적으로 돌아가게 만드는 쪽이다. 냉각 설비와 열관리 부품은 높은 집적도를 감당하는 역할을 하며, 기판·PCB·소재 업체는 칩과 보드를 실제 시스템으로 연결하는 위치에 있다.

예를 들어 전력 쪽에서는 HD현대일렉트릭과 LS ELECTRIC처럼 데이터센터용 전력기기 사업 연결을 공식 자료에서 확인할 수 있는 기업을, 냉각 쪽에서는 AI 데이터센터용 냉각 솔루션을 공개한 LG전자 등을 살펴볼 수 있다. 다만 같은 키워드로 묶여도 고객군과 적용처, 매출 비중이 다르다. 이 글은 매수·매도 추천이 아니라 AI 인프라와 기업을 읽기 위한 정보다.



투자 체크포인트

  • 글로벌 클라우드 기업과 데이터센터 운영사의 설비투자가 실제로 이어지는지
  • AI 서버의 랙 구조와 전력·냉각 방식이 어떻게 바뀌는지
  • 관련 기업이 고객사 인증, 신규 수주, 증설 계획을 공식 자료로 공개하는지
  • 매출 증가가 단순 물량 확대인지, 고사양 제품 비중 상승에 따른 수익성 개선인지
  • 전력망 인허가, 장비 납기, 고객사의 투자 일정 같은 병목이 생기지 않는지

특히 수주 공시나 실적 발표를 볼 때는 AI라는 단어가 들어갔는지만 보기보다, 어느 제품이 어느 고객군에 공급되는지와 해당 매출이 전체에서 차지하는 비중을 함께 확인해야 한다. AI 인프라 수요가 강해도 계통 연결, 인허가, 장비 공급 제약으로 프로젝트 일정은 늦어질 수 있다.



AI 데이터센터의 다음 병목은 어디일까

AI 데이터센터가 커질수록 돈의 흐름은 한 번에 이동하지 않는다. 처음에는 가속기와 메모리 주문이 눈에 띄고, 이후에는 전력 공급과 냉각 설계, 기판과 부품 조달처럼 시스템을 실제로 완성하는 영역이 중요해질 수 있다. 반대로 AI 투자 속도가 조절되면 이 공급망도 같은 시점에 같은 폭으로 움직이지 않을 가능성이 있다.

그래서 HBM 다음을 하나의 정답으로 고르기보다는, 데이터센터 안에서 다음 병목이 무엇인지 계속 보는 편이 낫다. 전기가 부족하면 전력 설비가, 열이 문제가 되면 냉각이, 연결과 집적도가 어려워지면 기판과 패키징이 주목받을 수 있다. AI 산업이 커질수록 반도체 바깥의 공급망까지 같이 봐야 하는 이유다.



Appendix. AI 데이터센터는 어떻게 구성될까

AI 데이터센터는 크게 연산 장비, 메모리와 저장장치, 네트워크, 전력 설비, 냉각 설비, 건물·전력망 인프라로 나눠볼 수 있다. HBM은 연산 장비의 성능과 연결되는 중요한 부품이지만, 서버가 실제로 24시간 돌아가려면 전력과 냉각, 기판과 보드까지 모두 맞물려야 한다.



출처 및 업데이트

이 글은 산업과 기업을 읽기 위한 정보이며 특정 증권의 매수·매도 추천이 아닙니다.

Beyond HBM: Where AI Data Center Spending Goes Next

At first, the AI trade looked simple. More AI meant more GPUs, and more GPUs meant more HBM.

That logic is still important. But a data center does not become useful the moment a chip leaves a factory.

It has to be powered, cooled, connected, packaged, installed, and kept running. That is where the AI story starts to look much wider than a single semiconductor category.

For technology investors, the more useful question may be this: when AI data centers expand, where does the money go after the GPU and HBM order is placed?

Server room at The National Archives in the United Kingdom
Photo: server room at The National Archives · EduVolunteer · CC BY 3.0 · source: https://commons.wikimedia.org/wiki/File:A_view_of_the_server_room_at_The_National_Archives.jpg


Key Takeaways

  • AI infrastructure spending begins with accelerators and memory, then continues through power equipment, cooling systems, networking, packaging materials, and data-center construction.
  • As computing density rises, electricity availability and heat removal become practical limits, not merely operating costs.
  • Advanced package substrates matter because they help connect increasingly complex chips to the rest of the system.
  • Different companies can benefit from the same AI buildout, but their results still depend on order timing, capacity, customer concentration, and margins.


The Key Question: What Happens After the GPU Order?

It is easy to understand why GPUs and HBM receive so much attention. They sit at the center of AI training and inference. Without powerful accelerators and fast memory, there is no modern AI cluster to speak of.

But a large AI cluster is not a box of chips. The chips need servers, the servers need racks, the racks need networking, and the entire facility needs a reliable flow of electricity and a way to move heat out of the building. The AI investment story can therefore spread in stages rather than arrive in one semiconductor order.

The point is not that HBM suddenly becomes less important. It is that HBM is one visible part of a much longer spending chain. NVIDIA’s FY2026 data-center revenue provides the scale of the accelerator buildout; the next question is whether the physical infrastructure can keep pace.



Power Has Become a Physical Constraint

Traditional data centers already consume significant electricity, but AI workloads change the conversation because they can concentrate much more computing power into a smaller footprint. The IEA projects data-center electricity consumption to rise from 485TWh in 2025 to about 950TWh in 2030, while AI-focused data-center demand could triple over the same period.

An AI server rack connected to power equipment, liquid cooling, and a semiconductor substrate in a data center
Original illustration of the power, cooling, and substrate layers behind an AI data center

That shifts attention toward equipment that once looked less glamorous: transformers, switchgear, uninterruptible power systems, power distribution units, busways, backup generation, and the grid connections outside the building. For a cloud provider, the question is not only whether it can buy more servers. It is whether enough power can be delivered to the site, on schedule, with the reliability required for an always-on service.

AI demand may be global, but power availability is local. Two projects with similar server budgets can move at very different speeds if one site has an easier path to interconnection, permitting, and electrical equipment. That difference matters when investors read capital-expenditure headlines.



Cooling Is No Longer a Side Detail

More power becomes more heat. That sounds obvious, but it is becoming one of the defining practical issues in AI infrastructure. Air cooling remains useful across much of the installed base, while higher-density deployments may call for improved airflow management, rear-door heat exchangers, direct-to-chip liquid cooling, or other liquid-based designs.

There is no single cooling architecture for every AI facility. The choice depends on workload, location, water availability, building layout, and an operator’s own engineering preferences. The durable point is that thermal management is moving closer to the center of the purchasing decision.

Cooling also affects construction, maintenance, energy use, and the speed at which additional computing capacity can be brought online. Vertiv’s 2026 results are one example of how suppliers describe the trend toward more complex, infrastructure-intensive AI deployments; they should not be read as proof that every cooling supplier will grow at the same rate.



Why Advanced Substrates Matter

Substrates are less visible than GPUs, but they are a useful reminder that advanced computing depends on specialized manufacturing below the finished chip. In semiconductor packaging, a substrate provides electrical pathways between a chip package and the circuit board. As chips become more complex, with more connections and tighter performance requirements, packaging and substrate technology become more demanding.

This matters in AI systems because high-performance processors, memory, and advanced packaging must work together under demanding electrical and thermal conditions. Ibiden’s announced investment in package substrates for AI and high-performance servers is a concrete example of the supply chain expanding beyond the silicon die.

The market does not reward every substrate supplier in the same way. Technology specifications, yield rates, customer qualification cycles, and capacity expansion all still matter. The broader lesson is simply that AI infrastructure is also a materials-and-manufacturing story.



Related Companies

This theme is better viewed as a map than as a single-stock story. At the chip level, accelerator designers, memory producers, foundries, and advanced-packaging specialists are the most visible participants. At the system level, server makers, networking suppliers, and optical-component companies turn silicon into an operating AI cluster.

Then comes the physical layer. Electrical-equipment manufacturers such as Eaton and Schneider Electric, thermal-management specialists such as Vertiv, data-center operators, engineering firms, and utilities can all sit somewhere along the path between an AI model and a functioning data center. These categories are not interchangeable: each has different lead times, customer dependencies, and margin structures.



Investment Watchpoints

  • Cloud-provider capital-expenditure plans and data-center construction updates
  • Utility interconnection timelines, permitting, and transformer or switchgear lead times
  • Server-rack density and stated cooling architecture
  • Package-substrate capacity additions, utilization, and yield commentary
  • The gap between an announced order and the period when it is recognized as revenue

Demand does not automatically become profit. A company can receive strong orders and still face component shortages, pricing pressure, project delays, or the cost of expanding capacity. This is not a buy-or-sell recommendation; it is a framework for following how AI capital spending moves through a supply chain.



What Could Slow the Story Down?

AI infrastructure spending can be powerful without moving in a straight line. Power shortages, permitting delays, grid upgrades, construction bottlenecks, component availability, and slower-than-expected returns on AI services can all affect the pace of deployment.

There is also a timing issue. Semiconductor orders, server shipments, and facility construction do not always happen in the same quarter. A strong chip cycle does not guarantee that every downstream supplier reports the same kind of growth at the same time.



Appendix. What Do Power, Cooling, and Substrates Actually Do?

Power infrastructure brings electricity from the grid into the facility and distributes it safely to computing equipment. Cooling systems remove the heat created by that equipment. Substrates help connect advanced semiconductor packages to the broader electronic system.

None of these areas may sound as exciting as a new AI-model launch. But without them, even the most advanced chips cannot operate at scale. AI is software at the surface, silicon at the core, and physical infrastructure underneath.



Sources and Update

This article is for general information and industry discussion only; it is not investment advice or a recommendation to buy or sell any security.

Japan Raises Rates to 1.25%, Why Should U.S. Tech Investors Care?


A rate decision in Tokyo can suddenly feel very close to a portfolio full of American tech stocks.

On September 18, the Bank of Japan announced an increase in its policy-rate guideline from around 1.0% to around 1.25%, effective September 24.

The connection I find interesting is the money behind a trade: a position in U.S. tech can be financed in another currency.

So how does a quarter-point move in Japan travel all the way to NVIDIA, Microsoft, and the Nasdaq?

Key Takeaways

  • The BOJ approved the increase by a 7–2 vote. The announcement and effective date are different: the new guideline starts on September 24. BOJ decision
  • Higher yen funding costs, especially alongside a strengthening yen, can put pressure on leveraged positions and encourage selling across markets.
  • Technology shares can feel that pressure through investor positioning and changing valuations. Lasting business damage requires a separate look at earnings and cash flow.
  • The outcome depends on expectations, currency moves, and U.S. financial conditions. An anticipated hike, a stable yen, and solid earnings can produce a much quieter story.


Key Driver: The Yen Carry Trade

The route starts with a simple attraction: borrow where financing is cheaper, then invest where the prospective return looks better. When the borrowing currency is the yen, this is commonly called a yen carry trade. Actual trades can use derivatives as well as loans, as the BIS explains.

That arrangement gets uncomfortable when financing becomes more expensive or the yen strengthens. For an unhedged investor, a dollar asset must eventually cover a liability measured in yen. A favorable stock return can be eaten away by the currency conversion.

Here is a hypothetical example. An investor borrows ¥15 million and exchanges it at ¥150 per dollar, receiving $100,000. Suppose the investment stays worth $100,000, but the exchange rate moves to ¥135 per dollar. Converting back produces only ¥13.5 million: a ¥1.5 million shortfall against the original principal, before interest and fees.

The stock has gone nowhere, yet the financing position has a problem. This example assumes no currency hedge or investment income; it is not a return calculation for an ordinary dollar-funded shareholder.

If investors face collateral calls or risk limits, they may reduce positions quickly. Readily traded U.S. stocks can become a source of cash even while the underlying companies continue operating normally.

We have seen this mechanism matter before. The BIS's account of August 2024 describes disappointing U.S. labor data interacting with leveraged equity and currency positions; volatility and margin requirements amplified the unwind. That episode involved several forces, so it offers a useful lesson about leverage rather than a timetable for another selloff. BIS, August 2024 analysis



Key Driver: Why Tech Can Feel the Pressure

Popular technology stocks can sit at the intersection of crowded positions and demanding growth expectations. If a leveraged portfolio needs cash, a liquid holding may be sold regardless of whether its latest product is selling well. That is a portfolio decision, with consequences for the share price.

Valuation adds another channel. Investors pay today for profits they expect in the future. If the return they require rises while those profit expectations stay unchanged, the present value falls. Businesses whose valuations depend heavily on distant cash flows can be especially sensitive to that calculation. The Federal Reserve's valuation framework separates the safe interest rate from the additional return investors require for taking risk.

The important bridge is U.S. financial conditions. A BOJ hike does not automatically lift the dollar discount rate applied to an American company.

Higher Japanese bond yields could make domestic bonds more attractive to Japanese investors and reduce demand for overseas bonds. The IMF's April 2026 assessment discusses that possibility, while noting that large institutional portfolios tend to adjust gradually.

At the same time, nervous investors may buy U.S. Treasuries. In the August 2024 turbulence, government bond yields fell as growth concerns and expectations of policy easing increased. BIS Quarterly Review

So watch the reason behind a yield move. Lower Treasury yields can support valuations, but weakening earnings expectations or a larger risk premium can offset that support. Tech shares respond to the combination.

Related Companies

NVIDIA — Nasdaq: NVDA. NVIDIA makes the discussion tangible because its business is closely tied to AI infrastructure investment. Its latest earnings release provides the business context. After any market funding shock, data-center revenue, gross margin, customer spending, and management's outlook would help assess whether the operating story has changed. Investor selling and weaker chip demand are different developments; the share price alone cannot tell us which one is occurring.

Microsoft — Nasdaq: MSFT. Microsoft's cloud and AI activities offer another angle. Its latest results provide a starting point for tracking Azure, infrastructure spending, and cash generation. Investors can compare the pace of investment with the revenue and cash flow it supports. A change in market appetite may affect the valuation immediately, while the business effect takes longer to establish through customer demand, spending plans, and financial results.

These companies illustrate what to examine. There is no claim here that either company, or a particular shareholder, funds its positions in yen.

Investment Watchpoints

The useful follow-up is to track how the decision passes through markets.

  • The yen's direction and speed. A rapid move can strain an unhedged position more abruptly than a modest change in annual financing costs. The announcement alone does not tell us where the currency goes next.
  • The expected policy path. Compare the BOJ's guidance with expectations for the Fed. Markets look ahead: the Fed's policy explainer describes how expected future policy influences longer-term rates. We should not label this decision a surprise without evidence of pre-meeting expectations.
  • Treasury yields, volatility, and credit spreads. Together they help show whether investors are reassessing rates, seeking safety, or demanding more compensation for risk.
  • Company updates. Orders, revenue, margins, cash flow, and investment plans help distinguish a change in the price investors will pay from a change in what a business can earn.

A stable yen and resilient company results would tell a different story from a fast currency move accompanied by tighter financing and earnings cuts. That is why the transmission deserves more attention than the headline rate alone.

Appendix: Reading the Yen and the Rate Gap

USD/JPY means the number of yen needed to buy one U.S. dollar. A move from 150 to 135 means the yen has strengthened: each dollar buys fewer yen. The numbers above are illustrative, not current market quotes.

The policy-rate gap is a useful starting point for comparing currencies, but it is not an investor's actual borrowing bill. Maturity, lender terms, collateral, and currency hedging all affect the economics. Fully hedged positions therefore behave differently from the simple unhedged example.

A 25-basis-point increase means 0.25 percentage point. Its effect on a portfolio depends on position size, financing arrangements, currency exposure, and how long the position remains open.

This article provides market information, not a recommendation to buy or sell securities.

Sources and Update

Updated September 18, 2026. Research cutoff: approximately 16:44 JST/KST (03:44 EDT), before the U.S. regular stock-market session opened. This article explains possible transmission channels; it does not report that day's U.S. market reaction.

A rate decision in Tokyo can suddenly feel very close to a portfolio full of American tech stocks.

On September 18, the Bank of Japan announced an increase in its policy-rate guideline from around 1.0% to around 1.25%, effective September 24.

The connection I find interesting is the money behind a trade: a position in U.S. tech can be financed in another currency.

So how does a quarter-point move in Japan travel all the way to NVIDIA, Microsoft, and the Nasdaq?

Key Takeaways

  • The BOJ approved the increase by a 7–2 vote. The announcement and effective date are different: the new guideline starts on September 24. BOJ decision
  • Higher yen funding costs, especially alongside a strengthening yen, can put pressure on leveraged positions and encourage selling across markets.
  • Technology shares can feel that pressure through investor positioning and changing valuations. Lasting business damage requires a separate look at earnings and cash flow.
  • The outcome depends on expectations, currency moves, and U.S. financial conditions. An anticipated hike, a stable yen, and solid earnings can produce a much quieter story.

Key Driver: The Yen Carry Trade

The route starts with a simple attraction: borrow where financing is cheaper, then invest where the prospective return looks better. When the borrowing currency is the yen, this is commonly called a yen carry trade. Actual trades can use derivatives as well as loans, as the BIS explains.

That arrangement gets uncomfortable when financing becomes more expensive or the yen strengthens. For an unhedged investor, a dollar asset must eventually cover a liability measured in yen. A favorable stock return can be eaten away by the currency conversion.

Here is a hypothetical example. An investor borrows ¥15 million and exchanges it at ¥150 per dollar, receiving $100,000. Suppose the investment stays worth $100,000, but the exchange rate moves to ¥135 per dollar. Converting back produces only ¥13.5 million: a ¥1.5 million shortfall against the original principal, before interest and fees.

The stock has gone nowhere, yet the financing position has a problem. This example assumes no currency hedge or investment income; it is not a return calculation for an ordinary dollar-funded shareholder.

If investors face collateral calls or risk limits, they may reduce positions quickly. Readily traded U.S. stocks can become a source of cash even while the underlying companies continue operating normally.

We have seen this mechanism matter before. The BIS's account of August 2024 describes disappointing U.S. labor data interacting with leveraged equity and currency positions; volatility and margin requirements amplified the unwind. That episode involved several forces, so it offers a useful lesson about leverage rather than a timetable for another selloff. BIS, August 2024 analysis

Key Driver: Why Tech Can Feel the Pressure

Popular technology stocks can sit at the intersection of crowded positions and demanding growth expectations. If a leveraged portfolio needs cash, a liquid holding may be sold regardless of whether its latest product is selling well. That is a portfolio decision, with consequences for the share price.

Valuation adds another channel. Investors pay today for profits they expect in the future. If the return they require rises while those profit expectations stay unchanged, the present value falls. Businesses whose valuations depend heavily on distant cash flows can be especially sensitive to that calculation. The Federal Reserve's valuation framework separates the safe interest rate from the additional return investors require for taking risk.

The important bridge is U.S. financial conditions. A BOJ hike does not automatically lift the dollar discount rate applied to an American company.

Higher Japanese bond yields could make domestic bonds more attractive to Japanese investors and reduce demand for overseas bonds. The IMF's April 2026 assessment discusses that possibility, while noting that large institutional portfolios tend to adjust gradually.

At the same time, nervous investors may buy U.S. Treasuries. In the August 2024 turbulence, government bond yields fell as growth concerns and expectations of policy easing increased. BIS Quarterly Review

So watch the reason behind a yield move. Lower Treasury yields can support valuations, but weakening earnings expectations or a larger risk premium can offset that support. Tech shares respond to the combination.

Related Companies

NVIDIA — Nasdaq: NVDA. NVIDIA makes the discussion tangible because its business is closely tied to AI infrastructure investment. Its latest earnings release provides the business context. After any market funding shock, data-center revenue, gross margin, customer spending, and management's outlook would help assess whether the operating story has changed. Investor selling and weaker chip demand are different developments; the share price alone cannot tell us which one is occurring.

Microsoft — Nasdaq: MSFT. Microsoft's cloud and AI activities offer another angle. Its latest results provide a starting point for tracking Azure, infrastructure spending, and cash generation. Investors can compare the pace of investment with the revenue and cash flow it supports. A change in market appetite may affect the valuation immediately, while the business effect takes longer to establish through customer demand, spending plans, and financial results.

These companies illustrate what to examine. There is no claim here that either company, or a particular shareholder, funds its positions in yen.

Investment Watchpoints

The useful follow-up is to track how the decision passes through markets.

  • The yen's direction and speed. A rapid move can strain an unhedged position more abruptly than a modest change in annual financing costs. The announcement alone does not tell us where the currency goes next.
  • The expected policy path. Compare the BOJ's guidance with expectations for the Fed. Markets look ahead: the Fed's policy explainer describes how expected future policy influences longer-term rates. We should not label this decision a surprise without evidence of pre-meeting expectations.
  • Treasury yields, volatility, and credit spreads. Together they help show whether investors are reassessing rates, seeking safety, or demanding more compensation for risk.
  • Company updates. Orders, revenue, margins, cash flow, and investment plans help distinguish a change in the price investors will pay from a change in what a business can earn.

A stable yen and resilient company results would tell a different story from a fast currency move accompanied by tighter financing and earnings cuts. That is why the transmission deserves more attention than the headline rate alone.

Appendix: Reading the Yen and the Rate Gap

USD/JPY means the number of yen needed to buy one U.S. dollar. A move from 150 to 135 means the yen has strengthened: each dollar buys fewer yen. The numbers above are illustrative, not current market quotes.

The policy-rate gap is a useful starting point for comparing currencies, but it is not an investor's actual borrowing bill. Maturity, lender terms, collateral, and currency hedging all affect the economics. Fully hedged positions therefore behave differently from the simple unhedged example.

A 25-basis-point increase means 0.25 percentage point. Its effect on a portfolio depends on position size, financing arrangements, currency exposure, and how long the position remains open.

This article provides market information, not a recommendation to buy or sell securities.

Sources and Update

Updated September 18, 2026. Research cutoff: approximately 16:44 JST/KST (03:44 EDT), before the U.S. regular stock-market session opened. This article explains possible transmission channels; it does not report that day's U.S. market reaction.