Mixture-of-experts models promise a neat idea: use only the specialists needed for each token, rather than running the entire model every time.
That can lower active compute. It does not make the model’s full collection of expert weights disappear.
Those weights still need to wait somewhere, and the moment they leave fast accelerator memory, the problem becomes less about storage capacity and more about timing.
Can the right expert arrive before the GPU needs it? That small question may decide whether offloading is a cost-saving design or a new source of delay.
Previous Tech(EN) post: Prefix Caching Saves AI Compute—But It Does Not Make the Memory Problem Disappear
Key Takeaways
- MoE activates only some experts for a token, but the full expert-weight pool must still be stored somewhere in the system.
- Keeping more weights resident in fast memory can make latency more predictable; offloading can increase usable capacity but adds a transfer deadline.
- A prefetch is useful only if the needed weights arrive before the relevant layer runs. A nominal cache hit that arrives late can still stall inference.
- Research on expert offloading is promising but conditional. Benchmark results, model changes, hardware configuration, and quality trade-offs need to be read together.
- NVIDIA, AMD, and Micron have different infrastructure exposure to the issue. That does not establish that they supply the cited research systems or that any one technique determines revenue or share performance.
Original illustration generated for this publication team; no third-party artwork, logos, or copied diagrams used.
A Sparse Model Still Has a Full Closet of Weights
An MoE model has multiple expert networks and a router that chooses a small subset for each token. The appeal is obvious: the model does not need to run every expert on every step.
But sparse activation is not the same thing as a small total model. A hotel may have hundreds of rooms even when only a few guests check in at once. Likewise, an MoE system may activate only a few experts while still needing access to a large pool of weights.
The simplest way to serve that model is to keep all the relevant weights resident close to compute. That favors predictable response time, but it consumes accelerator memory. At a larger model size or higher concurrency, holding everything nearby can become expensive or impossible within a chosen hardware configuration.
Offloading changes the arrangement. A serving system can keep a working set of likely experts in fast memory and store less active experts in a lower tier. The full model remains available, but not every part is immediately available at GPU-memory speed.
That is why the active-parameter headline needs a second question beside it: how large is the full weight pool, and how reliably can the next required expert be made ready?
Offloading Is Really a Deadline Problem
It is tempting to describe offloading as a simple capacity fix: put the weights that do not fit in HBM somewhere else. The harder part is bringing them back before the model reaches the layer that needs them.
Imagine a production line that needs a particular tool at the next station. The tool can be kept in a cheaper warehouse, but that helps only if it reaches the station before the work arrives. If it shows up late, the line waits.
MoE inference has a comparable sequence. The router selects experts as the model processes tokens and layers. A runtime may try to predict which expert will be needed next, start the transfer early, and evict a lower-value resident expert to make room. The system succeeds when the selected weights arrive before their compute deadline.
The September 11, 2026 SeqMoE preprint proposes this kind of predictive memory management, combining sequence prediction, deadline-aware prefetching, cache eviction, and a graph-compatible runtime. In its evaluation, the paper reports 45% expert residency, a 96.97% average hit rate, and 80.22% of full-load performance. Those are paper-specific results under the authors’ test conditions, not a general service-level guarantee or a hardware benchmark for every MoE workload.
This is the useful distinction: a high cache hit rate tells us that a requested expert was found in a planned location. An on-time hit tells us that it was ready early enough to avoid a stall. For an interactive AI service, the second measure is usually the one a user feels.
Why an SSD Is Not a Drop-In Substitute for HBM
HBM, server DRAM, and enterprise SSDs all store data, but they are not interchangeable places to run an MoE model. They differ in capacity, bandwidth, latency, connection path, and the software work required to use them well.
HBM is close to the accelerator and designed for high-bandwidth compute workloads. Server DRAM can offer a larger memory tier with different access characteristics. SSDs can provide much more economical capacity for colder data, but an expert weight stored there has to travel through a longer path before it can be used by the GPU.
The September 16, 2026 Edge0 preprint makes the timing challenge very clear. Its authors argue that merely placing experts on an SSD is insufficient because the next-layer routing decision normally arrives too late to hide the read. Their proposed answer is a prerouter that makes a one-token-ahead routing choice, plus a recovery LoRA trained to address quality loss associated with routing replacement and INT4 quantization.
That detail is important. Edge0 is not simply a lossless cache-management recipe. It changes part of the routing behavior and adds a quality-recovery component. The paper reports a 35B-class model serving at 20 tokens per second with 3 GiB of peak active memory on a single 24 GB system. That reported outcome belongs to its particular model, hardware, precision, and evaluation setup; it should not be read as proof that SSD-backed MoE serving is already a universal production solution.
Transfer speed alone also does not settle the question. Queueing, interconnect topology, host overhead, prefetch accuracy, and concurrency can change whether loading a weight beats leaving it resident or recomputing a different design choice. Peak specification numbers do not tell the whole serving story.
Some of the Memory Problem Can Move Into Model Design
There is more than one response to a growing expert pool. A runtime can move weights around more cleverly, but model designers can also reduce how much distinct expert material must be available at once.
The MoRE preprint, revised September 17, 2026, explores one such direction: sharing an expert pool across adjacent layers while keeping layer-specific routers. This is an architectural approach to parameter growth, not an expert-offloading runtime.
That separation matters for readers following AI infrastructure. A software cache, a storage tier, an interconnect strategy, and a redesigned model can all improve a capacity problem in different ways. Treating them as the same solution would hide the trade-offs.
MoRE is marked as accepted to COLM 2026, but its reported experiments cover models from 114 million to 1.15 billion parameters. That is meaningful research evidence, not enough to generalize its results to every frontier-scale or commercial serving environment.
Expert Parallelism and Expert Offloading Are Different Moves
There is another way to avoid putting every expert on one accelerator: distribute experts across more GPUs. That is expert parallelism, not offloading to a lower memory tier.
NVIDIA documents wide expert parallelism in TensorRT-LLM and describes a multi-GPU implementation in its Wide Expert Parallelism technical post. The idea is to reduce the number of expert weights each GPU must hold by spreading the model across a larger system.
That may relieve per-GPU memory pressure, but it replaces one constraint with others: inter-GPU communication, topology, system scale, and scheduling. It should not be confused with the SSD or host-memory offloading explored in the research papers above.
The same distinction applies to NVIDIA Dynamo. NVIDIA’s March 16, 2026 announcement describes Dynamo as production-grade inference software, and its public materials discuss KV and state-data management. That is relevant context for inference systems, but it is not evidence that Dynamo offloads MoE expert weights or that it implements SeqMoE or Edge0.
Related Companies: Different Exposure, Different Evidence
NVIDIA (Nasdaq: NVDA) is exposed through accelerators, networking, and inference software. Its public TensorRT-LLM and Wide Expert Parallelism materials show that large-scale MoE serving is a platform consideration. A broader MoE deployment cycle could make memory capacity, interconnect design, runtime support, and system utilization more important. It does not follow that NVIDIA supplies the cited papers, that a named customer uses their methods, or that efficiency changes GPU demand in one automatic direction.
AMD (Nasdaq: AMD) has relevant accelerator-memory attributes through its Instinct portfolio. AMD’s MI350 Series page, checked September 28, 2026, lists up to 288 GB of HBM3E and up to 8 TB/s of peak theoretical memory bandwidth per GPU. Those are hardware specifications, not an MoE-offload benchmark or proof that the system runs the research methods discussed here.
Micron (Nasdaq: MU) participates across memory and storage categories. Its June 24, 2026 fiscal Q3 release discusses HBM4 shipments, PCIe Gen6 SSD production, and a broader DRAM, NAND, and storage portfolio. Those products can occupy different infrastructure tiers, but the release does not show that Micron supplies the research systems or quantify a revenue contribution from MoE offloading.
The business path remains longer than a technical diagram: workload need, system design, qualification, customer deployment, shipment, product mix, and margin. Each step needs its own evidence.
Investment Watchpoints
For readers trying to follow the infrastructure implications, the more useful questions are often operational rather than promotional.
- Resident working set: How much expert weight must remain close to compute at the target model size, precision, and concurrency?
- On-time prefetch rate: What share of required experts reaches the accelerator before the relevant layer needs it, rather than merely appearing somewhere in the cache hierarchy?
- Measured transfer path: What do real bandwidth, latency, queueing, and topology measurements show—not just peak component specifications?
- Quality trade-off: Does the approach alter routing, quantization, or training objectives, and what happens to task quality as well as tokens per second?
- Commercial evidence: Are there named qualifications, deployments, shipments, product-mix disclosures, or filing commentary beyond a paper or technical demonstration?
There is a genuine two-sided demand question here. Better serving efficiency can reduce infrastructure required for one request, yet it can also enable lower-cost services, higher concurrency, larger models, or more total requests. The net effect on accelerator, HBM, DRAM, or SSD demand cannot be inferred from an efficiency claim alone.
This article is for industry context, not a recommendation to buy or sell any security.
Appendix: Where This Matters in Practice
MoE memory design becomes easier to see in an AI service that has a large model, many simultaneous users, and uneven expert reuse. One user’s request may repeatedly favor a small set of experts while another triggers a different group. The serving system has to decide which weights are worth keeping nearby and which can wait in another tier.
That is why “sparse” does not mean “memory no longer matters.” It means the system has a new choice: reduce active compute by selecting fewer experts, then manage the full set of expert weights so the selected few arrive on time.
Sources and Date Checked
- SeqMoE: Toward Full-Load Performance via Predictive and Graph-Compatible MoE Offloading, arXiv preprint submitted September 11, 2026; checked September 28, 2026.
- The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction, arXiv preprint first posted September 16, 2026 and revised September 17, 2026; checked September 28, 2026.
- MoRE: Mixture of Reused Experts, arXiv preprint revised September 17, 2026; checked September 28, 2026.
- NVIDIA TensorRT-LLM, product documentation, checked September 28, 2026.
- NVIDIA: Scaling Large MoE Models with Wide Expert Parallelism, October 20, 2025; checked September 28, 2026.
- NVIDIA Dynamo 1.0 announcement, March 16, 2026; checked September 28, 2026.
- AMD Instinct MI350 Series, product page, checked September 28, 2026.
- Micron fiscal Q3 2026 results, June 24, 2026; checked September 28, 2026.
댓글 없음:
댓글 쓰기
안녕하세요 :)