The answer is set by three things: the model’s weight footprint, the KV cache your context and concurrency produce, and how your serving framework spreads both across GPUs. An eight-GPU HGX B300 provides roughly 2.1 TB of aggregate HBM joined by NVLink, for the models whose working set does not fit on fewer cards.
Sizing an inference server is usually framed as a yes or no question: does the model fit. For frontier reasoning workloads that framing hides the decision that actually matters. What has to fit is the working set, the model weights plus a KV cache that grows with context length and concurrency, and what matters almost as much is where that working set physically lives across the GPUs and how often the serving path has to reach from one GPU to another. An eight-GPU HGX B300 is one answer to that problem, and it is the right answer for a specific class of model and load. Whether your workload is in that class, and what the node does and does not give you when it is, is something you can work out in advance from the model and your targets.
The HGX B300 places eight B300 SXM GPUs on one baseboard, joined by fifth-generation NVLink and NVSwitch, with roughly 2.1 TB of HBM3e in aggregate across the eight. That aggregate is the headline number, and its precise meaning matters: eight separate GPU memories that software can treat as one capacity budget because any GPU can read any other’s memory over NVLink. It is not a single uniform pool that every GPU reaches at local speed. What lives on each GPU, and how often a request crosses between them, is decided by the parallelism strategy and the serving framework, not by the hardware alone.
Key Takeaways
- You can decide up front whether you need an eight-GPU node. Size the weight footprint, the KV cache at your context and concurrency, and your prefill and decode targets, then compare against a smaller server before buying the largest configuration.
- The node gives ~2.1 TB of aggregate HBM across eight GPUs, not one uniform local pool. The eight HBM stacks are joined by NVLink and NVSwitch, and the serving framework and parallelism strategy decide what weights and cache live on each GPU.
- There are three memory tiers, not one. Local HBM is about 7.7 TB/s per GPU, NVLink is 1.8 TB/s GPU-to-GPU through NVSwitch, and the network is slower again; NVLink keeps a multi-GPU model off the network, it does not make every read local.
- Prefill and decode have different bottlenecks, and KV-cache size is architecture-dependent. Decode is usually memory-bandwidth-bound and prefill compute-bound; a Multi-head Latent Attention model like DeepSeek-R1 caches a small compressed latent, while a grouped-query model of similar size caches far more. Compute it from the model’s config.
- On single-tenant, fixed-cost OpenMetal hardware the node runs committed, with no per-GPU-hour meter. It is a build-to-requirement proof of concept scoped by the OpenMetal Engineering team, not a self-serve SKU, and there is no published availability date.
OpenMetal builds single-tenant GPU hardware to a requirement rather than from a fixed menu, and the catalog is actively expanding. The eight-GPU HGX B300 is one point on that spectrum; a workload that points to a different accelerator, a different GPU density, or a larger multi-node build is a design conversation with the OpenMetal Engineering team rather than a configuration to pick off a page.
Aggregate Memory, Placed by Software
Each B300 GPU has its own HBM, which it reads at roughly 7.7 TB/s. NVLink and NVSwitch let any GPU read any other GPU’s HBM at up to 1.8 TB/s, which is the figure NVIDIA publishes for fifth-generation NVLink GPU-to-GPU bandwidth. Those two numbers are both real and they are not the same number: a read from a GPU’s own HBM is several times faster than a read pulled across NVLink from a neighbor, and both are far faster than anything that crosses the network to another node.
This is why the ~2.1 TB is best described as aggregate capacity rather than a single pool. When a model is too large for one GPU, the serving framework splits it, by tensor parallelism within a layer, by expert parallelism across a Mixture-of-Experts layer’s experts, or by pipeline parallelism across layer groups, and it places the KV cache according to that split. Whether a given token’s computation stays on one GPU or has to gather data from several is a property of that placement and of the model’s communication pattern, not something the hardware decides for you. NVLink’s contribution is bounded and specific: for the cross-GPU traffic a parallel model does generate, it keeps that traffic on a 1.8 TB/s intra-node fabric instead of the network. It does not make cross-GPU access free, and a well-placed model minimizes how much of it happens in the first place.
Prefill and Decode Are Different Problems
Reasoning inference is not uniformly bound by one resource. It has two phases with different characteristics, and conflating them is the most common sizing error. Prefill processes the entire prompt at once, in parallel across its tokens, and is typically compute-bound: it is where the FLOPS matter, and on a long prompt it dominates time-to-first-token. Decode then generates the answer one token at a time, reading the weights and the growing KV cache on every step while doing comparatively little arithmetic per token, so it is typically bound by memory bandwidth and by how much of the working set stays resident.
Reasoning models are decode-heavy by construction, because test-time scaling spends quality on generating long chains of intermediate tokens, so steady-state throughput on a reasoning service usually tracks memory bandwidth and residency more than peak FLOPS. That is a tendency, not a law. The actual bottleneck shifts with batch size, context length, cache precision, the parallelism strategy, and the attention implementation in your serving framework, and on long prompts the compute-bound prefill phase can dominate the latency a user feels first. Size against both phases and against your own time-to-first-token and per-token targets, measured on your prompts, rather than assuming one number governs.
Sizing the Working Set: a Worked Example
The working set is weights plus KV cache, and both are computable from a model’s published configuration. DeepSeek-R1 is a useful worked example because it is a documented open reasoning model and because its architecture makes the point that KV-cache cost is not universal.
DeepSeek-R1 (verified from config.json, 2026-10-08): 61 layers, Multi-head Latent
Attention, Mixture-of-Experts (256 routed experts, 8 active per token),
~671B total parameters, FP8 weights, up to 163,840-token context.
Weights (memory follows TOTAL parameters, not the 37B active per token):
~671B params x 1 byte (FP8) ~= ~671 GB (approximate)
-> far larger than one 288 GB GPU; spread across the node by expert/tensor parallelism.
KV cache (MLA caches a compressed latent, NOT num_kv_heads x head_dim):
per token = layers x (kv_lora_rank + qk_rope_head_dim) x bytes_per_element
= 61 x (512 + 64) x 2 (FP16) ~= ~68.6 KiB per token
one request at 163,840 tokens ~= ~10.7 GiB (FP16 cache)
64 requests averaging 32K tokens ~= ~137 GiB (aggregate)The arithmetic shows two things, and both generalize into a sizing method. First, for this model the weights are the dominant term and the reason the workload needs node-scale memory: at roughly 671 GB they cannot sit on one GPU, so they are distributed across the node and reached over NVLink. Second, Multi-head Latent Attention keeps the KV cache small, roughly 10.7 GiB for a single maximum-context request, because MLA caches a compressed latent of 512 plus 64 values per layer instead of the full per-head key and value tensors. A grouped-query or multi-head model of comparable size caches many times more per token, and that is the regime where the KV cache can rival or exceed the weights. So the common claim that the KV cache dominates is true for some architectures and false for others; compute it from the config rather than assuming it.
Precision is two separate decisions, often conflated. NVFP4 is a weight and compute format: it shrinks the weight footprint and feeds Blackwell Ultra’s FP4 tensor cores. The KV cache is a different allocation with its own precision, chosen in the serving framework, commonly FP16, FP8, or INT8, and native FP4 compute does not automatically store the KV cache in FP4. If you intend to quantize the cache, confirm the format your serving stack actually supports for it rather than inferring it from the GPU’s compute capability.
Do You Need Eight GPUs? A Sizing Method
Work the decision in this order, against your own model and service-level objectives, before assuming the largest node.
- Weight footprint. Total parameters times bytes per parameter, using total parameters for a Mixture-of-Experts model. If this exceeds one or two GPUs, you need multi-GPU memory and a parallelism strategy.
- KV-cache capacity. Compute per-token cache from the model’s attention config, multiply by your target context length and concurrency, and add it to the weights. This is the number that grows with load.
- Prefill and decode targets. Separate time-to-first-token (prefill, compute) from per-token latency (decode, bandwidth), and check each against your service-level objective.
- Parallelism and communication. Decide how the framework will split the model and how much cross-GPU traffic the split implies. This determines how much the NVLink fabric, rather than local HBM, is in your path.
- Serving-framework support. Confirm the framework supports the model’s attention implementation, your parallelism mode, and your intended KV-cache precision.
- Compare against a smaller configuration. If the working set fits one or two GPUs at your context and concurrency, a smaller server is the better-matched and lower-cost substrate. The eight-GPU node is for the working sets that do not.
This is a method, not a benchmark. The figures it produces are capacity and bandwidth estimates from published specifications; actual throughput for your model and load is something to measure in a proof of concept, not to assert from a spec sheet.
Where the Node Earns Its Place
The eight-GPU node is the right substrate when the working set, weights plus the KV cache you intend to serve, exceeds what one or two GPUs hold, so the model must be distributed and kept on one tightly-coupled NVLink domain rather than split across the network. If it fits fewer GPUs at your targets, use fewer and spend less. The boundary does not disappear at the node edge either: a workload larger than one node spans multiple nodes over the network, where the inter-node interconnect, not NVLink, sets the ceiling, and the right multi-node topology is scoped to the workload with the OpenMetal Engineering team rather than assumed.
The eight-GPU HGX B300 is a configuration OpenMetal builds to a requirement as an engineered proof of concept, not a self-serve SKU you order today, and there is no published availability date. The proof of concept measures the working set and the throughput your reasoning service actually produces, at your context length and concurrency, against the node, rather than sizing it from a spec sheet.
What This Looks Like on OpenMetal
For a team weighing this substrate, here is how the node is actually delivered.
| Dimension | What you get |
|---|---|
| Product | 8-GPU NVIDIA HGX B300 node, built to requirement (HGX B300 hardware page) |
| Delivery model | Single-tenant, dedicated bare metal, delivered as an engineered proof of concept rather than a self-serve SKU; no published availability date |
| Access | Full root plus out-of-band management (IPMI/BMC); you bring your own OS image |
| Base configuration | Eight B300 SXM GPUs joined by NVLink and NVSwitch, roughly 2.1 TB of aggregate HBM3e (not one uniform pool), on a dual Intel Xeon 6 host |
| What you add, scoped to the workload | Host memory, local NVMe capacity, and multi-node cluster size are sized to your model and load; component pricing and lead time are quoted at the time of order |
| Local storage | Front E1.S NVMe for datasets, checkpoints, and model repositories, with OpenMetal storage nodes on the private network for larger sets |
| Network | Eight ConnectX-8 SuperNICs as the GPU scale-out fabric, plus a single-tenant OpenMetal private network with non-metered east-west traffic |
| Software boundary | You run the stack: the CUDA driver branch, NCCL, the container runtime, the serving framework, and your parallelism and KV-cache settings are yours |
| Validation path | The OpenMetal Engineering team scopes and validates the node or cluster against your reasoning workload as a proof of concept |
| Pricing path | Fixed monthly, included egress, no per-GPU-hour meter; contact OpenMetal for a quote |
Why Fixed-Cost Delivery Lets You Commit the Node
An always-on reasoning endpoint is, by design, the workload that keeps a node busy around the clock, and on a per-GPU-hour meter that steady high utilization is exactly the pattern the meter charges most for. OpenMetal delivers the node single-tenant at a fixed monthly cost with included egress and no per-GPU-hour meter, so the aggregate capacity is yours to keep full, and the hardware sits behind your own isolation boundary rather than sharing silicon with another tenant. The pricing model and the serving strategy agree: keep as much of the working set resident as the node holds, and run it committed.
What the Decisions Add Up To
Whether reasoning inference needs an eight-GPU B300 node is a sizing question, not a preference. Compute the weight footprint from total parameters, compute the KV cache from the model’s own attention config at your context and concurrency, separate the prefill and decode targets, and account for how your parallelism strategy uses the NVLink fabric. If the working set fits one or two GPUs, serve it on fewer and spend less. If it does not, the eight-GPU node gives you ~2.1 TB of aggregate HBM on one NVLink domain, delivered single-tenant at a fixed cost so you can run it full, and validated against your own workload in a proof of concept rather than promised from a datasheet.
Sources
- NVIDIA Blackwell Ultra datasheet (HGX B300 per-GPU and per-node memory, per-GPU HBM bandwidth, 1.8 TB/s fifth-generation NVLink, NVSwitch, FP4): https://resources.nvidia.com/en-us-blackwell-architecture/blackwell-ultra-datasheet (verified in session 2026-10-08)
- NVIDIA HGX platform (eight-GPU B300 baseboard, NVLink/NVSwitch): https://www.nvidia.com/en-us/data-center/hgx/ (verified in session 2026-10-08)
- DeepSeek-R1 model configuration, config.json on Hugging Face (layer count, Multi-head Latent Attention dimensions, Mixture-of-Experts counts, FP8 weights, context length): https://huggingface.co/deepseek-ai/DeepSeek-R1 (config verified in session 2026-10-08)
Request a Proof of Concept
The way to size this is against your own reasoning model, your context length, and your concurrency target. We can scope an HGX B300 node so you can measure the working set and the throughput your service actually produces under load. Tell us the model and the service-level objective and the OpenMetal Engineering team will help you plan the deployment.

































