The NVIDIA HGX B300 is NVIDIA’s Blackwell Ultra platform in its densest single-node form: eight B300 SXM GPUs wired together on one baseboard by fifth-generation NVLink and NVSwitch into a single, low-latency memory domain. Unlike a PCIe GPU server, where each card is a discrete accelerator with its own memory, the eight GPUs on an HGX B300 baseboard provide roughly 2.1 TB of aggregate fast HBM3e (eight separate GPU memories reachable over NVLink), so a frontier-scale model and a very large KV cache can stay resident across the whole node and every GPU can reach any other GPU’s memory at 1.8 TB/s. That is the architecture that makes the HGX B300 compelling for the hardest reasoning-inference and large-model training jobs, where the binding constraint is how much model plus working state you can keep in memory and how fast the GPUs can talk to each other.

OpenMetal builds this as a single-tenant, dedicated bare metal node. Like every OpenMetal deployment it carries fixed monthly pricing with included egress and no per-GPU-hour meter, so a node pinned at full utilization for a multi-week job costs the same as one sitting idle. Because the HGX B300 is a dense, high-power platform and a possibility built to a requirement rather than a stock SKU, the right starting point is a scoping conversation with the OpenMetal Engineering team.

Key Takeaways

  • Aggregate HBM across eight GPUs over NVLink, not one uniform pool. Fifth-generation NVLink and NVSwitch join the eight B300 GPUs into ~2.1 TB of aggregate fast memory at 1.8 TB/s GPU-to-GPU GPU-to-GPU. A reasoning model plus a large KV cache stays resident across the whole node, which is what keeps high-concurrency, long-context serving at low latency instead of paging against a single card’s ceiling.
  • Blackwell Ultra memory and native FP4 at node scale. Up to 144 PFLOPS of FP4 (sparse) per node and 62 TB/s of aggregate memory bandwidth make the HGX B300 a fit for large-model training, long-output reasoning, and real-time generation, not a default for every workload.
  • Scale-out is built in. Eight NVIDIA ConnectX-8 SuperNICs at up to 800 Gb/s plus two BlueField-3 DPUs per node mean the single-node design extends to multi-node clusters over a high-bandwidth fabric when a workload outgrows one node.
  • Fixed monthly pricing avoids the idle silicon tax. On a sustained training run or an always-on reasoning endpoint, the most expensive thing on a metered GPU-hour cloud is exactly the thing OpenMetal charges a flat rate for: a GPU kept busy. Here the marginal cost of running the node harder is zero.
  • Single-tenant bare metal, built to requirement. The whole node is yours, with no shared tenancy and full control. The configuration (GPU density, host memory, storage, cluster size) is scoped to your workload with the OpenMetal Engineering team rather than picked from a fixed menu.

NVIDIA HGX B300 node architecture: eight B300 SXM GPUs linked by NVSwitch into ~2.1 TB of aggregate HBM3e over NVLink, with a dual Intel Xeon 6 host and ConnectX-8 scale-out networking.

What it shows: the eight B300 GPUs share ~2.1 TB of aggregate HBM3e over NVLink over NVSwitch (270 GB usable per GPU), with the serving software placing weights and KV cache across the eight GPUs.

Server Configuration at a Glance

ComponentSpecification
GPUs8x NVIDIA B300 SXM (Blackwell Ultra) on an HGX baseboard
GPU memory288 GB HBM3e per GPU (physical). NVIDIA’s HGX B300 datasheet lists 270 GB usable per GPU and ~2.1 TB of fast memory per 8-GPU node
GPU memory bandwidth7.7 TB/s per GPU; ~62 TB/s aggregate per node
GPU-to-GPU interconnectFifth-generation NVIDIA NVLink at 1.8 TB/s via NVSwitch (aggregate memory reachable across all 8 GPUs, placed by software); ~14.4 TB/s NVLink-switch bandwidth
Tensor / precision supportFP4 (NVFP4, Blackwell-native), FP8/FP6, INT8, BF16/FP16, TF32, FP64
Node AI throughputUp to ~144 PFLOPS FP4 (sparse) / ~108 PFLOPS FP4 (dense) per node
Per-GPU max powerConfigurable up to 1,100 W
ProcessorDual Intel Xeon 6 (Granite Rapids) P-cores (for example, 64-core Xeon 6768P; 128 cores / 256 threads per node)
System memoryBuilt to requirement: up to 4 TB DDR5-6400 (1 DIMM per channel) or up to 8 TB DDR5-5200 (2 DIMMs per channel). Component pricing and lead time are quoted at the time of order
Boot storage2x M.2 NVMe
Data storageFront hot-swap E1.S NVMe bays, configured to the workload
Scale-out networking8x NVIDIA ConnectX-8 SuperNIC up to 800 Gb/s + 2x NVIDIA BlueField-3 DPU
Form factor / cooling8U, air-cooled
Power6x 6.6 kW (3+3 redundant) Titanium-level power supplies
TenancySingle-tenant, dedicated bare metal
AvailabilityBuild-to-requirement / proof of concept with the OpenMetal Engineering team. Not a self-serve order; no published availability date
PricingFixed monthly, included egress, no per-GPU-hour meter. 

Build to Your Requirement

The OpenMetal GPU catalog is actively expanding and is built to a requirement rather than picked from a fixed menu. The eight-GPU HGX B300 is one point on that spectrum. If your workload points to a different accelerator, a different GPU density per node, a different storage or memory build, or a larger multi-node cluster than what is shown here, that is a design conversation worth having. We evaluate the build against the workload rather than the other way around.

Scope a Build with Engineering

The Platform: NVIDIA HGX B300 (Blackwell Ultra)

The HGX B300 places eight B300 SXM GPUs on a single baseboard. Each GPU carries 288 GB of HBM3e physically (NVIDIA specifies 270 GB usable per GPU and roughly 2.1 TB of fast memory across the node), with 7.7 TB/s of bandwidth per GPU and about 62 TB/s aggregated across the node. The eight GPUs are joined by fifth-generation NVLink and NVSwitch, giving every GPU direct access to every other GPU’s memory at 1.8 TB/s. This is the defining difference from a PCIe GPU server: the node operates as one tightly-coupled domain of aggregate memory over NVLink, with the serving software placing weights and cache across the eight GPUs.

Blackwell Ultra adds native FP4 (NVFP4) alongside FP8/FP6, INT8, BF16/FP16, TF32, and FP64, and the node reaches up to roughly 144 PFLOPS of FP4 (sparse). For low-precision reasoning inference and mixed-precision training, that is a generational step over Hopper-class hardware. For precise, current per-GPU and per-node figures, see the NVIDIA sources at the end of this page, and always confirm whether a published memory or throughput number is per GPU, per node, or per rack before comparing it.

Host: Dual Intel Xeon 6 (Granite Rapids)

The GPU baseboard is paired with two Intel Xeon 6 (Granite Rapids) processors with P-cores. A typical build uses the 64-core Xeon 6768P for 128 cores and 256 threads per node, with Intel AMX and AVX-512 for CPU-side tokenization, data loading, preprocessing, and embedding pipelines that feed the GPUs. The exact processor is part of the scoping conversation.

Memory

Host memory is configured to the workload. On this platform that means up to 4 TB of DDR5-6400 at one DIMM per channel, or up to 8 TB of DDR5-5200 at two DIMMs per channel. The maximum-capacity configuration and the maximum-bandwidth configuration are not the same build: the 8 TB fill runs at the lower DDR5-5200 grade, while full DDR5-6400 bandwidth is a one-DIMM-per-channel configuration. We size host memory to stage training datasets and to keep data-loading pipelines ahead of eight GPUs. RAM capacity and lead time move with the memory market and are quoted at the time of order.

Storage

OpenMetal separates boot and data storage. The node boots from two M.2 NVMe drives, keeping the OS isolated from data, and data sits on front hot-swap E1.S NVMe bays sized to the workload. Fast local NVMe matters at this scale: loading and swapping frontier-size model weights is read-bound, and keeping that off the critical path reduces cold-start and model-switch latency. Drive capacity and endurance are configured to the requirement, and like RAM, component pricing and lead time are quoted at the time of order.

Networking and Scale-Out

Each node integrates eight NVIDIA ConnectX-8 SuperNICs at up to 800 Gb/s and two NVIDIA BlueField-3 DPUs. That is the GPU scale-out fabric: it is what lets a single-node HGX B300 extend into a multi-node cluster when a workload grows past one node. Multi-node cluster topology (node count, fabric, and storage design) is scoped with the OpenMetal Engineering team to the workload.
Separately from that compute fabric, every OpenMetal deployment sits on a single-tenant private network. East-west traffic between your nodes, whether GPU nodes exchanging data or pulling datasets and checkpoints from OpenMetal storage nodes, runs on that private network and is not metered, so distributed training and multi-node inference do not accrue internal-transfer charges. Public connectivity is billed on a 95th-percentile model with a generous included allotment rather than per gigabyte, which matters for always-on inference endpoints where responses stream continuously. Connectivity is provisioned to the build, and every deployment carries OpenMetal’s network SLA and included DDoS protection.

On the Roadmap. Higher-bandwidth OpenMetal network connectivity is on the roadmap, with configurations supporting up to 100 Gbps of throughput under evaluation. The throughput available for your deployment would be validated as part of the proof of concept.

Security and Tenancy

The HGX B300 runs as a single-tenant, dedicated bare metal node: physical isolation, not a shared hypervisor, which is the foundational property for protecting proprietary models and inference data. Confidential-computing scoping (Intel TDX with NVIDIA Confidential Computing) is a separate engineering conversation for this platform: NVIDIA’s validated confidential-GPU mode is one GPU per confidential VM, so a confidential configuration on a multi-GPU node is not a validated turnkey mode and would be scoped and validated with the OpenMetal Engineering team for a specific workload.

Because the node is yours alone, your model weights, prompts, and inference data never share a GPU, a CPU, or a memory bus with another tenant. For teams running proprietary or fine-tuned models, that physical boundary is the difference between renting time on shared accelerators and operating your own: there is no co-tenant to isolate from, no hypervisor in the data path, and the hardware sits inside your own network boundary.

Recommended Workloads

Long-Context, High-Concurrency Reasoning Inference

The ~2.1 TB of aggregate fast memory and native FP4 let a frontier-scale reasoning model and a very large KV cache stay resident across the whole node, so test-time-scaling and high-concurrency serving hold their latency target instead of paging against a single card’s memory ceiling. Serve with NVIDIA NIM, vLLM, or TensorRT-LLM.

Large-Model and Large-MoE Training

Fifth-generation NVLink and NVSwitch make the eight GPUs one tight training domain with 1.8 TB/s GPU-to-GPU bandwidth, which is what large-model and large-Mixture-of-Experts training need when gradients and expert activations move constantly between GPUs. Frameworks: PyTorch (FSDP), DeepSpeed, Megatron, NVIDIA NeMo.

Real-Time Generation and Multimodal

Blackwell Ultra throughput and FP4 suit real-time generative media and multimodal pipelines where both memory footprint and token rate are demanding.

Runs with Your OpenMetal Footprint

An HGX B300 node or cluster does not have to stand alone. It can attach to your existing OpenMetal Hosted Private Cloud or bare metal on the same single-tenant private network, so your GPU nodes, your general compute, and your storage all sit on one non-metered east-west fabric. For data-heavy training and serving, OpenMetal storage nodes (Ceph-backed clusters or dedicated NVMe) hold datasets, checkpoints, and model repositories next to the GPUs, so loading weights and streaming training data does not cross a metered link or leave your environment. The storage and private-cloud design are scoped to the workload alongside the GPU build.

Bare-Metal Control: Your Stack, Your Versions

The node is bare metal with full root access and out-of-band management (IPMI/BMC), and you bring your own OS image. Frontier GPU stacks are version-sensitive: the CUDA driver branch, NCCL, the container runtime, and the inference or training framework often have to be pinned to exact, compatible versions. On a single-tenant bare metal node you control that whole stack directly, with the full GPU passed through and no hypervisor in the path, rather than inheriting a managed cloud’s image and driver choices. That control is what lets you reproduce a validated configuration, match a model vendor’s reference stack, and keep it stable across a long training run or a production endpoint.

Scope Your HGX B300 Build

Tell us about your workload and the OpenMetal Engineering team will scope an HGX B300 node or cluster as a proof of concept. Share the models you plan to run, your latency and concurrency targets or your training scale, and your timeline.

  • Single 8-GPU HGX B300 node: ~2.1 TB of aggregate NVLink-connected HBM for frontier-scale reasoning and training
  • Multi-node HGX B300 cluster: scaled over the 800 Gb/s ConnectX-8 fabric, designed to the workload
  • Proof of concept: validate your models and serving stack before you commit
  • Custom configuration: host memory, storage, and GPU density built to requirement

All OpenMetal deployments include fixed monthly pricing, included egress, and a 99.96%+ network SLA.


Sources

External specifications on this page are the NVIDIA HGX B300 / Blackwell Ultra platform figures, verified against NVIDIA’s published materials (2026-10-08). Confirm live before publishing, as vendor figures can change.