Prefill is compute-bound, decode is memory-bandwidth-bound. Why splitting inference into two purpose-fit GPU pools beats one uniform fleet.
Category: Technical Articles
After weights load, the HBM left over is your KV-cache budget. Why the H200’s 141GB buys more context and concurrency than a 94GB H100.
Map MongoDB, Redis, Kafka, ClickHouse, and Kubernetes workers to OpenMetal SKUs by the resource each role saturates, then size the failure domain.
Yes, Llama 3.3 70B runs on a single OpenMetal H200 at FP8 with full 128K context. See the VRAM fit math, KV-cache budget, and vLLM setup.
An ordered Day-2 playbook for a single-tenant H200: full root and IPMI, owning the CUDA stack, boot-data isolation, and a node-bounded blast radius.
OpenMetal XL v5 keeps 64 cores but changes node, memory, I/O, power, AMX, and TDX readiness. Where v5 wins, and the one spec that regresses.
The v5 generation can be told as a cores-and-clocks story, but a significant change is bandwidth: the private fabric doubled to 40 Gbps, memory moved to DDR5-6400, and the lane budget grew to 88 PCIe 5.0 lanes.
All-NVMe OSDs, an isolated boot pool, a clean lane budget, and identical nodes: how OpenMetal’s v5 hardware makes Ceph behave predictably instead of needing tuning.
How Nomad uses CSI to consume OpenStack Cinder + Ceph block storage. Build scheduler-agnostic persistent storage on dedicated OpenMetal infrastructure.

































