A first-party FP4 checkpoint moves GPU selection from memory capacity to native tensor-core format support, and inverts the usual verdict. Based on the Inkling-Small.
Tag: LLM Inference
Prefill is compute-bound, decode is memory-bandwidth-bound. Why splitting inference into two purpose-fit GPU pools beats one uniform fleet.
Yes, Llama 3.3 70B runs on a single OpenMetal H200 at FP8 with full 128K context. See the VRAM fit math, KV-cache budget, and vLLM setup.

































