Running Kimi K3 on MI355X at Better Performance per Dollar Than B300
Ranked #1 on Hacker News with 75 points and 3 comments.
Running Kimi K3 at ~952 tok/s/node, AMD continues to prove its case as the winner in performance per dollar.
Over the past several months, weβve seen an explosion in the capabilities of open source models. With DeepSeek V4-Pro and GLM5.2 reaching near-Opus levels of intelligence, open source has emerged as a real, cost-efficient alternative to the closed source models weβve been married to.
But we have yet to see one like Kimi K3. Promising Fable/Sol levels of intelligence, Kimi K3 marks the start of a new era for open source.
But a smarter model means a bigger model β and these models are expanding in size just as fast as they are in capabilities. GLM5.2 has 753B parameters, DeepSeek V4-Pro 1.6T, and Kimi K3 weighs in at 2.8T (!!) parameters. Thatβs over 1.5TB of VRAM before allocating a KV cache for 1M tokens of context. Not even a B200 node (8 GPUs) can fit Kimi K3. That leaves you with limited options: serve on a node of B300s, which have 288GB of VRAM per GPU, or commit two B200 nodes (TP16) to serving Kimi.
But guess which other non-NVIDIA GPU has 288GB of VRAM? AMDβs MI355X. Can you tell we like these chips yet? At around ~2.4Γ cheaper per GPU on average versus a B300 and ~1.7Γ cheaper than a B200, the MI355X is a cost-efficient alternative to Blackwells with comparable hardware specs. The only problem with AMD is software support β slower kernels and less day-0 support on inference frameworks make serving frontier models on AMD a real engineering effort. Our claim at Wafer is that agents are improving at kernel and model optimization, closing this gap as we speak. But with AMD shipping day-0 support for Kimi K3, most of the work was already done for us.
The results are great: on a 1,024-token input / 400-token output benchmark, the MI355X reaches 952 tok/s/node and 118 tok/s single stream β over 3.8Γ the aggregate throughput per node and over 1.3Γ the single-stream decode of our TP16 B200 deployment (whose 498 tok/s is a 16-GPU, 2-node total β ~249/node). B300 nodes still win ~1.65Γ on aggregate throughput over the MI355X, but at 2.4Γ the price, the MI355X crushes the B300 on performance per dollar.