All notes

AI

Aug 2, 2026

Kimi K3 Runs on AMD MI355X with Better Cost Efficiency Than B300

Wafer AI benchmarked Kimi K3 on AMD MI355X hardware and reports better performance per dollar than NVIDIA B300, a meaningful data point for teams evaluating inference infrastructure.

Wafer AI published results running Kimi K3 on AMD's MI355X accelerator. The headline finding: cost-efficiency beats NVIDIA's B300 for this workload.

This matters for a specific reason. B300 is NVIDIA's current high-end inference card, and most production LLM deployments default to NVIDIA silicon without evaluating alternatives. A credible third-party result showing AMD competitive on a frontier Chinese model changes that calculus, at least for teams willing to do the integration work.

Kimi K3 is Moonshot AI's latest model. It sits in the frontier tier and has drawn attention for strong reasoning performance. Running it efficiently on non-NVIDIA hardware is non-trivial — ROCm support, kernel optimization, and memory bandwidth utilization all have to align. The team's work on MI355X suggests the gap in software tooling between AMD and NVIDIA is narrowing for inference use cases.

The practical implication for infrastructure decisions is about optionality. NVIDIA supply constraints have been a real bottleneck for teams scaling inference. If MI355X delivers better performance per dollar on workloads like Kimi K3, it becomes a viable primary target, not just a fallback.

For solo founders and small teams, the cost-per-token metric is the lever that controls how much you can serve before margin pressure forces architectural changes. A hardware path that improves that metric without sacrificing throughput is worth evaluating seriously.

The announcement does not claim MI355X wins on raw throughput or latency in isolation — the framing is specifically performance per dollar. Engineers should benchmark against their own traffic profiles and batch sizes before committing to an AMD-first inference stack. But the data point exists now, and it is from a team running real hardware, not simulation.