Packt · almost finished
GPU Engineering: AI Inference and System Design
Modern AI workloads don't fail at the prompt, they fail at the systems boundary. This book covers the anatomy of high-throughput inference serving. It starts where traffic hits: concurrency at the gateway versus the engine, multi-tenant isolation, and the divide between the control plane and the data plane. From there it works down to bare metal: interconnect fabric, silicon, compilers, and kernels.
Written for systems architects, infrastructure leads, and platform engineers who have to design the fleet rather than consume an endpoint.
Almost finished. Launch is end of 2026 or early 2027. The Maven course is a compressed version.