Nehodí se? Vůbec nevadí! Zboží můžete vrátit až do 30 dní
S dárkovým poukazem nešlápnete vedle. Obdarovaný si za dárkový poukaz může vybrat cokoliv z naší nabídky.
Až 30 dní na vrácení zboží
Master LLM Inference and Scale Your AI Infrastructure
In 2026, inference spend surpassed training spend across the tech industry. The engineers who can maximize tokens per second on H100, H200, and B200 GPU fleets are the most valuable specialists in AI. Inference at Full Throttle turns complex GPU performance engineering into a reproducible, highly practical discipline.
Written by the ChatVariety Team-an elite collective of ML infrastructure engineers and vLLM contributors-this book provides the exact mathematical formulas and production configurations needed to run large language models at extreme scale without breaking the bank.
What You Will Master:Stop wasting millions on sub-optimal cloud GPU allocations. Learn how to design, benchmark, and operate multi-tenant, high-throughput, and ultra-low-latency LLM serving architectures today.
Ahoj! Jsem Libroamiko, tvůj knižní rádce.
Jak ti můžu pomoct?