LIBRISTO
LIBROAMANTO
povinné
Staňte se součástí komunity milovníků knih z celého světa a získejte hromadu výhod. Založit účet zdarma
0
Doprava zdarma se Zásilkovnou nad 1 499 Kč
Kurýr DPD 69 Balíkovna 69 PPL kurýr 74 PPL box 39 Zásilkovna 39 Výdejní místo DPD 49 PPL shop 49 Balíkovna 49

Doprava zdarma při nákupu nad 1 499 Kč přes Zásilkovnu nebo PPL Box.

Inference at Full Throttle

LLM serving performance with vLLM, quantization, KV cache tuning and speculative decoding

Jazyk AngličtinaAngličtina
Kniha Brožovaná
Kniha Inference at Full Throttle ChatVariety Team
Libristo kód: 53520833
Nakladatelství Independently published, srpen 2026
Master LLM Inference and Scale Your AI InfrastructureIn 2026, inference spend surpassed training spe... Celý popis
? points 24 b Připravujeme Připravujeme Nové Nové
235
Očekávané naskladnění Naskladnění 17. 08. 2026

Až 30 dní na vrácení zboží

Master LLM Inference and Scale Your AI Infrastructure

In 2026, inference spend surpassed training spend across the tech industry. The engineers who can maximize tokens per second on H100, H200, and B200 GPU fleets are the most valuable specialists in AI. Inference at Full Throttle turns complex GPU performance engineering into a reproducible, highly practical discipline.

Written by the ChatVariety Team-an elite collective of ML infrastructure engineers and vLLM contributors-this book provides the exact mathematical formulas and production configurations needed to run large language models at extreme scale without breaking the bank.

What You Will Master:
  • GPU Memory Mathematics: Derive the exact KV cache formula from first principles to budget HBM memory perfectly.
  • Advanced Quantization: Deploy FP8, AWQ INT4, GPTQ, and MXFP4 based on real-world throughput and quality trade-offs.
  • Serving Stack Optimization: Fine-tune vLLM, SGLang, TensorRT-LLM, and TGI for enterprise workloads.
  • Ultra-Fast Decoding: Implement speculative decoding, EAGLE-class self-speculation, and KV cache prefix caching.
  • Distributed Scale: Combine tensor, pipeline, and expert parallelism (MoE) across multi-node GPU clusters.
  • Production Benchmarking: Avoid common traps by measuring TTFT, TPOT, and tail latency under realistic workloads.

Stop wasting millions on sub-optimal cloud GPU allocations. Learn how to design, benchmark, and operate multi-tenant, high-throughput, and ultra-low-latency LLM serving architectures today.

Herečka & Polyglotka
EWA KASP pro
Přehrát video
Ewa Kasp
Libristo má největší výběr cizojazyčné literatury. Proto své knihy kupuji tady.

Informace o knize

Plný název Inference at Full Throttle
Jazyk Angličtina
Vazba Kniha - Brožovaná
Datum vydání 2026
Počet stran 82
EAN 9798192412626
Libristo kód 53520833
Nakladatelství Independently published
Váha 123
Rozměry 152 x 229 x 4
Darujte tuto knihu ještě dnes
Je to snadné
1 Přidejte knihu do košíku a zvolte doručit jako dárek 2 Obratem vám zašleme poukaz 3 Kniha dorazí na adresu obdarovaného

Přihlášení

Přihlaste se ke svému účtu. Ještě nemáte Libristo účet? Vytvořte si ho nyní!

 
povinné
povinné

Nemáte účet? Získejte výhody Libristo účtu!

Díky Libristo účtu budete mít vše pod kontrolou.

Vytvořit Libristo účet
Knižní rádce Libroamiko
Ahoj, jsem Libroamiko, můžu pomoct?