The Inference Mechanic.
Tuning LLMs for speed, cost, and scale.
Training builds the engine. Inference is the car people actually drive. This is the manual for the people who keep it running: what the hardware allows, how a model is built and tuned, how one engine becomes a fleet, how to lock it down, and how far to the edge it can go. Every number in it is measured on real GPUs, not quoted.
By Purushottam Chaudhary — founder & CEO of QuickDial AI, the inference engineer behind a voice stack that runs under a cent a minute on commodity compute.
Under the hood, one system at a time.
The book follows the order a mechanic would: start at the metal, learn how the engine is made, tune it, then take it out on the road. Every chapter ends with a Dyno Test — a benchmark on real GPUs with the exact command and the measured numbers. Boxes throughout: Under the Hood explains why, Torque Spec gives the number to hit, Roadside Repair tells a production failure and its fix.
- 01
Hardware and Its Limits
From the machine you already own to a GPU server: the arithmetic of a token, the vendors and their bets, racks, power, and how to read the gauges. Read it free →
- 02
Model Architectures and What They Cost
Inside one block, the bills every architecture pays, shrinking the KV cache, mixtures of experts, System One decision models, and models that make pictures.
- 03
How Models Are Made
Pre-training, scaling laws bent by inference, fine-tuning, preference tuning, RL as an inference workload, adapters, merging and distillation.
- 04
The Tune-Up
Every optimization, one by one: batching and scheduling, paged memory, quantization, pruning, distillation, speculative decoding, kernels, parallelism, caching above the model.
- 05
Beyond Text: Modalities
Everything becomes tokens. Images read and made, audio listening and speaking, video, 3D and the physical world — and what each one costs.
- 06
From Engine to Fleet
Serving engines, packaging and shipping, Kubernetes for inference, and running the fleet: turning one fast box into capacity you can price.
- 07
Locks and Guardrails
Threats drawn plainly, guardrails inside and around the model, System One guards, prompt injection and agents, enterprise security for the serving layer.
- 08
The Edge
Model formats, small language models, edge hardware, running models on devices, edge and cloud together, and the future of generative AI at the edge.
- 09
On the Dyno
Proving the numbers on an RTX 4090: the rig, the models on the bench, the test catalogue and the results, test by test. If it isn't on the dyno, it isn't in the book.
- Engineers shipping LLM features who pay the inference bill and want it to go down.
- Platform and infrastructure teams deciding what runs on which hardware, and how many boxes.
- Founders and technical leaders choosing between rented APIs, their own GPUs and the edge.
The first hundred pages, open on the bench.
The cover, the full table of contents and all of Chapter 1 — Hardware and Its Limits — read right here in your browser, no download and no sign-up. When you reach the end, leave your name and email and you are first in line for the early release.
- Chapter 1 in full: the machine you own, bolting on a GPU, the arithmetic of a token, vendors, racks, power, and the gauges
- Six Dyno Tests with real measurements
- Works on phone, tablet and desktop
Cover · Contents · Chapter 1
The first 100 readers get it free.
Leave your name and email and you will hear from us once, when the book ships. The first hundred people on the list receive the EPUB and PDF at no charge; everyone after that gets the launch price.
- One email when it ships — nothing else, no newsletter
- EPUB and PDF, read on anything
- Unsubscribe with one click, any time
