The Inference Mechanic — Tuning LLMs for speed, cost, and scale — by Purushottam Chaudhary. Book cover. Free for the first 100 on the list
A service manual for large language models · 2026 · EPUB & PDF

The Inference Mechanic.

Tuning LLMs for speed, cost, and scale.

Training builds the engine. Inference is the car people actually drive. This is the manual for the people who keep it running: what the hardware allows, how a model is built and tuned, how one engine becomes a fleet, how to lock it down, and how far to the edge it can go. Every number in it is measured on real GPUs, not quoted.

By Purushottam Chaudhary — founder & CEO of QuickDial AI, the inference engineer behind a voice stack that runs under a cent a minute on commodity compute.

672pages · 9 chapters
370+figures and derivations
Measuredevery number, on real GPUs
100 ppfree to read right now
Inside the book

Under the hood, one system at a time.

The book follows the order a mechanic would: start at the metal, learn how the engine is made, tune it, then take it out on the road. Every chapter ends with a Dyno Test — a benchmark on real GPUs with the exact command and the measured numbers. Boxes throughout: Under the Hood explains why, Torque Spec gives the number to hit, Roadside Repair tells a production failure and its fix.

  1. 01

    Hardware and Its Limits

    From the machine you already own to a GPU server: the arithmetic of a token, the vendors and their bets, racks, power, and how to read the gauges. Read it free →

  2. 02

    Model Architectures and What They Cost

    Inside one block, the bills every architecture pays, shrinking the KV cache, mixtures of experts, System One decision models, and models that make pictures.

  3. 03

    How Models Are Made

    Pre-training, scaling laws bent by inference, fine-tuning, preference tuning, RL as an inference workload, adapters, merging and distillation.

  4. 04

    The Tune-Up

    Every optimization, one by one: batching and scheduling, paged memory, quantization, pruning, distillation, speculative decoding, kernels, parallelism, caching above the model.

  5. 05

    Beyond Text: Modalities

    Everything becomes tokens. Images read and made, audio listening and speaking, video, 3D and the physical world — and what each one costs.

  6. 06

    From Engine to Fleet

    Serving engines, packaging and shipping, Kubernetes for inference, and running the fleet: turning one fast box into capacity you can price.

  7. 07

    Locks and Guardrails

    Threats drawn plainly, guardrails inside and around the model, System One guards, prompt injection and agents, enterprise security for the serving layer.

  8. 08

    The Edge

    Model formats, small language models, edge hardware, running models on devices, edge and cloud together, and the future of generative AI at the edge.

  9. 09

    On the Dyno

    Proving the numbers on an RTX 4090: the rig, the models on the bench, the test catalogue and the results, test by test. If it isn't on the dyno, it isn't in the book.

Who it is for
  • Engineers shipping LLM features who pay the inference bill and want it to go down.
  • Platform and infrastructure teams deciding what runs on which hardware, and how many boxes.
  • Founders and technical leaders choosing between rented APIs, their own GPUs and the edge.
Read before you buy

The first hundred pages, open on the bench.

The cover, the full table of contents and all of Chapter 1 — Hardware and Its Limits — read right here in your browser, no download and no sign-up. When you reach the end, leave your name and email and you are first in line for the early release.

  • Chapter 1 in full: the machine you own, bolting on a GPU, the arithmetic of a token, vendors, racks, power, and the gauges
  • Six Dyno Tests with real measurements
  • Works on phone, tablet and desktop
Open the preview
Cover · Contents · Chapter 1
Purushottam Chaudhary
About the author

Purushottam Chaudhary

Founder & CEO, AI Inference Engineer · QuickDial AI

Puru built the speech recognition, the language model and the voices behind QuickDial's AgentBox, and engineered them to run on commodity CPUs at under a cent a minute. Before that: fifteen years shipping software inside GE Healthcare, S&P Global, Bristol Myers Squibb and State Street, a head-of-technology role at a blockchain platform company, and a generative-3D startup where he trained the models himself. Two papers in 2026 — one peer-reviewed in clinical machine learning, one on composable KV-cache segments for voice agents on commodity hardware. The book is the workshop manual he wished he had.

Waiting list · open now

The first 100 readers get it free.

Leave your name and email and you will hear from us once, when the book ships. The first hundred people on the list receive the EPUB and PDF at no charge; everyone after that gets the launch price.

  • One email when it ships — nothing else, no newsletter
  • EPUB and PDF, read on anything
  • Unsubscribe with one click, any time
Join the waiting list

Free for the first 100 subscribers. We keep your address for this book only — see the privacy note.