Inference Engineering --- how models run in production - full curriculum
How models actually run - the generation loop, prefill and decode, batching, KV cache, the serving stack, and the workloads that changed the job.
- 01
The Generation Loop
Understand how an LLM produces text token by token --- from the forward pass and logits through sampling and stopping to the serving stack that puts it all in production.
- 02
The Hardware Floor
Measure and model GPU performance from first principles --- throughput vs latency, FLOPs and bytes, the memory hierarchy, arithmetic intensity, the roofline model, and why decode is memory bound. By the end you can read a profiler trace and name exactly which regime your bottleneck lands in.
- 02.01Why the GPU WaitsVideo
- 02.02FLOPs and BytesVideo
- 02.03The Memory HierarchyVideo
- 02.04Arithmetic IntensityVideo
- 02.05The RooflineVideo
- 02.06Why Decode Is Memory BoundVideo
- 02.07Where the Memory WentVideo
- 02.08PrecisionVideo
- 02.09Reading the DeviceVideo
- 02.10Profiling a Forward PassVideo
- 02.11Launch OverheadVideo
- 02.12Naming Your BottleneckVideo
- 03
The KV Cache
Design and measure the single most important inference optimisation. From why caching exists through the arithmetic, the memory wall, attention variants built to shrink it, fragmentation and paged allocation, prefix caching and radix trees, to hit rate as the production metric that drives both cost and latency.
- 04
The Batch
Turn a model server into a scheduler. One request at a time wastes the machine; this module covers continuous batching, chunked prefill and prefill/decode disaggregation, then the latency numbers that say whether any of it worked.
- 04.01One Request at a TimeVideo
- 04.02Static BatchingVideo
- 04.03Continuous BatchingVideo
- 04.04The Scheduler LoopVideo
- 04.05Chunked PrefillVideo
- 04.06Prefill/Decode InterferenceVideo
- 04.07DisaggregationVideo
- 04.08Moving the KV CacheVideo
- 04.09TTFT, TPOT, ITLVideo
- 04.10Throughput vs LatencyVideo
- 04.11Goodput and SLOsVideo
- 04.12Benchmarking Your ServerVideo
- 05
Fewer Bytes, Fewer Steps
There are only two levers on decode: move fewer bytes per step, or take fewer steps per token. Quantisation, FlashAttention and speculative decoding are all one or the other, and each has a regime where it makes things worse.
- 05.01Two Ways to Go FasterVideo
- 05.02Number FormatsVideo
- 05.03Weight-Only QuantisationVideo
- 05.04Activation QuantisationVideo
- 05.05Calibration and QualityVideo
- 05.06KV Cache QuantisationVideo
- 05.07FlashAttentionVideo
- 05.08Speculative DecodingVideo
- 05.09Draft ModelsVideo
- 05.10EAGLE and MTPVideo
- 05.11When Speculation LosesVideo
- 05.12Stacking the WinsVideo
- 06
The Fleet
One GPU becomes many, and many users become one bill. Sharding a model across devices, serving mixture-of-experts, routing requests so the cache still hits, and arriving at a defensible cost per million tokens.
- 06.01When One GPU Is Not EnoughVideo
- 06.02Tensor ParallelismVideo
- 06.03Pipeline ParallelismVideo
- 06.04The InterconnectVideo
- 06.05Mixture of ExpertsVideo
- 06.06Expert ParallelismVideo
- 06.07Routing RequestsVideo
- 06.08Multi-Tenancy and LoRAVideo
- 06.09Autoscaling and Cold StartsVideo
- 06.10Choosing the HardwareVideo
- 06.11Cost per Million TokensVideo
- 06.12Self-Host or APIVideo
- 07
Serving Agents
An agent is not a chatbot. It is hundreds of turns against one enormous shared prefix, with a tool call stalling the loop in the middle of every one. That workload breaks the batching and cache assumptions modules 3 and 4 were built on, and turns the KV cache from a buffer into a storage tier.
- 07.01The Agent WorkloadVideo
- 07.02Session AffinityVideo
- 07.03Cache-Aware RoutingVideo
- 07.04The KV Memory HierarchyVideo
- 07.05KV Offload and ReloadVideo
- 07.06Tool Calls in the LoopVideo
- 07.07The Tool Schema TaxVideo
- 07.08Structured OutputVideo
- 07.09What Grammars CostVideo
- 07.10Long-Context PrefillVideo
- 07.11Context CompactionVideo
- 07.12Agents per MegawattVideo
- 08
Thinking Costs Tokens
A model that thinks before it answers moves the compute from training to serving, makes output length unpredictable, and drags reproducibility and the reinforcement-learning loop in behind it. Every number in modules 2 and 4 has to be recomputed for a request that emits thirty thousand tokens nobody reads.
- 08.01Test-Time ComputeVideo
- 08.02What a Reasoning Model EmitsVideo
- 08.03The Decode-Heavy WorkloadVideo
- 08.04Effort and BudgetsVideo
- 08.05Unpredictable Output LengthVideo
- 08.06Parallel SamplingVideo
- 08.07Verifiers in the Serving PathVideo
- 08.08Serving Two ModelsVideo
- 08.09SLOs for ThinkingVideo
- 08.10DeterminismVideo
- 08.11Rollouts Are InferenceVideo
- 08.12Cost per Solved TaskVideo
