---
title: "Prefill vs decode in LLM inference"
description: "A practical guide to prefill, decode, TTFT, inter-token latency, chunked prefill, and why one LLM request creates two serving workloads."
canonical_url: "https://fanout.sh/blog/prefill-vs-decode-llm-inference"
md_url: "https://fanout.sh/blog/prefill-vs-decode-llm-inference.md"
last_updated: "2026-08-03"
access: "public"
---

# Prefill vs decode in LLM inference

A practical guide to prefill, decode, TTFT, inter-token latency, chunked prefill, and why one LLM request creates two serving workloads.

- Author: Suraj Gaud

- Published: 2026-08-03

- Track: Inference engineering

- Access: Fanout Pro

- Tags: LLM inference, prefill, decode, latency, vLLM

The complete article body is available to Fanout Pro members on the canonical page.

---
This representation contains public Fanout content only. Protected Pro lessons, account data, billing, checkout, and pricing are not included.

Browse the public content map: https://fanout.sh/sitemap.md
