Topic: Inferencing

topic

Topic: Inferencing

seed (no source yet -> emerging at 1 source -> established at 2+ corroborating sources)

About this note

Status: seed (no source yet -> emerging at 1 source -> established at 2+ corroborating sources)

Living, cross-source synthesis on LLM inference / serving. Many sources feed this note; merge and de-duplicate as they arrive (architect persona). Every claim cited.

On this pageWhat this coversSynthesisKey claimsKey visualsOpen questions / conflictsSources feeding this topic

What this covers#

Running models efficiently: the prefill/decode phases, the KV cache, batching (static / continuous), quantization, speculative decoding, attention optimizations (e.g. paged attention), throughput vs. latency tradeoffs, and serving stacks (vLLM, TGI, TensorRT-LLM, llama.cpp).

Synthesis#

(empty - populated as sources are ingested.)

Key claims#

Claim Sources (cited) Confidence
(none yet) - -

Key visuals#

(Diagrams/plots/code across sources - e.g. a KV-cache or continuous-batching diagram - embedded with caption + citation.)

Open questions / conflicts#

Sources feeding this topic#