topic
Topic: Inferencing
seed (no source yet -> emerging at 1 source -> established at 2+ corroborating sources)
About this note
Status: seed (no source yet -> emerging at 1 source -> established at 2+ corroborating sources)
Living, cross-source synthesis on LLM inference / serving. Many sources feed this note; merge and de-duplicate as they arrive (architect persona). Every claim cited.
On this page
What this coversSynthesisKey claimsKey visualsOpen questions / conflictsSources feeding this topicWhat this covers#
Running models efficiently: the prefill/decode phases, the KV cache, batching (static / continuous), quantization, speculative decoding, attention optimizations (e.g. paged attention), throughput vs. latency tradeoffs, and serving stacks (vLLM, TGI, TensorRT-LLM, llama.cpp).
Synthesis#
(empty - populated as sources are ingested.)
Key claims#
| Claim | Sources (cited) | Confidence |
|---|---|---|
| (none yet) | - | - |
Key visuals#
(Diagrams/plots/code across sources - e.g. a KV-cache or continuous-batching diagram - embedded with caption + citation.)
Open questions / conflicts#
- (none yet)
Sources feeding this topic#
- (none yet)