Reference · GlossaryKV cacheExplaining why long contexts cost more and why caching prefixes helps latency.On this pageWhen to useWhen not toExample#When to useExplaining why long contexts cost more and why caching prefixes helps latency.#When not toAs a user-facing product feature name — say “prompt caching” instead.#ExampleAfter the system prompt is processed once, later tokens in the same session reuse cached KV blocks.Learn by doing → Learn by doing →Related termsPrompt cachingInferenceLatency← All terms