Glossary term
Semantic caching
Reusing a prior answer or result for a semantically similar request to reduce repeated model computation.
Distinct from exact-match caching — it catches requests that are worded differently but mean the same thing. One of the handful of AI-specific cost-optimization practices the FinOps Foundation names explicitly, alongside model tiering and routing.
Related research