GPTCache
A caching library that can reuse language-model responses using exact or semantic matches.
About GPTCache
Repeated requests may spend model time and tokens on answers that can safely be reused.
Who it’s for
- Developers with repetitive LLM workloads
- Teams evaluating response reuse
When to consider it
Consider GPTCache only after measuring repetition in your own workload. A useful cache needs a correct reuse policy as well as a good hit rate. Compare exact matching first, then evaluate whether semantic matching adds value.
Tradeoffs & limitations
- Similar questions can require different answers, especially with changing facts or user-specific context.
- Cache invalidation, isolation and compatibility with your model client need explicit tests.
- The maintainers no longer add adapters for new model APIs; use its generic get and set API instead.
A useful first evaluation
Build a test set of similar-looking requests that must not share answers. Measure false reuse as well as cache hits, and test expiry after underlying source information changes.
Suggested evaluation, not a report of tests we ran. Download the worksheet.
SOTA overview · Documentation-based assessment · Sources & review method
Updates
No updates shared yet.
Discussion
Newest firstAsk a question or share how you use GPTCache.
Keep it helpful. Community rules
Loading discussion…
Comparisons & guides
Sources & review method
Documentation-based assessment · Sep 30, 2026 · Prepared with AI assistance; not a hands-on benchmark.
Checked by SOTA · AI-assisted documentation review. Selection advice is our assessment; verify current requirements for your deployment.
Import history & original evidence
Imported source
Semantic cache with a Jev evaluator that uses Noul judgments to check whether a cached response can serve an incoming request.
Author and repository text is preserved where clear; missing languages are enriched automatically. Checks establish source-level Jev integration, not runtime, safety, or performance validation.