Submit projectSubmit

GPTCache

zilliztech / GPTCache

A caching library that can reuse language-model responses using exact or semantic matches.

Search & RetrievalClassification & Ranking

Save privately. Follow for reviewed updates in your SOTA inbox. Neither changes the ranking.

Your workspace

About GPTCache

Repeated requests may spend model time and tokens on answers that can safely be reused.

Who it’s for

  • Developers with repetitive LLM workloads
  • Teams evaluating response reuse

When to consider it

Consider GPTCache only after measuring repetition in your own workload. A useful cache needs a correct reuse policy as well as a good hit rate. Compare exact matching first, then evaluate whether semantic matching adds value.

Tradeoffs & limitations

  • Similar questions can require different answers, especially with changing facts or user-specific context.
  • Cache invalidation, isolation and compatibility with your model client need explicit tests.
  • The maintainers no longer add adapters for new model APIs; use its generic get and set API instead.

A useful first evaluation

Build a test set of similar-looking requests that must not share answers. Measure false reuse as well as cache hits, and test expiry after underlying source information changes.

Suggested evaluation, not a report of tests we ran. Download the worksheet.

SOTA overview · Documentation-based assessment · Sources & review method

Updates

No updates shared yet.

Discussion

Newest first

Ask a question or share how you use GPTCache.

Keep it helpful. Community rules

Loading discussion…

Comparisons & guides

Sources & review method

Documentation-based assessment · Sep 30, 2026 · Prepared with AI assistance; not a hands-on benchmark.

Checked by SOTA · AI-assisted documentation review. Selection advice is our assessment; verify current requirements for your deployment.

Editorial policy · Suggest a correction

Import history & original evidence

Imported source

Scope: Jev integration · Imported Sep 27, 2026 · MIT

Semantic cache with a Jev evaluator that uses Noul judgments to check whether a cached response can serve an incoming request.

Author and repository text is preserved where clear; missing languages are enriched automatically. Checks establish source-level Jev integration, not runtime, safety, or performance validation.