Semantic Caching

Save 30-50% on costs by caching semantically similar requests. When a request matches a previous response, CloudVera returns the cached result instantly.

Features

Embedding-Based Matching

Requests are compared using vector embeddings, not exact string matching. "What is the capital of France?" matches "Tell me France's capital city."

Configurable TTL

Set cache time-to-live per project. Default is 1 hour. Adjust based on how dynamic your use case is.

Cache Analytics

Track cache hit rate, tokens saved, and cost savings in real-time on the dashboard.

Per-Project Control

Enable or disable caching per project. Some projects (e.g., creative writing) may need fresh responses every time.