Semantic Caching
Save 30-50% on costs by caching semantically similar requests. When a request matches a previous response, CloudVera returns the cached result instantly.
Features
Embedding-Based Matching
Requests are compared using vector embeddings, not exact string matching. "What is the capital of France?" matches "Tell me France's capital city."
Configurable TTL
Set cache time-to-live per project. Default is 1 hour. Adjust based on how dynamic your use case is.
Cache Analytics
Track cache hit rate, tokens saved, and cost savings in real-time on the dashboard.
Per-Project Control
Enable or disable caching per project. Some projects (e.g., creative writing) may need fresh responses every time.