Plan your AI costs at scale
Choose your model, configure your query volume, and see exactly what your AI system will cost per month — with and without semantic caching.
100% client-side·No API calls·Zero server cost
Select Model
Configure Request Profile
Input tokens per request500 tokens
Output tokens per request300 tokens
Daily query volume10K queries/day
Optimisations
Semantic caching
Cache-read rate: $0.125/1M tokens (saves on repeated queries)
Prompt compressionNone
Shrinking your system prompt from 500 to 500 tokens — slide to simulate savings.
Gemini 2.5 ProGoogle
Input
$1.25/1M
Output
$10/1M
Context
2M
Cost / request$0.0036
Cost / 1K req$3.62
Monthly cost$1.1K300K requests
Annual estimate$13.1K
Model comparison — monthly cost
1Mistral Small 3.1$42.00
2Llama 3.3 70B$43.20
3Gemini 2.5 Flash-Lite$51.00
4DeepSeek V3$139.50
5GPT-5.4 nano$142.50
6Gemini 2.5 Flash$270.00
7GPT-5.4 mini$517.50
8Claude Haiku 4.5$600.00
Was this page helpful?
Sign in to cast your vote
Discussion
Sign in to share your feedback and join the discussion.

