AI Usage
Tokens by model, cost by feature, model latency and errors, and prompt caching savings.
AI usage blocks track what language models cost and how they perform. Costs come from token counts at per-million prices, with cached input tokens billed at their own, lower rate. Models keep their colour in a fixed order across blocks. Prices and model names in the examples are placeholders; pass your own.
Token Usage
Tokens per day stacked by model, with the total cost and the cost per million tokens.
- Tokens
- 203.7M
- Cost
- $1,515
- Per million tokens
- $7.44
- Reasoning
- Fast
- Embeddings
- Reasoning42.1M · $1,184
- Fast118M · $312
- Embeddings43.6M · $19
| Day | Reasoning | Fast | Embeddings |
|---|---|---|---|
| Sep 13 | 1.2M | 3.8M | 1.4M |
| Sep 14 | 1.2M | 3.8M | 1.4M |
| Sep 15 | 2.8M | 8.9M | 3.3M |
| Sep 16 | 2.8M | 9.2M | 3.4M |
| Sep 17 | 2.9M | 9.4M | 3.5M |
| Sep 18 | 3M | 9.7M | 3.6M |
| Sep 19 | 3.1M | 9.9M | 3.7M |
| Sep 20 | 1.2M | 3.8M | 1.4M |
| Sep 21 | 1.2M | 3.8M | 1.4M |
| Sep 22 | 3.3M | 10.7M | 3.9M |
| Sep 23 | 4.7M | 10.9M | 4M |
| Sep 24 | 4.8M | 11.2M | 4.1M |
| Sep 25 | 5M | 11.4M | 4.2M |
| Sep 26 | 5.1M | 11.7M | 4.3M |
Cost by Feature
Model spend by product feature, with each feature's input, cached and output tokens and its cost per thousand requests.
- Research agent$84260%94.8M tokens, 18.4K requests$46 per 1K requests64M input, 21M cached input and 9.8M output tokens.
- Chat assistant$39128%111.6M tokens, 212K requests$1.84 per 1K requests41M input, 58M cached input and 12.6M output tokens.
- Ticket summaries$14410%45.1M tokens, 96.5K requests$1.49 per 1K requests38M input, 4M cached input and 3.1M output tokens.
- Search embeddings$272%54M tokens, 1.5M requests$0.02 per 1K requests54M input, 0 cached input and 0 output tokens.
Model Performance
Requests, time to first token at the median and 95th percentile, output speed and error rate for each model.
| Model | Requests | TTFT p50 | TTFT p95 | Tokens/s | Errors |
|---|---|---|---|---|---|
| Reasoning | 41,200 | 1.24 s | 3.88 s | 58 | 0.40% |
| Fast | 318,000 | 310 ms | 790 ms | 164 | 0.12% |
| Vision | 12,600 | 880 ms | 2.94 s | 91 | 2.10%, above the 1% threshold |
TTFT is time to first token. Error rates over 1% are red.
Prompt Caching
The share of input tokens served from the cache, what it saved and what tokens would have cost without it.
- Saved by caching
- $224
- Spent on tokens
- $998
- would be $1,223 without the cache
- Input 197M
- Cached 83M
- Output 25.5M