Dashboardblocks
Components

AI Usage

Tokens by model, cost by feature, model latency and errors, and prompt caching savings.

AI usage blocks track what language models cost and how they perform. Costs come from token counts at per-million prices, with cached input tokens billed at their own, lower rate. Models keep their colour in a fixed order across blocks. Prices and model names in the examples are placeholders; pass your own.

Token Usage

Tokens per day stacked by model, with the total cost and the cost per million tokens.

Token usage
Tokens per day by model, last 14 days
Tokens
203.7M
Cost
$1,515
Per million tokens
$7.44
  • Reasoning
  • Fast
  • Embeddings
  • Reasoning42.1M · $1,184
  • Fast118M · $312
  • Embeddings43.6M · $19
Token usage: Tokens per day by model, last 14 days.
DayReasoningFastEmbeddings
Sep 131.2M3.8M1.4M
Sep 141.2M3.8M1.4M
Sep 152.8M8.9M3.3M
Sep 162.8M9.2M3.4M
Sep 172.9M9.4M3.5M
Sep 183M9.7M3.6M
Sep 193.1M9.9M3.7M
Sep 201.2M3.8M1.4M
Sep 211.2M3.8M1.4M
Sep 223.3M10.7M3.9M
Sep 234.7M10.9M4M
Sep 244.8M11.2M4.1M
Sep 255M11.4M4.2M
Sep 265.1M11.7M4.3M

Cost by Feature

Model spend by product feature, with each feature's input, cached and output tokens and its cost per thousand requests.

Cost by feature
Model spend by product feature, this month
  • Research agent$84260%
    94.8M tokens, 18.4K requests$46 per 1K requests
    64M input, 21M cached input and 9.8M output tokens.
  • Chat assistant$39128%
    111.6M tokens, 212K requests$1.84 per 1K requests
    41M input, 58M cached input and 12.6M output tokens.
  • Ticket summaries$14410%
    45.1M tokens, 96.5K requests$1.49 per 1K requests
    38M input, 4M cached input and 3.1M output tokens.
  • Search embeddings$272%
    54M tokens, 1.5M requests$0.02 per 1K requests
    54M input, 0 cached input and 0 output tokens.

Model Performance

Requests, time to first token at the median and 95th percentile, output speed and error rate for each model.

Model performance
Last 24 hours, all regions
ModelRequestsTTFT p50TTFT p95Tokens/sErrors
Reasoning41,2001.24 s3.88 s580.40%
Fast318,000310 ms790 ms1640.12%
Vision12,600880 ms2.94 s912.10%, above the 1% threshold

TTFT is time to first token. Error rates over 1% are red.

Prompt Caching

The share of input tokens served from the cache, what it saved and what tokens would have cost without it.

Prompt caching
Prompt caching across all requests, this month
Cache hit rateof input tokens↑ 5.6 pts on last month
Saved by caching
$224
Spent on tokens
$998
would be $1,223 without the cache
  • Input 197M
  • Cached 83M
  • Output 25.5M

On this page