Acuvo · measured, not marketed

Which provider route is honest about its cache

Everyone publishes which model is smartest. Nobody publishes which provider route actually caches your prompt, because that means paying several of them for the same real work at the same time and metering every call. We run a multi-provider chain and meter every call, so we can. A cached prompt token bills at about a tenth of a fresh one, which makes this the number that decides what a token costs — and it is on no pricing page.

Read this as a small sample, because it is one. Acuvo has no paying customers yet, so every request below is our own real work — building software through our own product — not a benchmark and not a customer’s traffic. It is honest and it is thin. Rows under 25 requests are shown and labelled rather than hidden, so this table cannot quietly become a list of our best results.

Together cached 94.4% of the prompt tokens we sent it. Relace cached 82.4% — a 1.1x spread. Same gateway, same weeks, same code. Neither provider publishes this number, and it is the one that decides what a token actually costs.

metered requests
1,000 of 12,745
prompt tokens
29.2M
output tokens
0.97M
window
2026-09-27 → 2026-09-28

This is the most recent 1,000 of 12,745 metered requests, not all of them — the database caps a single read and reports no error when it does. The dates above are this window’s, not the whole record’s.

By upstream — the leg that actually served

routerequestsprompt tokenscachedcold startsstanding
Together37813.16M94.4%6.1%measured
Relace40812.68M82.4%21.3%measured
(not an OpenRouter call)2043.16M74.5%27.0%measured
Makora70.21M45.1%57.1%too thin to publish (under 25 requests)
CoreWeave29910.0%100.0%too thin to publish (under 25 requests)
DeepInfra18300.0%100.0%too thin to publish (under 25 requests)

By model

routerequestsprompt tokenscachedcold startsstanding
deepseek/deepseek-v4-flash-073153715.90M81.1%21.6%measured
deepseek/deepseek-v4.1-flash38313.21M94.0%7.3%measured
deepseek/deepseek-v4-flash-vision-exp790.08M17.6%34.2%measured
deepseek/deepseek-chat10.02M0.0%100.0%too thin to publish (under 25 requests)

What we could not measure

Privacy

We publish no prompt, no completion and no customer identity on this page — only counts and rates about models and providers. We cannot make the wider promise: some providers we route to, DeepSeek among them, retain user inputs under their own terms, and that is their server and their policy, not something we control or can undo by writing it here.

Measured through acuvo-code, our CLI — npm i -g acuvo-code. Every figure on this page is a count or a ratio over our own metered calls, recomputed on load.