Simple, prepaid, and per token. Point your OpenAI-compatible
client at Pareta, use model: "auto", and pay for the tokens of
the answer — the routing, verification, and orchestration that produced it
are on us.
model: "auto" answer served
by Pareta's own models — one rate, no model tiers to pick.You pay for the answer, not the machinery. Behind one request Pareta may plan, route, verify, and re-check — none of that is billed. The meter charges the answer model's own token consumption, once per request.
When only a
frontier model clears the bar, auto says so and serves it: that
request is billed at the provider's list price instead of the token
rates. You see which happened on every response — the
X-Pareta-Billed header carries the exact amount, and the
dashboard reconciles to the penny.
Prepaid, with $30 free. New accounts start with a $30 credit and no card. Top up when you're ready; calls on an empty balance return a clean 402, never a surprise invoice.
The same battery behind the homepage numbers — production serving path, real metered billing:
| Task | Quality vs frontier | Cost vs frontier |
| Intent classification | 113% | 88× cheaper |
| Text embedding | 109% | 32× cheaper |
| ICD-10 medical coding | 106% | 138× cheaper |
| Invoice extraction | 103% | 142× cheaper |
| Document reranking | 101% | 31× cheaper |
| Contract field extraction | 96% | 20× cheaper |
Source: measured Pareta benchmark runs, July 2026, on the production serving path customers call. Frontier cost = usage tokens × vendor list price; Pareta = actual metered billing. Quality normalized to the strongest frontier model per task. Frontier set: GPT-5.5, Claude Opus 4.7/4.8, Sonnet 4.6, Gemini 3.5 Flash, OpenAI text-embedding-3-large. Token rates apply to answers served by Pareta-hosted models; frontier-served answers bill at provider list price.