Pricing
Each completed request is billed once. You pay for the model that produces the final response, not for routing, verification, retries, or non-serving attempts behind it.
One rate for every Pareta-hosted specialist, metered on the specialist's actual input and output tokens. No model tiers to choose.
When the specialist path cannot meet the workload's quality bar — or no Pareta specialist holds the bar for the workload — a frontier model serves the request. It is billed at that provider's current list price for the tokens it used, instead of the token rates.
Anything Pareta runs behind a request that does not produce the final response is included in whichever rate applies. One debit per completed request.
Prepaid, per request. Usage is deducted from your balance at the rates above, metered from the final answer model's actual token consumption. A call on an empty balance returns a 402, never a surprise invoice.
Two paths, one debit. A specialist-served answer bills at the token rates. A frontier-served answer bills at the provider's list price instead — not in addition. Every response carries the exact amount charged (see the request receipt below).
What runs, what is included, and which path is billed. The rates are the published rates above; the token counts of a real request set the actual amount.
auto.auto with a
json_schema response format.auto and each baseline on every item,
scores all of them the same way, and reports quality with
confidence intervals and cost per contender.auto calls at the rates above (token rates where a
specialist answered, list price where a frontier model did) plus
each frontier baseline's calls at its list price. On custom eval
sets graded by a judge panel, the grading calls are part of the
run total.Every completed response carries its bill in two headers, and the dashboard reconciles to them.
X-Pareta-BilledThe exact amount debited for
this request, in micro-dollars: the token rates if a Pareta
specialist produced the answer, the provider's list price if a
frontier model did.X-Pareta-Frontier-Would-Have-CostEstimated.
This request's usage tokens priced at a frontier model's list price
— a comparison figure, not an observed charge.Matched or beat the strongest frontier baseline on all five published benchmarks, with 30–140× lower serving cost. These are the same five runs behind the homepage numbers; dataset, path tested, and cost method are in the note at the foot of this page.
| Task · metric · baseline | Quality (strongest frontier = 100%) | Serving cost vs that frontier |
| Intent classification | 113% | 88× lower |
| Text embedding | 109% | 32× lower |
| ICD-10 medical coding | 106% | 138× lower |
| Invoice extraction | 103% | 142× lower |
| Document reranking | 101% | 31× lower |
Source: measured Pareta benchmark runs, May–July 2026. Published benchmarks: The path tested is stated per row on the benchmarks page — the benchmark harness for intent classification, the direct endpoints for embeddings and reranking, and the product path through auto for ICD-10 coding and invoice extraction. Frontier models were actually run on the same items; none of the figures include route-mix or escalation statistics. Frontier cost = their actual usage tokens × vendor list price. Pareta = actual metered billing at the token rates above. Quality normalized to the strongest frontier model per task. Frontier set: GPT-5.5, Claude Opus 4.7/4.8, Sonnet 4.6, Gemini 3.5 Flash, OpenAI text-embedding-3-large. Token rates apply to answers produced by Pareta-hosted specialists; frontier-served answers bill at provider list price.