Pricing

Pay for the model that answers

Each completed request is billed once. You pay for the model that produces the final response, not for routing, verification, retries, or non-serving attempts behind it.

Pareta specialist produces the final answer
$0.01per 1M input tokens
$0.05per 1M output tokens

One rate for every Pareta-hosted specialist, metered on the specialist's actual input and output tokens. No model tiers to choose.

Frontier model produces the final answer
List pricethe provider's current rate

When the specialist path cannot meet the workload's quality bar — or no Pareta specialist holds the bar for the workload — a frontier model serves the request. It is billed at that provider's current list price for the tokens it used, instead of the token rates.

Routing · verification · retries · non-serving attempts
Includednever a separate line

Anything Pareta runs behind a request that does not produce the final response is included in whichever rate applies. One debit per completed request.

How billing works

Prepaid, per request. Usage is deducted from your balance at the rates above, metered from the final answer model's actual token consumption. A call on an empty balance returns a 402, never a surprise invoice.

Two paths, one debit. A specialist-served answer bills at the token rates. A frontier-served answer bills at the provider's list price instead — not in addition. Every response carries the exact amount charged (see the request receipt below).

Three worked examples

What runs, what is included, and which path is billed. The rates are the published rates above; the token counts of a real request set the actual amount.

Example 1 · specialist meets the bar

Intent classification request

  1. Your application sends the request to the model ID auto.
  2. Pareta matches it to the deployed intent-classification specialist and runs it.
  3. The result is checked against the workload's quality bar: schema and structural checks on every specialist response, model-graded checks where the workload's measured quality calls for them.
  4. The result meets the bar and is returned.
BilledThe specialist's input tokens at $0.01 / 1M plus its output tokens at $0.05 / 1M.
IncludedRouting and verification.
Example 2 · frontier required to meet the bar

Invoice extraction request

  1. Your application sends the request to auto with a json_schema response format.
  2. The specialist result does not clear verification for this request, so the specialist path cannot meet the workload's quality bar.
  3. A frontier model produces the final response, which is returned to you.
BilledThe frontier model's usage tokens at that provider's current list price.
IncludedRouting, any specialist attempt, verification, and retries. The token rates are not charged on top.
Example 3 · an evaluation run

Benchmark on your own data

  1. You upload 5–50 representative items and name the frontier baselines to compare against.
  2. Pareta runs auto and each baseline on every item, scores all of them the same way, and reports quality with confidence intervals and cost per contender.
BilledOne debit for the run: the auto calls at the rates above (token rates where a specialist answered, list price where a frontier model did) plus each frontier baseline's calls at its list price. On custom eval sets graded by a judge panel, the grading calls are part of the run total.
Not billedScoring with Pareta's built-in scorers, items skipped for a transient infrastructure failure, and a run that fails. An empty balance is refused with a 402 before the run starts.

Request receipt

Every completed response carries its bill in two headers, and the dashboard reconciles to them.

Published benchmarks

Matched or beat the strongest frontier baseline on all five published benchmarks, with 30–140× lower serving cost. These are the same five runs behind the homepage numbers; dataset, path tested, and cost method are in the note at the foot of this page.

Task · metric · baseline Quality (strongest frontier = 100%) Serving cost vs that frontier
Intent classificationmacro-F1 · vs Claude Opus113%88× lower
Text embeddingnDCG@10 · vs OpenAI text-embedding-3-large109%32× lower
ICD-10 medical codingmicro-F1 · vs GPT-5.5106%138× lower
Invoice extractionF1 · vision · vs Claude Opus103%142× lower
Document rerankingnDCG@10 · vs Gemini Flash101%31× lower

Retrieval models

standalone rates · see Retrieval models
pareta-embed  $0.004 / 1M input tokens
pareta-rerank  $0.025 / 1,000 documents

Images & audio

unit-priced per generation / minute
Image generation, editing, speech-to-text and text-to-speech are metered per unit at the rates shown in the product — same prepaid balance, same receipts.
Start with $30 in credit, no card. Sign up, get a key, and benchmark the workload you already run before moving production traffic. Questions about volume or dedicated capacity: info@pareta.ai.
Benchmark your workload Read the production quickstart

Source: measured Pareta benchmark runs, May–July 2026. Published benchmarks: The path tested is stated per row on the benchmarks page — the benchmark harness for intent classification, the direct endpoints for embeddings and reranking, and the product path through auto for ICD-10 coding and invoice extraction. Frontier models were actually run on the same items; none of the figures include route-mix or escalation statistics. Frontier cost = their actual usage tokens × vendor list price. Pareta = actual metered billing at the token rates above. Quality normalized to the strongest frontier model per task. Frontier set: GPT-5.5, Claude Opus 4.7/4.8, Sonnet 4.6, Gemini 3.5 Flash, OpenAI text-embedding-3-large. Token rates apply to answers produced by Pareta-hosted specialists; frontier-served answers bill at provider list price.