Embeddings and reranking for production RAG

Two Pareta-hosted retrieval models for recurring production retrieval workloads. pareta-embed turns text into vectors through POST /v1/embeddings; pareta-rerank scores candidate documents against a query through POST /v1/rerank. Each is measured by nDCG@10 against its frontier baseline on Pareta’s published benchmarks. Call them directly from any retrieval stack with your Pareta API key.

pareta-embed

text embeddings · POST /v1/embeddings
Input
Text — one string or a list. Index passages as-is; embed search queries with input_type: "query".
Output
One unit-normalized vector per input, in input order (OpenAI-shaped response).
Metric
nDCG@10 vs OpenAI text-embedding-3-large
109%of the baseline nDCG@10 score
32×lower serving cost than the same baseline
$0.004 / 1M input tokens · metered per input token

pareta-rerank

document reranking · POST /v1/rerank
Input
A query and a list of candidate documents.
Output
Each document’s index and relevance score, most relevant first. Scores are calibrated, so a fixed threshold works as a keep/drop filter; top_n truncates the list.
Metric
nDCG@10 vs Gemini 3.5 Flash
101%of the baseline nDCG@10 score
31×lower serving cost than the same baseline
$0.025 / 1,000 documents · metered per document scored
Example

Call them directly

Standard REST endpoints with your Pareta API key. They work with any retrieval stack; no other Pareta usage is required.

curl https://api.pareta.ai/v1/embeddings \
  -H "Authorization: Bearer $PARETA_API_KEY" \
  -d '{"input": ["governed metrics catalog", "..."]}'

curl https://api.pareta.ai/v1/rerank \
  -H "Authorization: Bearer $PARETA_API_KEY" \
  -d '{"query": "termination clause", "documents": ["...", "..."]}'

Standalone retrieval vs. retrieval inside auto

Standalone retrieval workloads — search, RAG context selection, citation finding — call /v1/embeddings and /v1/rerank directly, as above. Retrieval-dependent workloads served through the model ID auto handle their retrieval step internally: you send the request to auto and do not call these endpoints yourself.

Both stages can be benchmarked on your own data: the text-embedding and document-reranking evaluation tasks score nDCG@10 against your graded relevance judgments.

Need dedicated capacity, or a model trained for your domain? Dedicated GPU capacity for either model, and custom rerankers and task models trained on your data, are available by arrangement. Write to .
Benchmark your workload Read the production quickstart
New accounts start with $30 in credit; no card required.
Published benchmarks: Pareta benchmark runs, July 2026, calling each retrieval model directly through the same serving bridge the /v1/embeddings and /v1/rerank endpoints use. They do not include end-to-end routing, verification, or frontier escalation. Metric: nDCG@10, normalized to the named baseline, which was run on the same items. Frontier cost = the baseline’s actual usage tokens × vendor list price; Pareta = actual metered billing.