Two Pareta-hosted retrieval models for recurring production
retrieval workloads. pareta-embed turns text into vectors
through POST /v1/embeddings; pareta-rerank scores
candidate documents against a query through POST /v1/rerank.
Each is measured by nDCG@10 against its frontier baseline on Pareta’s
published benchmarks. Call them directly from any retrieval
stack with your Pareta API key.
input_type: "query".top_n truncates the list.Standard REST endpoints with your Pareta API key. They work with any retrieval stack; no other Pareta usage is required.
curl https://api.pareta.ai/v1/embeddings \ -H "Authorization: Bearer $PARETA_API_KEY" \ -d '{"input": ["governed metrics catalog", "..."]}' curl https://api.pareta.ai/v1/rerank \ -H "Authorization: Bearer $PARETA_API_KEY" \ -d '{"query": "termination clause", "documents": ["...", "..."]}'
autoStandalone retrieval workloads — search, RAG context selection, citation
finding — call /v1/embeddings and /v1/rerank
directly, as above. Retrieval-dependent workloads served through the model
ID auto handle their retrieval step internally: you send the
request to auto and do not call these endpoints yourself.
Both stages can be benchmarked on your own data: the
text-embedding and document-reranking evaluation
tasks score nDCG@10 against your graded relevance judgments.
/v1/embeddings and /v1/rerank endpoints use.
They do not include end-to-end routing, verification, or frontier
escalation. Metric: nDCG@10, normalized to the named baseline, which was run
on the same items. Frontier cost = the baseline’s actual usage tokens ×
vendor list price; Pareta = actual metered billing.