A clinical-coding specialist distilled in-house and served from
the Pareta fleet. Send the clinical note through the same OpenAI-compatible
endpoint and model ID auto you use for your other workloads. By
default Pareta serves it with the ICD-10 specialist, checks the result
against the workload’s quality bar before returning it, and escalates to a
frontier model when needed to meet that bar. You are charged only for the
model that produces the final response.
response_format with a json_schema;
decoding is constrained to your schema, so codes arrive as a JSON array —
not parsed from prose.
There is no separate coding API. Point the OpenAI client at Pareta’s base
URL, set the model ID to auto, and declare the output contract
with response_format. Schema instructions given only in the
prompt are not enforced by the specialist; json_schema is the
supported contract — Pareta constrains decoding to it and validates the
response against it before returning
(structured outputs in the docs).
completion = client.chat.completions.create( model="auto", messages=[{"role": "user", "content": "Assign ICD-10-CM codes for the discharge summary below.\n\n" + chart_note}], response_format={ "type": "json_schema", "json_schema": { "name": "icd10_codes", "schema": { "type": "object", "properties": {"codes": {"type": "array", "items": {"type": "string"}}}, "required": ["codes"], "additionalProperties": False, }, }, }, ) codes = json.loads(completion.choices[0].message.content)["codes"]
auto through the eval API, including
routing and verification. Metric: micro-F1 vs GPT-5.5, run on the same
items. Frontier cost = GPT-5.5’s actual
usage tokens × vendor list price; Pareta = actual metered billing.
Verification is a quality gate, not a guarantee that an AI output can never
be wrong.