ICD-10 medical coding

A clinical-coding specialist distilled in-house and served from the Pareta fleet. Send the clinical note through the same OpenAI-compatible endpoint and model ID auto you use for your other workloads. By default Pareta serves it with the ICD-10 specialist, checks the result against the workload’s quality bar before returning it, and escalates to a frontier model when needed to meet that bar. You are charged only for the model that produces the final response.

106% of GPT-5.5 micro-F1 on ICD-10-CM coding — published benchmark (product path through auto), GPT-5.5 run on the same items · methodology
138× lower serving cost than the same GPT-5.5 run — its actual tokens at vendor list price vs Pareta’s actual metered billing
Structured JSON output Send response_format with a json_schema; decoding is constrained to your schema, so codes arrive as a JSON array — not parsed from prose.
$0.01 / $0.05 per 1M input / output tokens when the Pareta specialist produces the final response; frontier-served answers bill at the provider’s list price · pricing
Example

The same endpoint, a structured request

There is no separate coding API. Point the OpenAI client at Pareta’s base URL, set the model ID to auto, and declare the output contract with response_format. Schema instructions given only in the prompt are not enforced by the specialist; json_schema is the supported contract — Pareta constrains decoding to it and validates the response against it before returning (structured outputs in the docs).

completion = client.chat.completions.create(
    model="auto",
    messages=[{"role": "user", "content":
        "Assign ICD-10-CM codes for the discharge summary below.\n\n" + chart_note}],
    response_format={
        "type": "json_schema",
        "json_schema": {
            "name": "icd10_codes",
            "schema": {
                "type": "object",
                "properties": {"codes": {"type": "array", "items": {"type": "string"}}},
                "required": ["codes"],
                "additionalProperties": False,
            },
        },
    },
)
codes = json.loads(completion.choices[0].message.content)["codes"]
Benchmark it on your own charts. Upload de-identified charts in the dashboard and run the evaluation Pareta publishes: your data, the Pareta specialist and frontier models side by side, scored the same way.
Evaluating clinical data? Use de-identified records and review Pareta’s retention and subprocessor policies before sending production data.
Benchmark your workload Read the production quickstart
New accounts start with $30 in credit; no card required.
Published benchmark: Pareta benchmark run, 2026-07-06, on the product path — auto through the eval API, including routing and verification. Metric: micro-F1 vs GPT-5.5, run on the same items. Frontier cost = GPT-5.5’s actual usage tokens × vendor list price; Pareta = actual metered billing. Verification is a quality gate, not a guarantee that an AI output can never be wrong.