CortexCharts

Cortex Engine Benchmarks

Measured on the mistakes that matter.

These benchmarks score Cortex against de-identified primary-care charts with real coding and documentation gaps. The useful question is not which model sounds smartest. It is which engine catches more clinically relevant issues without becoming slow, expensive, or noisy.

The decision, in plain terms

01

The model in Cortex today earned its place here.

Claude Sonnet 5 — now running in Cortex — is right about 56% of what it flags, versus 37% for the previous model on the same review pipeline, at matching overall accuracy (46% vs 45%). Fewer wrong flags means fewer interruptions you learn to ignore — that is the trade we optimize for, and why it was promoted on 2026-07-01.

02

The latest sweep did not justify a switch.

Claude Fable 5 led the five new models at 50% overall accuracy and 58% trust-weighted accuracy. Its range still overlaps Cortex today, while a median chart took 22.9s versus 15.5s. That is not enough evidence to trade away speed for an uncertain gain, so Cortex stays on Claude Sonnet 5.

03

No model reaches your charts on hype.

Every candidate runs the same 10 de-identified real charts and is scored by an independent judge model. Promotion candidates are repeated three times to check that the result is stable. A model reaches Cortex only after it improves the current balance, especially by avoiding false alarms without missing too much.

04

Small test set, honest error bars.

10 charts is a small sample, so every score is published with its range. When two models' ranges overlap, we call it a tie instead of declaring a winner. Speed and cost are published too (1.1s per chart at the fast end, 0.11¢ at the cheapest) — the numbers vendors usually leave out.

Fastest

1.1s

GPT-OSS 120B (Cerebras)

Cheapest

0.11¢

OpenAI GPT-5.4 Nano

Best overall accuracy

50%

Claude Fable 5

Best new candidate

50%

Claude Fable 5

Model comparison

Compare the tradeoffs

Start with Cortex today and four useful reference points. Add any evaluated model when you need a deeper comparison.

5 of 15 models selected

Production always included

Accuracy, time, and cost

Circle size shows median cost per chart (smaller is better).

In Cortex today Compared
Selected benchmark model tradeoffsOverall accuracy is on the vertical axis, median seconds per chart is on the horizontal axis, and circle size represents median provider cost per chart. Higher accuracy, lower time, and lower cost are better. Use the model buttons below the chart to select a comparison.OVERALL ACCURACY (HIGHER IS BETTER)0%30%60%0s13s25sMEDIAN TIME PER CHART (LOWER IS BETTER)Cortex today · Claude Sonnet 5: 46% overall accuracy, 15.5s, 2.5¢ per chart1Claude Fable 5: 50% overall accuracy, 22.9s, 10.0¢ per chart2OpenAI GPT-5.6 Sol: 42% overall accuracy, 14.7s, 3.6¢ per chart3GPT-OSS 120B (Cerebras): 21% overall accuracy, 1.1s, 0.16¢ per chart4OpenAI GPT-5.4 Nano: 29% overall accuracy, 5.9s, 0.11¢ per chart5

Selected comparison

Open a row for precision, recall, and test details.

1

Cortex today · Claude Sonnet 5

The model reviewing charts now

Overall
46%
Trust-weighted
53%
Time
15.5s
Cost
2.5¢
Precision
56%
Recall
56%
95% confidence
26%–66%
Charts
10
Prompt
v0.5
Configuration
v0.7 · claude-sonnet-5

Scoped scoring: graded on specificity, missing-dx, audit-risk, em-level because this configuration is not designed to produce every benchmark category.

2

Claude Fable 5

Best new result; not a clear enough gain to switch

Overall
50%
Trust-weighted
58%
Time
22.9s
Cost
10.0¢
Precision
69%
Recall
50%
95% confidence
45%–54%
Charts
10
Prompt
v0.4
Configuration
v0.5 · claude-fable-5
3

OpenAI GPT-5.6 Sol

Very precise and responsive, but missed too many issues

Overall
42%
Trust-weighted
54%
Time
14.7s
Cost
3.6¢
Precision
96%
Recall
35%
95% confidence
26%–63%
Charts
10
Prompt
v0.4
Configuration
v0.5 · gpt-5.6-sol
4

GPT-OSS 120B (Cerebras)

Fastest tested, misses more

Overall
21%
Trust-weighted
28%
Time
1.1s
Cost
0.16¢
Precision
55%
Recall
20%
95% confidence
10%–32%
Charts
10
Prompt
v0.4
Configuration
v0.5 · gpt-oss-120b
5

OpenAI GPT-5.4 Nano

Cheapest tested

Overall
29%
Trust-weighted
34%
Time
5.9s
Cost
0.11¢
Precision
41%
Recall
24%
95% confidence
13%–48%
Charts
10
Prompt
v0.4
Configuration
v0.5 · gpt-5.4-nano

Benchmark method

How we test

The same real charts

Every candidate reviews the same 10 de-identified primary-care charts with real coding and documentation mistakes. Synthetic smoke tests are never included in the published scores.

A separate judge

An independent judge creates the expected issue list and checks each result. We report precision, recall, overall accuracy, time, and cost. Cortex never grades its own work.

Honest limits

Candidates run three times. Score ranges show how much the result could move, and overlapping ranges count as a tie. This is still a small, primary-care-only test set, so it guides product decisions. It is not a clinical validation claim.