These benchmarks score Cortex against de-identified primary-care charts with real coding and documentation gaps. The useful question is not which model sounds smartest. It is which engine catches more clinically relevant issues without becoming slow, expensive, or noisy.
The decision, in plain terms
01
The model in Cortex today earned its place here.
Claude Sonnet 5 — now running in Cortex — is right about 56% of what it flags, versus 37% for the previous model on the same review pipeline, at matching overall accuracy (46% vs 45%). Fewer wrong flags means fewer interruptions you learn to ignore — that is the trade we optimize for, and why it was promoted on 2026-07-01.
02
The latest sweep did not justify a switch.
Claude Fable 5 led the five new models at 50% overall accuracy and 58% trust-weighted accuracy. Its range still overlaps Cortex today, while a median chart took 22.9s versus 15.5s. That is not enough evidence to trade away speed for an uncertain gain, so Cortex stays on Claude Sonnet 5.
03
No model reaches your charts on hype.
Every candidate runs the same 10 de-identified real charts and is scored by an independent judge model. Promotion candidates are repeated three times to check that the result is stable. A model reaches Cortex only after it improves the current balance, especially by avoiding false alarms without missing too much.
04
Small test set, honest error bars.
10 charts is a small sample, so every score is published with its range. When two models' ranges overlap, we call it a tie instead of declaring a winner. Speed and cost are published too (1.1s per chart at the fast end, 0.11¢ at the cheapest) — the numbers vendors usually leave out.
Fastest
1.1s
GPT-OSS 120B (Cerebras)
Cheapest
0.11¢
OpenAI GPT-5.4 Nano
Best overall accuracy
50%
Claude Fable 5
Best new candidate
50%
Claude Fable 5
Model comparison
Compare the tradeoffs
Start with Cortex today and four useful reference points. Add any evaluated model when you need a deeper comparison.
5 of 15 models selected
Production always included
Accuracy, time, and cost
Circle size shows median cost per chart (smaller is better).
In Cortex today Compared
Selected comparison
Open a row for precision, recall, and test details.
Model
Overall
Trust-weighted
Time
Cost
Readout
Details
1
Cortex today · Claude Sonnet 5
In Cortex today
46%
53%
15.5s
2.5¢
The model reviewing charts now
2
Claude Fable 5
Latest candidate
50%
58%
22.9s
10.0¢
Best new result; not a clear enough gain to switch
3
OpenAI GPT-5.6 Sol
Latest candidate
42%
54%
14.7s
3.6¢
Very precise and responsive, but missed too many issues
4
GPT-OSS 120B (Cerebras)
Comparison
21%
28%
1.1s
0.16¢
Fastest tested, misses more
5
OpenAI GPT-5.4 Nano
Comparison
29%
34%
5.9s
0.11¢
Cheapest tested
1
Cortex today · Claude Sonnet 5
The model reviewing charts now
Overall
46%
Trust-weighted
53%
Time
15.5s
Cost
2.5¢
Precision
56%
Recall
56%
95% confidence
26%–66%
Charts
10
Prompt
v0.5
Configuration
v0.7 · claude-sonnet-5
Scoped scoring: graded on specificity, missing-dx, audit-risk, em-level because this configuration is not designed to produce every benchmark category.
2
Claude Fable 5
Best new result; not a clear enough gain to switch
Overall
50%
Trust-weighted
58%
Time
22.9s
Cost
10.0¢
Precision
69%
Recall
50%
95% confidence
45%–54%
Charts
10
Prompt
v0.4
Configuration
v0.5 · claude-fable-5
3
OpenAI GPT-5.6 Sol
Very precise and responsive, but missed too many issues
Overall
42%
Trust-weighted
54%
Time
14.7s
Cost
3.6¢
Precision
96%
Recall
35%
95% confidence
26%–63%
Charts
10
Prompt
v0.4
Configuration
v0.5 · gpt-5.6-sol
4
GPT-OSS 120B (Cerebras)
Fastest tested, misses more
Overall
21%
Trust-weighted
28%
Time
1.1s
Cost
0.16¢
Precision
55%
Recall
20%
95% confidence
10%–32%
Charts
10
Prompt
v0.4
Configuration
v0.5 · gpt-oss-120b
5
OpenAI GPT-5.4 Nano
Cheapest tested
Overall
29%
Trust-weighted
34%
Time
5.9s
Cost
0.11¢
Precision
41%
Recall
24%
95% confidence
13%–48%
Charts
10
Prompt
v0.4
Configuration
v0.5 · gpt-5.4-nano
Benchmark method
How we test
The same real charts
Every candidate reviews the same 10 de-identified primary-care charts with real coding and documentation mistakes. Synthetic smoke tests are never included in the published scores.
A separate judge
An independent judge creates the expected issue list and checks each result. We report precision, recall, overall accuracy, time, and cost. Cortex never grades its own work.
Honest limits
Candidates run three times. Score ranges show how much the result could move, and overlapping ranges count as a tie. This is still a small, primary-care-only test set, so it guides product decisions. It is not a clinical validation claim.