IndustriesPricingServicesBenchmarks
CyberTalk to salesStart
Start
Get started

GPT-6 Astra catches 9 more of every 100 critical alerts than your triage model at 10% lower cost

New release·GPT-6 Astra shipped to trusted access Oct 7, 10:00·Scored 640 cases by 11:12

Your alert triage runs GPT-5.6 Sol at $1.20 per alert and catches 78.4% of critical alerts
GPT-6 Astra catches 87.4% at $1.08 and now holds 5 of the 6 points on your frontier

Recall on critical alerts
87.4%▲ +9.0 pts
Cost per alert
$1.08▼ 0.9×
Estimated monthly cost at 150,000 alerts
$162k▼ −$18k
Claude Fable 5.1GPT-5.6 SolGPT-6.1 SolClaude Opus 5.5GPT-6 AstraProductionNew champion90%75%60%$6$4$2
Your frontierAlert triageFrontierGPT-6 AstraClaude Opus 5.5Others

What changed

All runs

21 models · 640 cases · avg@3
ModelRecallCost / alertp95Tokens / alertTurnsFrontier
GPT-6 Astra Max
91.8%
$4.60
98.4s
96,200
64
Yes
GPT-6 Astra Extra High
90.6%
$2.85
71.2s
61,800
46
Yes
GPT-6 Astra High
89.5%
$1.72
52.6s
39,400
34
Yes
Claude Opus 5.5 Max
88.9%
$6.80
142.3s
118,400
88
No
GPT-6 Astra Medium
87.4%
$1.08
34.8s
24,600
23
Yes
Claude Opus 5.5 High
86.8%
$3.10
78.5s
54,600
49
No
Claude Fable 5.1 Max
86.1%
$6.20
176.5s
142,000
104
No
Claude Fable 5.1 High
84.7%
$4.10
118.2s
89,500
72
No
Claude Opus 5.5 Medium
84.0%
$1.90
49.8s
33,200
32
No
GPT-5.6 Sol Extra High
83.2%
$3.30
88.1s
71,500
58
No
GPT-6.1 Sol High
82.4%
$2.40
63.9s
52,100
41
No
GPT-6 Astra Low
82.1%
$0.52
19.7s
11,800
13
Yes
Claude Fable 5.1 Medium
81.3%
$2.70
82.1s
56,300
51
No
GPT-5.6 Sol High
81.0%
$1.95
59.4s
44,800
39
No
GPT-6.1 Sol Medium
79.0%
$1.30
41.3s
31,700
28
No
GPT-5.6 Sol Medium
78.4%
$1.20
38.6s
28,900
27
No
Claude Opus 5.5 Low
77.5%
$0.85
26.4s
15,100
18
No
Gemini 3.8 Flash High
76.8%
$0.95
31.2s
68,400
39
No
GPT-6.1 Sol Low
73.6%
$0.70
22.8s
14,900
16
No
Gemini 3.8 Flash Medium
72.4%
$0.48
18.9s
39,100
24
Yes
GPT-5.6 Sol Low
71.6%
$0.62
21.5s
14,200
15
No

Where each model lands

Top right is best on both axes, faded dots are off the frontier

Expensive and fastCheap and fastExpensive and slowCheap and slow$0$1.9$70s53s180sCost per alert, cheaper to the rightLatency p95, faster upChampionProduction

Speed against cost

Latency p95 · cost per alert

Gemini 3.8 Flash Medium is the cheapest and fastest run, at $0.48 and 18.9s per alert.

GPT-6 Astra Low costs 4 cents more and catches 9.7 points more, at 82.1%. GPT-6 Astra Medium is the pick, catching 87.4% at $1.08 and 34.8s, both under production's $1.20 and 38.6s.

Slow and higher recallFast and higher recallSlow and lower recallFast and lower recall0s53s180s70%82.4%95%Latency p95, faster to the rightRecallChampionProduction

Recall against speed

Recall · latency p95

GPT-6 Astra High gives the best recall without the wait, 89.5% at a 52.6s p95.

GPT-6 Astra Max adds only 2.3 points and takes 45.8s longer. Claude Opus 5.5 Max takes 142.3s to reach 88.9%, still below GPT-6 Astra High.

Many tokens and higher recallFew tokens and higher recallMany tokens and lower recallFew tokens and lower recall045k150k70%82.4%95%Tokens per alert, fewer to the rightRecallChampionProduction

Token appetite

Recall · tokens per alert

GPT-6 Astra Medium is the most token-efficient top scorer, 87.4% on 24,600 tokens per alert.

Gemini 3.8 Flash High burns 68,400 tokens for 76.8%, nearly three times as many for 10.6 points less. GPT-6 Astra Medium uses 4,300 fewer tokens than production's 28,900.

Many turns and higher recallFew turns and higher recallMany turns and lower recallFew turns and lower recall03911070%82.4%95%Turns per alert, fewer to the rightRecallChampionProduction

Turns against recall

Recall · turns per alert

GPT-6 Astra Medium needs the fewest turns of any run above 85%, catching 87.4% in 23 turns.

That is 4 fewer turns than production for 9.0 points more recall. Claude Fable 5.1 Max needs 104 turns for 86.1%.

How this was scored

Cases
640
from your alert queue, CVE backlog, and attack replays
Runs
avg@3
3 runs per case, scores averaged
Grading
Ground truth
analyst labels, your test suites, and replayed attacks
Next run
Next release
OpenAI, Anthropic, and Google releases watched

Release history

DateReleaseChampionResult
Oct 7GPT-6 AstraGPT-6 AstraBetter on alert triage and vulnerability remediation
Sep 23GPT-6.1 SolGPT-5.6 SolSame on alert triage at 1.1× the cost
Sep 9GPT-5.6 SolGPT-5.6 SolBetter on alert triage and vulnerability remediation
Aug 12Claude Opus 5.5Claude Opus 5.5Better on detection engineering
Jul 21Gemini 3.8 FlashClaude Opus 5.1Slower on detection engineering
Jun 24Claude Fable 5.1Claude Opus 5.1Same on detection engineering at 2× the cost
Automatically updates with each model releaseRun this on your workflows →

Find your frontier with the most intelligent applied researcher

ProductStudioObserveEvaluateOptimizeIntegrations
GuidesLLM OpsLLM EvaluationEngineeringOptimize Apps
ResourcesLLM LeaderboardBenchmarksGlossaryBlog
© 2022–2026 Klu, Inc.
PrivacyTerms