GPT-6 Astra catches 9 more of every 100 critical alerts than your triage model at 10% lower cost
New release·GPT-6 Astra shipped to trusted access Oct 7, 10:00·Scored 640 cases by 11:12
Your alert triage runs GPT-5.6 Sol at $1.20 per alert and catches 78.4% of critical alerts
GPT-6 Astra catches 87.4% at $1.08 and now holds 5 of the 6 points on your frontier
- Recall on critical alerts
- 87.4%▲ +9.0 pts
- Cost per alert
- $1.08▼ 0.9×
- Estimated monthly cost at 150,000 alerts
- $162k▼ −$18k
What changed
All runs
21 models · 640 cases · avg@3| Model | Recall | Cost / alert | p95 | Tokens / alert | Turns | Frontier |
|---|---|---|---|---|---|---|
GPT-6 Astra Max | 91.8% | $4.60 | 98.4s | 96,200 | 64 | Yes |
GPT-6 Astra Extra High | 90.6% | $2.85 | 71.2s | 61,800 | 46 | Yes |
GPT-6 Astra High | 89.5% | $1.72 | 52.6s | 39,400 | 34 | Yes |
Claude Opus 5.5 Max | 88.9% | $6.80 | 142.3s | 118,400 | 88 | No |
GPT-6 Astra Medium | 87.4% | $1.08 | 34.8s | 24,600 | 23 | Yes |
Claude Opus 5.5 High | 86.8% | $3.10 | 78.5s | 54,600 | 49 | No |
GPT-6.1 Sol High | 82.4% | $2.40 | 63.9s | 52,100 | 41 | No |
GPT-5.6 Sol Medium | 78.4% | $1.20 | 38.6s | 28,900 | 27 | No |
Where each model lands
Top right is best on both axes, faded dots are off the frontier
Speed against cost
Latency p95 · cost per alertGemini 3.8 Flash Medium is the cheapest and fastest run, at $0.48 and 18.9s per alert.
GPT-6 Astra Low costs 4 cents more and catches 9.7 points more, at 82.1%. GPT-6 Astra Medium is the pick, catching 87.4% at $1.08 and 34.8s, both under production's $1.20 and 38.6s.
Recall against speed
Recall · latency p95GPT-6 Astra High gives the best recall without the wait, 89.5% at a 52.6s p95.
GPT-6 Astra Max adds only 2.3 points and takes 45.8s longer. Claude Opus 5.5 Max takes 142.3s to reach 88.9%, still below GPT-6 Astra High.
Token appetite
Recall · tokens per alertGPT-6 Astra Medium is the most token-efficient top scorer, 87.4% on 24,600 tokens per alert.
Gemini 3.8 Flash High burns 68,400 tokens for 76.8%, nearly three times as many for 10.6 points less. GPT-6 Astra Medium uses 4,300 fewer tokens than production's 28,900.
Turns against recall
Recall · turns per alertGPT-6 Astra Medium needs the fewest turns of any run above 85%, catching 87.4% in 23 turns.
That is 4 fewer turns than production for 9.0 points more recall. Claude Fable 5.1 Max needs 104 turns for 86.1%.
How this was scored
- Cases
- 640
- from your alert queue, CVE backlog, and attack replays
- Runs
- avg@3
- 3 runs per case, scores averaged
- Grading
- Ground truth
- analyst labels, your test suites, and replayed attacks
- Next run
- Next release
- OpenAI, Anthropic, and Google releases watched
Release history
| Date | Release | Champion | Result |
|---|---|---|---|
| Oct 7 | GPT-6 Astra | GPT-6 Astra | Better on alert triage and vulnerability remediation |
| Sep 23 | GPT-6.1 Sol | GPT-5.6 Sol | Same on alert triage at 1.1× the cost |
| Sep 9 | GPT-5.6 Sol | GPT-5.6 Sol | Better on alert triage and vulnerability remediation |
| Aug 12 | Claude Opus 5.5 | Claude Opus 5.5 | Better on detection engineering |
| Jul 21 | Gemini 3.8 Flash | Claude Opus 5.1 | Slower on detection engineering |
| Jun 24 | Claude Fable 5.1 | Claude Opus 5.1 | Same on detection engineering at 2× the cost |