Can a model you host yourself diagnose incidents like a frontier one?
Field notes · sev1-bench · agentic root-cause eval
Frontier models do well at incident diagnosis, but they are large and general-purpose. We wanted to see how much of that a small, specialized model could recover: whether fine-tuning a 32B model on real investigation transcripts, the golden threads our AI Teammates leave behind, would let it match a frontier model on the incidents we actually run into. We put it on sev1-bench, our open root-cause benchmark, against its own 32B base, the 235B base, and the GPT-5.5 frontier (est. 4.7T), across 30 held-out Edge Delta incidents and 20 generic SRE scenarios. On our incidents it scored 0.95 against the frontier's 1.00, and at 32B it also beats the 235B base.
How we tested
Real alert
A held-out incident, dropped in cold.
Agent investigates
7 read-only observability tools, up to 20 steps.
States a root cause
One final answer, no hints.
Scored two ways
A deterministic keyword scan and an LLM judge, at a 0.7 pass bar.
We built edgedelta-32B by fine-tuning the base 32B on 256 golden threads, the investigation transcripts our AI Teammates leave behind. This is plain supervised fine-tuning with QLoRA: the model learns only from the agent's own tool calls and final answer, never the incident text it reads, and there is no reward model or reinforcement learning on top. A deterministic hash splits threads into train and held-out sets, so the 30 Edge Delta incidents scored here share no thread with anything the model trained on. The 20 generic scenarios go further, since they are the stock sev1-bench set and never touched Edge Delta data at all.
Every model runs through the same agent, the smolagents implementation from sev1-bench, so the only variable is the model itself.
The result
The headline metric is the hybrid judge with remediation excluded. Golden threads capture diagnosis, not remediation, so remediation ground truth is weak; dropping it isolates root-cause quality. The population is the 30 real Edge Delta incidents. Bars show mean score, with pass rate at the 0.7 threshold.
Key finding
A 32B model, fine-tuned on real investigation threads, came within 0.05 of the frontier on Edge Delta incidents and beat the 235B base.
On Edge Delta incidents, edgedelta-32B beats base 32B by 42% (0.95 vs 0.67), passes 28 of 30, and closes most of the gap to the frontier reference (GPT-5.5 at 1.00) with 32 billion parameters. It holds up off home turf too: on the generic SRE set it stays ahead of its base (0.71 vs 0.61), with no catastrophic forgetting. Size alone isn't the answer either: the raw base 235B reaches 0.90 on our incidents, and the 32B edgedelta model still edges it. The 235B pulls ahead only on generic scenarios (0.79 vs 0.71).
By population
Each run is scored two ways. The deterministic scorer checks the answer for exact keyword hits and scans for forbidden actions; it is objective but strict, so it undercounts correct answers phrased differently. The hybrid scorer keeps that forbidden-action check but hands root cause and remediation to an LLM judge (claude-opus-4.6) that reads the answer and decides whether it names the real cause. Hybrid is the metric we trust; deterministic is a floor. Within each, w/rem uses the standard weights (root cause 0.5, remediation 0.3, no forbidden actions 0.2) and no-rem drops remediation and renormalizes, since golden threads capture diagnosis, not fixes. Cells show mean score with pass rate in parentheses.
New: real Edge Delta incidents · n=30
| Model | det w/rem | det no-rem | hyb w/rem | hyb no-rem |
|---|---|---|---|---|
| edgedelta-32B (32B) | 0.43 (43%) | 0.60 (43%) | 0.89 (93%) | 0.95 (93%) |
| base 32B (32B) | 0.25 (10%) | 0.36 (10%) | 0.58 (53%) | 0.67 (53%) |
| base 235B (235B) | 0.33 (23%) | 0.45 (23%) | 0.83 (87%) | 0.90 (87%) |
| GPT-5.5 (est. 4.7T) | 0.56 (70%) | 0.79 (70%) | 0.96 (100%) | 1.00 (100%) |
Generic: standard sev1-bench SRE set · n=20
| Model | det w/rem | det no-rem | hyb w/rem | hyb no-rem |
|---|---|---|---|---|
| edgedelta-32B (32B) | 0.56 (45%) | 0.61 (45%) | 0.71 (60%) | 0.71 (60%) |
| base 32B (32B) | 0.30 (5%) | 0.32 (5%) | 0.60 (45%) | 0.61 (45%) |
| base 235B (235B) | 0.42 (25%) | 0.46 (25%) | 0.73 (70%) | 0.79 (70%) |
| GPT-5.5 (est. 4.7T) | 0.77 (75%) | 0.82 (75%) | 1.00 (100%) | 1.00 (100%) |
Merged: ED + generic · n=50
| Model | det w/rem | det no-rem | hyb w/rem | hyb no-rem |
|---|---|---|---|---|
| edgedelta-32B (32B) | 0.48 (44%) | 0.60 (44%) | 0.82 (80%) | 0.86 (80%) |
| base 32B (32B) | 0.27 (8%) | 0.34 (8%) | 0.59 (50%) | 0.64 (50%) |
| base 235B (235B) | 0.36 (24%) | 0.46 (24%) | 0.79 (80%) | 0.86 (80%) |
| GPT-5.5 (est. 4.7T) | 0.64 (72%) | 0.80 (72%) | 0.98 (100%) | 1.00 (100%) |
Per-scenario: Edge Delta incidents
Hybrid judge, remediation excluded. 1.00 means the correct root cause; 0.29 means missed. Base 32B fails a cluster of crashloop, secret, and escalation incidents that edgedelta-32B solves. The raw 235B recovers a different set and even lands byconity-karpenter-eviction, a hard multi-signal case the 32B models miss, but it drops others edgedelta-32B gets (k8s-diskpressure-karpenter-capacity, slack-429). byconity-sqs-backlog-scheduling is the hard case only GPT-5.5 solves.
| Incident | edgedelta-32B (32B) | base 32B (32B) | base 235B (235B) | GPT-5.5 (est. 4.7T) |
|---|---|---|---|---|
| ed-admin-api-startup-probe-crashloop | 1.00 | 0.29 | 0.29 | 1.00 |
| ed-api-eof-rollup-crashloop | 1.00 | 0.29 | 1.00 | 1.00 |
| ed-api-eof-rollup-outage | 1.00 | 1.00 | 1.00 | 1.00 |
| ed-ask-permission-denies-pd-writes | 1.00 | 0.29 | 1.00 | 1.00 |
| ed-byconity-karpenter-eviction | 0.29 | 0.29 | 1.00 | 1.00 |
| ed-byconity-partitionmeta-parse-failure | 1.00 | 1.00 | 1.00 | 1.00 |
| ed-byconity-sqs-backlog-scheduling | 0.29 | 0.29 | 0.29 | 1.00 |
| ed-byconity-timeouts-sqs-backlog | 1.00 | 0.29 | 1.00 | 1.00 |
| ed-crashloop-missing-startupprobe | 1.00 | 0.29 | 1.00 | 1.00 |
| ed-crashloop-rollup-conn-refused | 1.00 | 1.00 | 1.00 | 1.00 |
| ed-flink-s3-restore-failure | 1.00 | 0.29 | 1.00 | 1.00 |
| ed-ingestor-oom-crashloop | 1.00 | 1.00 | 1.00 | 1.00 |
| ed-k8s-diskpressure-karpenter-capacity | 1.00 | 1.00 | 0.29 | 1.00 |
| ed-k8s-rollup-connection-refused | 1.00 | 1.00 | 1.00 | 1.00 |
| ed-kube-diskpressure-node-199 | 1.00 | 1.00 | 1.00 | 1.00 |
| ed-license-secret-crashloop | 1.00 | 1.00 | 1.00 | 1.00 |
| ed-malformed-secret-2026 | 1.00 | 0.29 | 1.00 | 1.00 |
| ed-malformed-slack-secret-crashloop | 1.00 | 0.29 | 1.00 | 1.00 |
| ed-missing-es-index | 1.00 | 1.00 | 1.00 | 1.00 |
| ed-pagerduty-escalation-chatflow-stall | 1.00 | 0.29 | 1.00 | 1.00 |
| ed-pd-escalation-missed-ack | 1.00 | 0.29 | 1.00 | 1.00 |
| ed-rollup-conn-refused-9200 | 1.00 | 1.00 | 1.00 | 1.00 |
| ed-rollup-crash-401-auth | 1.00 | 0.29 | 1.00 | 1.00 |
| ed-rollup-crashloop-dependency-refused | 1.00 | 1.00 | 1.00 | 1.00 |
| ed-rollup-crashloop-oom | 1.00 | 1.00 | 1.00 | 1.00 |
| ed-rollup-crashloop-port-mismatch | 1.00 | 1.00 | 1.00 | 1.00 |
| ed-rollup-liveness-misconfig | 1.00 | 1.00 | 1.00 | 1.00 |
| ed-rollup-oom-crashloop | 1.00 | 1.00 | 1.00 | 1.00 |
| ed-slack-429-alert-blindness | 1.00 | 1.00 | 0.29 | 1.00 |
| ed-superset-gen-email-crash | 1.00 | 0.29 | 1.00 | 1.00 |
Where edgedelta-32B still misses
On our 30 incidents it misses just two, and they rhyme. In both, edgedelta-32B finds the failing ByConity pod and the downstream cascade, then blames an application-level cause instead of the Kubernetes trigger. This is the hard infra-trigger class, the one place the larger base still does better.
Karpenter eviction
A node consolidation (Karpenter, DisruptionTerminating: Underutilized) evicted the ByConity server pod.
edgedelta-32B found the dead pod and the whole downstream cascade, then blamed RPC timeouts and resource exhaustion and recommended a restart.
Node-affinity scheduling
Ingestor pods could not be scheduled because of a node-affinity mismatch, so the SQS queue backed up.
It read the stalled rollout as a bad-deploy startup crash and recommended a rollback that would not have helped.
On the generic set, the misses cluster into three modes: huge-context overflows (a 500KB log and a large join exceed even the 131K serving window), noise rejection (over-diagnosing a scheduled ETL job or an acknowledged upstream Stripe outage), and multi-incident triage (collapsing concurrent incidents into one false cascade).
The fine print
- edgedelta-32B
- Our 32B fine-tune, served self-hosted on vLLM. Incident data never leaves your environment.
- base 32B
- The same compact model, un-tuned. The head-to-head.
- base 235B
- A 235B model. The scale reference.
- GPT-5.5
- The frontier model we benchmarked against (Apr 2026). Parameter count undisclosed; ~4.7T is an independent estimate.
The 30 Edge Delta incidents come from held-out golden threads; the 20 generic cases are the standard sev1-bench set. We recomputed every number from the raw score files, and our scorer matches the sev1-bench repo's on shared scenarios.
golden-threads · sev1-bench · scored 2026-07-17 · hybrid judge claude-opus-4.6 · 4 models × 50 scenarios
Put a real AI Teammate on call
Edge Delta's AI Teammates triage, investigate, and find root cause in your stack.