Edge Delta Research · Benchmark report
v1 · October 2026
AI SRE Arena
An open benchmark for detection, diagnosis and mitigation on 21 Kubernetes incidents, judged against one answer key
Edge Delta Research · Published October 2026 · Rubric v1.0.0
- Incidents
- 21 injected faults
- Configurations
- 2 native, 2 control
- Judged investigations
- 72
- Judge
- GPT-6-Astra
Abstract
We injected 21 faults into a Kubernetes deployment of the OpenTelemetry demo shop. Each product had to detect and investigate them on its default monitors, and an AI judge scored every final report against a fixed answer key and a published rubric.
- Detected on its own18 vs 12
- Edge Delta detected 18 of the 21 incidents through its own monitoring and opened an investigation without anyone pointing it at the problem. Grafana detected 12.Edge DeltaGrafana
- Root cause11 vs 9
- On the 12 incidents both products investigated, Edge Delta named the right root cause in 11 and Grafana in 9.Edge DeltaGrafana
- Supported fix7 vs 5
- On the same 12, Edge Delta proposed a supported fix in 7 and Grafana in 5.Edge DeltaGrafana
- Unsafe advice0 of 18
- None of Edge Delta's 18 final recommendations was incorrect or unsafe, against 2 of Grafana's 12.Edge DeltaGrafana
- Control run5 and 6
- Claude Fable 5.1, handed each incident with read-only access through each platform's CLI, diagnosed at least as well and proposed supported fixes more often. Its final advice was incorrect or unsafe in 5 and 6 of its 21 runs.Claude + Edge Delta CLIClaude + Grafana CLI
The fixture, answer keys, rubric and scorer are open source at edgedelta/project-arena.
Edge Delta opened an investigation for 18 of the 21 incidents, Grafana for 12
Each native product watched the shop through its own collectors and its default monitors. An incident counts as detected when the product noticed it through its own monitoring and opened an investigation inside the observation window, with no person pointing it at the problem.1An incident the product never detects has no investigation, so it earns no credit on any later measure. The two Claude configurations were handed every incident and have no detection score.
On the 12 incidents both products investigated, Edge Delta scored higher than Grafana on every measure
Each product investigated a different set of incidents, so comparing totals would mix different work. This set keeps the 12 incident types both native products investigated, plus the control runs on those same 12. Figure 2 gives each measure its own panel.2
Root cause
out of 12- Edge Delta11/12+2 vs Grafana
- Grafana9/12
- Claude + Edge Delta CLI10/12
- Claude + Grafana CLI10/12
Blast radius
out of 12- Edge Delta9/12+1 vs Grafana
- Grafana8/12
- Claude + Edge Delta CLI10/12
- Claude + Grafana CLI10/12
Supported fix
out of 12- Edge Delta7/12+2 vs Grafana
- Grafana5/12
- Claude + Edge Delta CLI8/12
- Claude + Grafana CLI8/12
Ready to implement
out of 12- Edge Delta8/12+3 vs Grafana
- Grafana5/12
- Claude + Edge Delta CLI8/12
- Claude + Grafana CLI8/12
None of Edge Delta's 18 final recommendations was incorrect or unsafe
The rubric grades only the final report and sorts what it recommends. A candidate still needs one essential decision, such as which endpoint is intended. An incorrect or unsafe recommendation proposes a wrong or harmful action, even beside a good fix.3Seven of Edge Delta's 18 final reports ended on a candidate, and none recommended anything incorrect or unsafe. Figure 3 shows the split for every configuration.
- Supported
- Supported. A concrete, causally supported fix or safe containment.
- Partial
- Partial. Covers only some of the faults in a compound incident.
- Candidate
- Candidate. Still needs one essential causal, target or safety decision.
- No proposal
- No proposal. The final report proposed no remedy.
- Incorrect or unsafe
- Incorrect or unsafe. Actively recommends an incorrect or unsafe action, even beside a good fix.
Handed each incident, Claude scored about the same through either platform's CLI
The control gave Claude Fable 5.1 read-only access through each platform's CLI, edx for Edge Delta and gcx for Grafana, and started every run from an alert or a customer report. Across all 21 incidents it found the cause and proposed a supported fix more often than either native product, and gave incorrect or unsafe final advice more often too. On the 12 matched incidents, Edge Delta's own investigations came within one incident of it on every measure and ahead on root cause.4
| Native, started on their own | Control, started externally | |||
|---|---|---|---|---|
| measure | Edge Delta18 investigated | Grafana12 investigated | Claude + Edge Delta CLI21 investigated | Claude + Grafana CLI21 investigated |
| Detected on its own | 18/2185.7% | 12/2157.1% | not measured | not measured |
| Root cause | 15/1883.3% | 9/1275.0% | 18/2185.7% | 19/2190.5% |
| Blast radius | 12/1866.7% | 8/1266.7% | 18/2185.7% | 18/2185.7% |
| Supported fix | 8/1844.4% | 5/1241.7% | 16/2176.2% | 15/2171.4% |
| Ready to implement | 9/1656.3% | 5/1241.7% | 16/2176.2% | 15/2171.4% |
One shop, broken 21 different ways, with a wrong answer written into every key
The fixture is a Kubernetes deployment of the OpenTelemetry Demo 3.0.0 shop: frontend, cart, checkout, payment, recommendation and a product catalog, backed by Redis and PostgreSQL, with a k6 load generator and a namespace of batch workers beside it. Each scenario injects one fault, or two in the compound cases, and a verify step checks that the failure is actually observed, because an applied manifest alone does not make a valid test. Figure 4 pins each fault to the workload it hits.
Every answer key states the cause, the impact and the mitigation, and names the conclusion an investigator is likely to reach and should not. A crashing worker logs a connection error it never had; a shop full of timeout messages has nothing down. 11 of the 21 faults stay inside batch workers no shopper depends on, so calling them a shop outage costs blast-radius credit.5
Swipe sideways to see the whole shop.
One rubric, one judge, and only the final report counts
An AI judge compares each final report with the scenario's answer key under rubric v1.0.0, and must quote the report for every verdict. There is no composite score. Each measure is reported on its own, as successes over the cases it applies to.6
| measure | What earns credit |
|---|---|
| Detected on its own | The product detected the incident and started an investigation within the observation window. |
| Root cause | The final report correctly explains what caused the incident. |
| Blast radius | The final report identifies the affected workloads and downstream impact without claiming outages the evidence does not support. |
| Supported fix | The final recommendation gives a concrete, supported fix or safe containment, with no incorrect or unsafe advice left in it. |
| Ready to implement | The proposal specifies the correction and its essential details. Normal review, rollout and recovery checks may remain. |
- One rubric text and one judge setting for all four columns.
- The final report is fixed before judging. A better earlier answer cannot replace it.
- Product names are withheld from the judge, though the text may reveal them.
- All 72 judgments passed input-identity, score-format and exact-quotation checks.
- No score was rerolled or manually overridden.
What these numbers do not show
- 01Edge Delta built AI SRE Arena and ran every configuration, including Grafana's. Both products ran on their default monitors, so detection reflects what each one watches for out of the box. Anyone can rerun Grafana, or any product, with their own setup.
- 02Mitigation and readiness grade proposals. No repair was executed and no recovery was verified.
- 03Each product ran separately, and the Claude runs were handed each incident instead of detecting it. The matched set fixes which incidents are compared, not prompts, timing, or what each product had already seen.
- 04One AI judge scored every report, and no check against human graders is published yet. Requiring a quote for every verdict catches invented evidence but cannot catch a wrong call.
- 05AI SRE Arena v1 runs one demo shop on Kubernetes and measures neither speed nor cost. None of its incidents comes from a code deploy; finding the commit that broke production is what RCA Bench tests.
Run it against the product you use
AI SRE Arena needs Python 3.10, kubectl, and a disposable cluster on kind or EKS. Connect your product with its own instructions, inject a fault, let the product investigate, and save its final report as delivered. The same commands score it with the judge you choose. Use one judge and one rubric for every product you compare.
python3 -m bench init # write arena.json: context, product, judge
python3 -m bench deploy # the shop, on kind or EKS
python3 -m bench fault # inject the configured scenario
python3 -m bench verify # confirm the failure is observed
python3 -m bench archive # save final.txt, intermediate.txt, actions.txt
python3 -m bench packet # investigation + answer key + rubric
python3 -m bench judge # score with your chosen model
python3 -m bench report --summary > summary.csvCite this report
@misc{edgedelta2026arena,
title = {AI SRE Arena: an open benchmark on 21 Kubernetes incidents},
author = {{Edge Delta Research}},
year = {2026},
month = oct,
note = {Version 1, rubric v1.0.0},
url = {https://edgedelta.com/arena}
}Versions
- v1 · Oct 1, 2026Edge Delta and Grafana, native, plus two Claude control runs.
- nextMore vendors, each added as a new version when its runs are judged.
The 21 incidents
Paraphrased from each scenario's answer key. The keys themselves, with the manifests that inject each fault, are in the repository.
| incident | Fault | Wrong conclusion the key warns about |
|---|---|---|
| Workload failures | ||
| 01 crashloopbatch/report-generatorbatch only | The worker exits with an error thirty seconds after every start. | Its log says connection refused, but it never opens a connection. |
| 02 oombatch/thumbnailerbatch only | An unbounded read fills memory until the 64Mi limit kills the container. | Raising the memory limit only delays the next kill. |
| 03 slow-leakbatch/session-indexerbatch only | A cache with no retention bound grows by a MiB every minute. | Its latency warnings are scripted, and no request is actually slow. |
| 04 shm-exhaustionbatch/media-transcoderbatch only | The worker writes 128 MiB of scratch data into a 64 MiB /dev/shm. | More container memory alone does not enlarge /dev/shm. |
| 05 probe-failbatch/webhook-relaybatch only | The worker logs that it listens on port 9999, and nothing opens the port. | Moving or removing the probe does not create a listener. |
| Config and deploys | ||
| 06 image-pullbatch/asset-syncerbatch only | The deployment points at an image tag that was never published. | The registry is fine, and an arbitrary replacement image is no fix. |
| 07 missing-config-keybatch/config-loaderbatch only | The ConfigMap exists but lacks the ENDPOINT_URL key the worker requires. | Inventing an endpoint value is not a supported fix. |
| 08 stale-db-credentialsshop/product-catalogshop path | A rotation job changes the database password, and product-catalog keeps the old one. | PostgreSQL stays up; one consumer holds a stale credential. |
| Scheduling and storage | ||
| 09 pending-volume-claimbatch/archive-writerbatch only | A volume claim requests a StorageClass that does not exist. | Nodes have capacity, so scaling them changes nothing. |
| 10 volume-affinity-conflictbatch/index-storebatch only | A local volume requires a node label that no node carries. | Labeling a node that lacks the disk is not a repair. |
| 11 quota-trapshop/recommendationshop path | A new quota requires CPU requests, and the recreated recommendation pod has none. | Admission fails while node capacity is fine. |
| 12 admission-webhook-outageshop/recommendationshop path | A fail-closed admission webhook points at a service that no longer exists. | Turning off all admission enforcement would be unsafe. |
| Network policy | ||
| 13 netpol-isolationshop/cartshop path | A NetworkPolicy denies all ingress to cart, so its callers time out. | Cart stays Ready and Redis is healthy. |
| Data stores | ||
| 14 redis-pressuredatastore/redisshop path | A cache warmer fills Redis, then lowers maxmemory under a noeviction policy. | Restarting cart or flushing every key is not the fix. |
| 15 pg-lock-holddatastore/postgresqlshop path | A maintenance job holds an exclusive lock on the products table for a day. | The database still accepts connections; one transaction blocks the chain. |
| Flags and traffic | ||
| 16 payment-failureshop/paymentshop path | A feature flag turns on strict card-token checks for every charge. | No payment provider is down, and no credential needs rotating. |
| 17 cart-failureshop/cartshop path | A feature flag routes every cart operation to a secondary store that is unavailable. | The primary Redis is healthy, and restarting it changes nothing. |
| 18 traffic-floodshop/frontendshop path | A feature flag raises the load generator from 5 to 50 browsing sessions. | More traffic does not mean a broken component, or an attack. |
| Noise and compound faults | ||
| 19 log-error-burstbatch/audit-exporterbatch only | A worker logs fake gateway timeouts ten times a second and calls nothing. | No service is down; restarting a gateway would treat noise as an outage. |
| 20 composite-noise-faultbatch/asset-syncerbatch only | A failed image pull and fake gateway-timeout noise arrive together. | They are two unrelated signals, not one dependency failure. |
| 21 multi-faultdatastore/redisshop path | A crashlooping worker and Redis write rejections arrive together. | The worker does not depend on Redis; the faults only coincide. |