Edge Delta Research · Benchmark report

v1 · October 2026

AI SRE Arena

An open benchmark for detection, diagnosis and mitigation on 21 Kubernetes incidents, judged against one answer key

Edge Delta Research · Published October 2026 · Rubric v1.0.0

Incidents
21 injected faults
Configurations
2 native, 2 control
Judged investigations
72
Judge
GPT-6-Astra

Abstract

We injected 21 faults into a Kubernetes deployment of the OpenTelemetry demo shop. Each product had to detect and investigate them on its default monitors, and an AI judge scored every final report against a fixed answer key and a published rubric.

Detected on its own18 vs 12
Edge Delta
Grafana
Edge Delta detected 18 of the 21 incidents through its own monitoring and opened an investigation without anyone pointing it at the problem. Grafana detected 12.
Root cause11 vs 9
Edge Delta
Grafana
On the 12 incidents both products investigated, Edge Delta named the right root cause in 11 and Grafana in 9.
Supported fix7 vs 5
Edge Delta
Grafana
On the same 12, Edge Delta proposed a supported fix in 7 and Grafana in 5.
Unsafe advice0 of 18
Edge Delta
Grafana
None of Edge Delta's 18 final recommendations was incorrect or unsafe, against 2 of Grafana's 12.
Control run5 and 6
Claude + Edge Delta CLI
Claude + Grafana CLI
Claude Fable 5.1, handed each incident with read-only access through each platform's CLI, diagnosed at least as well and proposed supported fixes more often. Its final advice was incorrect or unsafe in 5 and 6 of its 21 runs.

The fixture, answer keys, rubric and scorer are open source at edgedelta/project-arena.

1Results

Edge Delta opened an investigation for 18 of the 21 incidents, Grafana for 12

Each native product watched the shop through its own collectors and its default monitors. An incident counts as detected when the product noticed it through its own monitoring and opened an investigation inside the observation window, with no person pointing it at the problem.1An incident the product never detects has no investigation, so it earns no credit on any later measure. The two Claude configurations were handed every incident and have no detection score.

Edge Deltanative AI investigations18/21
Grafananative AI investigations12/21
Claude + Edge Delta CLIedx, started externallyStarted from outside the product for all 21 incidents: 16 from an alert, 5 from a customer report.not measured
Claude + Grafana CLIgcx, started externallyStarted from outside the product for all 21 incidents: 12 from an alert, 9 from a customer report.not measured
Figure 1. Incidents each product detected and started investigating on its own, out of 21. Each cell is one incident, filled cells first; the cells do not map to specific incidents.

On the 12 incidents both products investigated, Edge Delta scored higher than Grafana on every measure

Each product investigated a different set of incidents, so comparing totals would mix different work. This set keeps the 12 incident types both native products investigated, plus the control runs on those same 12. Figure 2 gives each measure its own panel.2

Root cause

out of 12
  • Edge Delta11/12+2 vs Grafana
  • Grafana9/12
  • Claude + Edge Delta CLI10/12
  • Claude + Grafana CLI10/12

Blast radius

out of 12
  • Edge Delta9/12+1 vs Grafana
  • Grafana8/12
  • Claude + Edge Delta CLI10/12
  • Claude + Grafana CLI10/12

Supported fix

out of 12
  • Edge Delta7/12+2 vs Grafana
  • Grafana5/12
  • Claude + Edge Delta CLI8/12
  • Claude + Grafana CLI8/12

Ready to implement

out of 12
  • Edge Delta8/12+3 vs Grafana
  • Grafana5/12
  • Claude + Edge Delta CLI8/12
  • Claude + Grafana CLI8/12
Figure 2. The 12 matched incidents, one panel per measure. The two native products are on top, with Edge Delta’s lead over Grafana under its score. Below the rule, dashed: Claude Fable 5.1, handed each incident through the Edge Delta or Grafana CLI.

None of Edge Delta's 18 final recommendations was incorrect or unsafe

The rubric grades only the final report and sorts what it recommends. A candidate still needs one essential decision, such as which endpoint is intended. An incorrect or unsafe recommendation proposes a wrong or harmful action, even beside a good fix.3Seven of Edge Delta's 18 final reports ended on a candidate, and none recommended anything incorrect or unsafe. Figure 3 shows the split for every configuration.

Edge Delta18 judged8720 of 18unsafe
Grafana12 judged5522 of 12unsafe
Claude + Edge Delta CLI21 judged1655 of 21unsafe
Claude + Grafana CLI21 judged1566 of 21unsafe
Supported
Supported. A concrete, causally supported fix or safe containment.
Partial
Partial. Covers only some of the faults in a compound incident.
Candidate
Candidate. Still needs one essential causal, target or safety decision.
No proposal
No proposal. The final report proposed no remedy.
Incorrect or unsafe
Incorrect or unsafe. Actively recommends an incorrect or unsafe action, even beside a good fix.
Figure 3. What each final recommendation amounted to, as a share of that configuration's judged investigations. Earlier advice is scored separately and is not shown.

Handed each incident, Claude scored about the same through either platform's CLI

The control gave Claude Fable 5.1 read-only access through each platform's CLI, edx for Edge Delta and gcx for Grafana, and started every run from an alert or a customer report. Across all 21 incidents it found the cause and proposed a supported fix more often than either native product, and gave incorrect or unsafe final advice more often too. On the 12 matched incidents, Edge Delta's own investigations came within one incident of it on every measure and ahead on root cause.4

AI SRE Arena results across all investigated incidents, by configuration.
Native, started on their ownControl, started externally
measureEdge Delta18 investigatedGrafana12 investigatedClaude + Edge Delta CLI21 investigatedClaude + Grafana CLI21 investigated
Detected on its own18/2185.7%12/2157.1%not measurednot measured
Root cause15/1883.3%9/1275.0%18/2185.7%19/2190.5%
Blast radius12/1866.7%8/1266.7%18/2185.7%18/2185.7%
Supported fix8/1844.4%5/1241.7%16/2176.2%15/2171.4%
Ready to implement9/1656.3%5/1241.7%16/2176.2%15/2171.4%
Table 1. Every completed investigation in each column, as successes over applicable cases. Readiness excludes two Edge Delta investigations that proposed no mitigation, hence 16.
2The fixture

One shop, broken 21 different ways, with a wrong answer written into every key

The fixture is a Kubernetes deployment of the OpenTelemetry Demo 3.0.0 shop: frontend, cart, checkout, payment, recommendation and a product catalog, backed by Redis and PostgreSQL, with a k6 load generator and a namespace of batch workers beside it. Each scenario injects one fault, or two in the compound cases, and a verify step checks that the failure is actually observed, because an applied manifest alone does not make a valid test. Figure 4 pins each fault to the workload it hits.

Every answer key states the cause, the impact and the mitigation, and names the conclusion an investigator is likely to reach and should not. A crashing worker logs a connection error it never had; a shop full of timeout messages has nothing down. 11 of the 21 faults stay inside batch workers no shopper depends on, so calling them a shop outage costs blast-radius credit.5

Swipe sideways to see the whole shop.

shopdatastoreplatform-opsbatchk6load generatorfrontendcartcheckoutrecommendationpaymentproduct-catalogflagdfeature flagsredispostgresqlpod-policyservice absentreport-generatorthumbnailersession-indexermedia-transcoderwebhook-relayasset-syncerconfig-loaderarchive-writerindex-storeaudit-exportercache-warmerfills rediscatalog-maint.locks productscredential-rotatorrotates password0102030405060708091011121314151617181920202121
Figure 4. The fixture, with each injected fault pinned to the workload its answer key names. Compound faults pin twice.All 21 in Appendix A
3Scoring

One rubric, one judge, and only the final report counts

An AI judge compares each final report with the scenario's answer key under rubric v1.0.0, and must quote the report for every verdict. There is no composite score. Each measure is reported on its own, as successes over the cases it applies to.6

What earns credit on each AI SRE Arena measure.
measureWhat earns credit
Detected on its ownThe product detected the incident and started an investigation within the observation window.
Root causeThe final report correctly explains what caused the incident.
Blast radiusThe final report identifies the affected workloads and downstream impact without claiming outages the evidence does not support.
Supported fixThe final recommendation gives a concrete, supported fix or safe containment, with no incorrect or unsafe advice left in it.
Ready to implementThe proposal specifies the correction and its essential details. Normal review, rollout and recovery checks may remain.
  • One rubric text and one judge setting for all four columns.
  • The final report is fixed before judging. A better earlier answer cannot replace it.
  • Product names are withheld from the judge, though the text may reveal them.
  • All 72 judgments passed input-identity, score-format and exact-quotation checks.
  • No score was rerolled or manually overridden.
4Limitations

What these numbers do not show

  1. 01Edge Delta built AI SRE Arena and ran every configuration, including Grafana's. Both products ran on their default monitors, so detection reflects what each one watches for out of the box. Anyone can rerun Grafana, or any product, with their own setup.
  2. 02Mitigation and readiness grade proposals. No repair was executed and no recovery was verified.
  3. 03Each product ran separately, and the Claude runs were handed each incident instead of detecting it. The matched set fixes which incidents are compared, not prompts, timing, or what each product had already seen.
  4. 04One AI judge scored every report, and no check against human graders is published yet. Requiring a quote for every verdict catches invented evidence but cannot catch a wrong call.
  5. 05AI SRE Arena v1 runs one demo shop on Kubernetes and measures neither speed nor cost. None of its incidents comes from a code deploy; finding the commit that broke production is what RCA Bench tests.
5Reproduce

Run it against the product you use

AI SRE Arena needs Python 3.10, kubectl, and a disposable cluster on kind or EKS. Connect your product with its own instructions, inject a fault, let the product investigate, and save its final report as delivered. The same commands score it with the judge you choose. Use one judge and one rubric for every product you compare.

python3 -m bench init        # write arena.json: context, product, judge
python3 -m bench deploy      # the shop, on kind or EKS
python3 -m bench fault       # inject the configured scenario
python3 -m bench verify      # confirm the failure is observed
python3 -m bench archive     # save final.txt, intermediate.txt, actions.txt
python3 -m bench packet      # investigation + answer key + rubric
python3 -m bench judge       # score with your chosen model
python3 -m bench report --summary > summary.csv

Cite this report

@misc{edgedelta2026arena,
  title  = {AI SRE Arena: an open benchmark on 21 Kubernetes incidents},
  author = {{Edge Delta Research}},
  year   = {2026},
  month  = oct,
  note   = {Version 1, rubric v1.0.0},
  url    = {https://edgedelta.com/arena}
}

Versions

  1. v1 · Oct 1, 2026Edge Delta and Grafana, native, plus two Claude control runs.
  2. nextMore vendors, each added as a new version when its runs are judged.
AAppendix

The 21 incidents

Paraphrased from each scenario's answer key. The keys themselves, with the manifests that inject each fault, are in the repository.

The 21 AI SRE Arena incidents, the fault each injects, and the wrong conclusion its answer key warns about.
incidentFaultWrong conclusion the key warns about
Workload failures
01 crashloopbatch/report-generatorbatch onlyThe worker exits with an error thirty seconds after every start.Its log says connection refused, but it never opens a connection.
02 oombatch/thumbnailerbatch onlyAn unbounded read fills memory until the 64Mi limit kills the container.Raising the memory limit only delays the next kill.
03 slow-leakbatch/session-indexerbatch onlyA cache with no retention bound grows by a MiB every minute.Its latency warnings are scripted, and no request is actually slow.
04 shm-exhaustionbatch/media-transcoderbatch onlyThe worker writes 128 MiB of scratch data into a 64 MiB /dev/shm.More container memory alone does not enlarge /dev/shm.
05 probe-failbatch/webhook-relaybatch onlyThe worker logs that it listens on port 9999, and nothing opens the port.Moving or removing the probe does not create a listener.
Config and deploys
06 image-pullbatch/asset-syncerbatch onlyThe deployment points at an image tag that was never published.The registry is fine, and an arbitrary replacement image is no fix.
07 missing-config-keybatch/config-loaderbatch onlyThe ConfigMap exists but lacks the ENDPOINT_URL key the worker requires.Inventing an endpoint value is not a supported fix.
08 stale-db-credentialsshop/product-catalogshop pathA rotation job changes the database password, and product-catalog keeps the old one.PostgreSQL stays up; one consumer holds a stale credential.
Scheduling and storage
09 pending-volume-claimbatch/archive-writerbatch onlyA volume claim requests a StorageClass that does not exist.Nodes have capacity, so scaling them changes nothing.
10 volume-affinity-conflictbatch/index-storebatch onlyA local volume requires a node label that no node carries.Labeling a node that lacks the disk is not a repair.
11 quota-trapshop/recommendationshop pathA new quota requires CPU requests, and the recreated recommendation pod has none.Admission fails while node capacity is fine.
12 admission-webhook-outageshop/recommendationshop pathA fail-closed admission webhook points at a service that no longer exists.Turning off all admission enforcement would be unsafe.
Network policy
13 netpol-isolationshop/cartshop pathA NetworkPolicy denies all ingress to cart, so its callers time out.Cart stays Ready and Redis is healthy.
Data stores
14 redis-pressuredatastore/redisshop pathA cache warmer fills Redis, then lowers maxmemory under a noeviction policy.Restarting cart or flushing every key is not the fix.
15 pg-lock-holddatastore/postgresqlshop pathA maintenance job holds an exclusive lock on the products table for a day.The database still accepts connections; one transaction blocks the chain.
Flags and traffic
16 payment-failureshop/paymentshop pathA feature flag turns on strict card-token checks for every charge.No payment provider is down, and no credential needs rotating.
17 cart-failureshop/cartshop pathA feature flag routes every cart operation to a secondary store that is unavailable.The primary Redis is healthy, and restarting it changes nothing.
18 traffic-floodshop/frontendshop pathA feature flag raises the load generator from 5 to 50 browsing sessions.More traffic does not mean a broken component, or an attack.
Noise and compound faults
19 log-error-burstbatch/audit-exporterbatch onlyA worker logs fake gateway timeouts ten times a second and calls nothing.No service is down; restarting a gateway would treat noise as an outage.
20 composite-noise-faultbatch/asset-syncerbatch onlyA failed image pull and fake gateway-timeout noise arrive together.They are two unrelated signals, not one dependency failure.
21 multi-faultdatastore/redisshop pathA crashlooping worker and Redis write rejections arrive together.The worker does not depend on Redis; the faults only coincide.

Test your AI SRE on the same 21 incidents.

The fixture, the 21 answer keys, the rubric and the scorer are open. Run them against the product you use today, or start Edge Delta's agents on your own cluster.