Edge Delta Benchmarks
Open, neutral benchmarks for AI incident reasoning and observability pipeline performance. Every model gets the same data and the same tools. We measure the reasoning.

Pipeline Performance Bench
4.6x faster
Can a pipeline handle petabytes of data?
HTTP log ingestion throughput compared across Edge Delta, Cribl, the OpenTelemetry Collector, and Fluentd under identical conditions.

Golden Threads Fine-Tune
0.95 vs 1.00 frontier
Can a small model you host yourself diagnose real incidents?
edgedelta-32B, fine-tuned on real Edge Delta investigation threads, matches a frontier model on our own incidents at a fraction of the size, running in your own environment.
Field notes
Same bug, four minds
One production panic paged four frontier-model agent teams at the same instant. Same diagnosis, same one-line fix, four different definitions of done. A behavioral teardown of a single real incident, model by model.
Read the teardown →Field notes
Small model, home turf
We fine-tuned Qwen3-32B on golden threads distilled from real Edge Delta investigations, then scored it against its base, a 235B model, and the GPT-5.5 frontier on 50 root-cause scenarios. A case study in specializing a small model for incident diagnosis.
Read the case study →

