Edge Delta Research · Buyer's guide
Revision 1 · October 2, 2026
The best AI SRE tools in 2026
15 products, compared from their vendors' own documentation on triggers, data access, permissions and price
Edge Delta Research · Published October 2, 2026 · Sources retrieved September 30 to October 2, 2026
- Products
- 15
- Sources
- 61
- Retrieved
- Oct 2026
- Revision
- 1
Summary
We read the product pages, documentation, pricing pages and release notes for 15 AI SRE products from September 30 to October 2, 2026, and recorded the same facts for each one, covering triggers, data access, permissions, price and the capabilities the vendor documents.
Definition. An AI SRE is an AI agent that performs the work of a site reliability engineer: it detects production issues, investigates them across logs, metrics, traces, and deployment context, identifies the root cause with supporting evidence, and remediates with approval and verification.
- Start without an alert9 of 15
- AWS DevOps Agent, Datadog Bits Investigation, Dynatrace SRE Agent, Edge Delta, Grafana Assistant Investigations, Microsoft Azure SRE Agent, New Relic Autopilot, Resolve AI and Traversal can begin from their own detection, a watcher or a schedule. The other 6 start when an alert rule fires, an incident is declared, or a person asks.
- Publish a price7 of 15
- AWS DevOps Agent, Datadog Bits Investigation, Dynatrace SRE Agent, Edge Delta, Grafana Assistant Investigations, Microsoft Azure SRE Agent and PagerDuty SRE Agent publish what the feature costs, in units that include AI credits, tokens, agent-seconds, Azure Agent Units, AI Actions, workflow-hours and a monthly plan with included credits. The rest sell it through sales or as an add-on without a public price, and Google's investigations are a free preview that needs Premium Support.
- Check the fix8 of 15
- AWS DevOps Agent, Datadog Bits Investigation, Edge Delta, Grafana Assistant Investigations, Microsoft Azure SRE Agent, PagerDuty SRE Agent, Splunk AI SRE and Traversal document checking whether the original problem was resolved after a change. The other vendors' documentation stops at recommending or making the change.
Edge Delta publishes this guide and is one of the 15 products in it. Entries are alphabetical, and each one links to the vendor pages it was written from.
Triggers, data access, permissions and price, product by product
Each cell restates what the vendor's own product pages, documentation and pricing pages say. Where a vendor publishes nothing on a point, the cell says that instead of guessing.
| Product | Starts from | Reads | May change | Pricing |
|---|---|---|---|---|
| AWS DevOps AgentGA Mar 2026 | Webhooks and integrations from ServiceNow, PagerDuty, Datadog, Dynatrace, New Relic, Splunk, Grafana and Slack, a manual request, or a scheduled custom agent | AWS accounts across Regions, Azure resources, CloudWatch, Datadog, Dynatrace, Splunk, New Relic, Grafana, GitHub, GitLab and your own MCP servers | Read-only by default; opt-in directed actions need an operator to approve each change in chat, and destructive tools are never called | $0.0083 per agent-second of active work, nothing while idle; a two-month free trial for new customers |
| Datadog Bits InvestigationGA Dec 2025 | Monitors with Auto-Investigate turned on, when they enter alert state, including monitors that Bits Detection creates and tunes itself in preview; a person can start one from Slack, a Synthetic test or a prompt, and workflows can trigger one | Datadog metrics, traces, logs, events, change tracking, RUM, profiles and GitHub code; Grafana, Dynatrace, Sentry and Splunk data in preview | Triage actions from chat; code fixes from Bits Code as pull requests a person opens; infrastructure actions such as Kubernetes patches, in preview, which guardrails can let run automatically, hold for approval or block | AI credits, about 6.5 per autonomous investigation; from $500 for 500 credits a month billed annually, or $1.30 a credit on demand |
| Dynatrace SRE AgentPreview | Error and availability problems that Dynatrace's causal AI has detected and traced to a root-cause entity; workflows can also run on events or a schedule | Logs, events, metrics and traces in Dynatrace, queried with DQL, plus the details of the problem it is working on | Writes its analysis and recommended fixes onto the problem; approved remediation, such as a rollback, runs through Dynatrace Workflows | No charge for the agentic AI yet; the workflow and the DQL queries it runs bill at rate-card prices, such as $0.03 per workflow-hour |
| Edge DeltaAvailable Oct 2025 | Its own monitors, including learned baselines on live telemetry, and scheduled loops; PagerDuty alerts and chat requests also start one | Logs, metrics, traces and events flowing through Edge Delta Telemetry Pipelines, plus connected tools such as Elastic, Kubernetes, AWS, Azure, GCP, GitHub, Jira and Linear | Starts at the Propose trust level: it reads freely and drafts changes, every write waits for a person's approval or a playbook you pre-approved, and destructive actions stay off until you raise the level | Free 14-day trial without a credit card; Pro from $20 a month, including $20 of credits, with unlimited investigations and spending limits |
| Google Gemini Cloud Assist investigationsPreview | A person starting one from the Investigations page, Logs Explorer, a Cloud Monitoring alert, chat or a product page; Proactive Mode for alerts is in private preview | Logs, metrics and configuration for Google Cloud resources in one project or App Hub application at a time, across 18 supported products | Nothing during the investigation; where a fix is supported it generates gcloud commands or Kubernetes manifests that the user reviews and runs | No charge during preview, listed under Gemini Code Assist Enterprise; no GA price published |
| Grafana Assistant InvestigationsGA Jul 2026 | An alert rule with an investigation attached, an incident or alert group in Grafana IRM, a Watcher's critical finding in preview, or a person in the app, Slack or the gcx CLI | Metrics, logs, traces and profiles in Grafana Cloud; code and deployment changes through a connected MCP server such as GitHub | Ends in a report with recommended next steps; write-capable MCP tools follow per-tool approval settings, and coding agents that open pull requests are in private preview | $2 per million tokens beyond included allowances, no per-investigation fee; Pro adds a $19 monthly platform fee and $20 per active AI user beyond three |
| Harness AI SREEnterprise plan | Alerts sent in by webhook and promoted to incidents by rules; incidents opened in the app or with /harness new in Slack | Deployments, pull requests, ServiceNow change records, alert data, and metrics and traces from affected services, plus Slack, Teams and Zoom conversations through AI Scribe; Datadog, Splunk, New Relic and Dynatrace send alerts by webhook | Runbooks that roll back, scale out or toggle a feature flag, started by a trigger or a responder; custom scripts run as Harness pipelines | Not published; sold as a module on the Enterprise plan |
| incident.io InvestigationsGA Aug 2026 | A declared incident, automatically or on conditions you set, or /inc investigate; an alert has to become an incident first | Telemetry queried on demand from tools such as Datadog, Grafana, New Relic, Splunk, Elasticsearch and AWS CloudWatch, plus GitHub or GitLab code and past incidents | Pull requests a person reviews and merges, and fields on the incident itself; it never merges or deploys | Paid add-on for the Pro and Enterprise plans; the add-on price is not published |
| Microsoft Azure SRE AgentGA Mar 2026 | Incidents from Azure Monitor, PagerDuty or ServiceNow, scheduled tasks, HTTP and log-query triggers, and chat | Azure resources, Azure Monitor, Application Insights, Log Analytics and Kusto; Datadog, Splunk, New Relic, Dynatrace and Elasticsearch through partner MCP connectors | Azure mitigations under its managed identity; in Review mode an administrator approves Azure infrastructure writes, in Autonomous mode they apply directly | $0.40 per agent-hour always on in US regions, plus token-metered usage, both billed in Azure Agent Units at $0.10 each |
| New Relic AutopilotGA Jul 2026 | An alert workflow with an AI agent destination, a question in the New Relic AI panel or Slack, or a Workflow Automation step, including after a deploy or on a schedule | New Relic telemetry across the accounts you select, Confluence and uploaded documents, Slack threads, and GitHub through a read-only MCP connection in preview | Recommends actions without taking them; remediation runs through New Relic Workflow Automation, which can pause for an approve or deny link | Not published; requires a Pro or Enterprise edition with the Advanced Compute add-on, billed by compute units |
| PagerDuty SRE AgentGA Oct 2025 | A responder asking on the incident, in the Operations Console or in Slack; in early access, an incident workflow or escalation policy can add it as a virtual responder that starts when an incident triggers | Incident history, runbooks and its own memory, plus 19 documented connectors including Datadog, Dynatrace, Grafana, New Relic, Splunk, AWS CloudWatch and GitHub | Recommends an incident workflow and runs it when a responder clicks Run, each workflow at most once per incident; a fully autonomous responder is announced for early access in the second half of 2026 | Included in PD Reliability Platform tiers from $2,800 a year and sold as an add-on for the Professional plan; each request uses 4 AI Actions |
| Resolve AISales-led | Alerts from PagerDuty and other alerting tools as they fire; background agents also run on a schedule or on events such as deploys | Observability, cloud, Kubernetes, code and CI/CD tools through 60+ integrations, with an optional Satellite, run on Kubernetes or AWS ECS, for private tools and Kubernetes data | Proposes alert silences and pull requests for a person to approve; in preview it can also dispatch, retry or cancel GitHub Actions runs, which wait for approval unless an admin sets the tool to Auto | Credits per chat, triage, investigation, incident or task, at rates set in your contract; no public rates |
| Rootly AI SREOn request | Alerts that match an auto-run investigation rule, account-level automatic investigation, or a person in the app or through @Rootly in Slack | Datadog, Grafana, New Relic, Honeycomb, Sentry, Dynatrace and Splunk, plus AWS, Azure, GCP, GitHub, GitLab and custom MCP tools | Reports a root cause and next steps; its GitHub and GitLab connections are read-only, and its docs warn that other write-capable connector tools run without an approval step | Not published for AI SRE; Rootly Incident Response and On-Call each start at $20 per user a month |
| Splunk AI SREGA Jun 2026 | A person opening a supported APM or Kubernetes alert, once the organization has opted in through Splunk Support, or choosing Run root cause analysis on the alert | Metrics, events, logs and traces in Splunk Observability Cloud; custom metrics are not supported | Nothing directly; remediation plans for Kubernetes alerts list commands an engineer runs, and an alpha built on Claude Managed Agents opens pull requests for review | Not published; Splunk directs customers to their representative |
| TraversalWorkers GA Sep 2026 | Incidents and alerts in Slack or Teams channels, where Workers decide when to step in; Custom Workers follow a written mission, on a schedule if you set one | Datadog, Dynatrace, Grafana, Splunk, Prometheus, Elasticsearch, CloudWatch, Sentry and others through read-only APIs, plus GitHub or GitLab; Kubernetes through a connector you run | Nothing directly, according to its docs; it recommends fixes and, over MCP, hands next steps to a coding agent you run. Its marketing describes self-healing remediation that the docs do not cover | Not published; sized to your environment after a walkthrough |
| Product | Starts on its own | Approval step | Checks the fix | Uses past incidents | Public price |
|---|---|---|---|---|---|
| AWS DevOps Agent | Documented | Documented | Documented | Documented | Documented |
| Datadog Bits Investigation | Documented | Documented | Documented | Documented | Documented |
| Dynatrace SRE Agent | Documented | Documented | Not found in public documentation | Not found in public documentation | Documented |
| Edge Delta | Documented | Documented | Documented | Documented | Documented |
| Google Gemini Cloud Assist investigations | Not found in public documentation | Documented | Not found in public documentation | Not found in public documentation | Not found in public documentation |
| Grafana Assistant Investigations | Documented | Documented | Documented | Documented | Documented |
| Harness AI SRE | Not found in public documentation | Documented | Not found in public documentation | Documented | Not found in public documentation |
| incident.io Investigations | Not found in public documentation | Documented | Not found in public documentation | Documented | Not found in public documentation |
| Microsoft Azure SRE Agent | Documented | Documented | Documented | Documented | Documented |
| New Relic Autopilot | Documented | Documented | Not found in public documentation | Documented | Not found in public documentation |
| PagerDuty SRE Agent | Not found in public documentation | Documented | Documented | Documented | Documented |
| Resolve AI | Documented | Documented | Not found in public documentation | Documented | Not found in public documentation |
| Rootly AI SRE | Not found in public documentation | Not found in public documentation | Not found in public documentation | Documented | Not found in public documentation |
| Splunk AI SRE | Not found in public documentation | Documented | Documented | Documented | Not found in public documentation |
| Traversal | Documented | Documented | Documented | Documented | Not found in public documentation |
- Starts on its own
- Can begin from a problem it detects itself, a watcher or a schedule, without an alert rule, a declared incident or a person asking.
- Approval step
- Documents that a person approves or carries out each change, or a setting that holds changes for approval, before anything in production changes.
- Checks the fix
- Documents checking, after a change, whether the original problem was resolved.
- Uses past incidents
- Draws on earlier incidents or a memory of past investigations when it works on a new one.
- Public price
- Publishes what the AI SRE feature costs, as rates, plan prices that include it, or usage units.
15 products, from their vendors' documentation
AWS DevOps Agent
Amazon Web Services · cloud provider
Generally available since March 31, 2026, after a preview from December 2025
AWS DevOps Agent investigates incidents that arrive from ticketing, paging and observability tools or from a manual request, across AWS accounts, Azure resources and connected tools. It produces a mitigation plan or a spec that a coding agent such as Kiro can implement. AWS previewed it in December 2025 and made it generally available on March 31, 2026.
Changes are off by default. With directed actions enabled, each approval covers one tool, operation and resource, for a single use or up to four hours, and CloudTrail records who approved it. A change attempted during an autonomous investigation fails instead of asking. Memory stores keep recurring root causes per alarm, standing directives and lessons from corrective feedback.
- Worth checking
- Cost follows agent time, about $30 per agent-hour, and AWS's own example of ten 8-minute investigations comes to $39.84. Queries it runs against services such as CloudWatch Logs Insights are billed separately, and each Agent Space runs three investigations at once by default.
Sources: AWS DevOps Agent, User guide, Directed actions, Pricing, GA announcement
Datadog Bits Investigation
Datadog · observability platform
Generally available since December 2025, launched as Bits AI SRE
Datadog's investigation agent forms hypotheses about a root cause and queries Datadog telemetry to confirm or rule out each one, then posts its findings to the Slack, case or on-call destinations already set on the monitor. Datadog made it generally available on December 2, 2025 as Bits AI SRE and renamed it Bits Investigation in mid-2026, alongside Bits Code and Bits Chat.
Datadog's March 2026 update put a typical investigation at three to four minutes. By default each monitor starts at most one automatic investigation in a rolling 24 hours, a limit admins can change, and warn, no-data, renotification and test events do not start one. Bits Infrastructure Operations, in preview, applies infrastructure fixes automatically where a guardrail allows it, for example in staging, and asks for approval elsewhere.
- Worth checking
- Investigations, chat messages (about 0.5 credits each) and code fixes (about 5) all draw on the same monthly credit bundle, and Datadog describes these figures as averages. Unused credits do not roll over, overage bills at the on-demand rate, and one-click infrastructure actions, guardrails, Bits Detection, Bits Infrastructure Operations and automatic memories are still in preview.
Sources: Bits Investigation docs, Bits Remediation docs, AI credits pricing, GA announcement, Bits Infrastructure Operations. Full comparison with Edge Delta
Dynatrace SRE Agent
Dynatrace · observability platform
Preview; the SRE Agent workflow template shipped in late July 2026, inside Dynatrace Intelligence, which launched in January 2026
At Perform in January 2026, Dynatrace introduced Dynatrace Intelligence, the successor to Davis AI, and added domain agents and agentic workflows to its platform. The SRE Agent, a workflow template that shipped in late July 2026, is one of eleven ready-made agentic workflows in the documentation, all marked preview. It picks up a problem after Dynatrace has run its own root-cause analysis, queries logs, events, metrics and traces with DQL, and writes its analysis and recommendations back onto the problem.
The default workflow is capped at five tool calls and a 1,000-word answer, and it does not apply a fix by itself. In July 2026 Dynatrace announced an Autonomous SRE Agent and a Cloud SRE Agent. The Cloud SRE Agents app hands problems to agents from AWS, Microsoft and Google, and its Hub listing describes it as a community-supported project.
- Worth checking
- The SRE Agent is marked preview in the documentation and coming soon in the Dynatrace Hub, so confirm what your tenant can use. We found no release note confirming that the Autonomous SRE Agent shipped after its August target.
Sources: Dynatrace Intelligence docs, SRE Agent in the Dynatrace Hub, January 2026 announcement, July 2026 announcement, SRE Agent docs, Agentic AI licensing FAQ, Rate card
Edge Delta
Edge Delta · telemetry pipelines and AI SRE
Available since October 2025, with a self-serve trial
Edge Delta is a telemetry-native AI SRE that continuously understands, investigates, and operates production. Its AI Teammates (SRE, Security Engineer, Software Engineer and Work Tracker) reason over the logs, metrics, traces and events moving through Edge Delta's Telemetry Pipelines, so baselines and service context already exist when an incident starts.
An investigation ends with a root cause and its evidence, and proposed fixes wait in an approvals queue. After a fix runs, the teammate compares telemetry from before and after the change, and the incident is saved to Production Memory for later investigations.
- Worth checking
- Detection on live data and the fullest context depend on telemetry passing through Edge Delta pipelines; data that stays in another backend is reached through connectors, so the teammates see what those tools return. Pro usage beyond the included $20 draws on credits whose cost per investigation is not published, and Pro keeps data and memory for 30 days.
Sources: Edge Delta AI SRE, Edge Delta pricing
Google Gemini Cloud Assist investigations
Google Cloud · cloud provider
Preview since June 2025; new investigations need Premium Support or account-team access since April 2026
The investigations feature in Gemini Cloud Assist analyzes logs, metrics and configuration for Google Cloud resources and returns root-cause hypotheses with suggested next steps. It entered public preview in June 2025, and since April 10, 2026 creating or running an investigation requires a Premium Support contract or access through your Google Cloud account team.
Investigations run with the permissions of the person who starts them, and the documentation states that this access is never used to change data. Proactive Mode, in private preview for Premium Support customers, investigates Cloud Alerting alerts and cost anomalies in the background and publishes results as Eventarc events.
- Worth checking
- It covers Google Cloud only, one project or App Hub application per investigation. Google advises against using it on data with residency requirements, because investigation data may be stored in any Google Cloud data center, and investigations stopped working inside VPC Service Controls perimeters on April 13, 2026.
Sources: Investigations docs, Release notes, Gemini Cloud Assist, Gemini pricing
Grafana Assistant Investigations
Grafana Labs · observability platform
Generally available since July 2026
Grafana Assistant became generally available in Grafana Cloud in October 2025, and its Investigations mode followed on July 29, 2026. An investigation explores metrics, logs, traces and profiles, builds hypotheses, and returns a report with findings, supporting evidence and recommended next steps.
By default, alert enrichment opens at most 50 new investigations every 20 minutes, and changes to the same alert group within six hours continue the existing investigation. An investigation can schedule its own re-checks, for example whether an error rate has recovered, and close them once the situation resolves. Metering for investigations began on October 1, 2026.
- Worth checking
- Investigations need a separate entitlement and are not available on self-managed Grafana, even when it connects to the Cloud Assistant backend. MCP tools set to auto-approve can write to or delete data in connected services without asking.
Sources: Investigations docs, GA release note, Grafana pricing, MCP server settings
Harness AI SRE
Harness · software delivery platform
Sold as an Enterprise plan module; no GA date published
Harness AI SRE combines incident management and on-call with AI agents. AI Scribe records what happens in Slack, Teams and Zoom calls, the RCA Change Agent ranks recent deployments, pull requests and ServiceNow change records as likely causes, and runbooks can roll back a deployment by running a Harness pipeline, scale out, or flip a feature flag.
On-call schedules, escalation policies, paging and a mobile app sit behind a feature flag, with import from PagerDuty, Opsgenie and xMatters. Alerts are enriched with similar past incidents and how they were resolved, AI post-mortems can be drafted when an incident closes, and ServiceNow change records are polled every five minutes.
- Worth checking
- Runbooks on automatic triggers run without a person, and Harness recommends running production rollbacks manually. Runbooks exposed to the AI investigator, which run automatically or with a person in the loop to feed root-cause analysis, are in early access, and correlation with feature flag, infrastructure and config changes sits behind a feature flag.
Sources: Harness AI SRE, Harness pricing
incident.io Investigations
incident.io · incident management
Generally available since August 5, 2026
Investigations is incident.io's AI SRE, sold as an add-on to its incident management platform. When an incident is declared, it queries connected telemetry, code and past incidents, then posts a root-cause hypothesis with a confidence level and linked evidence. It runs on Nexus, which incident.io calls its production intelligence model; each customer gets a separate instance, and the underlying models come from OpenAI, Anthropic and Google under zero-data-retention terms.
In September 2026 incident.io reported that the median time to the first accurate message in an incident channel fell from 6.7 to 3 minutes. An escalation step can hold alert-triggered paging until the first hypothesis exists, up to a time limit you set.
- Worth checking
- It works only on incidents, so an alert reaches it only after a workflow or a person turns it into one. The product page says sensitive data is redacted before it reaches model providers, while the documentation describes pattern-based redaction that stays off until you ask incident.io to enable it.
Sources: incident.io AI SRE, Triggering docs, incident.io pricing, GA changelog
Microsoft Azure SRE Agent
Microsoft · cloud provider
Generally available since March 10, 2026, after a preview from May 2025
Microsoft's agent takes incidents from Azure Monitor, PagerDuty or ServiceNow, runs scheduled and webhook-triggered tasks, and investigates across Azure and connected tools before proposing or applying a mitigation. Microsoft previewed it at Build in May 2025 and made it generally available on March 10, 2026.
By default the managed identity gets Reader and Monitoring Contributor roles, with a Privileged level offered at creation; delete and Key Vault commands are blocked outright, and administrators can add tool policies and hooks on top. Review mode pauses only Azure infrastructure writes, and new incident response plans and scheduled tasks default to Autonomous, including the quickstart plan it creates for high-severity alerts.
- Worth checking
- The always-on charge runs from creation until deletion, even while the agent is stopped, which comes to about $292 a month per agent in US regions before any investigation. A 30-day trial for new customers waives it for up to three agents.
Sources: Azure SRE Agent overview, Azure SRE Agent pricing, GA announcement
New Relic Autopilot
New Relic · observability platform
Generally available since July 2026; previewed as SRE Agent from February 2026
New Relic previewed an SRE Agent in February 2026 and made it generally available as Autopilot in July 2026. It attaches an analysis to alerts routed to it through a notification workflow and answers questions in the New Relic AI panel or in Slack, using telemetry from the New Relic accounts you select.
Memories, added in July 2026, are off by default and kept for 90 days unless you extend them to as long as a year. Each organization runs one Autopilot, limited to 100 requests an hour, and an investigation looks back over the last 24 hours by default.
- Worth checking
- Autopilot does not create alerts, so its coverage follows the alert conditions you already maintain in New Relic. The GitHub connection is in public preview, and Jira, named in the launch announcement, does not appear in the documentation.
Sources: New Relic Autopilot, Autopilot docs, GA release note, Memories docs
PagerDuty SRE Agent
PagerDuty · incident management
Generally available since October 30, 2025; called Paige in current docs
PagerDuty's SRE Agent, now called Paige in its documentation and pricing, works inside the incident. It pulls logs, metrics, runbooks and incident history from connected tools, suggests likely causes, and recommends an incident workflow for a responder to run. PagerDuty made it generally available on October 30, 2025.
Memory is kept per service. Paige saves what it learned from an incident once the incident resolves, keeps uploaded runbooks for later conversations, and a Memory API can view, update or redact what it holds. When an escalation policy triggers the agent, it works alongside the paged human and does not delay the page.
- Worth checking
- Each request, nudge or automatic trigger uses 4 AI Actions and per-tier allotments are not published, so estimate a month of usage before committing. The agent reads only the first 2,000 characters of an alert's custom details and of each incident note.
Sources: SRE Agent product page, SRE Agent docs, PagerDuty pricing, GA changelog
Resolve AI
Resolve AI · dedicated AI SRE
Sales-led; the trial button opens a demo request
Resolve AI was founded by Spiros Xanthos and Mayank Agarwal, who co-created OpenTelemetry and led Splunk Observability. It picks up alerts and incidents from the tools you already run and investigates them with teams of agents that pursue several hypotheses at once. In deep investigations, a separate verifier checks each finding against production evidence before it is shown.
Alerts run in one of three modes: triage, an adaptive mode that sets the depth from the alert's impact, runbooks and past handling, or a full investigation with proposed fixes. Background agents watch the deploy stream and flag regressions in error rate, latency or saturation before an alert fires. The company has raised more than $190 million, most recently at a $1.5 billion valuation in April 2026.
- Worth checking
- Each chat, triage, investigation, incident and background task draws credits at the rates in your contract. Reaching a monthly cap pauses automatic work, while work a person starts continues.
Sources: Resolve AI product overview, Mitigation actions docs, Billing and credit controls, Background agents, Agent Teams. Full comparison with Edge Delta
Rootly AI SRE
Rootly · incident management
Enabled per team by Rootly's account team; no launch date published
Rootly AI SRE investigates alerts and incidents through the observability, cloud and code tools connected to Rootly. It reports a root cause with a high, medium or low confidence level only when Rootly's checks on the recorded evidence pass; otherwise the result is labeled inconclusive, and teams on Rootly's six-stage flow can also see contributing factor or blocked.
It uses Claude Sonnet 4.6 by default with GPT-5 as a fallback, scrubs personal data from outgoing prompts, and runs inference in the US. Investigation rules can run manually or automatically, or be paused, and a 90-day preview shows which past alerts a rule would have matched.
- Worth checking
- Rootly's marketing says every change needs human sign-off, but its documentation says not to treat AI SRE as read-only, because connector and custom MCP tools can change data without pausing for approval. The AI SRE tab has no cancel control, and a run cannot be started from a workflow or the API.
Sources: Rootly AI SRE, AI SRE docs, Running an investigation, Rootly pricing
Splunk AI SRE
Splunk, a Cisco company · observability platform
Generally available since June 2026
Splunk's name for the product is AI SRE in Splunk Observability Cloud. When someone opens a supported alert, its troubleshooting agent ranks suspected root causes with a confidence level and evidence, and for Kubernetes alerts it adds a step-by-step remediation plan. The agent reached general availability in June 2026, first in the us1 realm.
Analysis runs once per user, so returning to an alert shows the earlier result immediately. In September 2026 Splunk added an alpha, built on Claude Managed Agents, that proposes a code fix, opens a pull request for review, confirms afterward that the fix worked, and keeps reviewer feedback in memory for later incidents.
- Worth checking
- Root-cause analysis covers APM service, business transaction and Kubernetes alerts. Alert grouping into incidents is in controlled availability in the us1 realm, and the troubleshooting agent is not offered in government realms.
Sources: Splunk AI SRE, Troubleshooting agent docs, September 2026 update
Traversal
Traversal · dedicated AI SRE
Sold since its June 2025 launch; Incident Workers generally available since September 2, 2026, Alert Workers in public beta
Traversal keeps a Production World Model of your system, rebuilt continuously from baselines, dependencies, changes and team knowledge. Its Causal Search Engine tests thousands of root-cause hypotheses in parallel and discards any that do not fit the dependencies, timing and spread of the failure. Workers join Slack or Teams channels to investigate incidents and triage alerts.
Traversal's documentation describes the platform as always read-only, though its AWS role is not limited to reads and write-capable MCP tools are yours to block. Its homepage presents self-healing as automated remediation, which the documentation does not describe. Named customers include American Express, DigitalOcean and Cloudways, and Traversal reports 38% lower MTTR at DigitalOcean.
- Worth checking
- Kubernetes and private data sources need the outbound-only Traversal Connector running in your environment, deployed with Helm on Kubernetes or as a container elsewhere. Alert Workers are still a beta you request through Traversal.
Sources: Traversal docs, Architecture, Workers GA announcement, Traversal pricing, Causal Search Engine. Full comparison with Edge Delta
Where to start a trial, by the stack you already run
| If | Start with | Because |
|---|---|---|
| Most of your telemetry already lives in Datadog | Datadog Bits Investigation | It reads data Datadog has already indexed, and Datadog publishes a per-investigation credit figure you can budget against. |
| You run Grafana Cloud with alert rules or Grafana IRM | Grafana Assistant Investigations | Alert rules and IRM webhooks can start investigations, and billing is by token with no per-investigation fee. |
| You run Dynatrace | Dynatrace SRE Agent | It starts from problems Dynatrace's causal AI has already detected; check that the preview is enabled for your tenant. |
| You monitor APM and Kubernetes in Splunk Observability Cloud | Splunk AI SRE | Root-cause ranking appears on supported alerts, and engineers keep every remediation step. |
| Most of your telemetry lives in New Relic | New Relic Autopilot | It attaches its analysis to alerts already routed through New Relic notification workflows. |
| Production runs mostly on one cloud provider | That cloud's agent: AWS DevOps Agent, Microsoft Azure SRE Agent or Gemini Cloud Assist investigations | Each reads its own cloud through the access model you already use. AWS bills per agent-second and Azure adds an always-on charge per agent, while Google's tool is a preview that needs Premium Support. |
| Paging and incidents run in PagerDuty, incident.io or Rootly | That vendor's agent | Each one works inside the incident you already open and reads your observability tools through connectors. |
| You deploy through Harness | Harness AI SRE | Its change agent ranks Harness deployments and pull requests as likely causes. |
| Your stack spans several observability vendors and you want a dedicated agent on top | Resolve AI or Traversal | Both work across the tools you already run. Resolve AI's write tools wait for approval by default, and Traversal's documentation describes it as read-only. |
| You want detection, investigation and fix checks on the telemetry itself | Edge Delta | Its teammates run on the telemetry pipeline, and Pro starts at $20 a month, including $20 of credits, after a 14-day trial. |
How this guide was put together, and what it does not show
- 01Edge Delta publishes this guide and sells one of the products in it. Products are listed alphabetically, and every entry links to the vendor pages it was written from.
- 02We read each vendor's product pages, documentation, pricing pages and release notes from September 30 to October 2, 2026. Where marketing pages and documentation disagreed, the entry follows the documentation and says so.
- 03The guide covers AI SRE products from established observability, incident management, cloud and software delivery vendors, plus Resolve AI and Traversal, the two dedicated AI SRE companies Edge Delta compares against in depth. Earlier-stage startups are not covered.
- 04A hollow mark in Figure 1 means we could not find the capability in public documentation. A product may do more than its documentation says. Features in preview, early access or alpha count as documented, and where a marketing page claims something the documentation contradicts, the mark follows the documentation.
- 05We have not run most of these products. The one hands-on measurement is AI SRE Arena v1, which Edge Delta built and ran: on 21 injected Kubernetes incidents, Edge Delta opened investigations for 18 and Grafana for 12.
- 06Prices are as published on the retrieval dates. Datadog's per-feature credit figures are averages, PagerDuty does not publish per-tier AI Action allotments, and Edge Delta does not publish the credit cost of an investigation.
Revisions
- October 2, 2026First publication.
What is an AI SRE?
An AI SRE is an AI agent that performs the work of a site reliability engineer: it detects production issues, investigates them across logs, metrics, traces, and deployment context, identifies the root cause with supporting evidence, and remediates with approval and verification.
Which AI SRE tools publish their pricing?
Of the 15 products in this guide, 7 publish a price for the AI SRE feature: Datadog bills AI credits at about 6.5 per autonomous investigation, Edge Delta Pro starts at $20 a month with $20 of included credits, Grafana charges $2 per million tokens beyond included allowances, AWS DevOps Agent charges $0.0083 per agent-second, Dynatrace does not charge for its agentic AI yet and bills the workflows and queries it runs at rate-card prices, Azure SRE Agent bills Azure Agent Units at $0.10 each in US regions, and PagerDuty includes its agent in PD Reliability Platform tiers from $2,800 a year. The others sell it through sales or as an add-on without a public price, and Google's investigations are a free preview that needs Premium Support.
Which AI SRE tools can make changes without a person approving them?
Azure SRE Agent's Autonomous mode applies Azure mitigations without waiting, and new response plans default to it. Datadog's Bits Infrastructure Operations, in preview, applies fixes automatically where guardrails allow it. Grafana MCP tools set to auto-approve run without prompting, Rootly's documentation warns that write-capable connector tools run without an approval step, and Harness runbooks on automatic triggers run without a person. Resolve AI's write tools wait for approval by default, though an admin can switch its preview GitHub Actions tools to Auto, and AWS DevOps Agent requires an operator to approve each change. incident.io changes code only through pull requests a person reviews, Traversal's documentation describes it as read-only, New Relic Autopilot and Splunk recommend steps that an engineer carries out, Google's investigations generate commands the user runs, and Edge Delta starts at the Propose trust level, where every write waits for a person's approval or a playbook you pre-approved.
Which AI SRE tools can start an investigation without an alert?
AWS DevOps Agent, Datadog Bits Investigation, Dynatrace SRE Agent, Edge Delta, Grafana Assistant Investigations, Microsoft Azure SRE Agent, New Relic Autopilot, Resolve AI and Traversal document starting from their own detection, a watcher or a schedule. The rest start from an alert rule, a declared incident or a person's request.
How is an AI SRE different from AIOps and observability tools?
Observability tools show humans what is happening, and AIOps tools rank what humans should look at first. An AI SRE performs the investigation itself and returns a conclusion a human can review.
How were the products in this guide chosen?
The guide covers AI SRE products from established observability, incident management, cloud and software delivery vendors, plus Resolve AI and Traversal. Edge Delta publishes the guide and is included, products are listed alphabetically, and each entry cites the vendor pages it was written from.