Edge Delta Research · Buyer's guide

Revision 1 · October 2, 2026

The best AI SRE tools in 2026

15 products, compared from their vendors' own documentation on triggers, data access, permissions and price

Edge Delta Research · Published October 2, 2026 · Sources retrieved September 30 to October 2, 2026

Products
15
Sources
61
Retrieved
Oct 2026
Revision
1

Summary

We read the product pages, documentation, pricing pages and release notes for 15 AI SRE products from September 30 to October 2, 2026, and recorded the same facts for each one, covering triggers, data access, permissions, price and the capabilities the vendor documents.

Definition. An AI SRE is an AI agent that performs the work of a site reliability engineer: it detects production issues, investigates them across logs, metrics, traces, and deployment context, identifies the root cause with supporting evidence, and remediates with approval and verification.

Start without an alert9 of 15
AWS DevOps Agent, Datadog Bits Investigation, Dynatrace SRE Agent, Edge Delta, Grafana Assistant Investigations, Microsoft Azure SRE Agent, New Relic Autopilot, Resolve AI and Traversal can begin from their own detection, a watcher or a schedule. The other 6 start when an alert rule fires, an incident is declared, or a person asks.
Publish a price7 of 15
AWS DevOps Agent, Datadog Bits Investigation, Dynatrace SRE Agent, Edge Delta, Grafana Assistant Investigations, Microsoft Azure SRE Agent and PagerDuty SRE Agent publish what the feature costs, in units that include AI credits, tokens, agent-seconds, Azure Agent Units, AI Actions, workflow-hours and a monthly plan with included credits. The rest sell it through sales or as an add-on without a public price, and Google's investigations are a free preview that needs Premium Support.
Check the fix8 of 15
AWS DevOps Agent, Datadog Bits Investigation, Edge Delta, Grafana Assistant Investigations, Microsoft Azure SRE Agent, PagerDuty SRE Agent, Splunk AI SRE and Traversal document checking whether the original problem was resolved after a change. The other vendors' documentation stops at recommending or making the change.

Edge Delta publishes this guide and is one of the 15 products in it. Entries are alphabetical, and each one links to the vendor pages it was written from.

1Comparison

Triggers, data access, permissions and price, product by product

Each cell restates what the vendor's own product pages, documentation and pricing pages say. Where a vendor publishes nothing on a point, the cell says that instead of guessing.

AI SRE tools compared by what starts an investigation, what the product reads, what it may change, and pricing
ProductStarts fromReadsMay changePricing
AWS DevOps AgentGA Mar 2026Webhooks and integrations from ServiceNow, PagerDuty, Datadog, Dynatrace, New Relic, Splunk, Grafana and Slack, a manual request, or a scheduled custom agentAWS accounts across Regions, Azure resources, CloudWatch, Datadog, Dynatrace, Splunk, New Relic, Grafana, GitHub, GitLab and your own MCP serversRead-only by default; opt-in directed actions need an operator to approve each change in chat, and destructive tools are never called$0.0083 per agent-second of active work, nothing while idle; a two-month free trial for new customers
Datadog Bits InvestigationGA Dec 2025Monitors with Auto-Investigate turned on, when they enter alert state, including monitors that Bits Detection creates and tunes itself in preview; a person can start one from Slack, a Synthetic test or a prompt, and workflows can trigger oneDatadog metrics, traces, logs, events, change tracking, RUM, profiles and GitHub code; Grafana, Dynatrace, Sentry and Splunk data in previewTriage actions from chat; code fixes from Bits Code as pull requests a person opens; infrastructure actions such as Kubernetes patches, in preview, which guardrails can let run automatically, hold for approval or blockAI credits, about 6.5 per autonomous investigation; from $500 for 500 credits a month billed annually, or $1.30 a credit on demand
Dynatrace SRE AgentPreviewError and availability problems that Dynatrace's causal AI has detected and traced to a root-cause entity; workflows can also run on events or a scheduleLogs, events, metrics and traces in Dynatrace, queried with DQL, plus the details of the problem it is working onWrites its analysis and recommended fixes onto the problem; approved remediation, such as a rollback, runs through Dynatrace WorkflowsNo charge for the agentic AI yet; the workflow and the DQL queries it runs bill at rate-card prices, such as $0.03 per workflow-hour
Edge DeltaAvailable Oct 2025Its own monitors, including learned baselines on live telemetry, and scheduled loops; PagerDuty alerts and chat requests also start oneLogs, metrics, traces and events flowing through Edge Delta Telemetry Pipelines, plus connected tools such as Elastic, Kubernetes, AWS, Azure, GCP, GitHub, Jira and LinearStarts at the Propose trust level: it reads freely and drafts changes, every write waits for a person's approval or a playbook you pre-approved, and destructive actions stay off until you raise the levelFree 14-day trial without a credit card; Pro from $20 a month, including $20 of credits, with unlimited investigations and spending limits
Google Gemini Cloud Assist investigationsPreviewA person starting one from the Investigations page, Logs Explorer, a Cloud Monitoring alert, chat or a product page; Proactive Mode for alerts is in private previewLogs, metrics and configuration for Google Cloud resources in one project or App Hub application at a time, across 18 supported productsNothing during the investigation; where a fix is supported it generates gcloud commands or Kubernetes manifests that the user reviews and runsNo charge during preview, listed under Gemini Code Assist Enterprise; no GA price published
Grafana Assistant InvestigationsGA Jul 2026An alert rule with an investigation attached, an incident or alert group in Grafana IRM, a Watcher's critical finding in preview, or a person in the app, Slack or the gcx CLIMetrics, logs, traces and profiles in Grafana Cloud; code and deployment changes through a connected MCP server such as GitHubEnds in a report with recommended next steps; write-capable MCP tools follow per-tool approval settings, and coding agents that open pull requests are in private preview$2 per million tokens beyond included allowances, no per-investigation fee; Pro adds a $19 monthly platform fee and $20 per active AI user beyond three
Harness AI SREEnterprise planAlerts sent in by webhook and promoted to incidents by rules; incidents opened in the app or with /harness new in SlackDeployments, pull requests, ServiceNow change records, alert data, and metrics and traces from affected services, plus Slack, Teams and Zoom conversations through AI Scribe; Datadog, Splunk, New Relic and Dynatrace send alerts by webhookRunbooks that roll back, scale out or toggle a feature flag, started by a trigger or a responder; custom scripts run as Harness pipelinesNot published; sold as a module on the Enterprise plan
incident.io InvestigationsGA Aug 2026A declared incident, automatically or on conditions you set, or /inc investigate; an alert has to become an incident firstTelemetry queried on demand from tools such as Datadog, Grafana, New Relic, Splunk, Elasticsearch and AWS CloudWatch, plus GitHub or GitLab code and past incidentsPull requests a person reviews and merges, and fields on the incident itself; it never merges or deploysPaid add-on for the Pro and Enterprise plans; the add-on price is not published
Microsoft Azure SRE AgentGA Mar 2026Incidents from Azure Monitor, PagerDuty or ServiceNow, scheduled tasks, HTTP and log-query triggers, and chatAzure resources, Azure Monitor, Application Insights, Log Analytics and Kusto; Datadog, Splunk, New Relic, Dynatrace and Elasticsearch through partner MCP connectorsAzure mitigations under its managed identity; in Review mode an administrator approves Azure infrastructure writes, in Autonomous mode they apply directly$0.40 per agent-hour always on in US regions, plus token-metered usage, both billed in Azure Agent Units at $0.10 each
New Relic AutopilotGA Jul 2026An alert workflow with an AI agent destination, a question in the New Relic AI panel or Slack, or a Workflow Automation step, including after a deploy or on a scheduleNew Relic telemetry across the accounts you select, Confluence and uploaded documents, Slack threads, and GitHub through a read-only MCP connection in previewRecommends actions without taking them; remediation runs through New Relic Workflow Automation, which can pause for an approve or deny linkNot published; requires a Pro or Enterprise edition with the Advanced Compute add-on, billed by compute units
PagerDuty SRE AgentGA Oct 2025A responder asking on the incident, in the Operations Console or in Slack; in early access, an incident workflow or escalation policy can add it as a virtual responder that starts when an incident triggersIncident history, runbooks and its own memory, plus 19 documented connectors including Datadog, Dynatrace, Grafana, New Relic, Splunk, AWS CloudWatch and GitHubRecommends an incident workflow and runs it when a responder clicks Run, each workflow at most once per incident; a fully autonomous responder is announced for early access in the second half of 2026Included in PD Reliability Platform tiers from $2,800 a year and sold as an add-on for the Professional plan; each request uses 4 AI Actions
Resolve AISales-ledAlerts from PagerDuty and other alerting tools as they fire; background agents also run on a schedule or on events such as deploysObservability, cloud, Kubernetes, code and CI/CD tools through 60+ integrations, with an optional Satellite, run on Kubernetes or AWS ECS, for private tools and Kubernetes dataProposes alert silences and pull requests for a person to approve; in preview it can also dispatch, retry or cancel GitHub Actions runs, which wait for approval unless an admin sets the tool to AutoCredits per chat, triage, investigation, incident or task, at rates set in your contract; no public rates
Rootly AI SREOn requestAlerts that match an auto-run investigation rule, account-level automatic investigation, or a person in the app or through @Rootly in SlackDatadog, Grafana, New Relic, Honeycomb, Sentry, Dynatrace and Splunk, plus AWS, Azure, GCP, GitHub, GitLab and custom MCP toolsReports a root cause and next steps; its GitHub and GitLab connections are read-only, and its docs warn that other write-capable connector tools run without an approval stepNot published for AI SRE; Rootly Incident Response and On-Call each start at $20 per user a month
Splunk AI SREGA Jun 2026A person opening a supported APM or Kubernetes alert, once the organization has opted in through Splunk Support, or choosing Run root cause analysis on the alertMetrics, events, logs and traces in Splunk Observability Cloud; custom metrics are not supportedNothing directly; remediation plans for Kubernetes alerts list commands an engineer runs, and an alpha built on Claude Managed Agents opens pull requests for reviewNot published; Splunk directs customers to their representative
TraversalWorkers GA Sep 2026Incidents and alerts in Slack or Teams channels, where Workers decide when to step in; Custom Workers follow a written mission, on a schedule if you set oneDatadog, Dynatrace, Grafana, Splunk, Prometheus, Elasticsearch, CloudWatch, Sentry and others through read-only APIs, plus GitHub or GitLab; Kubernetes through a connector you runNothing directly, according to its docs; it recommends fixes and, over MCP, hands next steps to a coding agent you run. Its marketing describes self-healing remediation that the docs do not coverNot published; sized to your environment after a walkthrough
Table 1. Vendor product pages, documentation and pricing pages, retrieved September 30 to October 2, 2026. Product names are trademarks of their owners.
Capabilities each vendor documents publicly
ProductStarts on its ownApproval stepChecks the fixUses past incidentsPublic price
AWS DevOps AgentDocumentedDocumentedDocumentedDocumentedDocumented
Datadog Bits InvestigationDocumentedDocumentedDocumentedDocumentedDocumented
Dynatrace SRE AgentDocumentedDocumentedNot found in public documentationNot found in public documentationDocumented
Edge DeltaDocumentedDocumentedDocumentedDocumentedDocumented
Google Gemini Cloud Assist investigationsNot found in public documentationDocumentedNot found in public documentationNot found in public documentationNot found in public documentation
Grafana Assistant InvestigationsDocumentedDocumentedDocumentedDocumentedDocumented
Harness AI SRENot found in public documentationDocumentedNot found in public documentationDocumentedNot found in public documentation
incident.io InvestigationsNot found in public documentationDocumentedNot found in public documentationDocumentedNot found in public documentation
Microsoft Azure SRE AgentDocumentedDocumentedDocumentedDocumentedDocumented
New Relic AutopilotDocumentedDocumentedNot found in public documentationDocumentedNot found in public documentation
PagerDuty SRE AgentNot found in public documentationDocumentedDocumentedDocumentedDocumented
Resolve AIDocumentedDocumentedNot found in public documentationDocumentedNot found in public documentation
Rootly AI SRENot found in public documentationNot found in public documentationNot found in public documentationDocumentedNot found in public documentation
Splunk AI SRENot found in public documentationDocumentedDocumentedDocumentedNot found in public documentation
TraversalDocumentedDocumentedDocumentedDocumentedNot found in public documentation
Figure 1. What each vendor documents publicly. A filled mark means the vendor's own pages describe the capability, including features in preview or alpha, and its documentation does not contradict them. A hollow mark means we could not find it, which is not proof the product lacks it.
Starts on its own
Can begin from a problem it detects itself, a watcher or a schedule, without an alert rule, a declared incident or a person asking.
Approval step
Documents that a person approves or carries out each change, or a setting that holds changes for approval, before anything in production changes.
Checks the fix
Documents checking, after a change, whether the original problem was resolved.
Uses past incidents
Draws on earlier incidents or a memory of past investigations when it works on a new one.
Public price
Publishes what the AI SRE feature costs, as rates, plan prices that include it, or usage units.
2The products

15 products, from their vendors' documentation

  • AWS DevOps Agent

    Amazon Web Services · cloud provider

    Generally available since March 31, 2026, after a preview from December 2025

    AWS DevOps Agent investigates incidents that arrive from ticketing, paging and observability tools or from a manual request, across AWS accounts, Azure resources and connected tools. It produces a mitigation plan or a spec that a coding agent such as Kiro can implement. AWS previewed it in December 2025 and made it generally available on March 31, 2026.

    Changes are off by default. With directed actions enabled, each approval covers one tool, operation and resource, for a single use or up to four hours, and CloudTrail records who approved it. A change attempted during an autonomous investigation fails instead of asking. Memory stores keep recurring root causes per alarm, standing directives and lessons from corrective feedback.

    Worth checking
    Cost follows agent time, about $30 per agent-hour, and AWS's own example of ten 8-minute investigations comes to $39.84. Queries it runs against services such as CloudWatch Logs Insights are billed separately, and each Agent Space runs three investigations at once by default.

    Sources: AWS DevOps Agent, User guide, Directed actions, Pricing, GA announcement

  • Datadog Bits Investigation

    Datadog · observability platform

    Generally available since December 2025, launched as Bits AI SRE

    Datadog's investigation agent forms hypotheses about a root cause and queries Datadog telemetry to confirm or rule out each one, then posts its findings to the Slack, case or on-call destinations already set on the monitor. Datadog made it generally available on December 2, 2025 as Bits AI SRE and renamed it Bits Investigation in mid-2026, alongside Bits Code and Bits Chat.

    Datadog's March 2026 update put a typical investigation at three to four minutes. By default each monitor starts at most one automatic investigation in a rolling 24 hours, a limit admins can change, and warn, no-data, renotification and test events do not start one. Bits Infrastructure Operations, in preview, applies infrastructure fixes automatically where a guardrail allows it, for example in staging, and asks for approval elsewhere.

    Worth checking
    Investigations, chat messages (about 0.5 credits each) and code fixes (about 5) all draw on the same monthly credit bundle, and Datadog describes these figures as averages. Unused credits do not roll over, overage bills at the on-demand rate, and one-click infrastructure actions, guardrails, Bits Detection, Bits Infrastructure Operations and automatic memories are still in preview.

    Sources: Bits Investigation docs, Bits Remediation docs, AI credits pricing, GA announcement, Bits Infrastructure Operations. Full comparison with Edge Delta

  • Dynatrace SRE Agent

    Dynatrace · observability platform

    Preview; the SRE Agent workflow template shipped in late July 2026, inside Dynatrace Intelligence, which launched in January 2026

    At Perform in January 2026, Dynatrace introduced Dynatrace Intelligence, the successor to Davis AI, and added domain agents and agentic workflows to its platform. The SRE Agent, a workflow template that shipped in late July 2026, is one of eleven ready-made agentic workflows in the documentation, all marked preview. It picks up a problem after Dynatrace has run its own root-cause analysis, queries logs, events, metrics and traces with DQL, and writes its analysis and recommendations back onto the problem.

    The default workflow is capped at five tool calls and a 1,000-word answer, and it does not apply a fix by itself. In July 2026 Dynatrace announced an Autonomous SRE Agent and a Cloud SRE Agent. The Cloud SRE Agents app hands problems to agents from AWS, Microsoft and Google, and its Hub listing describes it as a community-supported project.

    Worth checking
    The SRE Agent is marked preview in the documentation and coming soon in the Dynatrace Hub, so confirm what your tenant can use. We found no release note confirming that the Autonomous SRE Agent shipped after its August target.

    Sources: Dynatrace Intelligence docs, SRE Agent in the Dynatrace Hub, January 2026 announcement, July 2026 announcement, SRE Agent docs, Agentic AI licensing FAQ, Rate card

  • Edge Delta

    Edge Delta · telemetry pipelines and AI SRE

    Available since October 2025, with a self-serve trial

    Edge Delta is a telemetry-native AI SRE that continuously understands, investigates, and operates production. Its AI Teammates (SRE, Security Engineer, Software Engineer and Work Tracker) reason over the logs, metrics, traces and events moving through Edge Delta's Telemetry Pipelines, so baselines and service context already exist when an incident starts.

    An investigation ends with a root cause and its evidence, and proposed fixes wait in an approvals queue. After a fix runs, the teammate compares telemetry from before and after the change, and the incident is saved to Production Memory for later investigations.

    Worth checking
    Detection on live data and the fullest context depend on telemetry passing through Edge Delta pipelines; data that stays in another backend is reached through connectors, so the teammates see what those tools return. Pro usage beyond the included $20 draws on credits whose cost per investigation is not published, and Pro keeps data and memory for 30 days.

    Sources: Edge Delta AI SRE, Edge Delta pricing

  • Google Gemini Cloud Assist investigations

    Google Cloud · cloud provider

    Preview since June 2025; new investigations need Premium Support or account-team access since April 2026

    The investigations feature in Gemini Cloud Assist analyzes logs, metrics and configuration for Google Cloud resources and returns root-cause hypotheses with suggested next steps. It entered public preview in June 2025, and since April 10, 2026 creating or running an investigation requires a Premium Support contract or access through your Google Cloud account team.

    Investigations run with the permissions of the person who starts them, and the documentation states that this access is never used to change data. Proactive Mode, in private preview for Premium Support customers, investigates Cloud Alerting alerts and cost anomalies in the background and publishes results as Eventarc events.

    Worth checking
    It covers Google Cloud only, one project or App Hub application per investigation. Google advises against using it on data with residency requirements, because investigation data may be stored in any Google Cloud data center, and investigations stopped working inside VPC Service Controls perimeters on April 13, 2026.

    Sources: Investigations docs, Release notes, Gemini Cloud Assist, Gemini pricing

  • Grafana Assistant Investigations

    Grafana Labs · observability platform

    Generally available since July 2026

    Grafana Assistant became generally available in Grafana Cloud in October 2025, and its Investigations mode followed on July 29, 2026. An investigation explores metrics, logs, traces and profiles, builds hypotheses, and returns a report with findings, supporting evidence and recommended next steps.

    By default, alert enrichment opens at most 50 new investigations every 20 minutes, and changes to the same alert group within six hours continue the existing investigation. An investigation can schedule its own re-checks, for example whether an error rate has recovered, and close them once the situation resolves. Metering for investigations began on October 1, 2026.

    Worth checking
    Investigations need a separate entitlement and are not available on self-managed Grafana, even when it connects to the Cloud Assistant backend. MCP tools set to auto-approve can write to or delete data in connected services without asking.

    Sources: Investigations docs, GA release note, Grafana pricing, MCP server settings

  • Harness AI SRE

    Harness · software delivery platform

    Sold as an Enterprise plan module; no GA date published

    Harness AI SRE combines incident management and on-call with AI agents. AI Scribe records what happens in Slack, Teams and Zoom calls, the RCA Change Agent ranks recent deployments, pull requests and ServiceNow change records as likely causes, and runbooks can roll back a deployment by running a Harness pipeline, scale out, or flip a feature flag.

    On-call schedules, escalation policies, paging and a mobile app sit behind a feature flag, with import from PagerDuty, Opsgenie and xMatters. Alerts are enriched with similar past incidents and how they were resolved, AI post-mortems can be drafted when an incident closes, and ServiceNow change records are polled every five minutes.

    Worth checking
    Runbooks on automatic triggers run without a person, and Harness recommends running production rollbacks manually. Runbooks exposed to the AI investigator, which run automatically or with a person in the loop to feed root-cause analysis, are in early access, and correlation with feature flag, infrastructure and config changes sits behind a feature flag.

    Sources: Harness AI SRE, Harness pricing

  • incident.io Investigations

    incident.io · incident management

    Generally available since August 5, 2026

    Investigations is incident.io's AI SRE, sold as an add-on to its incident management platform. When an incident is declared, it queries connected telemetry, code and past incidents, then posts a root-cause hypothesis with a confidence level and linked evidence. It runs on Nexus, which incident.io calls its production intelligence model; each customer gets a separate instance, and the underlying models come from OpenAI, Anthropic and Google under zero-data-retention terms.

    In September 2026 incident.io reported that the median time to the first accurate message in an incident channel fell from 6.7 to 3 minutes. An escalation step can hold alert-triggered paging until the first hypothesis exists, up to a time limit you set.

    Worth checking
    It works only on incidents, so an alert reaches it only after a workflow or a person turns it into one. The product page says sensitive data is redacted before it reaches model providers, while the documentation describes pattern-based redaction that stays off until you ask incident.io to enable it.

    Sources: incident.io AI SRE, Triggering docs, incident.io pricing, GA changelog

  • Microsoft Azure SRE Agent

    Microsoft · cloud provider

    Generally available since March 10, 2026, after a preview from May 2025

    Microsoft's agent takes incidents from Azure Monitor, PagerDuty or ServiceNow, runs scheduled and webhook-triggered tasks, and investigates across Azure and connected tools before proposing or applying a mitigation. Microsoft previewed it at Build in May 2025 and made it generally available on March 10, 2026.

    By default the managed identity gets Reader and Monitoring Contributor roles, with a Privileged level offered at creation; delete and Key Vault commands are blocked outright, and administrators can add tool policies and hooks on top. Review mode pauses only Azure infrastructure writes, and new incident response plans and scheduled tasks default to Autonomous, including the quickstart plan it creates for high-severity alerts.

    Worth checking
    The always-on charge runs from creation until deletion, even while the agent is stopped, which comes to about $292 a month per agent in US regions before any investigation. A 30-day trial for new customers waives it for up to three agents.

    Sources: Azure SRE Agent overview, Azure SRE Agent pricing, GA announcement

  • New Relic Autopilot

    New Relic · observability platform

    Generally available since July 2026; previewed as SRE Agent from February 2026

    New Relic previewed an SRE Agent in February 2026 and made it generally available as Autopilot in July 2026. It attaches an analysis to alerts routed to it through a notification workflow and answers questions in the New Relic AI panel or in Slack, using telemetry from the New Relic accounts you select.

    Memories, added in July 2026, are off by default and kept for 90 days unless you extend them to as long as a year. Each organization runs one Autopilot, limited to 100 requests an hour, and an investigation looks back over the last 24 hours by default.

    Worth checking
    Autopilot does not create alerts, so its coverage follows the alert conditions you already maintain in New Relic. The GitHub connection is in public preview, and Jira, named in the launch announcement, does not appear in the documentation.

    Sources: New Relic Autopilot, Autopilot docs, GA release note, Memories docs

  • PagerDuty SRE Agent

    PagerDuty · incident management

    Generally available since October 30, 2025; called Paige in current docs

    PagerDuty's SRE Agent, now called Paige in its documentation and pricing, works inside the incident. It pulls logs, metrics, runbooks and incident history from connected tools, suggests likely causes, and recommends an incident workflow for a responder to run. PagerDuty made it generally available on October 30, 2025.

    Memory is kept per service. Paige saves what it learned from an incident once the incident resolves, keeps uploaded runbooks for later conversations, and a Memory API can view, update or redact what it holds. When an escalation policy triggers the agent, it works alongside the paged human and does not delay the page.

    Worth checking
    Each request, nudge or automatic trigger uses 4 AI Actions and per-tier allotments are not published, so estimate a month of usage before committing. The agent reads only the first 2,000 characters of an alert's custom details and of each incident note.

    Sources: SRE Agent product page, SRE Agent docs, PagerDuty pricing, GA changelog

  • Resolve AI

    Resolve AI · dedicated AI SRE

    Sales-led; the trial button opens a demo request

    Resolve AI was founded by Spiros Xanthos and Mayank Agarwal, who co-created OpenTelemetry and led Splunk Observability. It picks up alerts and incidents from the tools you already run and investigates them with teams of agents that pursue several hypotheses at once. In deep investigations, a separate verifier checks each finding against production evidence before it is shown.

    Alerts run in one of three modes: triage, an adaptive mode that sets the depth from the alert's impact, runbooks and past handling, or a full investigation with proposed fixes. Background agents watch the deploy stream and flag regressions in error rate, latency or saturation before an alert fires. The company has raised more than $190 million, most recently at a $1.5 billion valuation in April 2026.

    Worth checking
    Each chat, triage, investigation, incident and background task draws credits at the rates in your contract. Reaching a monthly cap pauses automatic work, while work a person starts continues.

    Sources: Resolve AI product overview, Mitigation actions docs, Billing and credit controls, Background agents, Agent Teams. Full comparison with Edge Delta

  • Rootly AI SRE

    Rootly · incident management

    Enabled per team by Rootly's account team; no launch date published

    Rootly AI SRE investigates alerts and incidents through the observability, cloud and code tools connected to Rootly. It reports a root cause with a high, medium or low confidence level only when Rootly's checks on the recorded evidence pass; otherwise the result is labeled inconclusive, and teams on Rootly's six-stage flow can also see contributing factor or blocked.

    It uses Claude Sonnet 4.6 by default with GPT-5 as a fallback, scrubs personal data from outgoing prompts, and runs inference in the US. Investigation rules can run manually or automatically, or be paused, and a 90-day preview shows which past alerts a rule would have matched.

    Worth checking
    Rootly's marketing says every change needs human sign-off, but its documentation says not to treat AI SRE as read-only, because connector and custom MCP tools can change data without pausing for approval. The AI SRE tab has no cancel control, and a run cannot be started from a workflow or the API.

    Sources: Rootly AI SRE, AI SRE docs, Running an investigation, Rootly pricing

  • Splunk AI SRE

    Splunk, a Cisco company · observability platform

    Generally available since June 2026

    Splunk's name for the product is AI SRE in Splunk Observability Cloud. When someone opens a supported alert, its troubleshooting agent ranks suspected root causes with a confidence level and evidence, and for Kubernetes alerts it adds a step-by-step remediation plan. The agent reached general availability in June 2026, first in the us1 realm.

    Analysis runs once per user, so returning to an alert shows the earlier result immediately. In September 2026 Splunk added an alpha, built on Claude Managed Agents, that proposes a code fix, opens a pull request for review, confirms afterward that the fix worked, and keeps reviewer feedback in memory for later incidents.

    Worth checking
    Root-cause analysis covers APM service, business transaction and Kubernetes alerts. Alert grouping into incidents is in controlled availability in the us1 realm, and the troubleshooting agent is not offered in government realms.

    Sources: Splunk AI SRE, Troubleshooting agent docs, September 2026 update

  • Traversal

    Traversal · dedicated AI SRE

    Sold since its June 2025 launch; Incident Workers generally available since September 2, 2026, Alert Workers in public beta

    Traversal keeps a Production World Model of your system, rebuilt continuously from baselines, dependencies, changes and team knowledge. Its Causal Search Engine tests thousands of root-cause hypotheses in parallel and discards any that do not fit the dependencies, timing and spread of the failure. Workers join Slack or Teams channels to investigate incidents and triage alerts.

    Traversal's documentation describes the platform as always read-only, though its AWS role is not limited to reads and write-capable MCP tools are yours to block. Its homepage presents self-healing as automated remediation, which the documentation does not describe. Named customers include American Express, DigitalOcean and Cloudways, and Traversal reports 38% lower MTTR at DigitalOcean.

    Worth checking
    Kubernetes and private data sources need the outbound-only Traversal Connector running in your environment, deployed with Helm on Kubernetes or as a container elsewhere. Alert Workers are still a beta you request through Traversal.

    Sources: Traversal docs, Architecture, Workers GA announcement, Traversal pricing, Causal Search Engine. Full comparison with Edge Delta

3Shortlists

Where to start a trial, by the stack you already run

Which AI SRE products to trial first in common situations
IfStart withBecause
Most of your telemetry already lives in DatadogDatadog Bits InvestigationIt reads data Datadog has already indexed, and Datadog publishes a per-investigation credit figure you can budget against.
You run Grafana Cloud with alert rules or Grafana IRMGrafana Assistant InvestigationsAlert rules and IRM webhooks can start investigations, and billing is by token with no per-investigation fee.
You run DynatraceDynatrace SRE AgentIt starts from problems Dynatrace's causal AI has already detected; check that the preview is enabled for your tenant.
You monitor APM and Kubernetes in Splunk Observability CloudSplunk AI SRERoot-cause ranking appears on supported alerts, and engineers keep every remediation step.
Most of your telemetry lives in New RelicNew Relic AutopilotIt attaches its analysis to alerts already routed through New Relic notification workflows.
Production runs mostly on one cloud providerThat cloud's agent: AWS DevOps Agent, Microsoft Azure SRE Agent or Gemini Cloud Assist investigationsEach reads its own cloud through the access model you already use. AWS bills per agent-second and Azure adds an always-on charge per agent, while Google's tool is a preview that needs Premium Support.
Paging and incidents run in PagerDuty, incident.io or RootlyThat vendor's agentEach one works inside the incident you already open and reads your observability tools through connectors.
You deploy through HarnessHarness AI SREIts change agent ranks Harness deployments and pull requests as likely causes.
Your stack spans several observability vendors and you want a dedicated agent on topResolve AI or TraversalBoth work across the tools you already run. Resolve AI's write tools wait for approval by default, and Traversal's documentation describes it as read-only.
You want detection, investigation and fix checks on the telemetry itselfEdge DeltaIts teammates run on the telemetry pipeline, and Pro starts at $20 a month, including $20 of credits, after a 14-day trial.
4Method and limits

How this guide was put together, and what it does not show

  1. 01Edge Delta publishes this guide and sells one of the products in it. Products are listed alphabetically, and every entry links to the vendor pages it was written from.
  2. 02We read each vendor's product pages, documentation, pricing pages and release notes from September 30 to October 2, 2026. Where marketing pages and documentation disagreed, the entry follows the documentation and says so.
  3. 03The guide covers AI SRE products from established observability, incident management, cloud and software delivery vendors, plus Resolve AI and Traversal, the two dedicated AI SRE companies Edge Delta compares against in depth. Earlier-stage startups are not covered.
  4. 04A hollow mark in Figure 1 means we could not find the capability in public documentation. A product may do more than its documentation says. Features in preview, early access or alpha count as documented, and where a marketing page claims something the documentation contradicts, the mark follows the documentation.
  5. 05We have not run most of these products. The one hands-on measurement is AI SRE Arena v1, which Edge Delta built and ran: on 21 injected Kubernetes incidents, Edge Delta opened investigations for 18 and Grafana for 12.
  6. 06Prices are as published on the retrieval dates. Datadog's per-feature credit figures are averages, PagerDuty does not publish per-tier AI Action allotments, and Edge Delta does not publish the credit cost of an investigation.

Revisions

  • October 2, 2026First publication.
5Questions

What is an AI SRE?

An AI SRE is an AI agent that performs the work of a site reliability engineer: it detects production issues, investigates them across logs, metrics, traces, and deployment context, identifies the root cause with supporting evidence, and remediates with approval and verification.

Which AI SRE tools publish their pricing?

Of the 15 products in this guide, 7 publish a price for the AI SRE feature: Datadog bills AI credits at about 6.5 per autonomous investigation, Edge Delta Pro starts at $20 a month with $20 of included credits, Grafana charges $2 per million tokens beyond included allowances, AWS DevOps Agent charges $0.0083 per agent-second, Dynatrace does not charge for its agentic AI yet and bills the workflows and queries it runs at rate-card prices, Azure SRE Agent bills Azure Agent Units at $0.10 each in US regions, and PagerDuty includes its agent in PD Reliability Platform tiers from $2,800 a year. The others sell it through sales or as an add-on without a public price, and Google's investigations are a free preview that needs Premium Support.

Which AI SRE tools can make changes without a person approving them?

Azure SRE Agent's Autonomous mode applies Azure mitigations without waiting, and new response plans default to it. Datadog's Bits Infrastructure Operations, in preview, applies fixes automatically where guardrails allow it. Grafana MCP tools set to auto-approve run without prompting, Rootly's documentation warns that write-capable connector tools run without an approval step, and Harness runbooks on automatic triggers run without a person. Resolve AI's write tools wait for approval by default, though an admin can switch its preview GitHub Actions tools to Auto, and AWS DevOps Agent requires an operator to approve each change. incident.io changes code only through pull requests a person reviews, Traversal's documentation describes it as read-only, New Relic Autopilot and Splunk recommend steps that an engineer carries out, Google's investigations generate commands the user runs, and Edge Delta starts at the Propose trust level, where every write waits for a person's approval or a playbook you pre-approved.

Which AI SRE tools can start an investigation without an alert?

AWS DevOps Agent, Datadog Bits Investigation, Dynatrace SRE Agent, Edge Delta, Grafana Assistant Investigations, Microsoft Azure SRE Agent, New Relic Autopilot, Resolve AI and Traversal document starting from their own detection, a watcher or a schedule. The rest start from an alert rule, a declared incident or a person's request.

How is an AI SRE different from AIOps and observability tools?

Observability tools show humans what is happening, and AIOps tools rank what humans should look at first. An AI SRE performs the investigation itself and returns a conclusion a human can review.

How were the products in this guide chosen?

The guide covers AI SRE products from established observability, incident management, cloud and software delivery vendors, plus Resolve AI and Traversal. Edge Delta publishes the guide and is included, products are listed alphabetically, and each entry cites the vendor pages it was written from.

Try Edge Delta on your own telemetry

Free for 14 days, no credit card required.