cloud_download
New version available
Click the blinking icon in the toolbar or the button below.
RCA Trace
0
Received
0
Buffering
0
Waiting
0
Escalated
0
Suppressed
hub Correlation Engine checking… Buffer:
Autonomous Rate
of resolved incidents
Avg MTTR · 24h
Signal Coverage
Filter:
0 selected
FIRST SEEN? LAST FIRED? OPEN FOR? CLUSTER? NAMESPACE? SERVICE? DIMENSION? DEVIATION? DURATION? LAYER? STATE? SUPPRESSION REASON? INCIDENT? INC STATE? AI?
monitor_heartSignal stream will appear here
Integrations
Connect data sources to improve RCA accuracy and workflow tools to extend autonomous actions. More connected = fewer blind spots.
hub
Kubernetes Cluster
The cluster the agent monitors — required before any other integration can be configured
expand_more
account_tree
Service Topology
topology.md — auto-generated per namespace during discovery. Used by the RCA engine to trace blast radius.
expand_more
Loading namespaces…
share
Cross-Namespace Resources
Shared infrastructure detected across namespaces — controls how the agent groups incidents from different tenants
expand_more
psychology
Agent Skills
YAML playbooks — define how the agent investigates and remediates each fault class
expand_more
psychology
databaseObservability Sources
Alertmanager
Replaced by CloudWatch EventBridge. Incidents now flow directly: CloudWatch Alarm → EventBridge rule → Lambda, eliminating the Alertmanager webhook dependency.
swap_horiz Replaced by EventBridge
AlertmanagerConfig removed · EventBridge active
Vector
Container metrics + log pipeline. Collects CPU, memory, and log data from pods and ships to CloudWatch — replaces Prometheus and Loki in this environment. Source label: vector.
check_circle Active · ships to CloudWatch
Metrics + logs pipeline active · signal_plane=k8s_oss
Prometheus
Golden signal metrics via Alertmanager. Not used in this environment — metric collection and alerting handled by Vector → CloudWatch pipeline.
swap_horiz Replaced by Vector
Not deployed · no Alertmanager dependency
Loki
Pod log aggregation. Not used in this environment — container logs are collected and shipped to CloudWatch by the Vector pipeline.
swap_horiz Replaced by Vector
Not deployed · logs available via CloudWatch Logs Insights
Datadog
Pull metrics and APM traces from Datadog as an alternative or complement to Prometheus. Supports hybrid-cloud setups.
Not connected
lock_open Unlocks APM trace correlation in RCA
OpenTelemetry
Distributed traces (OTLP). Enables span-level RCA — pinpoints the exact service boundary where a request fails.
Not connected
lock_open Unlocks span-level RCA evidence
cloudCloud Providers
Amazon Web Services
2/5 connected · EventBridge detection active · IAM drift monitoring live
CloudTrail
IAM change events, S3 access denied, API call history. Detects permission-drift incidents invisible to metrics — PutRolePolicy mutations surfaced via RCA.
check_circle Active — IAM drift detection live
PutRolePolicy events analysed · deny-s3-read detected
CloudWatch Alarms & Logs
Container Insights metric filters → CloudWatch Alarms → EventBridge → Lambda. Primary detection path for error rate, CPU throttling, and IAM drift — no Prometheus or Alertmanager required.
check_circle Active — 3 alarms monitoring
checkout-db-errors · checkout-cpu-high · checkout-iam-deny
Container Insights
Pod, container, node, and cluster-level CPU, memory, disk, and network metrics via CloudWatch. Installed via amazon-cloudwatch-observability EKS addon — no Prometheus required.
warning Not connected
lock_open Pod CPU/memory + node pressure in RCA
EKS Control Plane Logs
API server, controller-manager, scheduler, and audit logs from the EKS control plane in CloudWatch. Detects RBAC denials, API server errors, and cluster-level anomalies invisible in pod metrics.
warning Not connected
lock_open API server audit + RBAC events in RCA
RDS / CloudWatch Metrics
Database CPU, connections, read/write latency from CloudWatch. Enriches RCA with DB-layer evidence for connection and query failures.
warning Not connected
lock_open DB latency + connection evidence in RCA
ALB / Load Balancer
ALB request count, HTTP 5xx rate, and target response time from CloudWatch. Detects upstream load-balancer failures not visible in pod metrics.
warning Not connected
lock_open ALB 5xx rate + target health in RCA
AWS X-Ray
Distributed traces from the checkout service via OTel auto-instrumentation. Pinpoints which service hop introduced latency or errors during RCA.
warning Not connected
lock_open Distributed traces in RCA
DevOps Guru
ML-based anomaly detection for RDS, Lambda, and ALB. Enriches RCA with AWS-native insights.
Coming soon
Microsoft Azure
0/2 connected · Azure resources not correlated to incidents
Azure Monitor
RBAC deny events, AKS infra signals, and Log Analytics workspace queries. Permission-drift RCA for Azure resources.
warning Not connected — Azure RCA blind
lock_open Azure RBAC + AKS infra RCA
Entra ID / Azure AD
RBAC role assignment changes and conditional access events. Detects identity-plane permission drift.
Coming soon
Google Cloud Platform
0/2 connected · GCP resources not correlated to incidents
Cloud Logging
Cloud Audit Logs (IAM permission errors), GKE infra signals, and Cloud Monitoring metrics.
warning Not connected — GCP RCA blind
lock_open GCP IAM + GKE infra RCA
Cloud Monitoring
GCP metrics API — Cloud SQL, Cloud Run, GKE node metrics as native signals alongside Prometheus.
Coming soon
auto_awesomeWorkflow Tools
campaignNotifications & On-call
Slack
Post RCA summaries, blast-radius maps, and remediation approval requests to channels. Agent pings the on-call directly.
Not connected
lock_open Unlocks auto post-incident summaries
PagerDuty
Bi-directional: receive PD alerts as incidents, and auto-resolve PD incidents when the agent confirms fix.
Not connected
lock_open Unlocks auto-resolve PD incidents
OpsGenie
Alert enrichment with RCA context, and auto-close when agent resolves root cause.
Not connected
Microsoft Teams
Incident notifications and approve/reject cards delivered to Teams channels.
Not connected
merge_typeSource Control & GitOps
GitHub
Correlate incidents with recent commits. Auto-open PRs for config fixes. Rollback via revert PR.
Not connected
lock_open Unlocks commit correlation, auto-PR
ArgoCD
Detect failed rollouts, correlate sync failures with incidents, and trigger 1-click rollback to last healthy revision.
Not connected
lock_open Unlocks rollout correlation, 1-click rollback
CI/CD Webhook
Generic webhook for Jenkins, CircleCI, Buildkite — tag incident timelines with deployment events.
Not connected
terminalIDE & AI Interfaces
MCP Server
Expose the agent as an MCP server. Claude, Cursor, or any MCP client can call get_incidents, get_rca, approve_remediation from the terminal or IDE.
Not enabled
lock_open Unlocks Claude CLI + IDE-native incident management
Jira
Auto-create tickets for unresolved incidents with the full RCA report and blast-radius attached.
Not connected
lock_open Unlocks auto-ticket on escalation
ServiceNow
Enterprise ITSM — auto-open change requests before the agent executes remediations, satisfying change-control.
Not connected
psychologyKnowledge Base
gavelVerification Mode
Controls how the agent verifies that a remediation action actually resolved the incident. Deterministic — re-checks the metric that triggered the alert (fast, no LLM cost). LLM Judge — calls Claude Haiku to evaluate the before/after evidence snapshot and flag side effects or inconsistent actions. Both — runs deterministic first; only calls the judge when the metric result is ambiguous.
Applied per-namespace — stored in DynamoDB
info LLM Judge runs after a fault-class-specific stabilisation delay (15–60 s) and retries up to 3× on uncertain_recovering verdict before marking the incident REMEDIATION_FAILED.
verified_user
Cluster Pre-flight Check
Checking prerequisites for cluster integration…
refresh Running pre-flight checks…
Build a custom integration
Connect any tool not listed above — ITSM platforms, internal APIs, custom metrics endpoints. The AI assistant will guide you through generating the integration skill file.
info Mockup data — the metrics, charts, and trends on this page are illustrative. In production they are computed from resolved incidents using actual MTTR, SRE hours, and autonomous-action rates.
insights Insights
Business value, reliability trends, and AI-driven recommendations — updated after every incident.
Avg MTTR
4.2 min
vs 38 min baseline
▲ 89% faster
SRE Hours Saved
14.2 hrs
≈ $2,840 at $200/hr
▲ this period
Auto-Resolved
68%
of incidents, no operator action
▲ was 41% last month
Alert Noise Reduced
73%
raw signals suppressed
▲ suppression gate active
payments Business Value
MTTR trend, SRE cost savings, and observability consolidation — the numbers to share with leadership.
timer MTTR Trend (minutes)
38m 19m 0m baseline
7 days agonow: 4.2 min
savings Cost Savings Breakdown
Incidents auto-resolved 17
Avg time saved / incident 50 min
Total SRE hours saved 14.2 hrs
Estimated $ saved $2,840
hub Observability Cost Coverage
Signals the agent currently covers vs. what would require manual review or additional tooling.
Kubernetes metrics
92%
K8s logs (Vector)
88%
AWS infrastructure
74%
IAM / audit trail
55%
health_and_safety Operational Reliability
Fault class distribution, alert noise, blast radius prevented, and per-service reliability — where to focus next.
pie_chart Incidents by Fault Class
Network partition
10
Resource exhaustion
6
IAM / permission
5
Config drift
2
Network partition is the dominant fault class — topology.md coverage directly reduces RCA time here.
notifications_off Alert Noise Reduction
— Raw signals — Incidents opened
Raw signals (7d) 847
Incidents opened 25
Suppression rate 97.1%
shield Blast Radius Prevented
8 cascading failures
prevented this period
Incidents caught at source service before downstream services degraded — enabled by topology.md shared-service-first detection.
Services protected frontend, checkout, auth
Avg detection lead +2.4 min early
leaderboard Service Reliability (30d MTTR trend)
checkout 8.1 min ▲ worsening
postgres 5.3 min → stable
frontend 2.1 min ▼ improving
auth 1.8 min ▼ improving
checkout MTTR is rising — 3 network partition incidents in 14 days. See recommendations below.
psychology Agent Health
Is the agent improving? Autonomous rate trend, integration impact, and AI-driven recommendations based on incident patterns.
trending_up Autonomous Rate Trend
100% 50% 0% 80% target
30d ago: 41%now: 68%
68% auto-resolved 24% operator-approved 8% escalated
As memory.md and skills.md grow with each incident, the autonomous rate increases. Target: 80% within 90 days.
cable Integration Impact (this period)
Integration Usage Steps Hrs saved
topology.md
22 steps 5.2 hrs
Vector
25 steps 4.8 hrs
skills.md
14 steps 1.4 hrs
memory.md
10 steps 0.9 hrs
auto_awesome AI Recommendations — based on your incident history
warning
Add automated NetworkPolicy restore skill for checkout → postgres
checkout → postgres network partition has occurred 3× in 14 days. Each required manual egress rule restore (avg 12 min). A skill that detects the pattern and patches the NetworkPolicy automatically would have prevented 36 min of downtime.
Add to skills.md →
memory
Document postgres normal latency window in memory.md
2 of the last 5 RCA sessions queried postgres p99 latency and had no baseline to compare against — the agent fell back to generic thresholds. Adding a memory note ("postgres p99 > 50ms during nightly backup window is normal") would eliminate false-positive escalations during that window.
Add to memory.md →
cloud_queue
Connect CloudTrail to improve IAM incident confidence
The 2 IAM permission-drift incidents this period both completed with moderate confidence (52%, 58%) because the CloudTrail integration is not connected. Without audit logs, the agent cannot confirm which IAM change triggered the AccessDenied — it infers from patterns only. CloudTrail would raise confidence to 85%+.
Connect CloudTrail →
notifications
Connect Slack to reduce operator response lag on escalated incidents
8% of incidents were escalated and required operator intervention. Average time from escalation to operator acknowledgement: 9.4 min. Slack integration would send a rich-context notification the moment an incident is escalated, with the RCA summary attached — reducing ack time to under 2 min.
Connect Slack →
Active Incidents
smart_toy
AI SRE Assistant
System overview — ask about active incidents, patterns, or coverage
Resolved?
last 24h
Pending Approval?
awaiting review
Failed?
needs attention
Avg MTTR?
auto-remediated
Est. Cost Saved?
vs. manual triage
Tools
hub
Link to group
smart_toy AI SRE Assistant
No incident selected
Settings
Connection
API Endpoint ?
API Token ?
Active Namespace ?
Models
RCA Model ?
Assistant Model ?
RCA Behaviour
Max Tool Iterations ?
Observability
Prometheus URL ?
Loki URL ?
Remediation
Namespace Lock ?
Require Approval ?
Master switch — pause before any fix
Decision order (first match wins):
1. Hard blockers — protected namespace, blocked tool, no RCA, repeated failure → always blocked
2. Fault-class policy — per fault type override below → auto / pending / block
3. Require Approval ON → always pending (operator must approve)
4. Auto-approve class + confidence threshold met → auto-execute
5. Default → pending
Auto-Approve Classes ?
Protected Namespaces ?
Connect a cluster in Integrations to see namespaces.
Per Fault-Class Policy ?
network_partition ?
resource_exhaustion ?
iam_drift ?
bad_deployment ?
Block on Repeated Failures ?
Require review after a failed attempt
Connected Tools ?
Loading…
CloudTrail / IAM Drift
AWS Layer (CloudTrail) ?
Poll CloudTrail for IAM mutations
info Watched roles are auto-discovered from service account IRSA annotations in the topology snapshot. The list below is read-only — re-run topology discovery to update it.
CloudTrail Lookback Window ?
minutes per poll cycle
Signal Configuration
Namespace Mode ?
Signal Source ?
Observability Backend ?
Prometheus Rule Selector ?
K8s API Polling (K8sGPT) ?
NetworkPolicy mutations + pod failures — works on any namespace
Namespace Discovery ?
Opt out: kubectl label namespace <name> ai-sre-agent/monitored=false
Correlator Rules
These rules run in fixed priority order. Shared Resource Root and Attractor Routing always run.
Bad Deployment Detection ?
CrashLoop + RS mismatch + recent rollout → bad_deployment class
Same Fault Group ?
Same signal type across services in one window → single incident
Cross-Signal Dedup Policy ?
The correlator always runs 4 chained rules before this setting applies: shared-resource → bad-deployment → attractor routing → same-fault grouping.
Signal Confirmation
Confirmation Threshold ?
polls (~5 min to incident)
1 — fastest10 — balanced20 — strictest
CloudTrail signals (IAM drift) require 5 consecutive detections. Golden signal alerts (error rate, saturation) bypass this gate and fire immediately.
Signal Pipeline
Loading…
IAM Permissions ?
Attach to the Lambda execution role ai-sre-lambda-exec in AWS IAM.
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "SSMKubeToken",
      "Effect": "Allow",
      "Action": ["ssm:GetParameter","kms:Decrypt"],
      "Resource": [
        "arn:aws:ssm:*:*:parameter/sre/*",
        "arn:aws:kms:*:*:key/*"
      ]
    },
    {
      "Sid": "DynamoDB",
      "Effect": "Allow",
      "Action": [
        "dynamodb:GetItem","dynamodb:PutItem",
        "dynamodb:UpdateItem","dynamodb:DeleteItem",
        "dynamodb:Scan","dynamodb:Query",
        "dynamodb:DescribeStream","dynamodb:GetRecords",
        "dynamodb:GetShardIterator","dynamodb:ListStreams"
      ],
      "Resource": "arn:aws:dynamodb:*:*:table/ai-sre-*"
    },
    {
      "Sid": "Bedrock",
      "Effect": "Allow",
      "Action": ["bedrock:InvokeModel","bedrock:InvokeModelWithResponseStream"],
      "Resource": "arn:aws:bedrock:*::foundation-model/*"
    },
    {
      "Sid": "CloudTrailRead",
      "Effect": "Allow",
      "Action": ["cloudtrail:LookupEvents"],
      "Resource": "*"
    },
    {
      "Sid": "CloudWatchRead",
      "Effect": "Allow",
      "Action": [
        "cloudwatch:GetMetricData","cloudwatch:GetMetricStatistics",
        "logs:StartQuery","logs:GetQueryResults",
        "logs:FilterLogEvents","logs:DescribeLogGroups"
      ],
      "Resource": "*"
    },
    {
      "Sid": "IAMRead",
      "Effect": "Allow",
      "Action": [
        "iam:GetRolePolicy","iam:ListRolePolicies",
        "iam:ListAttachedRolePolicies",
        "iam:SimulatePrincipalPolicy","iam:GetRole"
      ],
      "Resource": "*"
    },
    {
      "Sid": "IAMRemediation",
      "Effect": "Allow",
      "Action": [
        "iam:DeleteRolePolicy",
        "iam:DetachRolePolicy"
      ],
      "Resource": "arn:aws:iam::*:role/ai-sre-sandbox-*"
    },
    {
      "Sid": "RDSELB",
      "Effect": "Allow",
      "Action": [
        "rds:DescribeDBInstances","rds:DescribeDBClusters",
        "elasticloadbalancing:DescribeTargetHealth",
        "elasticloadbalancing:DescribeLoadBalancers"
      ],
      "Resource": "*"
    }
  ]
}
Kubernetes RBAC ?
Apply to your cluster: kubectl apply -f agent-rbac.yaml then run setup-k8s-auth.sh to push the token to SSM.
apiVersion: v1
kind: ServiceAccount
metadata:
  name: ai-sre-agent
  namespace: kube-system
---
apiVersion: v1
kind: Secret
metadata:
  name: ai-sre-agent-token
  namespace: kube-system
  annotations:
    kubernetes.io/service-account.name: ai-sre-agent
type: kubernetes.io/service-account-token
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
  name: ai-sre-agent
rules:
- apiGroups: [""]
  resources: [pods,services,events,namespaces,nodes,configmaps]
  verbs: [get,list,watch]
- apiGroups: [""]
  resources: [pods]
  verbs: [delete]
- apiGroups: [apps]
  resources: [deployments,replicasets]
  verbs: [get,list,watch,update,patch]
- apiGroups: [networking.k8s.io]
  resources: [networkpolicies]
  verbs: [get,list,watch,update,patch]
- apiGroups: [autoscaling]
  resources: [horizontalpodautoscalers]
  verbs: [get,list,watch,update,patch]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
  name: ai-sre-agent
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: ClusterRole
  name: ai-sre-agent
subjects:
- kind: ServiceAccount
  name: ai-sre-agent
  namespace: kube-system
SSM Parameters ?
K8s Agent Token
/sre/k8s-agent-token
ServiceAccount token for kubectl access. Set by setup-k8s-auth.sh.
Backend Credentials ?
Alertmanager Webhook Token
/sre/alertmanager-token
Shared secret sent by Alertmanager in http_config.bearer_token.
aws ssm put-parameter --name /sre/alertmanager-token --value TOKEN --type SecureString --overwrite
Incidents API Token
/sre/incidents-api-token
Bearer token for the incidents REST API — used by this dashboard and the DataAgent CLI.
aws ssm put-parameter --name /sre/incidents-api-token --value TOKEN --type SecureString --overwrite
CloudFormation Stack ?
ai-sre-remediation-agent
us-west-2
Manages: 6 Lambda functions · API Gateway · DynamoDB tables · SQS queues · IAM roles · CloudWatch log groups
open_in_newOpen in AWS Console
sam build --use-container && sam deploy --region us-west-2 --no-confirm-changeset
80% coverage
Incident coverage: 80%
Unlock +18% more coverage
Connect CloudTrail to detect IAM permission-drift incidents the K8s agent can't see — policy changes that cause AccessDenied without any metric anomaly.
+IAM permission drift (CloudTrail)
·RDS connectivity failures coming soon
·S3 access denied errors coming soon
Dismiss for now
sync Integration Health Check
article POC Architecture
Future Direction
Chat-first interface — directional concept, not current roadmap. This illustrates the long-term UX inversion: the AI assistant becomes the primary interface; the incident table becomes a sidebar. The agent is proactive — it summarizes the situation without being asked. Generative UI components assemble inline for each question. The current dashboard (table-centered) remains the primary interface today.
Active · 3 incidents
smart_toy
DataAgent AI SRE
Chat-first · agent speaks first · generative UI inline
Discover EKS Clusters
Scan AWS regions for available clusters
REGIONS TO SCAN
Integration Setup Wizard
Custom Integration Wizard
System Health
Click "Run checks" to test connectivity
Pipeline Check
Injects a synthetic incident and verifies it flows through
CloudWatch alarm → EventBridge → anomaly detector → enricher → RCA engine
Tool Connectivity
Checks K8s API, Helm, CloudWatch, CloudTrail, Bedrock, and K8sGPT
terminalTerminal
Cluster Namespace
Digital Immune Agent  ·  Autonomous SRE Platform  ·  v1.0
kubectl  dataagent  ·  type help for all commands
Install local CLI:  curl -sSL https://demo.data-agent.co/install.sh | sh
$