0
Received
›
0
Buffering
›
0
Waiting
›
0
Escalated
›
0
Suppressed
—
Autonomous Rate
—
of resolved incidents
Avg MTTR · 24h
—
Signal Coverage
0 selected
| FIRST SEEN? | LAST FIRED? | OPEN FOR? | CLUSTER? | NAMESPACE? | SERVICE? | DIMENSION? | DEVIATION? | DURATION? | LAYER? | STATE? | SUPPRESSION REASON? | INCIDENT? | INC STATE? | AI? |
|
|
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
monitor_heartSignal stream will appear here | ||||||||||||||||
Integrations
Connect data sources to improve RCA accuracy and workflow tools to extend autonomous actions. More connected = fewer blind spots.
psychology
databaseObservability Sources
Alertmanager
Replaced by CloudWatch EventBridge. Incidents now flow directly: CloudWatch Alarm → EventBridge rule → Lambda, eliminating the Alertmanager webhook dependency.
swap_horiz Replaced by EventBridge
AlertmanagerConfig removed · EventBridge active
Vector
Container metrics + log pipeline. Collects CPU, memory, and log data from pods and ships to CloudWatch — replaces Prometheus and Loki in this environment. Source label:
vector.check_circle Active · ships to CloudWatch
Metrics + logs pipeline active · signal_plane=k8s_oss
Prometheus
Golden signal metrics via Alertmanager. Not used in this environment — metric collection and alerting handled by Vector → CloudWatch pipeline.
swap_horiz Replaced by Vector
Not deployed · no Alertmanager dependency
Loki
Pod log aggregation. Not used in this environment — container logs are collected and shipped to CloudWatch by the Vector pipeline.
swap_horiz Replaced by Vector
Not deployed · logs available via CloudWatch Logs Insights
Datadog
Pull metrics and APM traces from Datadog as an alternative or complement to Prometheus. Supports hybrid-cloud setups.
Not connected
lock_open Unlocks APM trace correlation in RCA
OpenTelemetry
Distributed traces (OTLP). Enables span-level RCA — pinpoints the exact service boundary where a request fails.
Not connected
lock_open Unlocks span-level RCA evidence
cloudCloud Providers
AWS
CloudTrail
IAM change events, S3 access denied, API call history. Detects permission-drift incidents invisible to metrics — PutRolePolicy mutations surfaced via RCA.
check_circle Active — IAM drift detection live
PutRolePolicy events analysed · deny-s3-read detected
CloudWatch Alarms & Logs
Container Insights metric filters → CloudWatch Alarms → EventBridge → Lambda. Primary detection path for error rate, CPU throttling, and IAM drift — no Prometheus or Alertmanager required.
check_circle Active — 3 alarms monitoring
checkout-db-errors · checkout-cpu-high · checkout-iam-deny
Container Insights
Pod, container, node, and cluster-level CPU, memory, disk, and network metrics via CloudWatch. Installed via
amazon-cloudwatch-observability EKS addon — no Prometheus required.warning Not connected
lock_open Pod CPU/memory + node pressure in RCA
EKS Control Plane Logs
API server, controller-manager, scheduler, and audit logs from the EKS control plane in CloudWatch. Detects RBAC denials, API server errors, and cluster-level anomalies invisible in pod metrics.
warning Not connected
lock_open API server audit + RBAC events in RCA
RDS / CloudWatch Metrics
Database CPU, connections, read/write latency from CloudWatch. Enriches RCA with DB-layer evidence for connection and query failures.
warning Not connected
lock_open DB latency + connection evidence in RCA
ALB / Load Balancer
ALB request count, HTTP 5xx rate, and target response time from CloudWatch. Detects upstream load-balancer failures not visible in pod metrics.
warning Not connected
lock_open ALB 5xx rate + target health in RCA
AWS X-Ray
Distributed traces from the checkout service via OTel auto-instrumentation. Pinpoints which service hop introduced latency or errors during RCA.
warning Not connected
lock_open Distributed traces in RCA
DevOps Guru
ML-based anomaly detection for RDS, Lambda, and ALB. Enriches RCA with AWS-native insights.
Coming soon
Az
Azure Monitor
RBAC deny events, AKS infra signals, and Log Analytics workspace queries. Permission-drift RCA for Azure resources.
warning Not connected — Azure RCA blind
lock_open Azure RBAC + AKS infra RCA
Entra ID / Azure AD
RBAC role assignment changes and conditional access events. Detects identity-plane permission drift.
Coming soon
GCP
Cloud Logging
Cloud Audit Logs (IAM permission errors), GKE infra signals, and Cloud Monitoring metrics.
warning Not connected — GCP RCA blind
lock_open GCP IAM + GKE infra RCA
Cloud Monitoring
GCP metrics API — Cloud SQL, Cloud Run, GKE node metrics as native signals alongside Prometheus.
Coming soon
auto_awesomeWorkflow Tools
campaignNotifications & On-call
Slack
Post RCA summaries, blast-radius maps, and remediation approval requests to channels. Agent pings the on-call directly.
Not connected
lock_open Unlocks auto post-incident summaries
PagerDuty
Bi-directional: receive PD alerts as incidents, and auto-resolve PD incidents when the agent confirms fix.
Not connected
lock_open Unlocks auto-resolve PD incidents
OpsGenie
Alert enrichment with RCA context, and auto-close when agent resolves root cause.
Not connected
Microsoft Teams
Incident notifications and approve/reject cards delivered to Teams channels.
Not connected
merge_typeSource Control & GitOps
GitHub
Correlate incidents with recent commits. Auto-open PRs for config fixes. Rollback via revert PR.
Not connected
lock_open Unlocks commit correlation, auto-PR
ArgoCD
Detect failed rollouts, correlate sync failures with incidents, and trigger 1-click rollback to last healthy revision.
Not connected
lock_open Unlocks rollout correlation, 1-click rollback
CI/CD Webhook
Generic webhook for Jenkins, CircleCI, Buildkite — tag incident timelines with deployment events.
Not connected
terminalIDE & AI Interfaces
MCP Server
Expose the agent as an MCP server. Claude, Cursor, or any MCP client can call
get_incidents, get_rca, approve_remediation from the terminal or IDE.Not enabled
lock_open Unlocks Claude CLI + IDE-native incident management
Jira
Auto-create tickets for unresolved incidents with the full RCA report and blast-radius attached.
Not connected
lock_open Unlocks auto-ticket on escalation
ServiceNow
Enterprise ITSM — auto-open change requests before the agent executes remediations, satisfying change-control.
Not connected
psychologyKnowledge Base
gavelVerification Mode
Controls how the agent verifies that a remediation action actually resolved the incident.
Deterministic — re-checks the metric that triggered the alert (fast, no LLM cost).
LLM Judge — calls Claude Haiku to evaluate the before/after evidence snapshot and flag side effects or inconsistent actions.
Both — runs deterministic first; only calls the judge when the metric result is ambiguous.
Applied per-namespace — stored in DynamoDB
info
LLM Judge runs after a fault-class-specific stabilisation delay (15–60 s) and retries up to 3× on uncertain_recovering verdict before marking the incident
REMEDIATION_FAILED.
Cluster Pre-flight Check
Checking prerequisites for cluster integration…
refresh
Running pre-flight checks…
⚙
Build a custom integration
Connect any tool not listed above — ITSM platforms, internal APIs, custom metrics endpoints. The AI assistant will guide you through generating the integration skill file.
info
Mockup data — the metrics, charts, and trends on this page are illustrative. In production they are computed from resolved incidents using actual MTTR, SRE hours, and autonomous-action rates.
insights
Insights
Business value, reliability trends, and AI-driven recommendations — updated after every incident.
payments
Business Value
MTTR trend, SRE cost savings, and observability consolidation — the numbers to share with leadership.
timer MTTR Trend (minutes)
7 days agonow: 4.2 min
savings Cost Savings Breakdown
Incidents auto-resolved
17
Avg time saved / incident
50 min
Total SRE hours saved
14.2 hrs
Estimated $ saved
$2,840
hub Observability Cost Coverage
Signals the agent currently covers vs. what would require manual review or additional tooling.
health_and_safety
Operational Reliability
Fault class distribution, alert noise, blast radius prevented, and per-service reliability — where to focus next.
pie_chart Incidents by Fault Class
Network partition is the dominant fault class — topology.md coverage directly reduces RCA time here.
notifications_off Alert Noise Reduction
— Raw signals
— Incidents opened
Raw signals (7d)
847
Incidents opened
25
Suppression rate
97.1%
shield Blast Radius Prevented
8
cascading failures
prevented this period
prevented this period
Incidents caught at source service before downstream services degraded — enabled by topology.md shared-service-first detection.
Services protected
frontend, checkout, auth
Avg detection lead
+2.4 min early
leaderboard Service Reliability (30d MTTR trend)
checkout
8.1 min ▲ worsening
postgres
5.3 min → stable
frontend
2.1 min ▼ improving
auth
1.8 min ▼ improving
checkout MTTR is rising — 3 network partition incidents in 14 days. See recommendations below.
psychology
Agent Health
Is the agent improving? Autonomous rate trend, integration impact, and AI-driven recommendations based on incident patterns.
trending_up Autonomous Rate Trend
30d ago: 41%now: 68%
68% auto-resolved
24% operator-approved
8% escalated
As memory.md and skills.md grow with each incident, the autonomous rate increases. Target: 80% within 90 days.
cable Integration Impact (this period)
Integration
Usage
Steps
Hrs saved
topology.md
22 steps
5.2 hrs
Vector
25 steps
4.8 hrs
skills.md
14 steps
1.4 hrs
memory.md
10 steps
0.9 hrs
auto_awesome
AI Recommendations — based on your incident history
Add automated NetworkPolicy restore skill for checkout → postgres
checkout → postgres network partition has occurred 3× in 14 days. Each required manual egress rule restore (avg 12 min). A skill that detects the pattern and patches the NetworkPolicy automatically would have prevented 36 min of downtime.
Add to skills.md →
Document postgres normal latency window in memory.md
2 of the last 5 RCA sessions queried postgres p99 latency and had no baseline to compare against — the agent fell back to generic thresholds. Adding a memory note ("postgres p99 > 50ms during nightly backup window is normal") would eliminate false-positive escalations during that window.
Add to memory.md →
Connect CloudTrail to improve IAM incident confidence
The 2 IAM permission-drift incidents this period both completed with moderate confidence (52%, 58%) because the CloudTrail integration is not connected. Without audit logs, the agent cannot confirm which IAM change triggered the AccessDenied — it infers from patterns only. CloudTrail would raise confidence to 85%+.
Connect CloudTrail →
Connect Slack to reduce operator response lag on escalated incidents
8% of incidents were escalated and required operator intervention. Average time from escalation to operator acknowledgement: 9.4 min. Slack integration would send a rich-context notification the moment an incident is escalated, with the RCA summary attached — reducing ack time to under 2 min.
Connect Slack →
smart_toy
AI SRE Assistant
System overview — ask about active incidents, patterns, or coverage
—
Resolved?
last 24h
—
Pending Approval?
awaiting review
—
Failed?
needs attention
—
Avg MTTR?
auto-remediated
—
Est. Cost Saved?
vs. manual triage
Tools