The Closed-Loop Incident Response Pattern: How AWS DevOps Agent Uses OpenSearch MCP to Investigate Itself

The Closed-Loop Incident Response Pattern: How AWS DevOps Agent Uses OpenSearch MCP to Investigate Itself

Cloud Platform Engineering

The pattern that makes this integration interesting isn’t just that DevOps Agent can query OpenSearch — it’s that the same system doing the alerting feeds the same system doing the investigation. When an anomaly detection monitor fires, the agent isn’t context-switching to a different tool. It’s querying the index that generated the alert.

The Mechanism

The six-stage closed loop works as follows:

  1. Applications emit logs and distributed traces to OpenSearch indices (typically application-logs-* and otel-traces-*)
  2. OpenSearch alerting monitors detect anomalies and publish to SNS
  3. SNS triggers a Lambda webhook forwarder
  4. The forwarder transforms notifications into HMAC-signed payloads and forwards them to DevOps Agent
  5. DevOps Agent queries OpenSearch indices via SearchIndexTool and ListIndexTool through MCP
  6. The agent correlates logs and traces with CloudTrail and CloudWatch data, then delivers a root cause analysis

AWS measures this at approximately five minutes from alert trigger to remediation recommendation.

The Three Hosting Paths

The OpenSearch MCP server can be deployed three ways, each suited to different operational contexts:

Self-managed on ECS (universal): Runs opensearch-mcp-server-py on Fargate behind an internal Network Load Balancer, accessible via VPC Lattice. Requires TLS certificates but works in all AWS regions — the choice when you need geographic flexibility or have existing ECS infrastructure.

Amazon Bedrock AgentCore (managed): One-click CloudFormation in select regions. Removes infrastructure management overhead at the cost of regional availability. The right choice when your team doesn’t want to own the MCP server runtime.

Built-in (OpenSearch 3.3+): Domains running version 3.3 or later expose an MCP endpoint directly at /_plugins/_ml/mcp — no separate server required. If you’re already on a current OpenSearch version, this is zero-infrastructure overhead.

The Authentication Layer

Security runs at two levels: IAM authenticates the caller, while OpenSearch’s fine-grained access control enforces index-level authorization. The separation matters — a compromised agent credential gets IAM-level access, not unrestricted OpenSearch access.

This is worth designing deliberately when you first configure the integration. Index-level authorization lets you scope which investigation data the agent can reach without granting blanket observability access.

The So What

The practical value here is convergence. Most incident response workflows span multiple tools: CloudWatch for metrics, OpenSearch or a third-party SIEM for logs and traces, a separate system for distributed traces, and PagerDuty or similar for alerting. The investigation bounces between systems while the engineer tries to hold the full picture in their head.

The DevOps Agent + OpenSearch MCP pattern flattens that: the same alerting fabric that caught the anomaly becomes the query target for the autonomous investigation. Engineers don’t get a better alert — they get an alert that already includes the first five minutes of investigation.

The five-minute metric is worth calibrating against your current mean time to investigate. If your team’s first-response protocol takes 15-20 minutes to scope an incident, having the correlations already in place when the pager fires is a structural change to the incident timeline — not just a workflow convenience.

Content created with AI assistance and reviewed for accuracy.

💬

Join the conversation

Stack Insiders is our free community for readers who want to go deeper — share resources, ask questions, and connect with others across every vertical we cover.

Join Stack Insiders →

Newsletter coming soon.

Curated digests across AI, biohacking, photography, travel, and more. Be the first to know when we launch.