On-Call Incident Agent

Agentic incident response·2026·Solo build

On-call engineers spend the first 10-15 minutes of every incident doing the same thing: checking logs, checking recent deploys, checking metrics, searching runbooks. This project automates that first pass. A CloudWatch alarm is delivered via SNS to a webhook that creates an incident record and starts an investigation; a hand-rolled tool-calling agent gathers evidence the way a human would and posts a structured summary to Slack: what happened, the likely root cause, every check it ran (including the ones that found nothing), and a next step grounded in the matched runbook. The on-call engineer approves or dismisses it with one click.

The constraint

The first minutes of an incident are spent on mechanical triage, not judgement. An agent can do that pass, but only if it shows its work. A confident root cause with no evidence trail is worse than none, and an agent with no budget will keep calling tools long after it has enough signal.

How it was built

  • –CloudWatch alarm → SNS → webhook that creates the incident record and kicks off the investigation
  • –Hand-rolled tool-calling loop on OpenAI function calling, no LangChain or LangGraph
  • –Four tools: CloudWatch Logs Insights, GitHub deploy history, CloudWatch Metrics, and RAG over internal runbooks (Milvus + OpenAI embeddings)
  • –Stops when it has enough signal or hits a call/time budget
  • –Structured Slack summary listing every check, including the ones that found nothing, with one-click approve/dismiss via Socket Mode
  • –The same four tools exposed as an MCP server, so the investigation logic works from any MCP client

Stack

TypeScript
Express
PostgreSQL
Prisma
Redis
OpenAI
AWS
Milvus
Slack API
MCP

The repository is private. Email sidhantsinghrathoreprsnl@gmail.com with your GitHub username and I'll add you.

Built with love by Sidhant Singh Rathore