/sre-incident

SRE Incident Responder & Root Cause Analysis

Log analyzer, downtime diagnostics, failover runbooks, and post-mortems.

/sre-incident Analyze high CPU usage logs and memory leak traceback to generate post-mortem

Where it installs

# .claude/commands/sre-incident.md
---
description: ...
argument-hint: [incident symptom, alert, or error]
allowed-tools: Read, Grep, Glob, Skill, TodoWrite, Task, ...
---

What it does

/sre-incident is a Cursor slash command. Type it in chat to run a saved workflow: SRE Incident Responder & Root Cause Analysis.

Diagnoses production downtime and latency spikes: analyzes structured logs, traces memory leaks, executes recovery runbooks, and generates post-mortems.

Log analyzer, downtime diagnostics, failover runbooks, and post-mortems. Unlike a skill, a command is something you invoke on purpose. The agent does not decide to run /sre-incident for you.

Why it exists

Founders retype the same multi-step prompt until it rots. /sre-incident exists so the pipeline, all 4 steps of it, is a file in .cursor/commands, versioned with the repo.

It ships in the Engineering Kit. It is wired to the Site Reliability Engineer (sre-engineer) and Root Cause Debugger (debugger) agents. Required skills: Observability, Telemetry & SRE Monitoring; Resilient Error Handling & Fault Tolerance.

When to use it

  • Use during production outages or performance degradations to isolate issues.
  • Use /sre-incident when you want that pipeline, not a freeform chat. If you only need one step, use a narrower command or a single agent.
  • Start a new chat. Do not run this command in a thread that just wrote marketing copy.

When not to use it

  • Do not run /sre-incident as a substitute for reading the diff. The command produces files; you still gate them.
  • Do not chain it into a 40-turn chat. Fresh context is part of the design.
  • Do not run it if you have not filled CURSOR.md. The pipeline will invent a stack.

Example workflow

  1. Parse server error logs and trace distributed latency spikes
  2. Identify failing upstream dependencies and resource exhaustion bottlenecks
  3. Formulate immediate mitigation and failover action steps
  4. Generate customer-ready incident post-mortem markdown report

Example usage

Type this in Cursor chat: /sre-incident Analyze high CPU usage logs and memory leak traceback to generate post-mortem

The command file tells the session which agents to adopt and which skills to read. You should see phase headers, not a single dump of code.

If a phase fails its gate, stop. Do not add 'just continue'.

Example output

  • Expected artifact: docs/incidents/post-mortem-YYYY-MM-DD.md

Best practices

  • Keep the prompt specific. /sre-incident Analyze high CPU usage logs and memory leak traceback to generate post-mortem is the shape: object, constraint, and outcome.
  • Let the listed agents work in order: sre-engineer → debugger.
  • Save outputs in the repo. Chat-only answers evaporate.
  • Engineering commands should leave tests or an audit note, not only implementation files.

Common mistakes

  • Typing /sre-incident with no object ('do the thing'). The pipeline will guess.
  • Re-running the command in the same chat after a failed gate instead of fixing the failing file.
  • Editing the command file to skip review so it 'goes faster'.
  • Skipping the Observability, Telemetry & SRE Monitoring skill that the command depends on.

Frequently asked questions

  • What does /sre-incident do in Cursor?
    Diagnoses production downtime and latency spikes: analyzes structured logs, traces memory leaks, executes recovery runbooks, and generates post-mortems.
  • When should I run /sre-incident?
    Use during production outages or performance degradations to isolate issues.
  • What is an example /sre-incident prompt?
    /sre-incident Analyze high CPU usage logs and memory leak traceback to generate post-mortem
  • Which skills does /sre-incident load?
    Observability, Telemetry & SRE Monitoring; Resilient Error Handling & Fault Tolerance
  • Is /sre-incident a Cursor skill?
    No. /sre-incident is a slash command you type. Skills are playbooks the agent may load. Use both: the command runs the workflow, the skills constrain how it writes.

Run /sre-incident from your own repo

AgenticKit installs 47 slash commands, 46 agents, and 61 skills. One command installation.