Observability, Telemetry & SRE Monitoring

.cursor/skills/sre-observability

Structured logging, OpenTelemetry tracing, Prometheus metrics, and alerts.

Where it installs

# .claude/skills/sre-observability/SKILL.md
---
name: sre-observability
description: ...
---
# identical content, skills are the same file in both editors

What it does

Observability, Telemetry & SRE Monitoring is a Cursor skill: a SKILL.md playbook that AgenticKit installs at .cursor/skills/sre-observability. Cursor's agent loads it when the task matches the skill description, instead of stuffing the same rules into every chat.

Observability standards for production SaaS apps, covering structured log formats, distributed tracing, error tracking with Sentry, and uptime checks.

Structured logging, OpenTelemetry tracing, Prometheus metrics, and alerts. The file is domain knowledge, not a slash macro. You do not type it. The agent reads it when the job is sre observability.

Why it exists

Default Cursor has no memory of how your team ships code, schema, and tests. Founders paste the same conventions into chat, then watch the model drift on the next turn. This skill exists so those conventions live on disk and load only when relevant.

It ships in the Engineering Kit. It is wired to the Site Reliability Engineer (sre-engineer) and Cloud Infrastructure Architect (cloud-architect) agents. It is wired to the /sre-incident: SRE Incident Responder & Root Cause Analysis and /audit: Codebase Quality & Security Auditor commands. That graph is the point: the skill is the standard, the agent is the role, the command is the trigger.

When to use it

  • Use Observability, Telemetry & SRE Monitoring when the work is specifically about structured logging, opentelemetry tracing, prometheus metrics, and alerts.
  • Load it before a long session that will touch this surface more than once. A one-line edit does not need the full playbook.
  • Use it on production SaaS work (multi-tenant apps, paid features, anything that will be reviewed) rather than a throwaway prototype.

When not to use it

  • Do not treat the skill as a replacement for a slash command. Skills teach. Commands run a pipeline.
  • Do not paste the whole SKILL.md into chat. If Cursor is not loading it, fix the description or invoke the matching command.
  • Do not use it as a generic 'write better code' rule. Scope is this domain only.

Example workflow

  1. Install the Engineering Kit so .cursor/skills/sre-observability lands in the repo.
  2. Open a new Cursor agent chat pointed at the files this skill governs.
  3. Ask the Site Reliability Engineer (sre-engineer) to read the skill, or run /sre-incident: SRE Incident Responder & Root Cause Analysis.
  4. Review the first artifact against the practices below. If it violates one, stop and correct the file. Do not prompt 'just finish it'.
  5. Commit the skill-guided files with the rest of the slice so the next session inherits the same standard.

Example usage

Example prompt: "Read .cursor/skills/sre-observability and apply it to this change. Do not invent extra conventions."

Or trigger the pipeline: /sre-incident: SRE Incident Responder & Root Cause Analysis. That command is written to load this skill.

Check the output against: Emit all application logs in structured JSON format with correlation IDs

Example output

  • The skill itself does not write a single output file. It changes what the agent is allowed to produce in code, schema, and tests.
  • When you run /sre-incident: SRE Incident Responder & Root Cause Analysis, expect repo files plus notes under docs/, not a chat-only answer.

Best practices

  • Emit all application logs in structured JSON format with correlation IDs
  • Instrument API routes with OpenTelemetry spans to trace distributed latency
  • Set up synthetic uptime monitors pinging health endpoints every 60 seconds
  • Configure alerts for 5xx error spikes and p99 response times exceeding 1000ms

Common mistakes

  • Ignoring "Emit all application logs in structured JSON format with correlation IDs" and hoping a later prompt will clean it up.
  • Copying the skill into .cursor/rules as always-on. That burns context and fights Cursor's load-on-match design.
  • Running two overlapping skills that contradict each other in the same turn.
  • Letting the agent skip tests or types because 'the skill is about architecture'.

Frequently asked questions

  • What is the Observability, Telemetry & SRE Monitoring Cursor skill?
    Observability standards for production SaaS apps, covering structured log formats, distributed tracing, error tracking with Sentry, and uptime checks. It lives at .cursor/skills/sre-observability after you install the Engineering Kit.
  • When should I use sre-observability instead of a Cursor rule?
    Use a rule for always-on or glob-scoped constraints. Use this skill for the full playbook that should load only when the task matches.
  • Which agents read sre-observability?
    Site Reliability Engineer (sre-engineer); Cloud Infrastructure Architect (cloud-architect). Those roles are told to open this file before they edit.
  • Which commands use sre-observability?
    /sre-incident: SRE Incident Responder & Root Cause Analysis; /audit: Codebase Quality & Security Auditor.
  • Does this skill work outside AgenticKit?
    Yes. A SKILL.md in .cursor/skills/sre-observability is a normal Cursor skill. AgenticKit is the packaged version plus the agent and command graph.

Add sre-observability and the other 48 Cursor skills

AgenticKit installs 61 skills, 46 agents, and 47 slash commands into .cursor/. One license, lifetime updates.