Deployment

Observability & Incident Response

Designs production telemetry, alerting, runbooks and incident command so problems are caught and handled fast.

What it does

Designs production telemetry, alerting, runbooks and incident command so problems are caught and handled fast.

When to use it

Use for structured logs, metrics, alerting design, on-call handoffs, or building real operational readiness.

Problems it prevents

Stops incidents being discovered by a customer complaint instead of an alert that actually fires on the real signal.

Typical workflow

  1. Define the outcomes worth observing
  2. Instrument safely, without leaking data
  3. Alert only on what needs action
  4. Run post-incident review

Example output

An alert that fires on a real user-facing signal, with a runbook the on-call person can actually follow at 3am.

What's included

  1. 01
    The complete prompt

    The full manual as a ready-to-save SKILL.md file, matched exactly to what ships in the T7 library.

  2. 02
    Install walkthrough

    Personal (~/.claude/skills/) and project (.claude/skills/) folder instructions, whichever fits how you work.

  3. 03
    Use notes

    How Claude invokes it, what a good result looks like, and what to check before you trust the output.

Installation

Save the manual as SKILL.md in ~/.claude/skills/observability-incident-response/ to use it on every project, or in .claude/skills/observability-incident-response/ to scope it to one repository. Then invoke it directly:

/observability-incident-response

Read the official Claude Code skills guide

Works well with

Save by bundling

Get Observability & Incident Response as part of Reliability Builder: five skills that work together as one toolkit, for £1.96 less than buying all five separately.

View Reliability Builder (£9.49)