backend-observability

内容来源:README.md(说明文档) · 原始地址 · 查看安装指南

原始内容

Backend Observability Plugin for Claude Code

A comprehensive Claude Code plugin for implementing production-ready observability in backend services. Covers the full spectrum from basic logging to advanced distributed tracing and SLO-based alerting.

Features

Commands

  • /instrument - Instrument backend services with observability
  • /audit - Audit existing observability implementation

Skills

Skill Description
instrumentation-planning Plan what to measure based on JTBD framework
request-tracing Implement distributed tracing with context propagation
database-observability Instrument database queries, pools, and transactions
cache-observability Monitor cache hit rates, latency, and memory
queue-observability Track message queues, consumer lag, and processing
error-handling Capture errors with context and classification
health-checks Implement liveness, readiness, and startup probes
slo-alerting Define SLIs/SLOs and burn-rate alerting

Agents

Agent Purpose
codebase-analyzer Analyze codebase architecture and existing telemetry
instrumentation-reviewer Review code for observability anti-patterns

Hooks

Real-time anti-pattern detection for:

  • Go
  • Python
  • TypeScript/JavaScript
  • Java/Kotlin
  • Rust
  • C#/.NET

Installation

# Add to Claude Code plugins
claude plugins add backend-observability

Or clone directly:

git clone https://github.com/nexus-labs/backend-observability ~/.claude/plugins/backend-observability

Quick Start

1. Analyze Your Codebase

/instrument --analyze

2. Plan Instrumentation

/instrumentation-planning

3. Implement Tracing

/request-tracing

4. Set Up Alerting

/slo-alerting

Reference Documentation

Methodology

  • RED Method - Rate, Errors, Duration for request-driven services
  • USE Method - Utilization, Saturation, Errors for resources
  • Four Golden Signals - Latency, Traffic, Errors, Saturation
  • JTBD for Backend - Job-based criticality classification
  • SLI/SLO Framework - Service level indicators and objectives

Platform Guides

  • Go (OTel, Prometheus, slog)
  • Python (FastAPI, SQLAlchemy, structlog)
  • Node.js (Express, Prisma, pino)
  • Java/Kotlin (Spring Boot, Micrometer)
  • Rust (Axum, tracing, SQLx)
  • .NET (ASP.NET Core, EF Core, Serilog)

Vendor Integration

  • OpenTelemetry (Collector, SDK)
  • Datadog (Agent, APM, Logs)
  • Grafana Stack (Prometheus, Loki, Tempo)
  • New Relic (APM, NRQL)
  • Honeycomb (Wide events, BubbleUp)
  • AWS CloudWatch (X-Ray, EMF, Logs Insights)
  • Jaeger (Distributed tracing)
  • Splunk (Observability Cloud)

Resilience Patterns

  • Circuit Breaker
  • Retry with Backoff
  • Timeout Management
  • Bulkhead Isolation
  • Rate Limiting
  • Distributed Tracing

Templates

Ready-to-use code for each platform (Go, Python, Node.js, Java, Rust, .NET):

  • HTTP middleware with tracing/metrics
  • Database instrumentation wrappers
  • Health check endpoints (liveness/readiness/startup)

Anti-Patterns Detected

The plugin detects common observability anti-patterns:

High Cardinality

// BAD: User ID in metric label
counter.WithLabelValues(userID).Inc()

// GOOD: Use attributes in traces instead
span.SetAttributes(attribute.String("user.id", userID))

Missing Context Propagation

# BAD: Context not propagated
requests.get(url)

# GOOD: Inject trace context
otel.inject(headers)
requests.get(url, headers=headers)

Unbounded Labels

// BAD: Full URL path as label
histogram.labels({ path: req.url }).observe(duration);

// GOOD: Use route pattern
histogram.labels({ route: req.route.path }).observe(duration);

Directory Structure

backend-observability/
├── plugin.json              # Plugin manifest
├── marketplace.json         # Marketplace metadata
├── CLAUDE.md               # Plugin instructions
├── commands/
│   ├── instrument.md       # /instrument command
│   └── audit.md           # /audit command
├── agents/
│   ├── codebase-analyzer/
│   └── instrumentation-reviewer/
├── skills/
│   ├── instrumentation-planning/
│   ├── request-tracing/
│   ├── database-observability/
│   ├── cache-observability/
│   ├── queue-observability/
│   ├── error-handling/
│   ├── health-checks/
│   └── slo-alerting/
├── hooks/
│   └── hooks.json          # Anti-pattern detection
└── references/
    ├── anti-patterns.md   # Common observability mistakes
    ├── methodology/       # RED, USE, 4GS, JTBD, SLI/SLO
    ├── platforms/         # Go, Python, Node.js, Java, Rust, .NET
    ├── vendors/           # OTel, Datadog, Grafana, etc.
    ├── patterns/          # Circuit breaker, retry, timeout, rate limiting
    └── templates/         # Middleware, database, health checks per platform

Best Practices Enforced

Metrics

  • Use semantic conventions for attribute names
  • Avoid high-cardinality labels
  • Group status codes (2xx, 4xx, 5xx)
  • Include units in metric names

Tracing

  • Propagate context across boundaries
  • Use route templates, not IDs
  • Record errors with context
  • Sample appropriately by importance

Logging

  • Use structured logging (JSON)
  • Include trace_id and span_id
  • Log at appropriate levels
  • Avoid logging sensitive data

Health Checks

  • Separate liveness and readiness
  • Include dependency checks
  • Set appropriate timeouts
  • Expose metrics endpoint

Contributing

See CONTRIBUTING.md for guidelines.

License

MIT - See LICENSE for details.