bmad-observability-agent

B-MAD Observability Agent

A comprehensive OpenTelemetry observability expert agent for B-MAD (Breakthrough Method for Agile AI Driven Development).

License: MIT B-MAD Documentation

🎯 What is this?

The B-MAD Observability Agent is an AI-powered expert that helps you build production-grade observability using OpenTelemetry. It provides:

🚀 Quick Start

Prerequisites

Installation

# Install via BMad CLI (requires BMad Method v6.3+ installed)
npx bmad-method install --custom-source https://github.com/henrikrexed/bmad-observability-agent

The installer will:

First Use

Invoke the O11y Engineer agent:

/o11y-engineer

Or run a specific capability directly:

/o11y-write-ottl        # Generate OTTL expressions
/o11y-generate-epics    # Create sprint-ready epics
/o11y-instrument-app    # Add OTel to your app
/o11y-redact-pii        # Configure PII redaction

🏗️ Architecture: The Observability Architect

The O11y Engineer acts as an Observability Architect — it designs and plans but does not implement code directly. It generates epics and stories for the BMAD agent team:

O11y Architect                    BMAD Agent Team
     │                                 │
     ├── Assess maturity               │
     ├── Design observability spec     │
     ├── Design collector pipeline     │
     ├── Generate epics ──────────────►│ Bob (Scrum Master) plans sprints
     │                                 ├──► Amelia (Developer) implements
     │                                 ├──► Murat (Test Architect) tests
     │◄── Quality gate (score ≥ 90) ──┤
     │                                 │
     └── Last Mile (direct):           │
         → Define SLI/SLO/KPI         │
         → dtctl apply dashboards     │
         → dtctl apply SLOs           │
         → dtctl apply workflows      │

Generated Epics

Running /o11y-generate-epics produces 6 epics across 4 sprints:

Sprint Epic Owner
1 Assessment & Observability Spec O11y Architect
1-2 Collector Pipeline (OTTL, PII, sampling) O11y Architect → Amelia
2 Custom Collector Distribution (OCB) O11y Architect → Amelia
2-3 Application Instrumentation (per-service) Amelia
3 Observability Test Suite Murat
4 Last Mile: SLI/SLO/KPI + Dynatrace O11y Architect (direct)

The Last Mile epic is a quality gate — it only starts after all tests pass (score ≥ 90).

📚 Documentation

Full documentation: https://henrikrexed.github.io/bmad-observability-agent/

🎯 Key Features

Intelligent Intent Detection

Ask natural questions via /o11y-engineer and get the right workflow:

You: "How do I know if my observability is good?"
Agent: Runs comprehensive quality checks and provides roadmap (QC menu)

You: "I need to create custom metrics"
Agent: Guides you through semantic convention design with Weaver (SC menu)

You: "My collector keeps crashing"
Agent: Identifies issues and provides fixes (DP menu)

Comprehensive Quality Checks

Select QC from the O11y Engineer menu to assess:

Score: 0-100 with actionable recommendations

Production-Grade Workflows

Capability Purpose Time
/o11y-engineer → QS Complete observability setup from scratch 2-4 weeks
/o11y-engineer → AM Maturity assessment + improvement roadmap 30 min
/o11y-engineer → CP Design OTel Collector pipeline 1-2 hours
/o11y-engineer → BD Build custom collector with OCB 2-4 hours
/o11y-engineer → SC Validate against semantic conventions 1 hour
/o11y-engineer → DD Create Dynatrace dashboard as code 30 min
/o11y-engineer → PD Build dashboard with discovered metrics (MCP) 15 min
/o11y-engineer → DB Build diagnostic notebook (MCP) 15 min
/o11y-engineer → SW AI-suggested automation workflows (MCP) 10 min

💡 Use Cases

For Homelab Enthusiasts

For Content Creators

For Production Environments

🤝 Multi-Agent Collaboration (BMAD)

This agent supports seamless handoff to other BMAD agents via /o11y-engineer:

# Generate handoff for next agent
Select HO from the O11y Engineer menu

# Create epics/stories for tracking
/o11y-generate-epics

# Get machine-readable status
Select SR from the O11y Engineer menu

# Sync from previous agent session
Select SS from the O11y Engineer menu

Handoff Output Example:

handoff:
  agent: "o11y-engineer"
  observability_status:
    overall_score: 78
    production_ready: false
  completed_actions:
    - action: "Configured OTel Collector"
      result: "success"
  pending_tasks:
    - task: "Add memory_limiter"
      priority: "critical"
  recommendations:
    immediate:
      - "Scale collector to 3 replicas"

🛠️ Agent Capabilities

OpenTelemetry Collector

Instrumentation

Semantic Conventions

Dynatrace Automation

With the Dynatrace MCP server, the agent can:

# MCP-powered capabilities (via /o11y-engineer menu)
PD  # Build dashboard with real metrics
DB  # Build troubleshooting notebook
SW  # Get AI-suggested automations

OTTL Transformations

Sensitive Data & PII Protection

Per-Language SDK Instrumentation

Sprint-Ready Epic Generation

📊 Example Output

/o11y-engineer → select QC

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
OBSERVABILITY QUALITY REPORT
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Overall Score: 78/100 ⚠️  NEEDS IMPROVEMENT

✅ PASSED (18 checks):
  ✓ Traces Present (15/15 pts)
  ✓ Metrics Present (15/15 pts)
  ✓ Collector HA - 3 replicas (10/10 pts)
  ...

⚠️  WARNINGS (5 checks):
  ! Semantic Convention Compliance - 87% (12/15 pts)
    Target: 95%+
    Fix: /o11y-engineer → SC

❌ FAILURES (3 checks):
  ✗ SLOs Not Configured (0/10 pts) 🚨 CRITICAL
    Fix: /o11y-engineer → DA

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
PRIORITY ACTIONS TO REACH 95+ (PRODUCTION-READY)
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

1. [CRITICAL] Configure SLOs (+10 pts)
   /o11y-engineer → DA
   Effort: 1 day

🤝 Contributing

Contributions welcome! Please read CONTRIBUTING.md first.

📝 License

MIT License - see LICENSE for details.

🙏 Acknowledgments

📺 Resources


Need help? Open an issue or reach out on Discord