The O11y Engineer provides comprehensive Dynatrace asset management through two complementary approaches: dtctl-based (configuration as code) and MCP-powered (AI-driven discovery).
CD (via /o11y-engineer)Generate dashboard YAML files deployed with dtctl apply.
CD
PD (via /o11y-engineer)Auto-discover metrics in your environment and build dashboards from real data.
PD
type: dashboard
name: "Service Overview"
content:
version: 7
variables: []
tiles:
- title: "Request Rate"
type: data
query:
type: dql
value: |
timeseries requests=sum(http_requests_total), interval:5m
visualization: lineChart
visualizationSettings:
chartSettings:
fieldMapping:
leftAxisValues: [requests]
timestamp: timeframe
querySettings:
maxResultRecords: 1000
defaultScanLimitGbytes: 500
maxResultMegaBytes: 1
defaultSamplingRatio: 10
enableSampling: false
davis: {}
layout:
x: 0
y: 0
w: 12
h: 6
=== “Success Rate”
```dql
timeseries {
success=sum(http_requests_total, filter:{status_code < 500}),
total=sum(http_requests_total)
}, interval:5m
| fieldsAdd success_rate=(success[]/total[])*100
```
=== “Latency Percentiles”
```dql
fetch spans
| filter service.name == "my-service"
| makeTimeseries {
p50=percentile(duration, 50),
p95=percentile(duration, 95),
p99=percentile(duration, 99)
}, interval:5m
| fieldsAdd p50_ms=p50[]/1000000, p95_ms=p95[]/1000000, p99_ms=p99[]/1000000
```
=== “Error Analysis”
```dql
fetch spans
| filter service.name == "my-service"
| filter status.code == "ERROR"
| summarize count=count(), by:{span.name, error.message}
| sort count desc
```
| Type | Use Case |
|---|---|
singleValue |
KPI display (error rate, latency) |
lineChart |
Time series trends |
areaChart |
Stacked area (capacity) |
barChart |
Categorical comparisons |
pieChart |
Proportional breakdown |
table |
Tabular data (traces, logs) |
honeycomb |
Entity health overview |
dtctl apply -f dashboard.yaml --dry-run # Validate first
dtctl apply -f dashboard.yaml # Deploy
Notebooks are interactive investigation documents combining markdown, DQL queries, and visualizations.
DN (via /o11y-engineer)DN
DB (via /o11y-engineer)Auto-discover data schema and generate contextually accurate notebooks.
DB
type: notebook
name: "Service Diagnosis"
content:
version: "7" # String, not integer
defaultTimeframe:
from: now()-2h
to: now()
sections: # Array of section objects
- id: "section-1"
type: markdown
markdown: |
# Service Diagnosis
Investigation notebook for troubleshooting.
- id: "section-2"
type: dql
state:
input:
value: |
fetch spans
| filter service.name == "my-service"
| filter status.code == "ERROR"
| sort start_time desc
| limit 50
timeframe:
from: now()-2h
to: now()
visualization: table
querySettings:
maxResultRecords: 1000
defaultScanLimitGbytes: 500
maxResultMegaBytes: 1
defaultSamplingRatio: 10
enableSampling: false
davis:
includeLogs: true
davisVisualization:
isAvailable: true
!!! note “Important Differences from Dashboards”
- content.version is a string ("7"), not an integer
- content.sections is an array, not an object
- Each section has a unique id
| Aspect | Dashboard | Notebook |
|---|---|---|
| Purpose | Ongoing monitoring | Investigation |
| Audience | Operations, on-call | Incident responders |
| Content | Fixed visualizations | Interactive queries |
| Lifecycle | Long-lived | Per incident or review |
Dynatrace workflows automate incident response, remediation, and reporting.
CW (via /o11y-engineer)CW
SW (via /o11y-engineer)Analyze your environment and recommend automation opportunities.
SW
title: "Auto-Remediation: High Error Rate"
description: "Detect high error rates and notify via Slack"
trigger:
event:
type: events
query: event.type == "PROBLEM"
tasks:
get_data:
action: dynatrace.automations:execute-dql-query
input:
query: |
fetch spans, from:-30m
| filter service.name == "my-service"
| filter status.code == "ERROR"
| summarize error_count=count(), by:{span.name}
process_data:
action: dynatrace.automations:run-javascript
input:
script: |
import { execution } from '@dynatrace-sdk/automation-utils';
export default async function ({ execution_id }) {
const ex = await execution(execution_id);
const result = await ex.result('get_data');
const errors = result.records.filter(r => r.error_count > 10);
return { high_error_spans: errors };
}
conditions:
states:
get_data: OK
predecessors:
- get_data
send_notification:
action: dynatrace.slack:slack-send-message
input:
channel: alerts-channel
message: |
High error rate detected:
conditions:
custom: ""
states:
process_data: OK
predecessors:
- process_data
| Type | Description | Example |
|---|---|---|
| Event-based | Problem or custom event | event.type == "PROBLEM" |
| Schedule-based | Cron schedule | "0 9 * * *" (daily 9 AM) |
| Manual | No trigger | Run on demand |
| Action | Description |
|---|---|
dynatrace.automations:execute-dql-query |
Execute DQL |
dynatrace.automations:run-javascript |
Run TypeScript |
dynatrace.automations:http-request |
HTTP request |
dynatrace.slack:slack-send-message |
Slack message |
dynatrace.jira:jira-create-issue |
Create Jira issue |
dynatrace.davis:davis-analyze |
Davis AI analysis |
| Capability | dtctl Command (via /o11y-engineer) |
MCP Command (via /o11y-engineer) |
|---|---|---|
| Dashboards | CD (Create Dashboard) |
PD (Project Discovery Dashboard) |
| Notebooks | DN (Diagnostic Notebook) |
DB (Dashboard Builder) |
| Workflows | CW (Create Workflow) |
SW (Suggest Workflows) |
| Discovery | Manual input | Discover services via MCP |
| Best for | CI/CD, GitOps | Interactive, prototyping |
dtctl get dashboard <id> -o yamldtctl apply -f dashboards/| Command | Description |
|---|---|
dtctl apply -f <file> |
Deploy a resource |
dtctl apply -f <file> --dry-run |
Validate without deploying |
dtctl get dashboards |
List all dashboards |
dtctl get dashboard <id> -o yaml |
Export dashboard |
dtctl get notebooks |
List all notebooks |
dtctl get workflows |
List all workflows |
dtctl execute workflow <id> |
Run workflow manually |
dtctl query "<dql>" |
Run DQL query |
dtctl config set-context <name> |
Configure environment |
dtctl config use-context <name> |
Switch environment |