Migrate logs, host metrics and APM from Datadog to Grafana Cloud
Send telemetry to both platforms with Grafana Alloy, rebuild dashboards and alerts, and verify coverage before retiring Datadog.
From Datadog to Grafana Cloud · Published 2026-10-08
Before you commit
- Hands-on effort
- 1–2 weeks for a few dozen hosts and dashboards; more for heavy APM or hundreds of monitors
- Elapsed time
- Run both platforms for at least two weeks, covering a deploy, a peak and an alert rehearsal.
- Examples cover
- Grafana Alloy 1.13+ · Datadog Agent 7.32+ for the OTLP example
- Content review
- 2026-10-09
- Integration testing
- Not recorded. Rehearse the commands and rollback on staging.
Prerequisites
- Read access to Datadog usage, dashboards and monitors, plus Grafana Cloud write credentials.
- An inventory of critical alerts, queries, custom metrics and on-call routing, with owners for each.
Don’t migrate yet if…
- Required integrations or alerts have not been rebuilt and verified.
- You cannot pay to send data to both platforms during the move or estimate the new metric and query costs.
Datadog bills per host and charges separately to ingest and index logs: $0.10/GB to ingest, then $2.50 per million events to index with 30-day retention (annual pricing). Infrastructure Pro is $15/host/month and APM starts at $31/host.
Grafana Cloud Pro is $19/month plus usage. Logs cost $0.50/GB with 30-day retention and metrics cost $6.50 per 1,000 active series. The first 50 GB of logs and 10,000 series each month are included. Start by checking which Datadog products account for your bill. Compare the logs-and-hosts estimate in the calculator, then add costs for the rest of your setup.
Installing the agent is the easy part. Most of the work is rebuilding dashboards and alerts, rewriting queries in PromQL and LogQL, and keeping the number of metric series your tags create under control.
0. Before you start
- Check your Datadog order form. Most Datadog contracts are annual commitments, often with a notice period before renewal. Find the renewal date and notice terms first. They set your deadline.
- Pull usage. In Datadog, open Plan & Usage. Write down infrastructure hosts, APM hosts, custom metrics, ingested log GB and indexed log events by retention. Only log GB and host count are calculator inputs; the others are needed for a full estimate. The calculator assumes 1 KB per event and indexes all ingested logs for 30 days, so it can overstate a filtered Datadog bill. Grafana’s modeled host cost uses Kubernetes Monitoring, not plain VM metrics or complete APM.
- Export monitors and dashboards. You need an API key and an application key with read access. Use your site’s API host (
api.datadoghq.eu,api.us3.datadoghq.com…) if you aren’t on US1.
export DD_API_KEY=... DD_APP_KEY=... DD_HOST=https://api.datadoghq.com
# All monitors in one file (no page parameter returns everything)
curl -s "$DD_HOST/api/v1/monitor" \
-H "DD-API-KEY: $DD_API_KEY" -H "DD-APPLICATION-KEY: $DD_APP_KEY" > monitors.json
# Dashboard list, then each dashboard's full JSON
mkdir -p dashboards
curl -s "$DD_HOST/api/v1/dashboard" \
-H "DD-API-KEY: $DD_API_KEY" -H "DD-APPLICATION-KEY: $DD_APP_KEY" \
| jq -r '.dashboards[].id' | while read -r id; do
curl -s "$DD_HOST/api/v1/dashboard/$id" \
-H "DD-API-KEY: $DD_API_KEY" -H "DD-APPLICATION-KEY: $DD_APP_KEY" > "dashboards/$id.json"
done
Keep these files. They are your spec for steps 5 and 6, and your backup once the Datadog account is gone.
- Rank them. Most organizations have dashboards nobody opens. Ask each team for its five most-used dashboards and the monitors that page someone, migrate those, and archive the rest as JSON.
1. Map the concepts
| Datadog | Grafana Cloud |
|---|---|
| Datadog Agent | Grafana Alloy (OpenTelemetry Collector distribution) |
| Infrastructure metrics, custom metrics | Grafana Cloud Metrics (Mimir), queried with PromQL |
| Log Management | Grafana Cloud Logs (Loki), queried with LogQL |
| APM / dd-trace | Grafana Cloud Traces (Tempo) via OpenTelemetry; Application Observability for RED metrics |
Tags (env:prod, service:checkout) |
Labels (env="prod", service="checkout") |
| Monitors | Grafana-managed alert rules |
Notification channels (@pagerduty-…, @slack-…) |
Contact points and notification policies |
| Dashboards | Grafana dashboards |
| Log indexes and exclusion filters | Filtering at the collector (loki.process) before data is sent |
Retention differs. Grafana Cloud Pro keeps metrics for 13 months and logs and traces for 30 days; the Free tier keeps everything for 14 days. Datadog keeps metrics for 15 months and sets log retention per index. If you have a 15-day or 90-day log index, decide what replaces it.
2. Install Alloy next to the Datadog Agent
Run both agents on every host for the whole migration. They don’t conflict, except on ports (see step 4). On Debian or Ubuntu (other systems):
sudo mkdir -p /etc/apt/keyrings
sudo wget -O /etc/apt/keyrings/grafana.asc https://apt.grafana.com/gpg-full.key
sudo chmod 644 /etc/apt/keyrings/grafana.asc
echo "deb [signed-by=/etc/apt/keyrings/grafana.asc] https://apt.grafana.com stable main" \
| sudo tee /etc/apt/sources.list.d/grafana.list
sudo apt-get update && sudo apt-get install alloy
In the Grafana Cloud portal, open your stack’s details page and copy the Prometheus remote-write URL and user ID, the Loki push URL and user ID, and create an access policy token with metrics:write and logs:write. Put them in /etc/default/alloy as environment variables, not in the config file.
A minimal /etc/alloy/config.alloy for host metrics and log files:
// Host metrics (node_exporter, built in)
prometheus.exporter.unix "host" {
disable_collectors = ["ipvs", "infiniband", "nfs", "nfsd"]
}
prometheus.scrape "host" {
targets = prometheus.exporter.unix.host.targets
forward_to = [prometheus.relabel.drop_unused.receiver]
}
prometheus.relabel "drop_unused" {
forward_to = [prometheus.remote_write.grafana.receiver]
rule {
source_labels = ["__name__"]
regex = "node_scrape_collector_.+"
action = "drop"
}
}
prometheus.remote_write "grafana" {
external_labels = { env = "prod" }
endpoint {
url = sys.env("GC_PROM_URL")
basic_auth {
username = sys.env("GC_PROM_USER")
password = sys.env("GC_TOKEN")
}
}
}
// Log files
local.file_match "app" {
path_targets = [{ __path__ = "/var/log/checkout/*.log", service = "checkout" }]
}
loki.source.file "app" {
targets = local.file_match.app.targets
forward_to = [loki.process.filter.receiver]
}
loki.process "filter" {
forward_to = [loki.write.grafana.receiver]
stage.json {
expressions = { level = "" }
}
stage.drop {
source = "level"
value = "debug"
drop_counter_reason = "debug"
}
}
loki.write "grafana" {
external_labels = { env = "prod" }
endpoint {
url = sys.env("GC_LOKI_URL")
basic_auth {
username = sys.env("GC_LOKI_USER")
password = sys.env("GC_TOKEN")
}
}
}
Apply it with sudo systemctl reload alloy. The alloy user must be able to read the log files. Start with two or three hosts, check the data in Grafana’s Explore view, then roll the package and config out with whatever deploys the Datadog Agent today (Ansible, Chef, a Helm chart).
3. Control cardinality and volume
Grafana Cloud bills metrics by active series, so every label value multiplies cost, as Datadog custom metrics do.
- Metric labels: don’t carry over every Datadog tag. Drop labels such as
user_id,request_id, pod UID or full URL with aprometheus.relabelrule usingaction = "labeldrop". Disable node_exporter collectors you don’t chart. - Loki labels should be few and low-cardinality:
service,env,host, maybelevel. Leave everything else (trace IDs, user IDs, paths) in the log line and filter it at query time. A high-cardinality label in Loki creates many small streams, which is slow to query. - Drop logs before they’re sent. The
stage.dropblocks above do the job of Datadog exclusion filters, except that Datadog still charged $0.10/GB to ingest excluded logs, while logs Alloy drops never reach Grafana Cloud and cost nothing. Debug logs, health-check access logs and noisy third-party libraries are the usual candidates.
4. Traces with OpenTelemetry (optional)
Two routes, depending on how much code you want to touch:
- Replace dd-trace with OpenTelemetry SDKs or auto-instrumentation. More code changes, but no Datadog libraries are left afterwards.
- Keep dd-trace and switch it to OTLP export. Datadog’s SDKs can export traces as OTLP to any receiver (in Preview; Java 1.62.0+, Python 4.8.0+, Node.js 5.98.0+, Go 2.8.0+, .NET 3.41.0+). Set
DD_TRACE_OTEL_ENABLED=true,OTEL_TRACES_EXPORTER=otlpandOTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318. Spans keep Datadog attribute names (http.status_code, nothttp.response.status_code), which matters when you build dashboards.
Either way, point the app at Alloy and have Alloy send to both backends during the dual-shipping period. Datadog Agent 7.32+ accepts OTLP if you enable it in datadog.yaml:
otlp_config:
receiver:
protocols:
grpc:
endpoint: localhost:4317
Then Alloy listens on HTTP 4318 only (the Agent has 4317) and forwards to both. The client_auth block needs Alloy 1.13 or later:
otelcol.receiver.otlp "apps" {
http { endpoint = "127.0.0.1:4318" }
output { traces = [otelcol.processor.batch.default.input] }
}
otelcol.processor.batch "default" {
output {
traces = [otelcol.exporter.otlphttp.grafana.input, otelcol.exporter.otlp.datadog_agent.input]
}
}
otelcol.auth.basic "grafana" {
client_auth {
username = sys.env("GC_OTLP_USER")
password = sys.env("GC_TOKEN")
}
}
otelcol.exporter.otlphttp "grafana" {
client {
endpoint = sys.env("GC_OTLP_URL") // https://otlp-gateway-prod-<region>.grafana.net/otlp
auth = otelcol.auth.basic.grafana.handler
}
}
otelcol.exporter.otlp "datadog_agent" {
client {
endpoint = "localhost:4317"
tls { insecure = true }
}
}
Alloy also has an otelcol.receiver.datadog component that accepts dd-trace’s native protocol on port 8126 without any app changes. It’s experimental (it needs --stability.level=experimental in CUSTOM_ARGS) and it conflicts with the Datadog Agent’s trace port, so it only fits after the Agent is gone.
5. Rebuild dashboards
There’s no general self-service converter. Grafana’s Datadog data source plugin (Cloud Pro and up) can query Datadog from Grafana, which helps you compare panels side by side, but it doesn’t translate dashboards. Rebuild your top dashboards by hand from the exported JSON, starting from Grafana’s prebuilt integration dashboards (the Linux node integration has host dashboards ready). Common translations:
| Datadog | Grafana Cloud |
|---|---|
avg:system.cpu.user{env:prod} by {host} |
100 * avg by (instance) (rate(node_cpu_seconds_total{mode="user", env="prod"}[5m])) |
avg:system.mem.used{*} by {host} |
node_memory_MemTotal_bytes - node_memory_MemAvailable_bytes |
sum:checkout.orders{env:prod}.as_count() (custom counter) |
sum(increase(checkout_orders_total{env="prod"}[$__interval])) |
Log search service:checkout status:error |
{service="checkout"} | json | level="error" |
Log search service:checkout "timeout" @http.status_code:504 |
{service="checkout"} |= "timeout" | json | http_status_code="504" |
| Log count by service, as a graph | sum by (service) (count_over_time({env="prod"} |= "error" [5m])) |
Datadog metric names use dots; Prometheus names use underscores, and counters usually end in _total. Write the mapping for your own custom metrics down once and share it.
6. Rebuild monitors as alert rules
Go through monitors.json and recreate the monitors that page people as Grafana-managed alert rules. A Datadog avg(last_5m):… > 90 becomes a PromQL query with a threshold of 90 and a pending period of 5 minutes. Log monitors become LogQL count_over_time rules.
Set up contact points first. PagerDuty (integration key) and Slack (webhook or bot token) are built in. Then add notification policies that route on labels such as team or severity, so you don’t have to set recipients rule by rule. Keep rules in a provisioning file or Terraform so they’re reviewed like code.
7. Verify by dual shipping
Run both platforms for at least two weeks, covering a deploy, a weekly peak and ideally a real incident:
- Put the same panel side by side (Datadog in one tab, Grafana in the other) for CPU, memory, request rate and error rate on a few hosts. Expect small differences from different collection intervals, not different shapes.
- Compare daily log counts per service. A big gap means a missed file path or an over-eager drop rule.
- Trigger each critical alert on purpose (stop a test service, fill a disk on a staging host) and confirm it reaches PagerDuty or Slack from Grafana.
- Check Grafana Cloud’s billing and usage dashboard after a week, and recheck the calculator with real numbers.
8. Rollback plan
Keep the Datadog Agent running and its monitors enabled until cutover. Mute the duplicate Grafana alerts, or route them to a test channel, until you trust them. Rolling back means unmuting Datadog monitors and pointing on-call there. No data is lost, because Datadog received everything the whole time.
9. Clean up
- Cut over on-call: turn on Grafana alert routing to PagerDuty, then mute and delete the Datadog monitors.
- Remove the Datadog Agent from your config management and uninstall it (
sudo apt-get remove datadog-agenton Debian/Ubuntu). Remove dd-trace from apps once they export through OpenTelemetry. - Revoke Datadog API and application keys, and delete them from your secrets store and CI.
- Send the cancellation in writing, as your order form requires. Datadog commitments usually run to the end of the term, so your last day of use and your last invoice may be months apart.
Found a command or version mismatch? Report a correction with the version and steps to reproduce it. Remove secrets and customer data first.