Datadog
by Datadog • • Infrastructure Monitoring Tools
Datadog is infrastructure monitoring software for DevOps and SRE teams watching cloud, container and on-premise hosts, with a free plan for small fleets.
Datadog puts every host, container and cloud service a platform team runs on one screen and flags the one that starts misbehaving. Datadog, Inc., a New York company traded on the NASDAQ as DDOG, builds and sells it with no parent above it. Datadog positions the wider platform as AI-powered observability and security; infrastructure monitoring is one line of that platform, and APM, log management and cloud security are sold beside it from the same pricing page. The Agent gathers host and container data from Linux, Windows, macOS and bare metal servers, and Autodiscovery detects containers as they start. The Host Map draws hosts as tiles coloured by a chosen metric and grouped by tags, and container views report CPU, memory and network use. The Kubernetes Explorer lists pods, nodes and deployments, with a link to a detailed dashboard. Custom dashboards use template variables to switch the hosts or environment in view, and anomaly monitors flag metrics that behave differently than they have in the past. Monitors send notifications to Slack, Microsoft Teams, PagerDuty, Jira and ServiceNow. Custom metrics arrive through DogStatsD or the HTTP API, the Agent scrapes Prometheus and OpenMetrics endpoints, and the AWS, Azure and Google Cloud integrations pull service metrics from each provider's monitoring API. The 1,000+ integrations come with out-of-the-box dashboards and monitors. Teams use Datadog to line up incidents against deployments with Change Tracking, which also records cloud resource changes in Preview. Serverless monitoring reports enhanced AWS Lambda metrics such as cold starts, timeouts and out-of-memory errors. Alerts reach on-call staff, chat channels and ticketing tools through the same monitors. Longer metric retention on the paid plans supports year-on-year comparisons of host behaviour. Datadog fits DevOps, SRE and platform teams that run a mix of cloud instances, containers and physical servers and want them in one view. It suits engineering groups that already send metrics from application code or Prometheus exporters, and AWS teams that want Lambda functions watched beside their hosts.
Features
- Included: Cloud provider metric integrations
- Included: Kubernetes cluster monitoring
- Included: Container-level metrics
- Included: Host map visualization
- Included: Tag-based grouping and filtering
- Included: Prebuilt dashboards for common services
- Included: Anomaly alerts on infrastructure metrics
- Included: Alert routing to chat and paging tools
- Included: Custom metrics via StatsD or an API
- Included: Prometheus endpoint scraping
- Included: Serverless function metrics
- Included: Configuration change tracking
- Included: Custom dashboard builder
- Included: Metric retention of at least 13 months
- Included: Auto-discovery of hosts and containers
- Included: On-premise and hybrid host support
- Included: Free tier for a handful of hosts
Additional Features
- Host Map tiles colored by metric and grouped by tag
- Tag filtering on graphs and dashboards
- Anomaly monitors on metrics
- Monitor notifications to Slack, Microsoft Teams, PagerDuty, Jira and ServiceNow
- DogStatsD and HTTP API custom metrics
- Prometheus and OpenMetrics endpoint scraping
- Metric integrations for AWS, Azure and Google Cloud
- Kubernetes pod, node and deployment views
Best for
- DevOps and SRE teams watching AWS instances and on-premise servers together
- Platform teams running Docker containers alongside virtual machines
- Engineering teams sending application metrics through DogStatsD or an HTTP API
- Teams with Prometheus exporters who want those metrics stored beside host data
- AWS teams tracking Lambda cold starts, timeouts and memory use
- Small teams trialling monitoring on up to 5 hosts at no charge
- On-call teams routing alerts to Slack, Microsoft Teams or PagerDuty
- Buyers who accept per-host pricing plus separate per-unit charges
Use cases
- Mapping host CPU load by availability zone
- Grouping hosts and dashboards by environment, team or region tag
- Alerting when a host metric departs from its usual pattern
- Sending monitor alerts to Slack, PagerDuty, Jira or ServiceNow
- Submitting application metrics through DogStatsD
- Scraping OpenMetrics endpoints into the same metric store
- Watching Lambda invocations, errors and cold starts
- Pulling AWS, Azure and Google Cloud service metrics into one place
- Comparing host metrics across a year of retained data
- Lining up incidents against deployments and configuration changes
- Trying out monitoring on a handful of hosts before buying
Screenshots & Videos
Explore Datadog in action