How to evaluate Datadog alternatives for machine data

Price searchable history and monitoring coverage, test an unplanned question over older data, and keep Datadog for the workflows your team still depends on.

Axiom · · 12 min read

Most evaluations of Datadog alternatives start after a budget conversation rather than a technical one. Logs and events have grown faster than the budget. Someone has already turned on exclusions or shortened retention, and the request is to price a replacement.

Comparing rates side by side does not show what you'll be able to look up three months from now. Retention, indexed percentage, and monitoring coverage do decide that.

This article covers what to measure, which questions to ask of logs, traces, metrics, and agent access, and how to run the comparison so both sides see the same workload.

Start with the evidence the current bill already dropped

Datadog is a formidable product.

The pressure that starts most evaluations is cost, not missing features. Logs and events have outgrown the indexing economics around them. By the time an alternative gets a serious look, most teams have already excluded a few sources, sampled a noisy service, or shortened retention to hold a number.

Start by listing which evidence survived the cost controls already in place, then use that list when you compare products.

Price searchable history, not only ingest

Ingestion and searchable history are billed separately

Datadog meters log ingestion by the gigabyte and Standard Indexing by the event. Our Datadog comparison page puts ingestion at roughly $0.10 per GB, then Standard Indexing at roughly $1.70 per million events for 15-day retention at annual list. The two meters stack, so the price of searchable history depends on how many events a gigabyte contains and what share of them get indexed.

Data loaded but left unindexed appears in a live tail rather than in the historical log explorer, which is the surface most investigations use. Retention is part of the outcome too. Fifteen days covers a lot of incident response, and it covers very little of a regression introduced two releases ago.

Pull the GB loaded, the indexed event count, the indexed percentage, and the retention path for each index, then price the workflow those numbers produce.

Each cost control also reduces what you can look up later

The standard levers are all rational. Teams exclude or sample events out of the index, convert logs to metrics to keep a trend, apply daily quotas, route older data to a cheaper tier or an archive, or filter at the collector.

Each one lowers spend by changing which evidence stays easy to search later. An aggregate can show that something changed, but it cannot restore the individual request, payload, or context that explains why. A daily quota protects the budget by stopping indexing once the cap is reached.

Test that pattern in the evaluation. When Datadog cost control starts deciding what you'll be able to know later, send that data workload to us.

A lower-cost log tier can be a different product

Long-retention log tiers are a genuine improvement over an archive-only strategy, and they keep logs searchable in place. They also carry a different contract. Query capacity is sized separately, and logs living only in that tier don't power monitors or Watchdog. Moving an index there to save money can also remove those logs from alerting.

We keep one lifecycle for loaded data. Everything loaded is queryable in APL right away, with no separate indexing SKU. Events settle into a schema-on-read event store that compresses heavily, and older data stays in the ordinary query path, feeding the same dashboards and monitors as fresh data.

We price all of it on one usage-based dial: data loaded, query compute, and compressed storage. Retention is then a question of how far back you want to see, not which log product each event ends up in. That's how we approach log management at petabyte scale.

Test each signal against a question nobody planned for

Traces: aggregate service health and individual evidence are different things

Datadog is strong at turning traces into operational service views, and it keeps aggregate APM metrics broadly available. The individual "span," the timed record of one step in a request's journey, follows a separate retention decision. Filters select which requests stay searchable once the live window closes.

We store each span as an event under the dataset's retention, with no head-based or tail-based sampling at ingest. Because those are traces that aren't sampled, you can search on an attribute nobody predicted would matter, then open the resulting trace.

This difference matters more if your support or engineering staff keep returning to specific requests weeks later than if you mainly use Datadog's mature APM workflow.

Metrics: confirm which pricing model is in the contract

Confirm which Datadog metric contract is active before comparing metrics. Datadog's metric-name model counts a metric name once, regardless of how many tag combinations sit under it. The familiar claim that Datadog charges for every tag combination no longer holds as a general rule. Indexed points, ingested points, and distribution multipliers still shape the total.

We price metric data through the same relationship as everything else: bytes loaded, compressed storage, and query work. MetricsDB treats "high cardinality" as something the store is built to handle rather than something to penalize. It places each exact series by its dimensions, so a demanding metric can spread across more ingest capacity. Cardinality still consumes resources, and it does not bill each series as its own commercial unit.

Agents inherit the evidence underneath them

We ship a Model Context Protocol server, or "MCP," and so does Datadog. Datadog's MCP server is sprawling, and so the useful comparison is which retained events, spans, and dimensions the agent can query.

An agent can't reason over an event that was excluded at the pipeline, a span that wasn't retained, or a dimension someone removed to control cardinality. Agents make broad history more valuable, because they will pursue more questions over longer periods than a human operator usually will. Axiom's MCP server reaches retained data through the same query primitives engineers use, on the same dial.

Give both systems the same unplanned production question over the same period, then compare what came back and what it cost.

Run both systems at once rather than planning a cutover

Most teams start by running both systems at once. Point the OpenTelemetry Collector, Vector, or Cribl at us, start with the highest-volume log source, and keep Datadog for highly specialized workflows while the economics prove out. Nothing has to be switched off. You can run logs, traces, metrics, and events on one event store beside Datadog for as long as it takes, then expand at the renewal seam.

Keep the two sides comparable while it runs: same events and fields retained, same time range, same aggregations, same monitoring requirement, same query concurrency. If one side is given an easier query mix, the numbers will not describe production.

Then pick one question nobody planned for, over data that's 90 days old, and follow it to an answer in both systems.

When Datadog still fits

Datadog still fits when you want one vendor to cover everything from service management, to software delivery, digital experience, and security, when APM, RUM, synthetics, and security breadth matter more than log economics, and when you use the integration catalog and turnkey dashboards heavily.

We're a fit when logs and events are what's breaking the bill, when you're indexing selectively or shortening retention to cope, and when every byte should stay queryable by engineers and agents on one predictable dial. For many teams the answer stays a mix of the two for a while, with the machine-data system of record on one side and the operational suite on the other.

Most Datadog-alternative evaluations start as a rate comparison after cost controls are already in place. A more useful test is whether those controls have already decided what you'll be able to know later, and whether a second system can keep that evidence in the ordinary query path.

Run that test by sending one high-volume log source to both systems, keep Datadog for the workflows you still depend on, and follow one unplanned question over older data in both systems.

FAQs

Related Reading

See the comparison on a live source

If cost control is already deciding what you'll be able to know later, a walkthrough on a live source will show ingest, retention, and query on the same data. Bring the highest-volume log stream and see Axiom on your data.