Agent Performance Comparison.

We gave the top five biggest observability platforms the same budget, the same data, and then let agents investigate different incidents on each platform.
The goal: which platform makes agents smart enough to find the true answer?

Using the OpenTelemetry Astronomy Shop Demo, we tested 10 fault scenarios, which we rated from easy to hardest. To fairly test the platforms, we adjusted the volume of data sent to each platform based on the public list prices of each and the volume of data generated by the demo application.

Method

The OpenTelemetry Astronomy Shop Demo generates telemetry. Each vendor receives the share of data we estimated the same reference spend would buy, based on its public list prices and the demo’s measured telemetry volumes. Coralogix retains all the telemetry.

We investigated each of the 10 scenarios five times on each of the five platforms: 250 investigations in total. Each repetition starts with an isolated agent and uses the same fixed telemetry window for that scenario.

The agents use the same model, system prompt, and scenario questions. Each has read-only access to the telemetry, with no hints about the fault.

Most scenarios introduce one fault. Five faults tests concurrent problems, and Optimize introduces no fault. Search relevance starts with a customer complaint; AI cost starts with a finance question. Agents are asked about customer and business impact, including money where the scenario provides that question.

Pricing calculations

We used an illustrative budget of $10,000 per month for each platform, based on public list prices and 15-day retention. Each platform is priced in its own billing units, with platform and seat fees included.

What data did we price?

We measured the demo’s telemetry rather than assuming generic event sizes. The sizing records include:

  • Logs: 126,475 records totaling 0.2062 GB in a six-hour clean-traffic window: approximately 1,631 bytes per log.
  • Traces: 561,949 spans in a six-hour clean-traffic window, averaging 2,874 bytes per span: approximately 1.62 GB.
  • Metrics: 9,833 active time series in a separate sizing measurement, with 64 collector self-telemetry series excluded from the comparison. Measured metric volume extrapolated to approximately 417.9 GB over a 30-day month before platform-specific cuts.

A time series is one metric and a unique combination of labels; a datapoint is one reading of that series. The same metric can therefore produce many series and repeated datapoints.

These are pricing-reference measurements, not the final volume ingested by every platform during the incidents. Telemetry changes with traffic, faults, instrumentation, and sampling. The reference mix used for the affordable-share calculation was 67% traces, 22% metrics, and 11% logs by volume.

List prices used in the model

We used the vendors’ published list prices, with ingestion, indexing, retention, and fixed fees included where applicable. The links below point to the official pricing pages; current offers may differ from the inputs used for this comparison. Prices captured September 2026.

Platform and sourcePricing inputs
CoralogixLogs $0.42/GB; traces $0.16/GB; metrics $0.06/GB. The model adds $0.0115/GB for customer-side Amazon S3 storage.
DatadogOn-demand ingestion at $0.10/GB plus $2.55 per million indexed log events or spans with 15-day retention, including applicable APM allowances. Custom metrics: $5 per 100 indexed metrics/month and $0.10 per 100 ingested metrics/month.
DynatraceLogs and traces: $0.20/GiB ingestion plus $0.0007/GiB-day for extra retention. Metrics: $0.15 per 100,000 datapoints. The model includes a query-cost reserve.
Grafana CloudLogs and traces: $0.40/GB write, $0.05/GB processing, and $0.10/GB retention at 30 days, prorated to 15 days in the model, for $0.50/GB total. Metrics: $6.50 per 1,000 active series/month.
New Relic$0.40/GB plus a $0.05/GB EU-region surcharge: $0.45/GB for logs, traces, and metrics, plus the seats listed below.

We converted GiB to decimal GB and used the measured log and span sizes to translate per-million-event charges into comparable volume costs. Metrics were priced in each platform’s native units: bytes, active series/custom metrics, or datapoints.

How that fits into the same budget

  1. Deduct platform and seat fees from the $10,000 monthly budget.
  2. Price the measured telemetry mix in each platform’s billing units, using 15-day retention.
  3. Allocate the available spend across logs, traces, and metrics, then adjust sampling, metric coverage, and collection frequency. The next section shows what we retained.
PlatformTracesLogsMetricsBlended affordable share
Coralogix100%100%100%100%
Datadog18.0%28.0%10.3%18.5%
Dynatrace82.3%100%5.7%41.1%
Grafana Cloud34.3%86.3%79.1%43.4%
New Relic38.1%95.9%15.9%38.9%

These percentages are calculated affordable shares, not the final sampling settings. Each platform’s budget is then allocated across traces, logs, and metrics. The full-data monthly cost estimates above divide the $10,000 reference budget by each platform’s blended affordable share and round to the nearest $1,000. They are illustrative estimates, rather than actual bills or quotes.

Additional pricing assumptions

  • New Relic: three Full Platform seats at $208/month in total: $10 for the first seat and $99 for each of the other two.
  • Grafana Cloud: a $19/month platform fee.
  • Datadog: an APM host fee of $48/month at the on-demand list price.
  • Dynatrace: a modeled 10% query reserve, assuming a typical small team’s dashboard and investigation usage.
  • Flat fees are deducted first from the $10,000/month budget before allocating the remaining spend to telemetry.

Data retained at each budget

To fit the estimated budgets, we combined the following actions. In our opinion, these are reasonable steps an engineering team would take to control telemetry costs:

  • Sample traces: prioritize error and slow traces, then retain a sample of routine traffic.
  • Reduce metric coverage: drop runtime metrics or less-used infrastructure detail.
  • Reduce metric frequency: report measurements less often where datapoints or volume drive cost.
  • Exclude lower-priority logs: use the demo’s service.criticality tag to inform service-level cuts, with additional successful proxy-log exclusions for Datadog.
TelemetryCoralogixDatadogDynatraceGrafana CloudNew Relic
Traces100%Errors + 45% of slow + 11% baselineErrors + slow + 38% baselineErrors + slow + 32% baselineErrors + slow + 33% baseline
Spans vs. Coralogix, normal traffic100%~8–10%~36–41%~29–31%~30–32%
LogsAll~51%All~93%~93%
JVM, .NET, and Go runtime metricsKeptDroppedDroppedDroppedDropped
Container and host metricsKeptDroppedCore only; every 3 minutesKeptEvery 2.5 minutes
Request, error, and latency metricsAll trafficAll traffic via APM statsAll trafficAll trafficAll traffic

Slow means a whole trace over 1.5 seconds. For Dynatrace, Grafana Cloud, and New Relic, we kept all error and slow traces. At Datadog’s list price, we chose to fit the budget by keeping all error traces and 45% of slow traces. Each baseline percentage applies to routine traces.

The demo’s service.criticality resource attribute classifies services as critical, high, medium, or low. This informed the log-cutting choices: for the Grafana Cloud, New Relic, and Datadog budgets, we chose to exclude logs from recommendation, ad, quote, and fraud detection. For Datadog, we also excluded successful proxy logs while keeping error responses.

For Dynatrace’s budget, we trimmed per-disk, per-interface, per-core, and container network and disk details, while keeping core container CPU, memory, limits, throttling, and uptime.

Results

Hover over an icon to see its exact score. You can also tap an icon or focus it with the keyboard. Use the switch to compare Cause + impact with Cause only.

How this was scored
ScenarioDifficultyCoralogixDatadogDynatraceGrafana CloudNew Relic
Delisted bestsellerEasy

A bestseller disappears from the catalog. Product pages and checkouts fail. This tests whether agents can connect an obvious error to affected shoppers and financial exposure.

We observed every agent finding the cause, as the failed checkouts and catalog logs were retained at every budget. Coralogix agents consistently measured the failure rate correctly. We observed sampled successful orders inflating the rates reported by several other vendors’ agents. In our scoring, Dynatrace and New Relic agents scored higher on the direct money question.

EasyCoralogix18/20Datadog14.5/20Dynatrace16.5/20Grafana Cloud13/20New Relic16.5/20
Recommendation cacheMedium

An unbounded recommendation cache exhausts memory and repeatedly restarts the service. Agents must distinguish the crash-loop symptom from its memory mechanism and measure failures without claiming lost orders.

In our scoring, Coralogix, Dynatrace, and New Relic agents tied on cause. Every Coralogix agent got the failure rates right, and we observed Grafana and Datadog agents also using unsampled request metrics effectively. New Relic agents explained the growing product list more often than Coralogix agents. The overall margin was narrow.

MediumCoralogix14.5/15Datadog12.5/15Dynatrace11/15Grafana Cloud13.5/15New Relic11.5/15
Kafka queue floodMedium

Checkout publishes duplicate order messages and fraud screening falls behind, while accounting keeps up. This tests a cause in one service, a symptom in another, and the value of orders left unscreened.

We observed every agent finding the cause from duplicate logs and slow consumer traces. Coralogix agents joined complete order and screening records to count the unscreened orders and measure their value. Other vendors’ agents estimated counts from queue metrics, and we observed some matching the money when asked directly.

MediumCoralogix18.5/20Datadog13/20Dynatrace13/20Grafana Cloud13.5/20New Relic13.5/20
CPU throttleMedium–hard

A tight product-catalog CPU limit slows lookups without producing errors. Agents must identify throttling, measure abandoned requests, and distinguish normal order-rate variation from attributable revenue loss.

In our scoring, Grafana agents scored highest, with accurate counts from unsampled request metrics and correct money answers. Coralogix, Datadog, and New Relic agents tied on cause, and Coralogix and New Relic agents tied for second overall. Two Coralogix agents claimed lost revenue from a change within normal variation.

Medium–hardCoralogix13/15Datadog12.5/15Dynatrace10/15Grafana Cloud14.5/15New Relic13/15
Search relevanceMedium–hard

A stricter relevance setting breaks multi-word searches while responses stay fast and successful. A customer complaint directs agents to investigate empty results, conversion, and orders that never happened.

We observed every agent finding the search change. Complete search-to-order traces let Coralogix agents measure the conversion drop and estimate lost orders. We found that sampled traces made counts and conversion harder for the other agents to establish. With an open investigation question, no agent found this fault.

Medium–hardCoralogix16/20Datadog10/20Dynatrace10/20Grafana Cloud11/20New Relic8/20
Ad garbage collectionHard

Repeated full garbage collection makes the ad service slow. Agents must prove the runtime mechanism, count failed or abandoned ad requests, and avoid attributing unrelated order variation to the fault.

We observed that slow traces were retained at every budget, but only the Coralogix data included JVM garbage-collection measurements. Every Coralogix agent proved the cause. We observed the other agents making partial inferences from container memory. In our scoring, Grafana agents scored higher on the money question because one Coralogix agent assumed USD.

HardCoralogix14.5/15Datadog5.5/15Dynatrace6.5/15Grafana Cloud10.5/15New Relic9.5/15
Early warningHard

Unclosed worker pools accumulate idle threads in fraud detection before any errors or slowdown appear. Agents must find the risk, explain it, forecast failure, and describe the exposure if screening stops.

Thread metrics revealed the mechanism. We found that container memory growth was enough to forecast the failure, but not to explain it.

HardCoralogix14.5/15Datadog0/15Dynatrace11/15Grafana Cloud10/15New Relic8.5/15
AI costHard

A prompt change drops the JSON-only instruction, causing silent AI retries and higher token spend while customers still receive answers. A finance question asks agents to explain the increase and quantify waste.

We observed every agent seeing token usage rise. Complete traces let Coralogix agents compare prompts and add up the price of every call. Dynatrace and New Relic agents found the cause but undercounted from sampled totals, and we observed Grafana agents using request metrics for accurate extra-call counts. Datadog agents also estimated dollar amounts from tokens. With an open question, no agent found the fault.

HardCoralogix10/10Datadog7.5/10Dynatrace7.5/10Grafana Cloud7.5/10New Relic7.5/10
Five faultsHardest

Payment failures, ad garbage collection, a recommendation cache crash loop, periodic cart latency, and an email memory leak occur together. Agents must explore without knowing the fault count, explain causes, and prioritize customer and business impact.

In our scoring, Coralogix agents got the most causes right, proved garbage collection, and each gave the exact failed-payment value per currency. We observed the Dynatrace agents finding every fault, but we found that, due to the data trimmed to fit the budget, they lacked GC proof and overstated payment failure rates from sampled traces. Two Coralogix agents missed the email leak.

HardestCoralogix28/30Datadog11.5/30Dynatrace24/30Grafana Cloud19.5/30New Relic18.5/30
OptimizeOpen-ended

The shop runs normally. Agents look for improvements and prioritize them by cost, reliability, and customer impact. We considered excessive collector CPU the most important finding; this is an opportunity-finding task, not a root-cause fault test.

Coralogix agents found an average of 8.4 real items, and every agent found the top item (collector CPU). The default icons summarize the opportunity-finding results. Cause only shows no fault to score.

Open-endedCoralogix8.4 itemsDatadog6.6 itemsDynatrace6.8 itemsGrafana Cloud7.0 itemsNew Relic8.0 items
BestStrongMixedWeak

Finding the cause and measuring impact

Average scores across nine scenarios, weighted equally. Optimize is excluded because it tests finding improvements rather than an injected fault, so it has no comparable cause or impact score. The guides sit at 80% because this is the threshold for a strong result in our scoring criteria.

False claims

False claims are statements the reviewing agent found contradicted the reference evidence, such as inflated failure rates or claims that available telemetry did not exist. Missing a finding was not automatically a false claim.

Recorded counts across the final 10 scenarios, with five investigations per platform in each. Fewer is better. These counts are separate from the scoreboard scores.

  1. Coralogix16
  2. Grafana Cloud17
  3. Datadog25
  4. New Relic36
  5. Dynatrace49

What the results show

We found that keeping complete telemetry helped agents explain why a problem happened and measure its impact more accurately.

In a scenario where searches returned no results, Coralogix agents could trace shoppers from search to checkout, measure how much less likely they were to place an order, and estimate the orders lost as a result. In a scenario assessing silent retries in an AI chatbot, by our scoring, Coralogix agents gave the most complete account of wasted spend. Having all the telemetry let them compare prompts, trace the repeated calls, and add up the cost of every AI call.

In our experience, more evidence was often the difference between seeing a symptom and proving its cause, or between estimating impact from a sample and measuring it from complete records.

We believe that with complete telemetry, agents can go beyond finding faults and build a clearer picture of key business metrics, including conversion, lost orders, and the cost of serving customers.

Scenarios and difficulty

Delisted bestseller

Easy · Missing product

A bestseller disappears from the catalog. Product pages and checkouts fail. This tests whether agents can connect an obvious error to affected shoppers and financial exposure.

Recommendation cache

Medium · Crash loop

An unbounded recommendation cache exhausts memory and repeatedly restarts the service. Agents must distinguish the crash-loop symptom from its memory mechanism and measure failures without claiming lost orders.

Kafka queue flood

Medium · Distributed backlog

Checkout publishes duplicate order messages and fraud screening falls behind, while accounting keeps up. This tests a cause in one service, a symptom in another, and the value of orders left unscreened.

CPU throttle

Medium–hard · Resource limit

A tight product-catalog CPU limit slows lookups without producing errors. Agents must identify throttling, measure abandoned requests, and distinguish normal order-rate variation from attributable revenue loss.

Search relevance

Medium–hard · Silent wrong results

A stricter relevance setting breaks multi-word searches while responses stay fast and successful. A customer complaint directs agents to investigate empty results, conversion, and orders that never happened.

Ad garbage collection

Hard · Runtime slowdown

Repeated full garbage collection makes the ad service slow. Agents must prove the runtime mechanism, count failed or abandoned ad requests, and avoid attributing unrelated order variation to the fault.

Early warning

Hard · Predicted failure

Unclosed worker pools accumulate idle threads in fraud detection before any errors or slowdown appear. Agents must find the risk, explain it, forecast failure, and describe the exposure if screening stops.

AI cost

Hard · Silent AI retries

A prompt change drops the JSON-only instruction, causing silent AI retries and higher token spend while customers still receive answers. A finance question asks agents to explain the increase and quantify waste.

Five faults

Hardest · Concurrent faults

Payment failures, ad garbage collection, a recommendation cache crash loop, periodic cart latency, and an email memory leak occur together. Agents must explore without knowing the fault count, explain causes, and prioritize customer and business impact.

Optimize

Open-ended · No injected fault

The shop runs normally. Agents look for improvements and prioritize them by cost, reliability, and customer impact. We considered excessive collector CPU the most important finding; this is an opportunity-finding task, not a root-cause fault test.

Scoring criteria

The three main questions

  1. Can you tell me what is going on in my system right now?
  2. What is the customer and business impact of this?
  3. How much money has this cost us so far, and how much is at risk?

These were the standard investigation questions. Specialized scenarios adapted them for forecasting, optimization, customer complaints, or AI spending.

How an answer earns points

Each response was assessed against the known fault and the scenario’s reference measurements. We scored the cause and the relevant impact components separately:

  • 1 point — correct: identifies the right mechanism or gives a supported impact measurement with the correct scope, time window, and denominator.
  • ½ point — partial: gets part of the answer right, but leaves the mechanism unproven or gives an incomplete or approximate impact estimate.
  • 0 points — wrong or missed: misses the finding, gives the wrong explanation, or makes an unsupported impact claim.

What does 15/20 mean?

15/20 means 15 of the 20 available points across five investigations: 75% of the maximum score.

Why are some scenarios out of 15 and others out of 20?

  • 15 points: cause, operational or predicted impact, and the explicit money answer — three components × five investigations. These scenarios had disruption or future risk, without measured financial loss attributable to the incident.
  • 20 points: the same three components, plus financial impact volunteered before the explicit money question — four components × five investigations. These scenarios had measurable financial consequences, such as failed checkouts, lost conversion, or orders left unscreened. The extra component rewards connecting the incident to money without a specific prompt.

The totals follow each scenario’s rubric. Compare platforms within the same scenario.

For example, in Delisted bestseller, Coralogix earned 5 for cause + 5 for affected counts + 4 for volunteered financial impact + 4 for financial impact after the money question = 18/20. This describes the quality and consistency of the answers to that incident; it is not a count of incidents solved.

  • Other totals: AI cost scores cause and financial impact for 10 points. Five faults has 25 possible cause points plus 5 money points, for 30 total.
  • Cause only: shows diagnosis points alone. Cause + impact: adds the scored counts and money components.
  • Optimize: reports the average number of valid improvement findings, rather than points out of a fault-test maximum. Cause only shows a dash.
  • False claims: recorded separately, without subtracting them from these totals. The results section lists the total counts across all 10 scenarios.

Reading the icons

  • Star: highest score in that scenario and view, including ties. A leading score can still be below 80%.
  • Check: at least 80% of available points. Dot: 55% to below 80%. Cross: below 55%.
  • Exact result: hover, focus, or tap an icon to see its points and maximum.

Caveats

  • These results are not intended to establish a universal vendor ranking. They merely cover one demo application and the tested data configurations, CLIs, and agent skills. They do not isolate data volume from other platform differences and capabilities.
  • The five investigations share one telemetry window per scenario and platform, so they measure how consistently agents investigate the same incident. We removed fault-flag telemetry and messages announcing injected faults to reduce demo-specific clues. Agents were instructed to rely on retrieved evidence, avoid identifying the demo, and assess the incident and its customers as they would in production. Prior familiarity with the demo may still influence answers.
  • Public list prices differ from negotiated contracts. The affordable shares depend on this application’s measured telemetry mix and the pricing assumptions above.
  • Keeping all errors and slow traces would push modeled spend over budget during incidents (up to ~2.6×). We allowed this, which gives the sampled cases more data than a strict cap would, and we expect it favored them.
  • The AI assistant uses a simulated model with realistic token counts and prices. Its costs are fractions of a cent; the finding concerns the retry mechanism and relative waste.

Telemetry generated from the OpenTelemetry Astronomy Shop Demo. Monthly costs are modeled from public list prices and are not quotes or actual bills. Results describe the tested data configurations.