Delisted bestseller
Easy · Missing product
A bestseller disappears from the catalog. Product pages and checkouts fail. This tests whether agents can connect an obvious error to affected shoppers and financial exposure.
We gave the top five biggest observability platforms the same budget, the same data, and then let agents investigate different incidents on each platform.
The goal: which platform makes agents smart enough to find the true answer?
Using the OpenTelemetry Astronomy Shop Demo, we tested 10 fault scenarios, which we rated from easy to hardest. To fairly test the platforms, we adjusted the volume of data sent to each platform based on the public list prices of each and the volume of data generated by the demo application.
The OpenTelemetry Astronomy Shop Demo generates telemetry. Each vendor receives the share of data we estimated the same reference spend would buy, based on its public list prices and the demo’s measured telemetry volumes. Coralogix retains all the telemetry.
We investigated each of the 10 scenarios five times on each of the five platforms: 250 investigations in total. Each repetition starts with an isolated agent and uses the same fixed telemetry window for that scenario.
The agents use the same model, system prompt, and scenario questions. Each has read-only access to the telemetry, with no hints about the fault.
Most scenarios introduce one fault. Five faults tests concurrent problems, and Optimize introduces no fault. Search relevance starts with a customer complaint; AI cost starts with a finance question. Agents are asked about customer and business impact, including money where the scenario provides that question.
We used an illustrative budget of $10,000 per month for each platform, based on public list prices and 15-day retention. Each platform is priced in its own billing units, with platform and seat fees included.
We measured the demo’s telemetry rather than assuming generic event sizes. The sizing records include:
A time series is one metric and a unique combination of labels; a datapoint is one reading of that series. The same metric can therefore produce many series and repeated datapoints.
These are pricing-reference measurements, not the final volume ingested by every platform during the incidents. Telemetry changes with traffic, faults, instrumentation, and sampling. The reference mix used for the affordable-share calculation was 67% traces, 22% metrics, and 11% logs by volume.
We used the vendors’ published list prices, with ingestion, indexing, retention, and fixed fees included where applicable. The links below point to the official pricing pages; current offers may differ from the inputs used for this comparison. Prices captured September 2026.
| Platform and source | Pricing inputs |
|---|---|
| Coralogix | Logs $0.42/GB; traces $0.16/GB; metrics $0.06/GB. The model adds $0.0115/GB for customer-side Amazon S3 storage. |
| Datadog | On-demand ingestion at $0.10/GB plus $2.55 per million indexed log events or spans with 15-day retention, including applicable APM allowances. Custom metrics: $5 per 100 indexed metrics/month and $0.10 per 100 ingested metrics/month. |
| Dynatrace | Logs and traces: $0.20/GiB ingestion plus $0.0007/GiB-day for extra retention. Metrics: $0.15 per 100,000 datapoints. The model includes a query-cost reserve. |
| Grafana Cloud | Logs and traces: $0.40/GB write, $0.05/GB processing, and $0.10/GB retention at 30 days, prorated to 15 days in the model, for $0.50/GB total. Metrics: $6.50 per 1,000 active series/month. |
| New Relic | $0.40/GB plus a $0.05/GB EU-region surcharge: $0.45/GB for logs, traces, and metrics, plus the seats listed below. |
We converted GiB to decimal GB and used the measured log and span sizes to translate per-million-event charges into comparable volume costs. Metrics were priced in each platform’s native units: bytes, active series/custom metrics, or datapoints.
| Platform | Traces | Logs | Metrics | Blended affordable share |
|---|---|---|---|---|
| Coralogix | 100% | 100% | 100% | 100% |
| Datadog | 18.0% | 28.0% | 10.3% | 18.5% |
| Dynatrace | 82.3% | 100% | 5.7% | 41.1% |
| Grafana Cloud | 34.3% | 86.3% | 79.1% | 43.4% |
| New Relic | 38.1% | 95.9% | 15.9% | 38.9% |
These percentages are calculated affordable shares, not the final sampling settings. Each platform’s budget is then allocated across traces, logs, and metrics. The full-data monthly cost estimates above divide the $10,000 reference budget by each platform’s blended affordable share and round to the nearest $1,000. They are illustrative estimates, rather than actual bills or quotes.
To fit the estimated budgets, we combined the following actions. In our opinion, these are reasonable steps an engineering team would take to control telemetry costs:
service.criticality tag to inform service-level cuts, with additional successful proxy-log exclusions for Datadog.| Telemetry | Coralogix | Datadog | Dynatrace | Grafana Cloud | New Relic |
|---|---|---|---|---|---|
| Traces | 100% | Errors + 45% of slow + 11% baseline | Errors + slow + 38% baseline | Errors + slow + 32% baseline | Errors + slow + 33% baseline |
| Spans vs. Coralogix, normal traffic | 100% | ~8–10% | ~36–41% | ~29–31% | ~30–32% |
| Logs | All | ~51% | All | ~93% | ~93% |
| JVM, .NET, and Go runtime metrics | Kept | Dropped | Dropped | Dropped | Dropped |
| Container and host metrics | Kept | Dropped | Core only; every 3 minutes | Kept | Every 2.5 minutes |
| Request, error, and latency metrics | All traffic | All traffic via APM stats | All traffic | All traffic | All traffic |
Slow means a whole trace over 1.5 seconds. For Dynatrace, Grafana Cloud, and New Relic, we kept all error and slow traces. At Datadog’s list price, we chose to fit the budget by keeping all error traces and 45% of slow traces. Each baseline percentage applies to routine traces.
The demo’s service.criticality resource attribute classifies services as critical, high, medium, or low. This informed the log-cutting choices: for the Grafana Cloud, New Relic, and Datadog budgets, we chose to exclude logs from recommendation, ad, quote, and fraud detection. For Datadog, we also excluded successful proxy logs while keeping error responses.
For Dynatrace’s budget, we trimmed per-disk, per-interface, per-core, and container network and disk details, while keeping core container CPU, memory, limits, throttling, and uptime.
Hover over an icon to see its exact score. You can also tap an icon or focus it with the keyboard. Use the switch to compare Cause + impact with Cause only.
| Scenario | Difficulty | Coralogix | Datadog | Dynatrace | Grafana Cloud | New Relic |
|---|---|---|---|---|---|---|
Delisted bestsellerEasyA bestseller disappears from the catalog. Product pages and checkouts fail. This tests whether agents can connect an obvious error to affected shoppers and financial exposure. We observed every agent finding the cause, as the failed checkouts and catalog logs were retained at every budget. Coralogix agents consistently measured the failure rate correctly. We observed sampled successful orders inflating the rates reported by several other vendors’ agents. In our scoring, Dynatrace and New Relic agents scored higher on the direct money question. | Easy | Coralogix18/205/5 | Datadog14.5/205/5 | Dynatrace16.5/205/5 | Grafana Cloud13/205/5 | New Relic16.5/205/5 |
Recommendation cacheMediumAn unbounded recommendation cache exhausts memory and repeatedly restarts the service. Agents must distinguish the crash-loop symptom from its memory mechanism and measure failures without claiming lost orders. In our scoring, Coralogix, Dynatrace, and New Relic agents tied on cause. Every Coralogix agent got the failure rates right, and we observed Grafana and Datadog agents also using unsampled request metrics effectively. New Relic agents explained the growing product list more often than Coralogix agents. The overall margin was narrow. | Medium | Coralogix14.5/155/5 | Datadog12.5/153/5 | Dynatrace11/155/5 | Grafana Cloud13.5/154.5/5 | New Relic11.5/155/5 |
Kafka queue floodMediumCheckout publishes duplicate order messages and fraud screening falls behind, while accounting keeps up. This tests a cause in one service, a symptom in another, and the value of orders left unscreened. We observed every agent finding the cause from duplicate logs and slow consumer traces. Coralogix agents joined complete order and screening records to count the unscreened orders and measure their value. Other vendors’ agents estimated counts from queue metrics, and we observed some matching the money when asked directly. | Medium | Coralogix18.5/205/5 | Datadog13/205/5 | Dynatrace13/205/5 | Grafana Cloud13.5/205/5 | New Relic13.5/205/5 |
CPU throttleMedium–hardA tight product-catalog CPU limit slows lookups without producing errors. Agents must identify throttling, measure abandoned requests, and distinguish normal order-rate variation from attributable revenue loss. In our scoring, Grafana agents scored highest, with accurate counts from unsampled request metrics and correct money answers. Coralogix, Datadog, and New Relic agents tied on cause, and Coralogix and New Relic agents tied for second overall. Two Coralogix agents claimed lost revenue from a change within normal variation. | Medium–hard | Coralogix13/155/5 | Datadog12.5/155/5 | Dynatrace10/152.5/5 | Grafana Cloud14.5/154.5/5 | New Relic13/155/5 |
Search relevanceMedium–hardA stricter relevance setting breaks multi-word searches while responses stay fast and successful. A customer complaint directs agents to investigate empty results, conversion, and orders that never happened. We observed every agent finding the search change. Complete search-to-order traces let Coralogix agents measure the conversion drop and estimate lost orders. We found that sampled traces made counts and conversion harder for the other agents to establish. With an open investigation question, no agent found this fault. | Medium–hard | Coralogix16/205/5 | Datadog10/205/5 | Dynatrace10/205/5 | Grafana Cloud11/205/5 | New Relic8/205/5 |
Ad garbage collectionHardRepeated full garbage collection makes the ad service slow. Agents must prove the runtime mechanism, count failed or abandoned ad requests, and avoid attributing unrelated order variation to the fault. We observed that slow traces were retained at every budget, but only the Coralogix data included JVM garbage-collection measurements. Every Coralogix agent proved the cause. We observed the other agents making partial inferences from container memory. In our scoring, Grafana agents scored higher on the money question because one Coralogix agent assumed USD. | Hard | Coralogix14.5/155/5 | Datadog5.5/150/5 | Dynatrace6.5/151/5 | Grafana Cloud10.5/152.5/5 | New Relic9.5/152/5 |
Early warningHardUnclosed worker pools accumulate idle threads in fraud detection before any errors or slowdown appear. Agents must find the risk, explain it, forecast failure, and describe the exposure if screening stops. Thread metrics revealed the mechanism. We found that container memory growth was enough to forecast the failure, but not to explain it. | Hard | Coralogix14.5/155/5 | Datadog0/150/5 | Dynatrace11/152.5/5 | Grafana Cloud10/152.5/5 | New Relic8.5/152.5/5 |
AI costHardA prompt change drops the JSON-only instruction, causing silent AI retries and higher token spend while customers still receive answers. A finance question asks agents to explain the increase and quantify waste. We observed every agent seeing token usage rise. Complete traces let Coralogix agents compare prompts and add up the price of every call. Dynatrace and New Relic agents found the cause but undercounted from sampled totals, and we observed Grafana agents using request metrics for accurate extra-call counts. Datadog agents also estimated dollar amounts from tokens. With an open question, no agent found the fault. | Hard | Coralogix10/105/5 | Datadog7.5/104/5 | Dynatrace7.5/105/5 | Grafana Cloud7.5/103/5 | New Relic7.5/105/5 |
Five faultsHardestPayment failures, ad garbage collection, a recommendation cache crash loop, periodic cart latency, and an email memory leak occur together. Agents must explore without knowing the fault count, explain causes, and prioritize customer and business impact. In our scoring, Coralogix agents got the most causes right, proved garbage collection, and each gave the exact failed-payment value per currency. We observed the Dynatrace agents finding every fault, but we found that, due to the data trimmed to fit the budget, they lacked GC proof and overstated payment failure rates from sampled traces. Two Coralogix agents missed the email leak. | Hardest | Coralogix28/3023/25 | Datadog11.5/308/25 | Dynatrace24/3020/25 | Grafana Cloud19.5/3015/25 | New Relic18.5/3014/25 |
OptimizeOpen-endedThe shop runs normally. Agents look for improvements and prioritize them by cost, reliability, and customer impact. We considered excessive collector CPU the most important finding; this is an opportunity-finding task, not a root-cause fault test. Coralogix agents found an average of 8.4 real items, and every agent found the top item (collector CPU). The default icons summarize the opportunity-finding results. Cause only shows no fault to score. | Open-ended | Coralogix8.4 items— | Datadog6.6 items— | Dynatrace6.8 items— | Grafana Cloud7.0 items— | New Relic8.0 items— |
Average scores across nine scenarios, weighted equally. Optimize is excluded because it tests finding improvements rather than an injected fault, so it has no comparable cause or impact score. The guides sit at 80% because this is the threshold for a strong result in our scoring criteria.
False claims are statements the reviewing agent found contradicted the reference evidence, such as inflated failure rates or claims that available telemetry did not exist. Missing a finding was not automatically a false claim.
Recorded counts across the final 10 scenarios, with five investigations per platform in each. Fewer is better. These counts are separate from the scoreboard scores.
We found that keeping complete telemetry helped agents explain why a problem happened and measure its impact more accurately.
In a scenario where searches returned no results, Coralogix agents could trace shoppers from search to checkout, measure how much less likely they were to place an order, and estimate the orders lost as a result. In a scenario assessing silent retries in an AI chatbot, by our scoring, Coralogix agents gave the most complete account of wasted spend. Having all the telemetry let them compare prompts, trace the repeated calls, and add up the cost of every AI call.
In our experience, more evidence was often the difference between seeing a symptom and proving its cause, or between estimating impact from a sample and measuring it from complete records.
We believe that with complete telemetry, agents can go beyond finding faults and build a clearer picture of key business metrics, including conversion, lost orders, and the cost of serving customers.
Easy · Missing product
A bestseller disappears from the catalog. Product pages and checkouts fail. This tests whether agents can connect an obvious error to affected shoppers and financial exposure.
Medium · Crash loop
An unbounded recommendation cache exhausts memory and repeatedly restarts the service. Agents must distinguish the crash-loop symptom from its memory mechanism and measure failures without claiming lost orders.
Medium · Distributed backlog
Checkout publishes duplicate order messages and fraud screening falls behind, while accounting keeps up. This tests a cause in one service, a symptom in another, and the value of orders left unscreened.
Medium–hard · Resource limit
A tight product-catalog CPU limit slows lookups without producing errors. Agents must identify throttling, measure abandoned requests, and distinguish normal order-rate variation from attributable revenue loss.
Medium–hard · Silent wrong results
A stricter relevance setting breaks multi-word searches while responses stay fast and successful. A customer complaint directs agents to investigate empty results, conversion, and orders that never happened.
Hard · Runtime slowdown
Repeated full garbage collection makes the ad service slow. Agents must prove the runtime mechanism, count failed or abandoned ad requests, and avoid attributing unrelated order variation to the fault.
Hard · Predicted failure
Unclosed worker pools accumulate idle threads in fraud detection before any errors or slowdown appear. Agents must find the risk, explain it, forecast failure, and describe the exposure if screening stops.
Hard · Silent AI retries
A prompt change drops the JSON-only instruction, causing silent AI retries and higher token spend while customers still receive answers. A finance question asks agents to explain the increase and quantify waste.
Hardest · Concurrent faults
Payment failures, ad garbage collection, a recommendation cache crash loop, periodic cart latency, and an email memory leak occur together. Agents must explore without knowing the fault count, explain causes, and prioritize customer and business impact.
Open-ended · No injected fault
The shop runs normally. Agents look for improvements and prioritize them by cost, reliability, and customer impact. We considered excessive collector CPU the most important finding; this is an opportunity-finding task, not a root-cause fault test.
These were the standard investigation questions. Specialized scenarios adapted them for forecasting, optimization, customer complaints, or AI spending.
Each response was assessed against the known fault and the scenario’s reference measurements. We scored the cause and the relevant impact components separately:
15/20 means 15 of the 20 available points across five investigations: 75% of the maximum score.
The totals follow each scenario’s rubric. Compare platforms within the same scenario.
For example, in Delisted bestseller, Coralogix earned 5 for cause + 5 for affected counts + 4 for volunteered financial impact + 4 for financial impact after the money question = 18/20. This describes the quality and consistency of the answers to that incident; it is not a count of incidents solved.
AI cost scores cause and financial impact for 10 points. Five faults has 25 possible cause points plus 5 money points, for 30 total.Telemetry generated from the OpenTelemetry Astronomy Shop Demo. Monthly costs are modeled from public list prices and are not quotes or actual bills. Results describe the tested data configurations.