Source
Data is ingested from any external source using Kafka Connect, which produces events and state to Kafka topics and KTables.
The total Coralogix estate over the past 7 days: scale, complexity and query volume.
A scaled account ingesting 40 TB a day, 1.2 PB queryable and 483 million timeseries of metrics.
A mid size account ingesting 300 GB a day, 9 TB queryable and 847,667 timeseries of metrics.
Events estimated at 1 KB each.
Learn more: DataPrime documentation DPXL expressions
Learn about an architecture that handles billions of daily queries and hundreds of petabytes of data, supporting hundreds of industry leading products and driving unprecedented customer savings.
Take the tour
What sets Coralogix apart is its data layer: a single Telemetry Lake. You ingest unsampled, full-fidelity data, keep it in your own cloud object storage, and query it whenever you like, whatever the signal, use case or reader. Here’s how it works.
Most platforms fix a field’s type the first time they see it. When the type changes, as it does whenever a team ships new code, you get a mapping exception, and the event is rejected or the field goes unindexed.
DataPrime builds an evolving schema instead. It tracks every field’s type over time and supports multi-type fields, so a field can be a number one day, a string the next and an object after that. Every event is kept, and every version of the field stays queryable.
Go indexless with Coralogix and you’re done with mapping exceptions and schema management for good.
Customers want more than open source. They want portable standards that add value to their business without tying it to a vendor.
We contribute to OpenTelemetry, from its Kubernetes Operator and Helm charts to the OpenTelemetry eBPF Instrumentation (OBI), and we’re contributing more all the time.
Observability use cases are measured in seconds, so we started by rethinking ingestion. Streama, our Kafka-based processing engine, transforms, enriches, parses, alerts on and generates metrics from data as it streams, without waiting on indexing or storage. Alerts fire faster, transformations happen instantly, and the platform costs less to run, a saving we pass on to customers.
Because analysis happens before storage, Streama can make routing decisions for you: generate a metric from a log and drop the original, and you keep the signal without paying to store the noise. That’s the power of separating storage from analysis.
Data is ingested from any external source using Kafka Connect, which produces events and state to Kafka topics and KTables.
Events flow to Kafka for stream analysis and are automatically parsed, enriched, and clustered using machine learning algorithms.
Data and insights go to any external destination once they’ve passed through the stream analysis engine.
Stream processing is CPU-bound work, so Coralogix scales ingestion and processing independently of storage and indexing. When traffic spikes, Streama adds capacity in seconds and releases it when traffic falls, so we keep pace with customer demand without waiting on storage.
Alerting runs in the stream, with no dependency on storage. Alerts analyse data as it moves through the stream, meaning instant alert triggers and no indexing lag. This architecture also means that in the event of query or ingest lag, alerts operate just fine.
This drives millions of alerts, with no impact on performance or cost.
Coralogix writes every event to your cloud object storage by default, and DataPrime queries it there directly. For the small subset of events that need millisecond ingestion and the fastest queries, you can also write to OpenSearch.
Data lands in wide Parquet, a columnar format we’ve tuned for observability. You route data into your own datasets and choose the bucket each one writes to.
Your data rests in your cloud object storage, the most scalable place to keep telemetry, and nowhere else. You pay Coralogix for processing, not retention.
It fits into your wider data strategy, with no vendor lock-in and no punitive migration charges.
Logs, spans, SIEM events, RUM sessions and AI sessions are stored as Parquet, and metrics in an open, machine-readable format, in your own object storage. Any engine that reads Parquet can read your telemetry: Apache Spark, Trino, Athena, DuckDB, your warehouse, a notebook.
So Coralogix storage is not a dead end. It is a nexus: one lake that the rest of your data architecture consumes from directly, with no export step and no second copy.
Coralogix uses dynamic materialization: the fields you query most stay materialized for speed, and the ones you rarely touch move into a lower-footprint encoding. It weighs each field’s cardinality and how often it’s read, so your storage footprint is never larger than your queries need.
It tracks thousands of fields at once, so however broad your access patterns, the data you need is in the right encoding. Fields that aren’t materialized are still fully queryable: DataPrime decodes them on the fly.
DataPrime is one language for all your telemetry, with schema-on-read querying, complex joins, window functions, unions and hundreds of native functions and commands. The engine plans every query and runs it in parallel, so even complex analyses scan terabytes in seconds.
And it never needs rehydration: any part of your data, however old, is queryable at any time.