perspectives

How to think about security data lakes in 2026

Why the difference between a queryable lake and a rehydration-bound archive matters more in 2026 than price per gigabyte.

Axiom · · 11 min read

Security teams have spent the last two years hearing that the answer to rising telemetry volume is a lake. Cloud audit trails, VPC flow logs, EDR events, WAF hits, and identity provider activity keep growing faster than any SIEM's ingest budget. Vendors across the industry have responded with a new storage tier that sits underneath the detection engine, and that is a reasonable response to a real problem. What matters in 2026 is what happens to the data once it lands in that tier, and that is where the architectures start to diverge in ways worth understanding before a platform team commits to one.

What a security data lake is

A "security data lake" is a storage layer built to hold large volumes of raw security telemetry economically, and for longer than a traditional SIEM index would allow, with the SIEM's detection rules, correlation logic, and case management sitting above it as a separate application. The lake is the data plane. The SIEM is the application layer. This is a deliberate split, and it is a good one in principle: it lets a team retain a full year of authentication logs or flow records without paying SIEM-tier indexing costs on all of it, while keeping the detection rules that fire alerts running against a smaller, higher-value working set.

The split between data plane and application

The split only pays off if the two layers stay connected in practice, not just in the architecture diagram. A lake that holds the data but requires a separate step to bring it back into a searchable state before an analyst can query it has reintroduced the exact bottleneck the lake was supposed to remove. Two products can both call themselves a security data lake and behave differently once an analyst needs data that is 90 days old at two in the morning.

The pattern the industry keeps repeating

A common pattern among incumbent SIEM vendors is what we would describe as lake-plus-promotion. Telemetry lands in a low-cost tier, and when a team needs to search it during an investigation, that data gets promoted, restored, or reindexed into the SIEM's active search layer first. The lake, in this design, functions less like an extension of the query path and more like a second operational copy of the data, one that has to be created on demand before it is useful. That promotion step might be fast in some cases and slow in others, but the underlying pattern is the same: querying old security data becomes a two-stage process instead of one.

This is a reasonable place to land when the search index and the storage tier were built as separate systems. The result is that "cheap storage" and "a place we can search" turn out to be two different promises, and a team evaluating a lake needs to know which one it is buying.

The 2026 question: query path or archive

For the year ahead, either the lake stays in the ordinary query path, or it is an archive someone has to rehydrate before the SOC can use it. Price per gigabyte gets most of the attention in vendor comparisons, but a lake that is inexpensive to store and expensive in analyst time during an investigation has moved the cost from the invoice to the incident timeline. A rehydration ticket that takes twenty minutes, or twenty hours, is the same problem sampling causes, under another name.

Under budget pressure, many organizations quietly filter, sample, or drop high-volume sources like VPC flow logs, WAF events, and cloud API activity before that data ever reaches the SIEM, a pattern we have described in detail when looking at security log normalization and retention economics. The intent is cost control. The effect is a blind spot, because the record that would reconstruct an incident may not exist. Sampling has the same failure mode as a slow-to-rehydrate lake: both trade completeness of the historical record for a lower running cost, and both bills come due at the moment a team can least afford them, mid-investigation. We have written before about the risks of sampling security telemetry, particularly for low-and-slow attacks that rely on being the one request out of a hundred that got dropped.

What changes when detections and cases stay put

Keep the SIEM a security team already trusts. The detection rules, the case management workflow, the escalation paths, and the analyst muscle memory built up over years of running a SIEM are not the part of the stack under pressure. What is under pressure is the volume of raw telemetry sitting underneath those detections, and the retention window a team can afford to keep searchable.

Coexistence is a practical starting point. The SIEM keeps doing what it does well: correlation, alerting, and case workflow, while a lake underneath handles the volume and the retention. Security teams do not need to relearn how they investigate an incident. They need the data behind that investigation to reach back further in time without turning every older query into a support ticket.

Where Axiom fits in this split

Axiom is the modern machine data platform, and its position in this split is straightforward. Axiom's Splunk comparison lists Axiom under Security / SIEM as log search, and we built our architecture around keeping data in the ordinary query path rather than moving it into a second, promoted copy. Our event store architecture, with no rehydration step stores settled security event data in a compact, immutable form that stays queryable without a customer-triggered rehydration workflow. A security team can reach back through months of VPC flow logs or authentication events the same way it would query the last five minutes, because the underlying system treats old and new data as one query path rather than two separate tiers with different rules.

That same continuity extends to how those events get queried in the first place. Security telemetry sits alongside application logs, events, traces, and metrics under petabyte-scale event data, kept exact. The pressure to sample security data almost always traces back to the same data center architecture problem that ingestion at scale creates for any high-volume telemetry source, security or otherwise.

Two other pieces of the architecture matter specifically for security and compliance teams evaluating a lake. First, self-serve RBAC, audit log, and SSO for security teams are table stakes for any team that will use this data plane to hold regulated or sensitive security telemetry, and they should not require a professional services engagement to configure. Second, as SOC investigation work increasingly involves agents querying the same telemetry a human analyst would, agent-native investigation with full access scope becomes relevant. An agent working an incident needs to reach the same data a person can reach, not a narrower slice gated behind a separate indexing tier. Underneath all of it, encryption, tenant isolation, and audit logging by default are the baseline a security data plane has to meet before the query performance conversation matters.

A short list of questions to ask any vendor

A few concrete questions cut through most of the marketing language around security data lakes:

  • What happens to data the moment it lands: is it queryable immediately, or does it wait for a promotion or reindexing job?

  • Does querying data from six months ago use the same interface and the same wait time as querying data from six minutes ago?

  • What gets sampled or filtered before it reaches the lake, and who made that decision?

  • Can an agent or automated investigation reach the same historical data a human analyst can, or is agent access gated to a narrower tier?

  • What does the vendor's own architecture documentation say happens between ingest and the point where a query can run against that data?

Incumbent SIEMs have been adding lake-tier storage. What is harder to verify from the outside is which of those tiers behave as an extension of the query path and which behave as an archive with a promotion step in front of it. That is the question worth asking directly, rather than assuming the term "data lake" answers it on its own.

FAQs

Is a security data lake meant to replace our SIEM?

No. In practice, a security data lake is a storage layer that sits underneath a SIEM, not a substitute for one. The SIEM still owns detection rules, correlation logic, and case management. The lake handles the volume and retention window that SIEM-tier indexing was never designed to hold affordably.

Why do some security data lakes feel slow during an investigation?

Many incumbent lakes were added onto an existing SIEM architecture as a separate storage tier, which means older data often needs to be promoted, restored, or reindexed into the active search layer before it is queryable. That promotion step can add minutes or hours to a query that should feel instant, especially mid-incident.

What is the difference between a lake and an archive?

An archive requires a deliberate action, such as a restore job or a rehydration request, before the data can be searched. A lake that lives in the ordinary query path lets an analyst run the same query against data from six months ago as against data from six minutes ago, with no separate step in between.

Does keeping raw security telemetry in a lake mean we have to give up sampling?

It means sampling becomes a choice rather than a forced cost-control measure. Many teams sample high-volume sources like VPC flow logs or WAF events specifically because full retention was too expensive to search. A lake that stays in the query path at scale removes that pressure. Writing to a block format that compresses machine data more efficiently than any known open format helps too. Even before a query runs, each lake gigabyte goes further with Axiom, so the decision to sample can be made on security grounds rather than budget ones.

How do we tell whether a vendor's lake is in the query path?

Ask what happens to data the moment it lands: is it queryable immediately, or does a background job need to run first? Also ask whether querying six-month-old data uses the same interface and response time as querying data from the last five minutes. If the answer involves a ticket, a restore window, or a separate reindexing step, it is functioning as an archive.

Where does agent-driven investigation fit into this?

As more SOC workflows involve an agent querying telemetry alongside a human analyst, the agent needs the same reach into historical data that a person has. If older security data sits behind a separate promotion step, an agent working an incident hits the same bottleneck an analyst would, just automated.

Do compliance requirements change when telemetry moves into a lake?

The expectations do not lower because the storage tier changed. Role-based access control, audit logging, and single sign-on should be available on the lake itself, not bolted on later, since it is now holding the same regulated security data the SIEM used to hold alone.

Related Reading

See how Axiom handles security telemetry at scale

The fastest way to test any of this is against your own data. Book a demo to see how Axiom keeps security event data in the ordinary query path, from the last five minutes to years back, without a rehydration step in between.