What is a security data lake?
A "security data lake" is a high-volume, full-fidelity store for security telemetry that stays queryable for hunting and investigation, rather than a compressed, sampled summary of it. The term gets used loosely, so it helps to be precise about what the lake does and does not do. A SIEM correlates events into alerts, manages cases, and gives analysts a workflow for triage and response. A security data lake holds the raw and normalized events behind those alerts, at a scale and cost that lets teams keep months or years of history instead of weeks.
Most organizations already run a SIEM, and it stays useful for exactly what it was built for: detection logic, alert routing, and case management. The gap shows up when SIEM ingestion costs force a choice between shrinking the data going in or shrinking how long it stays. A security data lake sits beside the SIEM as the event store and query layer, holding the full volume of telemetry the SIEM can't economically retain, while the SIEM keeps doing detection and case work on the subset it needs. On Axiom's Splunk comparison, the Security / SIEM row for Axiom is log search, and this piece treats that as a starting constraint, not a limitation to work around.
Step 1: Decide what gets ingested and how it gets normalized
The first decision is source coverage. A useful security data lake typically ingests VPC flow logs, WAF events, cloud API activity such as CloudTrail or Azure Activity Log, EDR telemetry, and authentication logs from identity providers. These sources rarely arrive in a shared shape. Field names differ, timestamps use different formats, and severity levels mean different things depending on the vendor.
Two approaches handle this. Upfront normalization standardizes field names, timestamps, and event categories before the data lands in storage, so every source looks the same downstream at query time. OCSF (the Open Cybersecurity Schema Framework) is the open schema built for this, and Axiom publishes its OCSF schema on GitHub. Schema-on-read instead preserves the raw event and applies structure at query time, trading some upfront engineering for flexibility when a new source shows up or a field turns out to matter later than expected.
Most security data lakes end up using a mix: normalize the fields analysts query constantly, like source IP, user identity, and event outcome, while keeping the raw payload available for anything normalization missed. Axiom supports both sides of that mix: schema-less ingest keeps the raw event, and virtual fields apply structure at query time, so a field that matters later does not need a re-ingest.
The pattern of normalizing security logs from multiple sources while preserving the original event is a practical template for this step. The post's example is correlating VPC Flow, WAF, application auth, and GuardDuty around one incident, without discarding the raw record.
Step 2: Pick an event store that keeps full history affordable
Once ingest and normalization are settled, the store has to hold the volume without forcing a sampling decision. Sampling is the default cost lever for high-volume telemetry, and it carries the sharpest security downside of any shortcut in this list. Sampling security events can drop exactly the low-and-slow connection patterns an attacker relies on, the kind of activity that only looks meaningful once every event in the sequence is visible. A security data lake is built to avoid that trade-off, so the event store needs compression and query economics that make full-fidelity retention affordable rather than aspirational.
This is where the architecture choice matters most. A full-fidelity event store that keeps compressed history in the normal query path, with no rehydration step required as data ages, removes the operational tax that usually pushes teams toward shorter retention windows or sampled ingestion. Running that store yourself is possible, but it means owning cluster sizing, compaction, and capacity planning for a workload that grows with attacker dwell time, not on a predictable schedule. A fully managed event store removes that operational toil, letting the team that owns the security data lake spend its time on queries and coverage instead of keeping a cluster alive.
Step 3: Build the query path analysts will use
If analysts never query it, the store is an expensive archive. The query path has to match how security analysts work: starting from a hypothesis, narrowing across time and fields, and following a thread from one anomalous event to the next. That's an ad hoc, iterative pattern, closer to a series of questions than a fixed dashboard.
APL, Axiom's piped, sequential query language, is built for that kind of investigation. Each stage of a query filters, groups, or transforms the result of the stage before it, which maps onto how a hunt proceeds: start broad, filter to a suspect IP range, group by user, pivot to a related source. Dashboards still have a role for standing metrics and known-bad indicators. But the query path an analyst or their agent peer reaches for during an active investigation needs to support unplanned questions, beyond the ones anticipated when the dashboard was built. The SIEM keeps its own case management and alert workflow: the data lake's query path is where the follow-up investigation happens once a case is open.
Step 4: Set retention as a policy decision
Retention windows in most stacks get set by what the budget allows, then get treated as a fact of the architecture rather than a choice. A security data lake works better when retention starts from an actual requirement, whether that's the audit and investigation window a compliance framework demands or typical dwell times for the industry, and the store is sized to match.
This only works if retention is configurable rather than fixed at the platform layer. Axiom's configurable retention and usage-based pricing let teams set the window that matches their compliance and investigation needs, without paying for a separate indexing tier or negotiating a new SKU every time the requirement changes. Governance controls matter here too. Self-serve RBAC, audit log, and SSO keep access to a security data lake auditable in the same way the telemetry it stores is. The store itself also needs to meet the bar the data deserves: encryption, audit logging, and tenant isolation are required for a system holding security telemetry.
Step 5: Plan coexistence with the existing SIEM
The rollout that works in practice does not touch the SIEM's detection logic or case management at all. It starts with the sources that are either too high-volume to justify SIEM ingestion cost or effectively orphaned today. VPC flow logs, cloud audit trails, and raw EDR telemetry are common candidates, and they land in the security data lake first. The SIEM keeps every detection rule, alert, and case workflow it already runs, pointed at the sources it already ingests.
This is deliberately the same shape as Axiom's broader migration pattern. On Axiom's Splunk comparison, the Security / SIEM row for Axiom is log search, and the practical path is to land beside the existing SIEM on the workloads it was never economically sized for, prove the model on that slice, then expand coverage at the natural renewal seam rather than forcing a cutover. Analysts get a wider window of full-fidelity history to pull from during an investigation, those SIEM jobs stay exactly where they are, and the two systems divide the work along the line each one is built for. For Splunk environments, Axiom for Splunk keeps that split: security telemetry lands in Axiom, and Splunk keeps detection and response.
A reference architecture in one pass
Telemetry from cloud, network, identity, and endpoint sources gets normalized enough to query consistently, while raw payloads stay available for anything normalization missed. That data lands in a fully managed event store sized for full-fidelity retention rather than sampled snapshots. Analysts query it through APL, moving from hypothesis to evidence in the same session, while the SIEM's dashboards and alerts keep running against the sources it already covers. Retention is set against compliance and investigation requirements, not against whatever the budget happens to allow. The SIEM stays the system of record for detections and cases, and the data lake becomes the system of record for the evidence behind them.
Getting started without a rebuild
Standing up a security data lake next to an existing SIEM is an ingest and query addition, not a replacement project. The first useful step is landing one high-volume source in the lake this quarter and running a real investigation against it.
FAQs
Does a security data lake replace our SIEM?
No. The pattern described here is coexistence. The SIEM keeps detection and case management, since that's what it's built for. The security data lake takes on the sources and retention windows the SIEM was never sized to hold economically, and gives analysts a wider window of full-fidelity history to query during an investigation.
Which sources should we land in the data lake first?
Start with the highest-volume or most orphaned sources: VPC flow logs, cloud audit trails like CloudTrail or Azure Activity Log, and raw EDR telemetry. These are usually the sources driving SIEM ingestion cost the hardest, or the ones getting dropped entirely once budget runs out. Landing them in the lake first proves the model without touching any existing detection rule.
How is a security data lake different from just archiving logs to cold storage?
Cold storage trades queryability for cost. A security data lake keeps the same data queryable at high volume, without a rehydration step before an analyst can search it. If pulling data out of storage requires a restore job before a hunt can start, that's an archive, not a data lake.
Why does sampling matter more for security telemetry than for other logs?
Sampling assumes that dropping some percentage of events preserves the signal that matters. For low-and-slow attacker behavior, the signal is often the sequence of small, individually unremarkable events, and sampling can drop exactly the events that would have made the pattern visible. That risk is specific to security use cases in a way it isn't for, say, routine application metrics.
How long should we retain security telemetry?
Long enough to cover the audit window a compliance framework requires and the dwell time patterns typical for the industry, whichever is longer. Retention should be set against that requirement rather than against whatever a fixed budget happens to allow, which is why configurable retention matters more here than in general-purpose log management.
What does Axiom provide in this architecture?
Axiom is the event store and query layer: ingest and normalization support, a fully managed store built for full-fidelity retention without rehydration, and APL as the query language analysts use during investigations. Axiom does not provide detection rules, alert correlation, or case management. Those stay with the SIEM.
Related reading
Try it on your own security telemetry
The fastest way to evaluate this pattern is to land one high-volume source, VPC flow logs or an EDR feed are common starting points, and run a real hunt against it. Book a demo to see the ingest path, APL, and retention controls against your own data.
