Most security tools fail quietly in the same place: they never received the event that mattered. Before you can talk about AI-driven detection, automated investigation, or a security copilot, you need security data collection that actually reaches every corner of your environment: cloud control planes, Linux and Windows servers, containers, SaaS apps, and network devices.
This is a problem for anyone running or buying a security platform, because collection gaps are invisible until an incident. Logs that were never ingested look exactly like activity that never happened. Teams discover the gap when an attacker's path runs through a SaaS app or a DNS request nobody was recording, or when the cloud bill reveals what the alerts did not.
This article explains what "every source" really covers across five domains, the engineering trade-offs you have to settle (agent versus agentless, push versus pull, volume and cost, coverage and reliable delivery), how thorough collection differs from the superficial kind, what good looks like, a step-by-step way to build the layer in phases, a concrete before-and-after example, and realistic cost and timeline expectations. It is written for engineering leads, security teams, and founders scoping an AI security platform or auditing the visibility they already have.
You Cannot Detect What You Cannot See
Every clever thing an AI security platform does — the enrichment, the correlation, the conversational copilot, the automated response — depends on one unglamorous fact: it has the data. If an event never gets collected, no amount of AI further up the stack can reason about it. It simply didn't happen, as far as the platform is concerned.
This is why the data collection layer is the foundation the whole AI-native security platform stands on. It is also the layer most likely to be quietly incomplete. A platform that ingests AWS CloudTrail but misses the SSH logins on the servers, or watches the servers but ignores the SaaS apps where half the company's data actually lives, has blind spots — and attackers are drawn to blind spots the way water finds the crack in a wall.
The goal of this module is deceptively simple to state and genuinely hard to deliver: collect security-relevant events from every possible source, reliably, at a cost you can live with. This piece walks through what "every source" really means, the engineering tensions you have to resolve to get there, and how to tell a thorough collection layer from a superficial one.
What "Every Source" Actually Covers
Coverage is not one checklist. It's several overlapping domains, each with its own log formats, access methods, and failure modes. A serious collection layer reaches into all of them.
Cloud providers
For most modern businesses, the cloud control plane is the crown jewels. The events that matter here are less about traffic and more about who did what: IAM changes, new access keys, role assumptions, and permission grants. On AWS that means CloudTrail as the backbone, plus VPC Flow Logs for network movement, S3 access logs, RDS and Lambda logs, and the findings that GuardDuty and Security Hub already produce. Azure and GCP have their equivalents — activity logs, audit logs, and their native threat-detection findings.
Pulling in the provider's own findings (GuardDuty, Security Hub) is worth calling out: you are not just collecting raw events, you are collecting the cloud's own opinions about them, which the platform can then correlate with everything else.
Linux and Windows servers
Servers are where a foothold turns into a breach. On Linux, the high-value signals are SSH logins, sudo usage, cron jobs, auditd records, kernel messages, and syslog — plus the application logs that sit on top: nginx, Apache, MySQL, PostgreSQL, Redis, and Docker. On Windows, it's the Windows Event Logs, PowerShell activity, RDP sessions, Active Directory events, service and registry changes, and Defender output. A failed-password line or an unexpected sudo at 3am is often the first honest sign that something is wrong.
Containers and Kubernetes
Containerised estates move fast and leave thin trails. At the Docker level you want container creation and deletion, image pulls, known image vulnerabilities, and container network traffic. At the Kubernetes level, the events that matter are pod creation, secrets access, RBAC changes, node status, ingress logs, and — critically — the API server logs, which are the audit trail for the entire cluster. A pod that suddenly reads a secret it never touched before is exactly the kind of signal that only exists if you collected it.
SaaS applications
This is the domain teams most often forget, and it's often where the sensitive data actually lives. Microsoft 365 and Google Workspace hold email and files. Slack and Zoom hold conversations. Dropbox and Salesforce hold documents and customer records. GitHub, GitLab, and Bitbucket hold your source code and deployment keys. Each exposes an audit API — sign-ins, sharing changes, admin actions, repository access — and collection here is almost always agentless, pulling from those APIs on a schedule.
Network devices
Finally, the network layer: firewalls, VPN concentrators, routers, switches, proxies, DNS resolvers, and IDS/IPS sensors. DNS logs in particular are underrated — a lot of malware phones home over DNS, and the request to a suspicious domain is often visible long before anything else fires.
The Engineering Tensions You Have to Resolve
Listing the sources is the easy part. The reason collection is hard is that every source forces a set of trade-offs, and getting them wrong is expensive to fix later.
Agent-based vs agentless
The first fork in the road. An agent is a small piece of software you install on the server or endpoint. It sees everything locally — processes, file changes, kernel events — and can collect deep signals that no remote API exposes. The cost is operational: you have to deploy it, keep it updated, and account for the resources it consumes on every host.
Agentless collection pulls from an API or a log endpoint the source already exposes — CloudTrail, the Microsoft 365 audit API, a firewall's syslog stream. Nothing to install, which makes it fast to onboard and impossible to break the host. The cost is depth: you only see what the API chooses to expose, and often at a delay.
In practice a good platform is not dogmatic about this. Cloud and SaaS are almost always agentless — there's no host to put an agent on, and the APIs are the right interface. Servers and endpoints usually justify an agent, because the deep signals (a suspicious process spawning, a config file changing) simply aren't available any other way. The right answer is a blend, chosen per source.
Push vs pull
Related but distinct. In a push model, the source sends events to the platform — a server's agent streams logs out, a firewall forwards syslog. In a pull model, the platform reaches out on a schedule and asks the API for what's new. Push gives you lower latency and works well for high-volume streaming sources. Pull is simpler to secure (the platform initiates the connection, so you're not opening inbound ports on every host) and is the natural fit for SaaS APIs. Most real deployments run both, and the collection layer has to handle each cleanly.
Log volume and cost
This one blindsides people. Security logs are enormous. VPC Flow Logs alone can dwarf everything else. A mid-sized estate can easily generate hundreds of gigabytes a day, and if you're paying a per-gigabyte ingest fee — the pricing model of most incumbent SIEMs — costs scale in a way that quietly punishes you for having good visibility. The engineering answer is to filter at the edge (drop the genuinely useless noise before it travels), tier your storage (hot store for recent data, cheap object storage for older retention), and choose a log store that doesn't fall over at real volume. This is where the choice of the log processing pipeline downstream — and a columnar store like ClickHouse — earns its keep.
Coverage and reliable delivery
The last two tensions are about trust. Coverage is the discipline of knowing what you're not collecting. A collection layer should track which sources are configured and, more importantly, surface the ones that aren't — an unmonitored server is worse than a known gap because nobody's thinking about it. Reliable delivery is the promise that once an event is generated, it actually arrives. That means buffering when the pipeline is busy, retrying on failure, and detecting silence — if a source that normally sends 10,000 events an hour goes quiet, that's either a broken collector or an attacker who disabled logging, and both need an alert.
Thorough vs Superficial Collection
The difference between a collection layer that protects you and one that looks like it does comes down to details that are invisible in a demo.
| Dimension | Superficial collection | Thorough collection |
|---|---|---|
| Cloud | CloudTrail only | CloudTrail + VPC Flow + S3/RDS/Lambda + GuardDuty findings |
| Servers | Syslog forwarding | Agent capturing SSH, sudo, auditd, process and file events |
| SaaS | Not collected | Audit APIs for M365, Workspace, GitHub, Slack |
| Containers | Ignored | Docker + Kubernetes API server and RBAC events |
| Network | Firewall allow/deny | Firewall + DNS + VPN + IDS/IPS |
| Coverage tracking | None — you hope it's complete | Explicit inventory of configured vs missing sources |
| Delivery | Best-effort | Buffered, retried, with silence detection |
| Cost control | Pay per GB, no filtering | Edge filtering + tiered storage |
Neither column is a demo you can tell apart in five minutes — both light up dashboards. The difference shows the day an incident happens: whether the event you need to reconstruct what happened was actually captured, or whether it fell into a gap nobody was tracking.
Benefits of Thorough Security Data Collection
Good collection rarely shows up in a demo, but almost every downstream capability depends on it.
Detection that covers the real attack path
Attacks rarely stay in one domain. A phished SaaS account leads to source-code access, which leads to a leaked cloud key, which leads to new infrastructure. When every step is collected, detection rules and AI correlation can see the chain rather than isolated fragments. Thorough collection is what turns three weak signals in three tools into one strong alert, early enough to act on.
Investigations that can actually be completed
When something does go wrong, the first question is what happened and in what order. Analysts and investigation agents can only answer that from the events that were captured. With cloud, host, SaaS, and network events in one place, an incident timeline can be rebuilt in hours instead of days of requesting exports from different teams, and gaps in the story are far less likely to stay unexplained.
Tampering becomes visible
Attackers often try to disable logging early. A layer with silence detection treats a source going quiet as an event in its own right. That turns one of the quietest intrusion techniques into an alert, and it also catches the mundane case of a broken collector before it leaves a weeks-long hole in the record.
Predictable costs without sacrificing visibility
Edge filtering and tiered storage break the link between seeing more and paying more per gigabyte. Teams no longer face the choice of switching off useful sources to stay within budget. Cost becomes something engineered and tuned, rather than a reason to accept blind spots.
Faster onboarding of new systems
A collection layer built around a common schema and a clean ingestion abstraction absorbs a new SaaS app or cloud account as configuration. Security keeps pace with the business instead of lagging behind every new tool the company adopts, which is often where unmonitored data ends up living.
Honest answers about coverage
Explicit coverage tracking lets a security team say exactly which sources are monitored and which are not. That is useful for internal risk discussions, customer security questionnaires, and auditors, and it replaces the vague confidence that comes from dashboards full of activity.
Security Data Collection Use Cases
These scenarios show where a complete collection layer makes the difference, and which sources each one depends on.
Catching a leaked cloud access key
A developer's credentials are exposed and an attacker uses an access key to create resources. CloudTrail records the API calls, GuardDuty may flag the behavior, and VPC Flow Logs show the new traffic. With those sources collected and watched, the unusual IAM activity raises an alert when it happens, rather than surfacing weeks later on the cloud bill. The concrete example below walks through this exact path.
Spotting a compromised SaaS account
An employee's Microsoft 365 or Google Workspace account signs in from an unfamiliar location, then creates a forwarding rule or shares files externally. Only the SaaS audit APIs reveal those actions. Pulling them on a schedule and correlating sign-ins with admin and sharing events lets the platform flag the takeover before sensitive documents leave the organization.
Detecting a foothold on a server
An internet-facing Linux host shows a run of failed SSH logins, then a successful one, followed by an unexpected sudo and a new cron job. Syslog forwarding alone might show the logins, but a host agent capturing auditd, process, and file events shows what the intruder did next. That sequence is what separates a noisy scanner from an active compromise.
Finding malware calling home over DNS
A workstation infected through a malicious attachment starts making regular requests to an unusual domain. Nothing else has fired yet. Collected DNS resolver logs make that request visible, and enrichment downstream can match the domain against threat intelligence. This is one of the earliest signals available, but only for teams that treat DNS as a first-class source.
Watching for secrets access in Kubernetes
A pod that has never read a particular secret suddenly does, or an RBAC binding grants a service account broader rights. Kubernetes API server audit logs are the only record of these events. Collecting them lets the platform flag lateral movement inside a cluster that would otherwise look like normal container churn. Because containers are short-lived, these audit records are often the only trace left once a compromised pod has been replaced.
Security Data Collection Best Practices
If you're evaluating or specifying this layer, these are the practices that separate a foundation you can build on from one you'll have to rip out.
- Breadth by design. The architecture assumes new source types will keep arriving and absorbs them without a rewrite. A new SaaS app or a new cloud service should be a configuration exercise, not an engineering project.
- A blend of agent and agentless. Deep host signals where they matter, API pulls where that's the right interface — chosen per source, not applied dogmatically.
- Explicit coverage tracking. The platform knows what it's collecting and what it isn't, and treats an unmonitored source as a visible gap rather than silence.
- Reliable delivery with silence detection. Events are buffered and retried, and a source going unexpectedly quiet raises an alert — because disabling logging is a classic early move in an intrusion.
- Cost-aware ingestion. Filtering at the edge and tiered storage, so good visibility doesn't become a bill that pressures you into collecting less.
- Clean handoff downstream. Collected events land in a form the log processing pipeline and detection engine can immediately work with, rather than a mess the next layer has to untangle.
- Least-privilege collector credentials. Collectors hold read access to some of the most sensitive audit trails in the company. Give each one a dedicated, narrowly scoped identity, rotate its credentials, and log its own activity, so the collection layer does not become an attacker's shortcut.
- An owner and a test event for every source. Each source should have a named person responsible for it and a known event, such as a failed login or a new access key, that can be generated on demand to prove the pipeline still works end to end after any change. Run those tests after collector upgrades and permission changes, not just at onboarding.
A Concrete Example
Picture a 40-person logistics firm running a modest but real estate: a handful of AWS accounts, a dozen Linux servers, Microsoft 365 for email and files, GitHub for code, and a firewall at the edge.
Before a proper collection layer: they had CloudTrail flowing into an S3 bucket nobody read, and their servers wrote logs locally. When a developer's laptop was compromised and the attacker used a leaked access key to spin up resources, the evidence existed — in CloudTrail — but it was never being watched, and the M365 sign-in from an unfamiliar country that preceded it was never collected at all. They found out from their cloud bill.
After: CloudTrail, VPC Flow Logs, and GuardDuty findings feed the platform in near real time; a lightweight agent on each server streams SSH and sudo events; the M365 and GitHub audit APIs are pulled on a schedule; and the firewall forwards DNS and connection logs. The same attack now surfaces as a sequence — an anomalous M365 sign-in, followed by GitHub key access, followed by unusual IAM activity in CloudTrail — that correlates into a single alert hours before any resources spin up. Nothing changed about the attack. What changed is that every step of it was being collected.
Cost and Timeline Reality
Honest framing: the collection layer is unglamorous work, and it's where a surprising amount of the effort in a security platform actually goes. Onboarding a source is rarely just "point it at the API." Each one has its own auth model, its own quirks, its own rate limits and edge cases, and the long tail of sources is where the time disappears.
A useful first cut — cloud collection for one provider plus Linux server agents, delivered reliably into a processing pipeline — is typically a matter of a few weeks for a focused team. Broad coverage across cloud, servers, containers, SaaS, and network, with proper coverage tracking and silence detection, is a longer effort measured in months and never quite "finished," because new sources keep appearing.
The expensive mistake here is the same one that haunts the rest of the platform: architecting the first version in a way that can't absorb new source types without a rewrite. Getting the ingestion abstraction right at the start is cheap. Retrofitting it once ten source-specific collectors have hard-coded their own assumptions is not. The same design discipline that matters when securing any AI agent applies here — decide the hard structural questions early, when they're still cheap to change.
Building the Collection Layer Step by Step
Trying to collect everything on day one is how these projects stall. A phased build gets useful visibility early and keeps the architecture honest.
- Inventory your sources. List every cloud account, server group, cluster, SaaS app, and network device. Mark which ones hold sensitive data or admin access. This list becomes your coverage baseline.
- Define a common event schema. Decide early how events are normalised (timestamp, actor, action, target, source) so each new collector maps into the same shape. Open standards such as the OpenTelemetry data model are a reasonable reference even if you do not adopt them wholesale.
- Start with the control plane. Cloud audit logs and identity sign-ins usually give the highest signal per gigabyte, so they come first.
- Add host agents to critical servers. Begin with internet-facing and privileged hosts, then widen.
- Connect the SaaS audit APIs. Email, file storage, and source control are the usual priorities because they hold data and credentials.
- Turn on silence detection and coverage reporting. Every source gets a baseline volume and an owner, and missing sources show up as visible gaps.
- Add edge filtering and tiered retention. Once you know what normal volume looks like, drop proven noise and move older data to cheaper storage.
Each phase should end with a test: generate a known event (a failed SSH login, a new access key, a file shared externally) and confirm it arrives, is parsed correctly, and is searchable downstream.
Common Security Data Collection Mistakes
Collection layers usually fail through gaps that nobody notices until an incident forces the question.
Collecting logs that nobody watches
Sending CloudTrail to a storage bucket feels like coverage, but if nothing reads, parses, and alerts on those events, the evidence only helps after the damage is done. The logistics firm in the example above had exactly this setup. Collection means delivery into a pipeline that processes the data, not just retention somewhere cheap.
Forgetting SaaS and DNS
Teams naturally start with cloud and servers because they feel like infrastructure, and treat SaaS and DNS as optional extras. Yet email, file storage, and source control hold much of the sensitive data, and DNS often carries the earliest sign of malware. Leaving them out creates blind spots precisely where many attacks begin.
Hard-coding each collector
Building a bespoke collector for each source, with its own schema and assumptions, works for the first few sources and becomes a maintenance burden after that. Every new source turns into an engineering project, and normalizing events downstream gets harder. A common event schema and a shared ingestion abstraction are cheap at the start and expensive to retrofit.
Cutting sources to control cost
When per-gigabyte ingest fees climb, the tempting fix is to switch off the noisiest sources, often flow logs or DNS. That trades a budget problem for a visibility problem. Filtering proven noise at the edge and moving older data to cheaper storage usually gets costs under control without blinding the platform.
Treating silence as good news
A source that stops sending events is often read as a quiet day. Without baselines and silence alerts, a broken collector or a deliberately disabled log can go unnoticed for weeks. Every source should have an expected volume and an owner who is told when it drops.
Related guides
- Building an AI-native security operations platform: the full architecture
- The log processing pipeline: parse, normalise, enrich at scale
- Building the detection engine
- Cloud security posture management
- AI agent security: what business owners need to know
- Our AI development services
We Build the Foundation Right
Data collection is the layer that decides whether everything above it works. We treat it as core engineering, not plumbing to rush through — a blend of agent and agentless collection chosen per source, explicit coverage tracking so you know what you're not seeing, and reliable delivery with silence detection built in from the start.
If you're planning a security platform, or you suspect your current visibility has gaps you can't quite name, we're happy to map out what a complete collection layer would look like for your specific estate.
Talk to us about your platform — no commitment, just a conversation.
Frequently Asked Questions
What is the data collection layer in a security platform?
It's the module responsible for gathering security-relevant events from every source in your environment — cloud providers, servers, containers, SaaS applications, and network devices — and delivering them reliably to the rest of the platform for processing and analysis. It's the foundation everything else depends on: enrichment, detection, and the AI copilot can only reason about events that were actually collected. If a source isn't being collected, it's a blind spot no amount of downstream intelligence can compensate for.
Should I use agent-based or agentless collection?
Both, chosen per source. Agentless collection — pulling from APIs like CloudTrail or the Microsoft 365 audit log — is the right fit for cloud and SaaS, where there's no host to install anything on and nothing to break. Agent-based collection makes sense for servers and endpoints, where deep signals like process activity and file changes simply aren't exposed by any remote API. A platform that insists on only one approach is either missing depth on your servers or adding operational overhead where an API would have done the job.
Which log sources are most commonly missed?
SaaS applications and DNS logs, in our experience. Teams tend to focus on cloud and servers because those feel like "infrastructure," and forget that Microsoft 365, Google Workspace, GitHub, and Slack hold much of the sensitive data and are common initial attack vectors. DNS logs are the other blind spot — a lot of malware communicates over DNS, and a request to a malicious domain is often the earliest visible signal of a compromise, but only if you're collecting it.
How much do security logs cost to collect and store?
More than most people expect, and the pricing model matters as much as the volume. A mid-sized estate can generate hundreds of gigabytes of security logs a day, and incumbent SIEMs typically charge per gigabyte ingested — which means good visibility quietly becomes an expensive habit. A well-designed collection layer controls this by filtering genuinely useless noise at the edge before it travels, and tiering storage so recent data sits in a fast store while older retention moves to cheap object storage. The store you choose downstream has a large effect on the total bill.
How do I know if a log source has stopped sending data?
Through silence detection, which a thorough collection layer builds in deliberately. Each source has a normal baseline volume, and the platform watches for unexpected drops — if a source that usually sends thousands of events an hour suddenly goes quiet, that raises an alert. This matters for two reasons: a silent source might be a broken collector you need to fix, or it might be an attacker who disabled logging to cover their tracks, which is a classic early move in an intrusion. Silence should never pass unnoticed.
Can we start with a few sources and add more later?
Yes, and that's the sensible way to do it — provided the architecture is built to absorb new source types without a rewrite. A good first phase is often one cloud provider plus Linux server agents, delivered reliably into the processing pipeline, which already gives real visibility. The critical constraint is the ingestion abstraction: it has to treat "a new source" as a configuration exercise rather than an engineering project. Get that right at the start and expanding coverage is easy; get it wrong and every new source becomes a fight.
Does a small business need a full security data collection layer?
Not on day one. A small company with one cloud account, a few servers, and Microsoft 365 or Google Workspace gets most of the value from collecting cloud audit logs, identity sign-ins, and SaaS admin events, then adding server agents for anything internet-facing. That scope is affordable and already catches the common attack paths: stolen credentials, leaked keys, and risky sharing. The important part is choosing an architecture that can add sources later without rework, so coverage can grow as the business and its attack surface grow.
Conclusion
An AI security platform can only reason about the events it receives, which makes the collection layer the real ceiling on what it can detect. The hard part is not listing sources; it is settling the trade-offs for each one, keeping ingestion costs from pushing you toward collecting less, and knowing at all times which sources are missing or have gone silent.
The practical lessons are consistent. Blend agent and agentless collection per source instead of picking one dogmatically. Treat SaaS audit logs and DNS as first-class sources, because they are where attacks often start. Build silence detection and coverage reporting in from the first phase, since a quiet source can mean a broken collector or an attacker covering their tracks. Most of all, get the ingestion abstraction right early, because retrofitting it after a dozen custom collectors is expensive.
A good next step is to write down every source in your environment and mark which ones you are actually collecting today. The gaps on that list are your priorities. If you would like help designing a collection layer that can grow with your estate, book a call with our engineering team.
