Select logs by investigation value
Prioritise identity events, privileged role use, control-plane changes, key operations, public exposure, network flows, workload security events and sensitive data access. Enable detailed data events selectively where volume and privacy allow.
For each source, document the owner, latency, expected fields, retention and a test that proves collection still works.
- Authentication and authorization
- Administrative and policy changes
- Network and DNS activity
- Storage and database access
- Runtime and application security events
Protect the collection path
Send logs to an account or project where workload administrators cannot modify them. Encrypt transport and storage, restrict readers and alert on collection changes.
Plan for regional or provider outages. Buffer critical events and monitor gaps. A silent pipeline failure creates false confidence.
- Independent log archive
- Immutable retention where required
- Collection health checks
- Controlled investigation access
Enrich before alerting
A resource identifier is not enough. Add service owner, environment, data class, internet exposure and business criticality. Join identity events with role and device context.
Enrichment lets the same technical signal be prioritised differently for a test sandbox and a payment system.
- Asset and owner catalogue
- Identity and HR context with privacy controls
- Known change windows
- Threat and vulnerability context
Build a detection service, not a rule pile
Each detection needs a threat hypothesis, data dependency, severity logic, runbook, owner and review date. Measure whether it produces useful investigations, not merely how often it fires.
Feed incident findings into new or improved detection. Retire rules that cannot be operated or justified.
- Detection-as-code review
- Regular simulation and validation
- False-positive learning
- Coverage mapped to attacker behaviour
A practical 30-day field plan
Week one — Generate. Identity, control plane, network, data and workload events. Put one accountable owner in the room, agree which business service or decision is in scope, and record the assumptions the team is making. Resist the urge to begin with a technology purchase; the first deliverable is a shared view of the problem and the authority to change it.
Week two — Protect. Central collection, integrity and access controls. Walk through the current process with the people who operate it. Compare the written design with real access paths, data flows, exceptions and on-call practice. Mark every point where an owner is missing or where the team cannot produce evidence that a control works.
Week three — Enrich. Asset owner, criticality, user and threat context. Choose a narrow pilot that can be observed safely. Define the expected result, the rollback path and the person who may accept a trade-off. Capture operational friction as product feedback; controls that are difficult to use will eventually be bypassed.
Week four — Act. Detection, investigation, containment and learning. Review the pilot with engineering, operations, security and the service owner. Close urgent gaps, assign longer work to a funded backlog and set the next evidence review. The month should end with a repeatable operating rhythm, not a one-time presentation.
Evidence worth keeping
Good evidence is understandable outside the team that created it. Keep a concise record that connects the decision, owner, technical implementation and observed result. Screenshots can support evidence, but configuration, logs, test output and approved records are stronger because another person can reproduce or challenge them.
- Authentication and authorization — owner, current state, last validation and any open exception
- Independent log archive — owner, current state, last validation and any open exception
- Asset and owner catalogue — owner, current state, last validation and any open exception
- Detection-as-code review — owner, current state, last validation and any open exception
- Decision log showing who approved residual risk and when it will be reviewed
- Test or exercise result with the actual outcome, not only a pass label
Questions for the leadership review
Use these questions to keep the discussion connected to operating risk rather than tool activity. A useful answer names a person, a service and evidence.
- Who is accountable for the cloud detection outcome when teams disagree about delivery and risk?
- Which critical service or customer promise would be affected by a failure in this area?
- What evidence would tell us the design is working in production today?
- Which exception creates the largest concentration of access, dependency or recovery risk?
- What would the team contain first, and how would it restore a trustworthy service?
- Which improvement can be completed in the next 30 days without waiting for a large programme?
Common questions
How long should cloud logs be retained?
Set retention from investigation needs, legal obligations, storage cost and privacy. Different sources may need different periods.
Should every log go to the SIEM?
No. Keep a protected archive, then route the sources and fields needed for active detection and investigation.
