Assign core roles
The incident commander owns coordination and decision tracking. A technical lead directs investigation and containment. A communications lead manages internal and external updates. A scribe maintains the timeline, actions and evidence.
Add legal, privacy, HR, supplier and business-service representatives when the facts require them. Keep the core group small enough to decide.
- Incident commander
- Technical investigation lead
- Communications lead
- Scribe and evidence coordinator
- Business service owner
Create a reliable operating rhythm
Open a controlled working channel, record known facts and assumptions, set the next objective and schedule updates. Use short situation reports: impact, scope, actions, decisions needed and next review time.
Separate the technical room from the executive briefing if necessary. Investigators should not repeatedly stop to recreate status.
- Clear severity and declaration criteria
- Decision log
- Action owner and deadline
- Regular situation report
Control containment and recovery
Containment can interrupt revenue, safety or evidence, so pre-authorise common actions and identify who decides unusual ones. Record the reason and expected effect.
Recovery requires acceptance criteria from the service owner, security validation and close monitoring. Do not declare victory when systems merely restart.
- Identity and endpoint isolation
- Network or service restrictions
- Known-good recovery path
- Heightened monitoring after restore
Close with accountable learning
Hold a blameless review focused on conditions, decisions and system changes. Distinguish immediate fixes from longer architectural work.
Assign actions to normal delivery plans with executive visibility for material risk. Test that important fixes work.
- Timeline and contributing conditions
- Detection and response gaps
- Control and architecture improvements
- Owner, due date and verification
A practical 30-day field plan
Week one — Declare. Set severity, commander and working channels. Put one accountable owner in the room, agree which business service or decision is in scope, and record the assumptions the team is making. Resist the urge to begin with a technology purchase; the first deliverable is a shared view of the problem and the authority to change it.
Week two — Stabilise. Scope, contain and protect critical operations. Walk through the current process with the people who operate it. Compare the written design with real access paths, data flows, exceptions and on-call practice. Mark every point where an owner is missing or where the team cannot produce evidence that a control works.
Week three — Recover. Restore safely and communicate status. Choose a narrow pilot that can be observed safely. Define the expected result, the rollback path and the person who may accept a trade-off. Capture operational friction as product feedback; controls that are difficult to use will eventually be bypassed.
Week four — Learn. Review causes, decisions and corrective work. Review the pilot with engineering, operations, security and the service owner. Close urgent gaps, assign longer work to a funded backlog and set the next evidence review. The month should end with a repeatable operating rhythm, not a one-time presentation.
Evidence worth keeping
Good evidence is understandable outside the team that created it. Keep a concise record that connects the decision, owner, technical implementation and observed result. Screenshots can support evidence, but configuration, logs, test output and approved records are stronger because another person can reproduce or challenge them.
- Incident commander — owner, current state, last validation and any open exception
- Clear severity and declaration criteria — owner, current state, last validation and any open exception
- Identity and endpoint isolation — owner, current state, last validation and any open exception
- Timeline and contributing conditions — owner, current state, last validation and any open exception
- Decision log showing who approved residual risk and when it will be reviewed
- Test or exercise result with the actual outcome, not only a pass label
Questions for the leadership review
Use these questions to keep the discussion connected to operating risk rather than tool activity. A useful answer names a person, a service and evidence.
- Who is accountable for the incident response outcome when teams disagree about delivery and risk?
- Which critical service or customer promise would be affected by a failure in this area?
- What evidence would tell us the design is working in production today?
- Which exception creates the largest concentration of access, dependency or recovery risk?
- What would the team contain first, and how would it restore a trustworthy service?
- Which improvement can be completed in the next 30 days without waiting for a large programme?
Common questions
Who should be incident commander?
A trained coordinator with authority and calm judgement; the role can rotate and should be practised.
When should executives join?
When business risk or decisions require them. Give concise briefings without pulling them into every technical task.
