An incident response team is the defined set of roles—typically an incident commander, security analysts, infrastructure owners, legal counsel, and a communications lead—who execute an organization’s incident response plan. Triage is the first job the team performs.
Key takeaways
- An incident response team is a set of roles, not a headcount or a department. One person can hold several roles.
- The incident commander is the load-bearing role. It exists so decisions don’t stall in committee while an attacker moves.
- A responsible, accountable, consulted, informed (RACI) matrix is the fastest way to remove ambiguity about who decides versus who executes.
- Triage—classifying severity and scope—is the first thing the team does and the step that determines how much of the plan activates.
- Small organizations don’t hire an incident response (IR) team. They assign roles to existing staff and fill the 24×7 gap with a managed detection and response (MDR) provider.
- Team size scales with risk and complexity, not with a formula.
Every organization has an incident response team, whether or not anyone has written it down. The question is whether the roles were assigned in advance or sorted out during the incident, and that difference shows up directly in response time. A defined team means the person who can isolate a compromised host already knows they have that authority, and the person deciding whether to notify regulators isn’t also running log queries. If you’re building the underlying process first, start with incident response and the lifecycle it follows, then come back to who owns each step.
What roles make up an incident response team?
Six roles cover most incidents. Match them to people, not to job titles.
Incident commander. Declares the incident, sets and revises severity, makes final calls, and owns communication to leadership. This person coordinates—they don’t investigate. The moment the commander is deep in a packet capture, nobody is running the incident. Name at least two backups, because incidents don’t respect on-call schedules.
Technical lead. Directs investigation and containment, decides what to isolate and in what order, and coordinates the analysts and system owners doing hands-on work.
Security analysts. Investigate, scope, gather evidence, and execute containment actions. In a managed model, this is largely the MDR provider’s SOC.
Infrastructure or application owner. Executes changes on the affected systems. This role exists because the security team usually can’t—and shouldn’t—unilaterally change production.
Legal counsel. Assesses regulatory exposure and notification obligations, and manages privilege over the investigation record. Bring counsel in early, not once you’ve decided it’s reportable.
Communications lead. Handles internal updates, customer messaging, and external statements. Keeping this separate from the commander prevents the person making technical decisions from also drafting customer emails.
One more role that’s easy to skip and shouldn’t be: a scribe who maintains the timeline. The record they produce is what your after-action report, your regulatory filing, and your insurer all rely on.
How do you build a RACI matrix for incident response?
A RACI matrix resolves the specific ambiguity that costs the most time during an incident: the difference between the person who does the work and the person who owns the decision.
A simplified version:
| Activity | Incident commander | Technical lead | Infra owner | Legal | Comms |
|---|---|---|---|---|---|
|
Declare the incident |
A/R | C (Consulted) | I (Informed) | I | I |
|
Set severity |
A/R | C | I | C | I |
|
Isolate a host |
A (Accountable) | R (Responsible) | C | I | I |
|
Take a production system offline |
A/R | C | R | C | C |
|
Determine if reportable |
A | I | I | R | C |
|
Notify customers |
A | I | I | C | R |
|
Own the after-action report |
A/R | C | C | C | I |
Two rules make this useful rather than decorative. First, exactly one A per row. If two people are accountable for taking production offline, nobody is. Second, build the matrix around the decisions that will be contested at 2am, not around every task—a 40-row RACI is a document nobody opens.
Fill it in with names and mobile numbers, then walk it through a tabletop exercise. That’s where you find out that your only named infrastructure owner is on parental leave.
What is triage in incident response?
Triage is the first thing the team does after a potential incident surfaces: classify its severity, scope, and business impact to decide how much of the plan to activate.
It answers four questions in order:
- Is this real? Validate the alert or report. Most alerts aren’t incidents.
- What’s the scope? Which systems, accounts, and data are involved, and is it still spreading?
- What’s the business impact? Revenue systems, regulated data, and customer-facing services change the severity calculus.
- What tier of response does this warrant? A single phished credential and confirmed lateral movement are not the same event.
Effective triage prevents both failure modes. Under-triage means an incident runs for hours as a routine ticket. Over-triage means you wake nine people and take a system offline for something a single account reset would have fixed—and it trains people to ignore the next escalation.
The classification should follow your severity scale with concrete triggers, and it should be revisable. Scope almost always grows, and severity should grow with it. NIST’s SP 800-61r3 makes a related point worth internalizing: incident response works better treated as a continuous risk management function than as a set of tasks that begins when an alert fires. Teams that classify once and never revisit tend to keep responding to the incident they first thought they had.
Should you build an in-house team or use outside support?
The honest answer? Almost nobody builds a genuinely 24×7 in-house team, and most who try are solving the wrong problem.
Round-the-clock coverage is a staffing problem before it’s a security one. Covering every hour of every week, with enough depth to absorb vacation, sick leave, and turnover, takes a team several times larger than most organizations expect—and that’s before tooling, detection engineering, or a threat intelligence function. The roles are hard to fill and harder to retain, particularly on the overnight shifts where response speed matters most.
What splits cleanly is this:
| Keep in-house | Get from a provider |
|---|---|
|
Incident commander and decision authority |
24×7 monitoring and triage |
|
Business risk decisions |
Detection engineering and tuning |
|
Legal and regulatory determination |
Alert investigation at volume |
|
Customer and internal communication |
Initial containment execution |
|
System and application ownership |
Threat hunting |
|
Post-incident business remediation |
Cross-environment threat intelligence |
The left column can’t be outsourced, because it depends on knowing your business, your risk tolerance, and your customers. The right column is where scale genuinely helps—an MDR provider covering many environments sees attacker techniques that no single organization would, and staffs the overnight hours that break in-house rotations.
The failure mode to avoid is a fuzzy seam between the two columns. Write down which side owns each decision, or you’ll discover the answer mid-incident.
How should team size scale with organization size?
There’s no formula, and anyone offering a ratio is guessing. Size follows risk, regulatory exposure, and environmental complexity.
- Small organization, low complexity. Three to five named roles held by existing IT and leadership staff. One person may be commander and comms lead. Monitoring and response execution come from a provider.
- Mid-size, regulated or multi-cloud. Five to eight named roles with real backups, a dedicated security lead as commander, and named counsel. The provider handles the SOC function.
- Enterprise. A dedicated IR function, often 10 or more people, with separate detection engineering and threat hunting, plus a retained IR firm for major incidents.
What matters more than headcount at every size: are the roles named, do the backups exist, and has the team practiced together? A five-person team that has run three tabletop exercises will outperform a 12-person team that has never met.
Incident response team best practices
- Name backups for every role, especially incident commander. One deep is not a plan.
- Keep the incident commander out of hands-on technical work during an incident.
- Assign a scribe every time. The timeline is your defensible record.
- Bring legal in early, before the reportable determination, so privilege is intact.
- Fill the RACI with names and mobile numbers, not job titles.
- Define the seam with your provider explicitly—what they contain without asking, and what always escalates to you.
- Practice the team, not just the plan. Test the team with a tabletop exercise at least twice a year.
- Rotate who plays commander in exercises so the role doesn’t depend on one person.
Expel’s take
The teams that respond well aren’t the biggest ones; they’re the ones where nobody has to ask who decides.
Expel functions as the SOC layer of our customers’ incident response teams—24×7 monitoring, triage, investigation, and pre-authorized containment. Detections land in Expel Workbench™, where automated triage and enrichment handle the correlation work, and Ruxie™, our AI SOC manager, assembles the investigative context so a named analyst starts from a working picture instead of a bare alert. Human-led, AI-powered. An analyst makes every containment call, and a customer can see the reasoning behind it. That’s how we keep a 13-minute mean time to respond (MTTR).
What we won’t do is take the incident commander’s seat, and customers shouldn’t want us to. Deciding whether to pull a revenue system offline during a sales quarter, whether a breach is reportable, and what to tell customers requires knowing the business. The engagements that go smoothly are the ones where that boundary was written down before the first incident, usually in a tabletop exercise. The ones that don’t are where “the MDR provider handles response” was assumed to include decisions it never could.
Build a small internal team that owns decisions. Get the round-the-clock function from someone who staffs it as their whole job. If you’d rather not write scenarios from scratch to test it, our Oh Noes! tabletop kit is free and ready to run.
Frequently asked questions
What roles make up an incident response team?
A typical incident response team includes an incident commander who makes final calls, security analysts who investigate and contain, an infrastructure owner who executes technical changes, legal counsel for regulatory exposure, and a communications lead for messaging.
What is an incident commander?
The incident commander is the single decision-maker during an incident response, responsible for coordinating the team, prioritizing actions, and communicating status to leadership. The role exists so decisions aren’t delayed by committee during a time-sensitive event.
What is triage in incident response?
Triage is the first step an incident response team performs after a potential incident is reported: quickly classifying its severity, scope, and business impact to decide how much of the response plan to activate. Effective triage prevents both under- and over-reacting.
Can a small company build an incident response team without hiring new staff?
Yes. Small companies typically assign IR responsibilities to existing IT and leadership staff on a RACI matrix, then fill 24×7 monitoring and rapid-response gaps with a managed detection and response (MDR) provider, giving round-the-clock coverage without a dedicated in-house SOC.
How big should an incident response team be?
Team size scales with organization size and risk rather than a fixed formula. A small business might need three to five named roles covering command, technical response, and communications, while an enterprise may run a dedicated team of 10 or more.

