Est.
FeaturesLong read

Severity Classification Frameworks for Shared IR Teams

Clear severity definitions help small teams triage faster without waking the wrong person.

Reporter · · 11 min read
Cover illustration for “Severity Classification Frameworks for Shared IR Teams”
Features · September 11, 2026 · 11 min read · 2,457 words

Severity classification determines who gets woken up, how fast a team moves, and whether a security incident gets contained in an hour or spreads for a week. For a lean team where a founder or an IT generalist is making the triage call at 11 PM, the frameworks built for dedicated security operations centers simply don't hold up. This piece breaks down how to build a classification system that actually works when the person on call has four other jobs.

Most severity frameworks assume a dedicated analyst on shift, trained on the taxonomy, backed by a second-shift reviewer. Strip that away and hand the same framework to a founder who also handles payroll, or an IT generalist who's also patching servers, and the wheels come off fast. One team's "critical" becomes another team's "we'll look at it Monday." Nobody wants to be the person who paged the CEO over a false alarm, so real emergencies drift toward lower ratings until containment stops being an option. And when everything gets stamped urgent because nobody trusts the tiers, alert fatigue sets in, and nothing gets treated as urgent at all.

The incentive problem sits underneath all of it: lower severity ratings carry less burden. Fewer notifications, no executive call, no documented response window. That means any manual classification system, left alone, creates quiet pressure to under-report. This isn't a training gap. It's what happens by default when the tiers are ambiguous and a tired person is the one applying them.

What severity classification is actually doing, and what it isn't

Severity measures business impact. Plain and simple: how much is this incident hurting operations, data, customers, or the company's standing with regulators, right now, at this moment.

That's a different question from priority. Priority is about what gets worked first given the hands available today. A low-severity incident can be high priority because a compliance deadline is bearing down on it. A high-severity incident can drop in priority temporarily if containment is already holding the line. Mixing up the two leads to two failure patterns: teams over-mobilize on something that's already handled, or they under-mobilize on something that's actively spreading because it didn't look urgent on first glance.

Severity also isn't technical complexity, and this trips up more non-specialists than anything else. A misconfigured firewall rule is technically trivial to explain but can be operationally catastrophic if it exposes a customer database. A sophisticated intrusion attempt might set off every alarm bell in the SIEM and still be fully contained within minutes. The alert volume or the cleverness of the attacker tells a decision-maker almost nothing about how bad the day is about to get.

The anchor has to be business impact, full stop, not attacker sophistication and not how loud the tool alert is. CREST's 2025 Cybersecurity Incident Management Guide makes this point directly: classification should weigh potential business impact, regulatory exposure, and reputational risk, not just the technical vector involved. NIST SP 800-61r3, finalized in April 2025, backs the same idea from a different angle, folding incident response into broader enterprise risk management instead of treating it as a technical sideshow that lives apart from the rest of the business.

How many severity levels a lean team can realistically operate

Three to five tiers cover the vast majority of organizations. Past five, classification gets harder to apply under pressure without buying any real operational benefit. More tiers sound more precise on paper; in practice, they just give a tired responder more places to get stuck.

A three-tier model, Critical / High / Low, fits smaller or simpler environments well. It's fast to apply when someone's half-asleep and staring at an alert at 2 AM. It's easy to train a non-specialist on in twenty minutes rather than two hours. And for a team that isn't handling a high volume or wide variety of incidents, three tiers are usually all the precision that's needed.

A five-tier model earns its keep when the team handles real volume and real variety. PagerDuty's five-level model (SEV 1 through SEV 5) treats SEV-1 and SEV-2 as major incidents requiring formal, coordinated response, while SEV-4 and SEV-5 route as normal-hours queue items that don't need to wake anybody. PagerDuty's own data suggests that classifying incidents properly under this kind of model can cut resolution times by as much as 40%. The tradeoff is upfront cost: five tiers demand a lot more pre-decision work to define each one clearly enough that nobody's guessing.

For most small and mid-sized businesses without a dedicated security team, four plain-language tiers hit the sweet spot:

  • Critical, could stop operations or expose sensitive information quickly
  • High, may allow outside access to important data or systems
  • Medium, could be used as a stepping stone or cause operational confusion
  • Low, unlikely to cause serious impact alone, but worth fixing

The design principle that matters most here isn't the number of tiers. It's whether the definitions are written so a non-security employee can apply them without picking up the phone to ask what they mean.

The inputs that drive the classification decision at triage

Triage pulls from a handful of sources, and a lean team needs to know what to look at without hunting: SIEM or EDR alerts flagging affected systems, a user reporting something odd (a phishing email, an unexpected lockout, weird behavior on a shared drive), an external notification from a vendor or even law enforcement, and whatever proactive monitoring shows about what's affected and what an adversary looks to be doing right now.

From there, the person making the call needs answers to a specific set of questions. How many systems or users are touched, and is that number climbing? Is sensitive or regulated data potentially exposed? Is this actually stopping the business from operating, revenue, production, customer access? Is there any sign of active adversary movement, lateral movement between systems, data leaving the network, something planting itself for later? And what's the regulatory exposure if this turns out to be the real thing?

A single suspicious login attempt and ransomware actively spreading across file shares are not the same category of problem, even though both might trigger an alert that looks similar on a dashboard. The inputs have to separate scope and trajectory, not just flag that something happened.

Recoverability belongs in this mix too, and it's often overlooked. An incident one person can close out in an hour is a structurally different animal from one that needs a full backup restoration or a vendor engagement to resolve. That difference should shape the tier, not just the response plan.

None of this is a one-time verdict. Classification gets revisited as new information comes in. What starts as a Medium at 9 AM might be a Critical by 9:20, once it's clear the scope is growing rather than shrinking.

What each severity level must specify before an incident occurs

Every tier needs four things nailed down long before anyone needs them: who gets notified and in what order, who holds the incident commander role for that tier, what the response time target actually is (based on who's really available to staff it, not a number that sounds good in a slide deck), and what activation looks like, full team, a subset, or one person handling it solo.

A Critical tier, for instance, might mean an immediate page to the on-call technical lead and incident commander, executive notification inside a defined window, and a decision on external communication made in that same window. A High tier might mean internal stakeholders get looped in, with no public statement unless things get worse. A Low tier might just be a ticket assigned to whoever's next up during business hours, no escalation unless someone re-evaluates the tier later.

One detail gets missed constantly: the coordination channel itself. If email is how the team talks during an incident, and email is what got compromised, the whole plan collapses at the exact moment it's needed. Critical and High tiers need a named, pre-established out-of-band channel, a Signal group, a separate Teams tenant with its own MFA, something that doesn't depend on the system that might be down.

CREST's 2025 guidance recommends a RACI matrix built around incident type and severity level. For a lean team, that means naming who's Responsible and who's Accountable, by name or by role title, before anything happens, not while it's happening. A small IR team, at minimum, should have an incident commander, a technical lead, an identity lead (someone who owns Microsoft 365 and Entra ID specifically, since identity compromise sits behind so many incidents), a communications lead, and an executive sponsor.

The real test of whether any of this works: if the most junior person on the team is first on the scene at 3 AM, can they read the tier specification and know exactly what to do next, without calling anyone to ask?

Severity drift, the failure mode that quietly degrades every framework

Severity drift happens because lower tiers carry less weight, fewer pages, no executive notification, no documented obligation to respond fast. Given that setup, people rationally, if unconsciously, classify things a notch lower than they should. This isn't really about culture. It's what any manual classification system does over time if nobody builds in a counterforce.

Drift shows up in a few recognizable patterns. Escalation gets delayed because nobody wants to be the one who woke the CTO for nothing, so genuine emergencies sit unreported until containment is no longer realistic. Sometimes it runs the other way too: if Critical means "everything stops," teams get reluctant to declare it even when the situation calls for exactly that. And once severity gets applied inconsistently across incidents, MTTR data (mean time to resolution, broken down by tier) stops meaning anything, because the tiers themselves weren't consistent to begin with.

A few countermeasures actually hold up. Build a "declare early" norm into the framework itself: over-classifying early is fine and won't be penalized, but under-classifying late carries a real cost, and everyone should know that going in. Add a mandatory re-check, say thirty minutes into response, where the initial tier gets explicitly confirmed or bumped up. And when the call is genuinely a toss-up between two tiers, the rule should always be to start at the higher one.

Research on the 2025 Cost of a Data Breach Report found that organizations with tested incident response plans and AI or automation tooling saved an average of $2.66 million per breach compared to those without. Late or wrong severity classification is exactly what erodes that kind of readiness before it ever gets a chance to pay off.

Adapting the framework when responsibility is split across roles or external partners

Lean teams generally run one of three models, and classification has to bend to fit each one.

In-house only works fine when one person owns the on-call function and the escalation paths are written down and tested. It gets fragile fast when that same person is also handling infrastructure, user support, and backups, with nobody covering nights or vacation.

Co-managed setups split the load: internal staff keep the environmental knowledge and daily relationships, while an outside partner brings senior judgment, monitoring, and extra capacity once an incident outgrows what the internal team can handle alone. Classification here has to be a shared decision from the start, who owns the first call, who's allowed to bump the tier up, who has the authority to bring the outside team in.

Fully outsourced models suit organizations running with limited IT headcount or heavier compliance obligations. Even here, the severity framework needs to be agreed with the provider ahead of time, not figured out live during an incident.

The handoff is where things usually break. If internal staff and an outside provider are each working from their own severity definitions, the argument over classification happens right at the moment of escalation, burning exactly the time the whole framework exists to save. A shared severity definition document, agreed on, signed off, and actually tested in a tabletop exercise, is close to a prerequisite for co-managed or outsourced arrangements. CREST's 2025 guidance notes that the severity call should determine whether full team activation is needed and whether outside support gets engaged, and that only works cleanly if the external partner's activation criteria line up with the internal tier definitions. Get that wrong, and a team either under-mobilizes on something spreading fast, or over-mobilizes on a false alarm and pays a credibility cost with everyone who got paged for nothing.

Testing and maintaining the framework so it holds under real pressure

A framework nobody has tested is a guess dressed up as a plan. The first real incident is the worst possible moment to discover that a tier definition is vague or that the emergency contact list still has someone's old phone number on it.

Tabletop exercises are the main way to find the cracks before they matter. Run a scenario, have the team classify it, and let the disagreements surface, then fix the tier definitions based on what didn't line up, before an incident forces the issue. The scenarios worth running are the ones where the right classification isn't obvious, not another easy Critical or easy Low. After each exercise, update any tier definition that produced different answers from different people.

Blameless post-incident reviews double as an audit of the classification itself. Every review should ask directly: was the initial call right, and if not, what was missing or misread? Framing this without blame matters more than it sounds like it should, because the moment a review starts assigning fault for a bad classification, people stop surfacing their own uncertainty next time, and that's exactly what feeds drift. The useful questions are things like: why did someone reasonable wait before escalating? Why didn't the alert reach the right person fast enough? Why did the whole recovery hinge on one person remembering where the backups live?

The framework itself needs upkeep too. Contact lists need review on a set schedule, because people change roles and phone numbers, and an outdated pager contact at the Critical tier is a failure waiting for its moment. Tier definitions should get revised whenever a new kind of incident exposes a gap nobody planned for. And the out-of-band channel, the Signal group or the separate Teams tenant, only works if it's actually been tested and everyone knows how to get into it, not just referenced in a document nobody's opened since it was written.

The real measure of whether any of this holds up: a junior team member, facing a real incident for the first time, follows the tier specification and reaches the right people inside the response window, without improvising a single step.

Sources

  1. Cyber Security Incident Management Guide
  2. Incident Severity Classification: Best Practices to Speed Resolution

More in Features