When it happens: incident response
What to do when something happens in the network that should not have: the response phases and preparation (log server, contacts, network map, configuration backups), roles and escalation, where the first signal comes from and how to triage it, containment by quarantine and its price, investigation from a timeline and flows including the real blast radius, remediating the cause instead of the symptom, safe recovery with the N-1 test, and a blameless post-mortem.
Lesson 1: Before it happens
The phases of incident response
Before you start: this course builds on basics from the beginner courses (addresses and mask, gateway, VLANs, ports, DNS, firewall) – if any feels unfamiliar, check those first. It also follows on from the courses Monitoring, logging and detection, Zero Trust and microsegmentation and When the network is down. An incident is a situation where something happened in the network that should not have: a device behaves differently than it should, someone reached a place they should not have, data left the building. Incident response is not one person’s heroic feat at midnight – it is a procedure that fits on one page and can be walked through. That is why it is split into phases: preparation (before it happens), detection (you find out something is going on), containment (you stop the spread), remediation (you remove the cause), recovery (you bring service back) and lessons learned (you change the network so it does not repeat). The phases are not bureaucracy. They exist because in a crisis people tend to do things in the wrong order: delete, reboot and ‘tidy up’ before they know what actually happened. Keep the order and each step has a purpose and does not destroy the ground for the next one. And one more thing people underestimate: the phases overlap and loop back. During the investigation you often find it is still spreading and you return to containment. That is normal – what matters is knowing which phase you are in and what its goal is.
Step by step
- Preparation is the phase that happens before anything occurs: log collection, a network map, contacts. Without it the other phases do not work.
- Detection: something is happening and you learn of it – from a sensor alert, a log, or a user. Only here does the incident begin to exist.
- Containment stops the spread. Only then come remediation (removing the cause) and recovery (return to service).
- The last phase is lessons learned: what to change so the same does not repeat. Phases readily loop back – that is normal.
What to prepare in advance
The quality of an incident response is decided not when the incident happens, but months earlier. Four things save you the most time in a crisis and it is wise to have them ready. First, a log server: one place where devices send their events. Without it you have no timeline and the investigation shrinks to guesswork – and what was never written down cannot be added afterwards. Second, contacts: who is called at night, who talks to management, who to the connectivity provider, who to the application vendor. Hunting for a phone number while production stands still is a pure waste. Third, a network map: what is where, what is critical and what talks to what. In the simulator its image is the topology check and the connections overview – knowing them from calm times lets you recognize a deviation in a crisis. And fourth, configuration backups of the devices: firewall rules, VLANs, routing. When it turns out someone changed a setting, you want something to compare against and something to restore. The common denominator: all four can be arranged calmly, cheaply and without adrenaline – and not one of them can be added during a crisis. The readiness test is simple: imagine your main server has been compromised since yesterday and ask what you would start with. If the answer is ‘I do not know where to look’, you know what to add.
Step by step
- A log server comes first. Devices send it events, so you have something to build an incident timeline from.
- A network map: what is where, what is critical, what normally talks to what. The topology check and connections overview keep it current.
- Configuration backups answer the question ‘did someone change this?’ and, above all, give you something to restore.
- Contacts, a map, logs, backups – all are arranged calmly. In a crisis you can add none of them.
Who does what
In a small company people often say ‘IT will handle the incident’. But even one person needs to know which role they are currently in, because the roles interfere with each other. There are three and they can be described very briefly. The one who decides says what may be done: whether a server is disconnected, whether production is halted, whether the police are called. That is a decision about business impact, not a technical one – so it should not be made by whoever is elbow-deep in the log. The one who investigates collects evidence and assembles the timeline; they need calm and above all must not be cleaning up at the same time. The one who communicates keeps management, users and possibly customers informed – and shields the investigator from questions like ‘will it be long?’. When you are alone on an incident, switch roles deliberately and in turns, not all at once. Part of roles is escalation: a threshold agreed in advance, beyond which you call further. For example: the incident lasts over an hour, it touches personal data, it reaches more than one segment, or you are unsure of the scope. Escalation is not an admission of incompetence but an ordinary working step – and a threshold set in advance spares you awkward deliberation in the middle of the night. Remember a simple rule: whoever investigates does not decide; whoever decides does not speak for everyone; and one person always speaks outward.
Step by step
- Who decides: may the server be disconnected? Is service halted? That is a decision about business impact, not a technical detail.
- Who investigates collects evidence and assembles the timeline. They need calm – and must not clean up at the same time.
- Who communicates keeps management and users informed and shields the investigator from ‘will it be long?’.
- Escalation is a threshold set in advance: over an hour, personal data, several segments, unknown scope. Then you call further.
Lesson 2: Detection: the first signal
Where the first report comes from
Incidents rarely begin the way people imagine – with a big red sign on a screen. They begin in four far more ordinary ways and it is worth knowing all of them, because each has a different speed and a different reliability. First, a user: ‘something is flashing here’, ‘a strange email arrived’, ‘the computer is slow’. It is the most common source and also the least precise – but do not underestimate it, because users often notice things no tool sees. Second, an IDS/IPS sensor: it recognizes a known attack pattern and raises an alert. It is fast and specific, but sees only what passes under its nose and what its signatures cover. Third, monitoring and logs: a server that stopped answering, a sharp rise in traffic, repeated failed logins. That signal is slower but shows context. And fourth, a decoy (honeypot): a device nobody has any legitimate reason to talk to. When a connection appears at it, it is almost never a false alarm – which makes a decoy the cleanest signal you can have in a network. The practical conclusion: do not rely on a single source. A sensor is silent about what its signatures miss, a user will not notice quiet data exfiltration, and monitoring shows only the consequence. Only together do they form a net an incident gets caught in – and the sooner it is caught, the smaller the damage.
Step by step
- Most often a user speaks up first: ‘a strange email arrived’, ‘the computer is slow’. Imprecise, yet often ahead of the tools.
- An IDS/IPS sensor recognizes a known attack pattern and raises an alert. Fast and specific – but limited to its signatures.
- Monitoring and logs show a server that stopped answering, or repeated failed logins. Slower, but with context.
- A decoy has no legitimate reason for traffic, so a connection there is the cleanest signal. Relying on one source does not pay.
An alert vs. an incident
Between ‘an alert arrived’ and ‘we have an incident’ there is a stretch of road called triage. It is a quick sorting: what is a false alarm, what is a trifle you handle in passing, and what is a real incident that starts the whole procedure. Without triage one of two bad things happens. Either every alert becomes an alarm, the team burns out and within a month nobody looks at alerts – that is alert fatigue, familiar from the Monitoring, logging and detection course. Or, the other way round, nothing becomes an alarm and a real breach sinks among hundreds of lines. Triage rests on three questions you may well write on paper. Is it real? Is there confirmation from a second source, or is it a single isolated line? How far does it reach? One device, one segment, or is it already showing up in several places? What is at stake? A test computer, or a database of personal data? Only the answers decide whether the procedure starts and, above all, how fast. A useful shortcut: severity is not the same as loudness. An alert labeled critical may be noise, and a line at the ‘informational’ level may be the first step of a breach. So in triage you look at what happened and what it concerns, not at the color the tool used in its message.
Step by step
- An alert arrived. That alone is no incident – between a report and an incident lies triage, a quick sorting.
- First question: is it real? Does a second source confirm it, or is it one isolated line with no backing?
- Second and third questions: how far does it reach (one device or several places) and what is at stake (a test PC or a database).
- Severity is not loudness: a critical alert is often noise, and an ‘informational’ line may be a breach’s first step.
Verify before you raise the alarm
Before you declare an incident, verify that what you think is happening really is. It sounds like a needless delay, yet it is the cheapest step of the whole procedure: a false alarm costs the whole company time and, worse, blunts the willingness to react next time. Verification has three simple steps and you can do all of them in the simulator. First, a second source: does a log on another device confirm the alert? The sensor reports a suspicious connection – do you see it in the connections overview too, in who talks to whom? If so, it is real traffic, not a lone pattern match. Second, time: when did it first appear and is it still going on? A one-off event from last week is something entirely different from a connection arising every minute. Third, context: is there a plain explanation? A new application, a planned update, a new colleague, a test run. A great many ‘attacks’ dissolve the moment you learn a new backup job started on Friday. But beware the opposite mistake, which is more insidious: verification must not become a reason to do nothing. When you have a serious suspicion plus confirmation from a second source, do not postpone containment because you want absolute certainty. A sensible rule: verify for minutes, not hours. A short check protects you from a false alarm; a long analysis only gives the attacker time.
Step by step
- The sensor reports a suspicious connection. So far it is a hypothesis, not a finding – a false alarm costs the whole company time.
- A second source: do the log and the connections overview confirm the same? If so it is real traffic, not a lone pattern match.
- Time and context: is it still going on, or was it once? And is there a plain explanation – a new backup, an update, a test?
- But verification must not become a reason to do nothing. A short check guards against a false alarm; a long one gives the attacker time.
Lesson 3: Containment: stop the spread
Isolating a device and its price
Containment has a single goal: stop the spread. It is not a fix nor a cleanup, it is a stopper. The most direct tool the network gives you is quarantine – in the simulator you find it under the right mouse button as Isolate. The device stays physically in place, but the network cuts it off: it does not talk to neighbors, does not talk outward, and nobody talks to it. From the incident’s standpoint it is a very strong step, because it immediately shrinks the blast radius – how far one can get from the compromised device (a term you know from the Switching and VLANs course). But it has a price and it is fair to name it right away. First, the user stops working: isolate the accountant’s computer and the accountant does nothing. Isolate a server and everything depending on it stands still. Second, isolation is visible: an attacker sitting on the device knows at once that you found them and may react – erasing traces or jumping elsewhere if they are already elsewhere. And third, most often overlooked: isolation is not remediation. A cut-off device still contains whatever compromised it, and the moment you return it to the network without a fix, it starts again. Quarantine is therefore an excellent first step, not a last one. It stops the clock and gives you room to find out what happened – and that is exactly what containment is about: buying time without destroying the ground for the investigation.
Step by step
- The compromised PC gets further from here – to the server and outward. How far it reaches is its blast radius.
- Quarantine: right button → Isolate. The device stays in place, but the network cuts it off from its surroundings and the internet.
- The rest of the network runs on and the blast radius shrank. But the price is real: an isolated user or server does not work.
- Careful: isolation is not remediation. The device still contains what compromised it – unfixed, it restarts on return to the network.
When to isolate and when to watch
Quarantine is a strong tool, but not always the right one. There is a second option that seems odd at first: let it run and watch. Choosing between ‘isolate now’ and ‘watch a while longer’ is one of the few genuinely hard decisions of an incident, so it is worth knowing what guides it. Isolate at once when it may spread within minutes (encrypting malware working across shared drives), when sensitive data is at stake, when you already see active exfiltration, or when you simply lack the capacity to watch the situation and understand it at the same time. Fast stopping is almost always better than slow deciding. Watch a while when observation truly brings what you cannot get otherwise: everywhere the attacker reached, who their second target is, how they got in. That makes sense with a slow, quiet breach where one isolated device is not the whole story anyway – and only when you have the means to watch and someone really attends to it. Watching without monitoring is not a strategy, only inaction with a nicer name. The difference between the two choices fits in one sentence: isolation saves your network, observation saves your knowledge. When you have no choice, take saving the network. And if you decide to watch, give it a firm time limit and say in advance what must appear for you to isolate immediately.
Step by step
- When it spreads within minutes or sensitive data is at stake, isolate at once. Fast stopping beats slow deciding.
- Do you see active exfiltration? Then you do not wait for perfect understanding – the data is leaving right now.
- With a quiet breach watching can make sense: it shows everywhere the attacker reached. But only with monitoring and someone watching.
- Isolation saves the network, watching saves knowledge. With no choice, take the network – and give watching a firm time limit.
What isolation does not solve
The most common mistake after successful containment is relief. The device is isolated, the alerts went quiet, so it is done – and this is exactly where incidents come back to life. Isolation solves the spread, but leaves three things untouched. First, the cause: you still do not know how it got there. If it was an email attachment, nothing stops someone opening the same attachment on another computer tomorrow. If it was a vulnerable service exposed outward, it is still sitting out there. Second, the other devices: an attacker rarely stops at the first machine. Before you isolated it, they may long have been elsewhere – so containment is followed by a check whether the same pattern shows up elsewhere, typically in the connections overview and in the log. And third, persistent access: an added account, a changed firewall rule, a redirected DNS record – changes the attacker made while they still controlled the device, before the isolation. Quarantining one computer does not remove them and they keep working. The practical consequence: after isolation do not take a break but go straight on to the investigation – precisely because quarantine stopped the clock and you now have room. And one more unpleasant lesson better read here than lived: a compromised device does not return to the network merely because it went quiet. It returns after remediation – cleaned or reinstalled, with changed passwords and a check that it will not start again. That is what the Remediation and recovery lesson is about.
Step by step
- The device is isolated and the alerts went quiet. Here people feel relief – and exactly here incidents come back to life.
- The cause stayed untouched: the exposed vulnerable service is still out there, the same attachment can be opened elsewhere tomorrow.
- An attacker rarely stops at the first machine. Check whether the same pattern shows up elsewhere – in the connections overview and the log.
- And persistent access remains: an added account, a changed rule, a redirected record. After isolation go straight on to the investigation.
Lesson 4: Investigation
A timeline from the logs
In practice, investigating an incident looks far less dramatic than expected: you assemble a timeline. You take individual records from syslog, from the sensor and from alerts, order them by time and look for where the story begins. The timeline has three points you always want to find. The first occurrence: the oldest record belonging to the incident. It is almost always earlier than you thought – the alert came on Tuesday, but the first odd connection is from Friday. The turning point: the moment suspicious behavior became something harmful (data sent, a setting changed, another device touched). And the last activity: the newest record, from which you tell whether the incident is still live. One rule helps while assembling: stick to time, not to impressions. When every line has an exact time, what anyone remembers stops mattering. That is exactly why the Monitoring, logging and detection course insists so much on one log server – without a shared timeline, times from different devices align badly. Expect holes in the timeline too. A device that sent no logs, moments when syslog was flooded, events lost on the way. Note the hole and do not pretend nothing was there – unrecorded does not mean it did not happen. The timeline is the most valuable output of the whole investigation: from it follow the scope, the remediation and what later goes into the post-mortem.
Step by step
- Look for the first occurrence: the oldest record belonging to the incident. It tends to be earlier than you thought – the alert is only a consequence.
- The turning point is when suspicious behavior became harmful: data sent, a setting changed, another device touched.
- The last activity tells you whether the incident is still live. Stick to time, not impressions – the time on the line does not lie.
- There will be holes – devices without logs, lost events. Note them: unrecorded does not mean it did not happen.
How it spread
The timeline says when. The other half of the investigation answers where through – and for that you look at traffic as flows: who talked to whom, for how long and how much moved. In the simulator you find this in the connections overview, which aggregates individual connections, so instead of a thousand lines you see pairs of ‘this one talks to that one’. For an investigation it is an ideal tool, because a breach almost always shows up as a new pair that was not there before: a work computer talking to a database it never touched; a DMZ server calling inward into the LAN; one station talking out to an address nobody else visits. Add your picture of what is normal and you have the answer to how far it got. The second thing flows give you is the blast radius: the real one, not the theoretical one. The theoretical radius you could compute in advance – ‘from the DMZ one can reach the database’. The real one shows in the flows: did the attacker actually look at the database, or only try and hit a rule? The difference between ‘could have’ and ‘actually was’ is decisive for the scope, because it decides which systems remediation covers and what must be reported onward. The practical procedure is simple: start at the compromised device and walk the links outward – who it talked to, and who they talked to. When the list stops growing, you know the incident boundary.
Step by step
- The connections overview aggregates traffic into pairs of ‘who talks to whom’. This is a normal flow that always belongs.
- A breach shows up as a new pair: a computer talking to a server it never touched. That is the lead you follow.
- A second typical lead: one station talks out to an address nobody else visits. Is something being sent right now?
- Flows give you the real blast radius: where the attacker actually got, not where they theoretically could. That sets the remediation scope.
What to write down
During an incident you work fast and only a fraction stays in your head. So you write things down from the start – and surprisingly it is not extra paperwork but the cheapest way to save yourself work later. Two different things get written. First, findings: what you saw, where and at what time. ‘14:02 – a new connection from PC to the server in the connections overview, not there before.’ Second, your own steps: what you did and when. ‘14:09 – PC isolated.’ The second part is underrated and later missed the most, because without it, two days on, you cannot tell what the attacker caused and what you did yourself while cleaning up. Around evidence there is a notion called chain of custody – a chain in which every record makes clear who took it, when and what happened to it since. The exact rules are a matter for lawyers and forensic specialists; you only need to grasp the principle and act on it: records are copied, not edited, stored away from the compromised device, and it is noted who took them. If the incident later becomes a legal matter or an insurance claim, this decides whether your material is worth anything. And one thoroughly practical piece of advice: do not start with cleanup. A reboot, a reinstall or deleting a suspicious file are irreversible steps that often destroy exactly what you would need an hour later. Record first, clean up after. The simulator prepares the base for you: the Syslog panel has a build report button – it generates the event timeline, isolated devices and observed connections, and you add your own steps.
Step by step
- Write down findings: what you saw, where and when. ‘14:02 – a new connection from PC to the server, not there before.’
- And above all your own steps: ‘14:09 – PC isolated.’ Without it, two days on, you cannot tell the attacker from your own cleanup.
- Evidence is copied, not edited, and stored away from the compromised device. It is noted who took it and when.
- Do not start with cleanup. Reboots, reinstalls and deletes are irreversible and often destroy just what you will need in an hour.
Lesson 5: Remediation and recovery
Remove the cause, not the symptom
Remediation is the phase where you remove what made the incident possible. It sounds obvious, yet in practice the symptom is commonly removed instead of the cause – because the symptom is visible and the cause is not. The typical examples are easy to spot: the server hung, so it was rebooted; the odd connection was blocked by a firewall rule; the compromised computer was reinstalled. In all three cases what bothered you stopped being visible. But the question ‘how did it get there?’ stayed unanswered, and while it does, the incident returns – possibly on another device and with another symptom. A simple question helps tell them apart; ask it at every remediation step: if the attacker tried the same again, would they get through? If yes, you were fixing a symptom. Causes in a network tend to be boring and repetitive: a service exposed that should not have been; a missing update; a shared or default password; a rule too broad, permitting more than it should; missing segmentation that let one machine reach everywhere. The remediation is then equally boring: close, patch, narrow, separate. Remediation also includes revoking persistent access mentioned under containment: added accounts, changed rules, redirected records. And beware one trap: do not fix blindly before you know what happened. Turning off half the network ‘just in case’ is an intervention too – irreversible for the investigation and painful for operations.
Step by step
- The server hung, so it was rebooted; the odd connection was blocked. It is out of sight – but that is a symptom, not the cause.
- The question ‘how did it get there?’ stayed. The exposed service, the missing update or the broad rule are still there.
- Remediation is as boring as the cause: close, patch, narrow, separate – and revoke the persistent access the attacker set up.
- At each step ask: if the attacker tried the same again, would they get through? If yes, you were fixing a symptom.
Back into service safely
Recovery is the return to normal – and it is the phase where the incident’s last big mistake is made: everything comes back at once and unchecked. A sensible procedure has four steps. First, bring things back in parts, in order of importance. Turn everything on at once and if something breaks you do not know what – bring them back one by one and you know immediately. Second, watch while doing it: after each return check the logs, the alerts and the connections overview. This is exactly where you learn whether the remediation was complete: if the old ‘who talks to whom’ pair reappears, it was not done. Third, verify the network is resilient again. During the incident you may have unplugged, blocked or rerouted something – and it easily happens that the network runs, but on a single path only. For this the simulator has the N-1 test: it removes elements one by one and shows what stops working after each failure. If after an incident the N-1 test shows new single points of failure, recovery is not over. The topology check serves the same purpose, finding forgotten unplugged cables or orphaned devices. And fourth, check that it really works for the users, not just for you on a diagram. Call the person who reported the incident and let them try what did not work. The sentence ‘it works on my side’ is the weakest possible proof of recovery.
Step by step
- Bring things back in parts, by importance. Turn everything on at once and if something breaks you do not know what – one by one you see it at once.
- After each return check the logs and the connections overview. If the old suspicious pair reappears, the remediation was incomplete.
- Run the N-1 test and the topology check: the network runs, but is a single path or an unplugged cable left behind?
- Finally let the user who reported it try. ‘It works on my side’ is the weakest possible proof of recovery.
When an incident is closed
Incidents have a peculiar property: they often never formally end. Service runs, nobody complains, so people simply stop thinking about it – and a month later nobody remembers what was actually solved and what was not. So closing an incident is said out loud and has its conditions. There are four and none can be skipped. First, the cause is removed, not merely covered – you can answer how it got in, and that way is closed. Second, the scope is known: you know which devices and data were concerned, and where you are unsure you wrote it down. Third, service runs normally and someone other than you verified it. And fourth, there is something to hand on: a record exists with the timeline, a list of your steps and what remains to be done. If any condition fails, the incident is not closed – it is merely quiet, which is not the same. A special category is an incident where not everything could be established. That is common and no disgrace; what matters is closing it honestly: write what you know, what you do not and what risk follows. Then it is decided whether the investigation continues or the residual risk is accepted and monitoring strengthened. Closing an incident has one more function, very important for a team: it is the moment when firefighting stops and thinking starts – the entry into the last phase, lessons learned.
Step by step
- First condition: the cause is removed, not merely covered. You can say how it got in, and that way is closed.
- Second: the scope is known – you know which devices and data were concerned. Where you are unsure, you wrote it down.
- Third: service runs normally and someone other than you verified it. Fourth: a record exists that can be handed on.
- When not everything could be established, close it honestly: what you know, what you do not and what residual risk follows.
Lesson 6: Lessons learned
A blameless post-mortem
The last phase is the one most often skipped, and yet the only one that improves the next incident. A post-mortem is a short meeting after closure, going through what happened and what follows. The key word is blameless: you do not look for who caused it, but for what made it possible. This is not kindness to people, it is a practical measure. When blame is sought, people stop reporting mistakes and next time you learn of an incident much later, that is, more expensively. When causes are sought, people speak openly and you find things nobody would otherwise tell you. A post-mortem has a short outline that fits half a page. What happened – the timeline in a few points. How we found out – and above all, whether it could have been found earlier; this is the most valuable question of the meeting. What we did – the steps, including those that did not work. What we will change – concrete measures, each with one name and one date. Without the last point a post-mortem is just talk. Two things have no place in one: general conclusions like ‘we must be more careful’ (that is not a measure, it is a wish) and a list of everything that could theoretically be improved. Three changes that actually happen beat thirty points nobody ever opens. The When the network is down course adds one thing: even an incident that ended well deserves half an hour of review.
Step by step
- What happened: the timeline in a few points, assembled from logs. Not a novel – points showing the sequence suffice.
- How we found out – and above all, could it have been found earlier? That is the meeting’s most valuable question.
- What we did, including steps that did not work. Those are the hardest to learn and nowhere else are they discussed.
- What we will change: concrete measures, each with one name and one date. Without that a post-mortem is just talk.
What to change in the network
Measures from a post-mortem almost always fall into three groups – worth knowing in advance, because they hint where to look. The first is segmentation. The question: if the same happened again, would the attacker get as far? If yes, a boundary is missing. Splitting into VLANs, separating servers from user computers, a special segment for things nobody normally works with – all of that shrinks the blast radius and is the most effective change you can make after an incident. The second is rules. Walk those the incident exploited and look for permits that are too broad: ‘from anywhere’, ‘to any port’, rules without a description, rules nobody knows the reason for. Ask each the question you know from the Network security course: who exactly needs this and why. The third is monitoring. Here you ask differently: could it have been found earlier? The answer is usually to add logging where it was missing, to send to the log server from a device that was forgotten, or to build a decoy – a device without legitimate traffic, where every connection is a real signal. Whatever you pick, hold to one rule: few changes, but finished. After an incident there is an urge to rebuild the whole network, but half-done segmentation is worse than none. Pick three things, write who and by when for each – and next time you start from a better position.
Step by step
- Question number one: would the attacker get as far a second time? If yes, a boundary is missing in the network.
- Segmentation: VLANs, separating servers from users, a special segment for what nobody normally works with.
- Rules: walk the broad ones – ‘from anywhere’, ‘any port’, no description. Ask each who exactly needs it.
- Monitoring: add logging where it was missing, or build a decoy. And hold the rule: few changes, but finished.
A readiness checklist
To close the course, let us sum the whole response into a list you can walk in a few minutes and that tells you where you stand today, not in three months. Do you collect logs in one place? Verify the firewall, the sensor and the servers send there too – and that it really arrives; syslog rides on UDP 514 and gets lost quietly. Do you recognize what is normal? Open the connections overview and walk the ‘who talks to whom’ pairs. If you see them for the first time during an incident, you have nothing to compare against. Do you know who decides and whom you call? Contacts and escalation thresholds belong on paper, not in one person’s head. Can you isolate a device? Try quarantine as a rehearsal and notice everything it stops – so you do not learn that first in a crisis. Do you have configuration backups and know how they are restored? A backup you never tried to restore is only a hope. Does the network survive one element failing? Run the N-1 test and the topology check; single points of failure are found in calm times. And the last, uncomfortable question: when did you last try this? A procedure nobody ever walked is not a procedure but a document. A half-hour rehearsal – ‘the server is compromised, now what?’ – finds a missing contact and missing logs more reliably than any audit. Incident response is not trained during an incident.
Step by step
- Do you collect logs in one place? And do they arrive from the firewall and the sensor too? Syslog rides on UDP 514 and gets lost quietly.
- Do you know what is normal? Walk the pairs in the connections overview. Seeing them first during an incident leaves nothing to compare.
- Can you isolate a device? Try quarantine as a rehearsal and notice everything it stops – not first in a crisis.
- And the last question: when did you last try this? Incident response is not trained during an incident, but calmly.