OT incident response: a ransomware playbook for plants that can't just stop
Build an OT incident response plan that respects process safety — roles, containment without shutdowns, backups that restore, and tested playbooks.
IT incident response assumes you can isolate, power off, and rebuild. A refinery, a glass furnace, or a paper machine cannot simply be powered off — an uncontrolled shutdown endangers people, destroys equipment, and costs more than the ransom. OT incident response therefore needs its own playbook: one that puts process safety first, contains without stopping unless stopping is the safe action, and restores control systems from known-good sources. This guide outlines that playbook.
Organize before the incident
Define OT-specific roles. The IT SOC cannot decide whether to trip a unit; operations owns that call. The response team needs an OT lead (usually from controls or operations engineering), an operations liaison with authority to order controlled shutdowns, IT security for forensics and enterprise containment, vendor contacts for control-system support, and a single incident commander. Put names, backups, and out-of-hours contacts in a printed call tree — ransomware encrypts the SharePoint page with the phone numbers.
Map what matters. You cannot protect or restore what you have not inventoried. The brownfield OT asset inventory is the foundation: every controller, workstation, and network device with its firmware, configuration backup status, and process criticality. Mark the systems whose loss forces shutdown versus those that degrade gracefully — that distinction drives every containment decision.
Pre-authorize the hard decisions. Agree in advance, in writing, with plant management: under what conditions operations will island the plant (sever OT from IT/corporate networks), who can order it, and what production loss that implies. Negotiating this during an incident wastes the hours containment needs.
Contain without breaking the process
Islanding beats shutdown. The first containment move is usually to sever the conduits between OT and everything else — pull the uplinks at defined zone boundaries — while the process keeps running under local control. Engineer these disconnect points in advance: labeled, tested, and operable without the network (a firewall rule change that requires the management network you just lost is not a disconnect plan).
Never reimage a running controller to "clean" it. Forensics on control systems means capturing configurations, logs, and network traffic — not wiping the PLC that is holding the process stable. Quarantine infected Windows assets (engineering stations, HMIs, historians) by network isolation, fail over to standby systems where they exist, and keep the process on local/manual control per operating procedures.
Communicate on out-of-band channels. Assume email, Teams, and VoIP are compromised or down. Incident communications run on phones, radios, runners, and a pre-established group chat on personal devices — decided and tested before the incident.
Restore from known-good, then learn
Restoration order follows process need: safety systems verified first, then control infrastructure, then supervisory and historian layers. Every restore needs known-good sources: offline, tested backups of controller programs, HMI projects, historian configurations, and workstation images — with restoration actually rehearsed, because a backup nobody has restored is a hypothesis. Change all shared and vendor credentials after containment; attackers harvest them first.
Afterward, run the blameless post-mortem within two weeks: timeline, what detected it, what contained it, what failed, and corrective actions with owners. Feed the findings into monitoring (SIEM use cases for OT, anomaly detection on control traffic), segmentation improvements, and the next tabletop exercise. Run tabletops at least annually — a playbook the team has never walked through is documentation, not capability.
Cite this page: OT incident response: a ransomware playbook for plants that can't just stop
, Shopfloor, 2026-10-04. https://shopfloor.space/articles/ot-incident-response-playbook/