OT Backup and Restore Runbook: PLCs, HMIs, SCADA, Historians and Network Devices

On this page

When a controller fails, a server’s disk dies or ransomware reaches the plant network, recovery speed depends on one thing: usable, recent, tested backups of every OT component, plus the software and licences needed to restore them. This runbook is a template to adapt to your site’s systems and procedures.

OT Backup and Restore Cycle: Inventory, Back up, Store safely, Test restore, Review
A backup is only proven when a restore has been tested.

1. Purpose

Ensure that every OT system can be restored to a known-good state within an agreed recovery time, after hardware failure, corruption, human error or a cyber incident.

2. Scope

Asset class Examples
Controllers PLCs, safety PLCs, DCS controllers, drives with parameters, robot controllers
HMIs and SCADA Operator panels, SCADA servers and clients, alarm and report configurations
Historians and MES servers Configuration, databases, archives
Engineering workstations Engineering software, projects, licences
Network and security devices Switches, routers, firewalls, remote access gateways
Infrastructure Virtualisation hosts, domain controllers, time servers, certificate services

3. Prerequisites

  • Asset inventory with owner, criticality, firmware/software versions and backup method for every asset
  • Recovery targets per asset: how much data loss is acceptable (RPO) and how fast it must be restored (RTO)
  • Access rights and approved tools for each platform
  • Secure backup storage, including an offline or immutable copy
  • Engineering software installers and licences (correct versions) archived

4. What to back up

Asset Back up Notes
PLC / DCS controller Program/project file from the engineering tool, uploaded from the running controller, hardware configuration, retentive data and recipes where needed Compare the online program with the archived version before trusting either
Safety PLC Program, signature/checksum, safety configuration documentation Follow safety change management; record checksums
Drives and servo Parameter sets Often forgotten; needed for replacement units
HMI panels Runtime and project files Keep the project file, not only the runtime
SCADA servers Project/configuration, tag databases, alarm configuration, scripts, user roles; full system image for fast recovery Application-level export plus image
Historian Configuration, tag database, archives or database backups Plan for archive size; test partial restores
MES / databases Database backups with logs, application configuration Follow database vendor practice
Network devices Running configurations of switches, routers, firewalls Store after every change
Engineering workstations Images or documented rebuild procedures Include licences and dongles
Documentation Drawings, I/O lists, network diagrams, IP registers, certificates Needed to rebuild anything

5. Risks and precautions

  • Online uploads from controllers are normally safe, but follow vendor guidance; avoid actions that change controller mode.
  • Credentials and secrets in backups (passwords, keys, certificates) must be protected.
  • Malware: scan media and store backups where a network compromise cannot reach them (offline or immutable storage).
  • Version mismatch: a project that needs an unavailable software version is not restorable.

6. Execution (routine backup)

  1. Confirm the change log since the last backup.
  2. Back up each asset using the approved method.
  3. Record: asset ID, date, method, software version, checksum or file hash, operator.
  4. Compare controller programs with the previous version; investigate unexpected differences (they may indicate undocumented changes).
  5. Copy to primary backup storage and to the offline/immutable copy.
  6. Update the backup register.

Frequency: after every change, plus scheduled backups (for example monthly for controllers, daily for servers and databases), according to criticality.

7. Restore procedure (general)

  1. Assess: identify the failed asset and confirm the cause (hardware, corruption, cyber). In a cyber incident, follow the incident response plan before restoring anything.
  2. Select the correct backup version (last known good, before the incident).
  3. Prepare replacement hardware with compatible firmware.
  4. Restore configuration or program using the approved tool.
  5. Validate (see below) before returning to production.
  6. Document the restore and any deviations.
OT Restore Procedure: Assess, Select, Prepare, Restore, Validate & document
Test restores regularly; a backup is only proven by a restore.

8. Validation

  • Controller or system runs without faults
  • Program/checksum matches the intended version
  • I/O, communication and interfaces working (SCADA, historian, MES)
  • Safety functions tested according to safety procedures where affected
  • Operators confirm displays, alarms and functions

9. Restore testing

Untested backups are assumptions. At least annually, and after major changes:

  • Restore a controller program to a spare or simulator and compare
  • Restore a SCADA or historian server to an isolated test environment
  • Restore network device configurations to spare devices
  • Time the restore and compare with the recovery target

10. Rollback and escalation

  • If a restore fails validation, return to the previous state (if safe) and escalate to the system owner and vendor support.
  • Keep vendor support contacts and contract numbers in the runbook.

11. Documentation and sign-off

Item Responsible
Backup register up to date System owner
Restore test records OT engineering
Exceptions and remediation OT lead
Approval of runbook Operations and OT management

Before you apply this in a plant: this article is for education. Always check the current edition of the relevant standards, the manufacturer's documentation for your exact product and version, and your site's procedures. Safety-related work needs qualified personnel. See our editorial policy.

Written by Bhargava Reddy Kapireddy

Bhargava has 16 years of hands-on experience with MES, SCADA, DCS, PLC and industrial data systems across power generation, oil and gas, pharmaceuticals and process manufacturing. He founded MFG Tech Hub to share practical, vendor-neutral automation knowledge.

More about the author → How we write and review articles