Bad Tag Quality and Communication Failures: Troubleshooting PLC to SCADA, Historian and MES Data

On this page

“Bad quality”, “comm failure”, question marks on the HMI, flat lines in the historian, MES counts that stop: these are daily problems in integrated plants. Because data passes through several layers (device → PLC → driver or OPC server → SCADA → historian → MES), the fault can be anywhere. This guide gives a layer-by-layer method to find it quickly.

Tracing Bad Tag Quality Along the Data Path: Field device, PLC / DCS, OPC server / driver, SCADA, Historian, MES / reports
Check value, quality and timestamp hop by hop; the first hop that differs holds the fault.

Understand the data path first

Draw the path for the affected tag, for example:

Transmitter → PLC analog input → PLC tag → OPC UA server / driver → SCADA server → historian collector → historian → MES / reports

Then ask at each hop: Is the value correct and is its quality good here? The first hop where the answer changes is where the fault is.

What quality codes tell you

Quality / status Typical meaning First place to look
Good Value received normally If the value is still wrong, check scaling, addressing and the source
Bad – not connected / comm failure The driver or server cannot reach the source Network, device state, driver configuration
Bad – configuration error / node unknown The address or NodeId does not exist Tag addressing, PLC program changes, namespace changes
Bad – device failure / sensor failure Source reports a fault (for example a transmitter outside 4–20 mA) Field device and wiring
Uncertain / last known value Value is old, substituted or outside limits Communication interruptions, stale data handling
Stale (no updates) Value hasn’t changed for longer than expected Subscriptions, deadbands, a frozen source

OPC UA uses detailed status codes (for example BadNotConnected, BadNodeIdUnknown, BadCommunicationError). See OPC UA Explained.

What Tag Quality Codes Mean: Good, Bad: not connected, Bad: config error, Bad: device failure, Uncertain, Stale
The quality code points to the layer where the problem is.

Symptom 1: One tag is bad, others from the same device are good

Likely causes: wrong address or NodeId, tag deleted or renamed in the PLC program, data type mismatch, array index out of range, field device fault on that channel.

Diagnostic steps:

  1. Check the value and quality of the tag in the PLC (online with the engineering tool).
  2. If it is bad in the PLC: go to the field (channel diagnostics, wiring, transmitter). See Instrument Troubleshooting.
  3. If it is good in the PLC: check the tag address in the OPC server/driver and in SCADA; browse the server again after program downloads.

Resolution: correct the address or data type; restore the tag in the PLC or update the mapping.

Prevention: symbolic addressing where available; change management that includes SCADA/historian tag impact when PLC programs change.

Symptom 2: All tags from one device are bad

Likely causes: device offline or in STOP, network path broken, IP changed, connection limit reached, driver stopped, certificate problem (secure protocols), firewall change.

Diagnostic steps:

  1. Is the device powered and running? Check its status LEDs and diagnostics.
  2. Ping the device from the SCADA/OPC server and test the application port; check the switch port. See OT Network Troubleshooting.
  3. Check the driver or OPC server status and its log (timeouts, refused connections, too many connections).
  4. For OPC UA: check certificates and security settings. See OPC UA Certificate Errors.
  5. Check recent changes: firewall rules, IP addresses, device replacements, firmware updates.

Resolution: restore the network path, device or configuration; restart the driver only after understanding why it failed.

Prevention: monitor device connection status as alarms; document IP addresses and firewall flows; limit concurrent connections per device.

Symptom 3: Many devices go bad at the same time

Likely causes: switch or network failure, SCADA/OPC server overload or crash, time-out settings too tight under load, broadcast/multicast storm, server patching or reboot, virtual machine host problem.

Diagnostic steps: check network switches and redundancy status, server CPU/memory, service status and event logs, and whether the failures coincide with scheduled tasks (backups, antivirus scans, patching).

Prevention: capacity monitoring, redundant servers and networks, scheduled tasks outside critical periods, tested failover. See Industrial Network Design for OT Engineers.

Symptom 4: Quality is good but the value is wrong

Likely causes: scaling or engineering unit mismatch, off-by-one Modbus addressing, word order of 32-bit values, signed/unsigned interpretation, wrong tag mapped.

Diagnostic steps: compare raw value in the device, PLC value, server value and SCADA value; check scaling at each layer (transmitter range, PLC scaling, SCADA scaling, which should be done in only one place). For Modbus, see Modbus RTU and Modbus TCP Explained.

Prevention: a single, documented place for scaling; tag lists with ranges and units reviewed during commissioning.

Symptom 5: SCADA is fine, but the historian has gaps or flat lines

Likely causes: historian collector disconnected, store-and-forward disabled or buffer full, exception or compression settings too coarse, collector reading from a different server, time synchronisation errors.

Diagnostic steps: check collector status and buffer levels, compare raw historian values with SCADA trends, review exception/compression settings, check timestamps and clocks.

Resolution: reconnect and allow buffers to empty (backfill); adjust settings for affected tags.

Prevention: buffer monitoring, alarms for stale historian tags, appropriate deadbands. See Industrial Historians Explained.

Symptom 6: MES counts or states stop updating

Likely causes: interface service stopped, handshake between MES and PLC stuck (request not acknowledged), counter rollover not handled, message queue backlog, MES server or database issues.

Diagnostic steps: check the interface or connector status and error queues, the handshake flags and sequence numbers in the PLC, and MES logs.

Prevention: cumulative counters with rollover handling, handshake timeouts with alarms, interface monitoring dashboards. See MES Integration Guide.

Symptom 7: Data from MQTT or IIoT platforms is stale

Likely causes: edge gateway disconnected, broker unavailable, duplicate client IDs, expired certificates, Sparkplug birth sequence missed.

Diagnostic steps and prevention: see MQTT and Sparkplug B, which covers state management, rebirths and broker troubleshooting.

Verification after any fix

  • The value and quality are correct at every layer of the data path.
  • Historian data after the fix is continuous; check whether backfilled data arrived for the outage period.
  • Alarms for communication failure have cleared and will trigger again if the fault recurs (test if possible).
  • Downstream reports and MES calculations for the affected period are reviewed and corrected if necessary.

Prevention checklist

  • Data path documentation for important tags (source → consumers)
  • Communication status alarms per device and per interface
  • Stale-data detection in SCADA, historian and MES
  • Store-and-forward at every hop that supports it
  • Scaling done in one documented place
  • Change management that includes impact on SCADA, historian and MES tags
  • Capacity monitoring of servers and networks
  • Certificate and time synchronisation monitoring

Frequently asked questions

What does bad quality mean in SCADA?

It means the SCADA system does not have a trustworthy value for the tag, usually because communication with the source failed, the address is invalid, or the source reported a fault. The displayed value should not be used for decisions until the cause is fixed.

Why is my tag value frozen but the quality good?

The source may be sending the same value (a frozen sensor, a PLC tag no longer written by logic), or a deadband may be filtering small changes. Check the value in the PLC and the device itself.

How do I find where data is lost between the PLC and the historian?

Follow the data path hop by hop and compare value, quality and timestamp at each layer. The first layer where they differ from the previous one contains the fault.

Key takeaways

  • Trace the data path hop by hop and compare value, quality and timestamp at each layer.
  • The pattern (one tag, one device, many devices, historian only, MES only) points to the layer at fault.
  • Prevent problems with communication alarms, stale-data detection, buffering and change management.

Before you apply this in a plant: this article is for education. Always check the current edition of the relevant standards, the manufacturer's documentation for your exact product and version, and your site's procedures. Safety-related work needs qualified personnel. See our editorial policy.

Written by Bhargava Reddy Kapireddy

Bhargava has 16 years of hands-on experience with MES, SCADA, DCS, PLC and industrial data systems across power generation, oil and gas, pharmaceuticals and process manufacturing. He founded MFG Tech Hub to share practical, vendor-neutral automation knowledge.

More about the author → How we write and review articles