Planning Firmware Updates for IIoT Gateways: Maintenance Windows, Rollback and Measurement Data Loss

Industrielles IIoT Gateway bei einem Firmware Update mit schematischer Darstellung von Wartungsfenster, Wiederherstellung und lokaler Messdatenpufferung.
→ Product category: IIoT solutions

An IIoT gateway has been reliably transmitting pressure, temperature and level data to a central platform for months. New firmware is then installed to address security vulnerabilities and improve communication. After the restart, the device is accessible again, the cloud connection is established and the dashboard displays current measurements. At first glance, the update appears to have been successful.

However, subsequent analysis reveals a gap in the measurement data history. No values were recorded for several minutes during the restart. Other measurements were received but carry the time of their subsequent transmission rather than the original acquisition time. In addition, the previous scaling of one data point has changed. The connection is working again, but the quality of the historical data has been compromised.

Firmware updates for IIoT gateways are therefore not merely an IT security task. They affect the entire industrial measurement data chain. In addition to the update itself, maintenance windows, local data buffering, timestamps, configuration, recovery and the required availability of alarm functions must be considered.

This technical article explains how firmware updates for industrial edge gateways can be prepared, implemented and verified in a controlled manner. It focuses on preventing undetected measurement data loss, establishing a reliable rollback plan, safely resuming data transmission and clearly distinguishing device availability from actual data quality.

Table of Contents

  1. Define the Update Objective and Measurement Data Requirements
  2. Consider the Complete IIoT Data Chain
  3. Which Functions a Firmware Update Can Interrupt
  4. Identify Gateway Inventory, Dependencies and Update Risks
  5. Plan Maintenance Windows and Recovery Times Realistically
  6. Test Firmware in a Test Environment Before Rollout
  7. Back Up Configuration, Measurement Data and Access Information
  8. Secure Firmware Authenticity and Update Transmission
  9. Prepare Reliable Rollback and Recovery Procedures
  10. Distinguish Measurement Data Loss, Transmission Failure and Data Gaps
  11. Design Local Data Buffering and Store-and-Forward
  12. Calculation Example: Buffer Capacity and Recovery Time
  13. Preserve Timestamps and Clock Synchronisation
  14. Preserve Data Models, Units and Scaling
  15. Handle MQTT, Delivery Acknowledgements and Duplicate Messages
  16. Maintain Local Alarms and Safe Plant Operation
  17. Verify Fieldbus, Sensor Connections and Communication After Restart
  18. Roll Out Firmware Updates to Multiple Gateways in Stages
  19. Perform a Controlled Firmware Update Procedure
  20. Verify Measurement Data Quality and Operational Readiness After Updating
  21. Systematically Diagnose Typical Update Errors
  22. Suitable IIoT Solutions from ICS Schneider
  23. Conclusion: Treat Firmware Updates as Controlled Changes to the Measurement Data Chain
  24. Frequently Asked Questions About Firmware Updates for IIoT Gateways

1. Define the Update Objective and Measurement Data Requirements

A firmware update can serve different purposes. Common objectives include closing known security vulnerabilities, correcting errors, improving communication protocols or providing additional functions. In an IIoT gateway, such changes can affect both IT connectivity and the processing of industrial measurement data.

Before updating, it is therefore necessary to determine which gateway functions are affected. A security update for a communication library differs from a firmware version that simultaneously changes the Modbus driver, database structure or MQTT messages.

The operational importance of the measurement data is equally significant. In a slow-changing level monitoring application, a short, documented data gap may be acceptable. Time-critical process monitoring, continuous quality recording or systems with important alarm functions may require considerably stricter conditions.

Planning should therefore define the maximum permissible interruption of the online connection, acceptable loss of measured values and required recovery time separately. For example, an interrupted cloud connection does not automatically mean that locally recorded measurements have been lost.

Two operational parameters are particularly useful: The Recovery Time Objective (RTO) describes the target maximum recovery time. The Recovery Point Objective (RPO) describes the maximum tolerable data loss relative to a recovery point. Both values must be defined according to the application.

An RPO of zero means that no relevant measurement data may be lost as a result of the event being considered. This objective can only be achieved if the architecture provides uninterrupted acquisition or subsequent recovery of measurement data during the update interruption. An existing cloud connection alone is insufficient.

2. Consider the Complete IIoT Data Chain

In an industrial IIoT application, the gateway is positioned between the field level and higher-level IT or OT systems. It reads measured values from sensors, transmitters or controllers, processes them and makes them available to other applications.

A typical arrangement may look as follows:

Sensor / Transmitter → Fieldbus or Analogue Input → Edge Gateway → Local Processing and Data Buffer → MQTT / HTTPS / OPC UA → SCADA, Historian or Cloud

The actual architecture may differ. Some systems already acquire measurement data in a PLC. Others perform time-based acquisition only at the gateway. Still others use separate local data loggers or an industrial database.

For firmware planning, it is essential to identify which tasks are actually performed by the device being updated. If the gateway only forwards data, an interruption may have different consequences from those of a gateway that polls the sensors itself while also performing local alarm calculations.

The IIoT solutions described on the ICS website connect measuring instruments to higher-level systems via RS-485/Modbus RTU, HART, IO-Link, OPC UA or Ethernet, for example. Scaling, timestamping, local alarm processing and data buffering can be performed at the edge.

The gateway is therefore a functional component of the measurement chain. Changing its firmware can affect how data is read, processed, temporarily stored and transmitted. Successfully installing a new version does not prove that the complete measurement task continues to operate unchanged.

3. Which Functions a Firmware Update Can Interrupt

Depending on the device architecture, a firmware update may restart individual applications or require a complete gateway reboot. Different functions may become temporarily unavailable during this process.

It is important to distinguish between the device that originally acquires the measurements and the device that transmits them. If sensor acquisition continues during a gateway restart, the data may be retrieved later. However, if acquisition itself is interrupted, actual measurement values may be missing unless an independent recording system is available.

Possible Effects of Firmware Updates on the Industrial Measurement Data Chain
Affected Function Possible Effect Necessary Countermeasure
Fieldbus polling Sensor values are not polled during the restart Assess the acquisition interruption; provide an independent history or upstream buffering if necessary
Local database Buffered data is temporarily unavailable or becomes corrupted if stored improperly Ensure persistent, restart-safe buffering and suitable data backup
MQTT or HTTPS connection Measured values temporarily fail to reach the broker or server Provide store-and-forward and controlled retransmission
Time service The system time may initially be incorrect after restarting Check clock synchronisation and flag unreliable timestamps
Driver or data model Addresses, data types, units or scaling change Compare configurations and verify reference measurement points
Local alarm function Limit monitoring may be unavailable during restart Provide an independent protective function or an approved alternative operating arrangement
Communication certificates The connection is not re-established after the update Check certificates, keys, validity and connection parameters beforehand
Remote management The gateway can no longer be accessed remotely after a fault Provide local recovery access and a defined rollback procedure

The table shows that not every update interruption necessarily results in permanent data loss. The decisive factors are which function fails and whether another component continues performing the required task.

Systems in which the same gateway acquires, buffers and processes measurements while also generating alarms are particularly critical. If this device fails completely, several functions may be affected simultaneously.

4. Identify Gateway Inventory, Dependencies and Update Risks

Before a firmware rollout, it is necessary to know which devices are affected. This requires a clearly defined inventory including device identifier, manufacturer, model, hardware revision, installed firmware and associated plant areas.

The interfaces used are equally important. For example, a gateway may poll several Modbus RTU transmitters, receive data from a PLC via OPC UA and simultaneously transmit information to a cloud platform using MQTT.

A firmware change may therefore affect different communication partners. If protocol libraries or drivers are updated, their compatibility with the connected devices and systems must be verified.

Dependencies involving certificates, VPN connections, time synchronisation and central management services also belong in the inventory. A gateway may start correctly from a technical perspective yet fail to transmit measurement data because required authentication no longer works after the update.

For each gateway, the measurement points it serves and the processes depending on those values should also be documented. This allows devices with relatively low operational importance to be distinguished from gateways with high requirements for availability and data integrity.

Update prioritisation should consider both security risk and operational risk. An update addressing a serious security vulnerability may be urgent. Nevertheless, an appropriate procedure must be selected to prevent the update from unexpectedly interfering with process functions.

For industrial OT systems, NIST SP 800-82 Rev. 3 describes a risk-based, documented approach to updates. IEC TR 62443-2-3 also addresses patch management in industrial automation and control systems. The specific implementation must be appropriate for the installation.

5. Plan Maintenance Windows and Recovery Times Realistically

A maintenance window is the period during which a planned change may be performed under agreed operating conditions. For an IIoT gateway, this period must include considerably more than the actual firmware installation.

Planning includes the final check of the original configuration, any necessary backup or emptying of data buffers, installation, restart, restoration of communication connections and verification of measurement data transmission.

Time for a possible return to the previous version must also be considered. If the entire maintenance window is consumed by installation, insufficient time remains for recovery and subsequent testing if a fault occurs.

An example time budget could allocate 10 minutes for backup and preparation, 15 minutes for the update, 5 minutes for restart, 20 minutes for functional testing and 30 minutes as a recovery reserve. This results in a planned maintenance window of 80 minutes. These durations are planning assumptions only and are not universally applicable manufacturer specifications.

The actual duration depends on factors including file size, transmission path, storage medium, device performance, number of interfaces and type of update. Connection quality must also be considered for devices using cellular communication.

A reliable maintenance window includes a clearly defined abort point. If successful commissioning has not been demonstrated by that time, recovery is initiated according to the established procedure.

For measurement points distributed worldwide, different time zones, production schedules and the availability of service technicians must also be considered. A convenient time for central IT may coincide with a critical operating period at a remote site.

The maintenance window must therefore be coordinated with the responsible OT operators. Necessary protective measures for safety-related functions must not be interrupted without an approved alternative protection concept.

6. Test Firmware in a Test Environment Before Rollout

Where possible, a firmware update should first be installed on a representative test device. The test environment must reproduce the essential characteristics of the intended application. These include hardware revision, previous firmware version, communication protocols and important configuration settings.

A gateway connected to a single sensor via Ethernet in a test laboratory only partially represents a production installation with several RS-485 devices, cellular communication and a local database.

The test environment should verify that all required services start correctly after the update. These may include fieldbus polling, MQTT connectivity, the time service, local measurement data buffering and any alarm functions.

An interrupted update process can also form part of an approved robustness test. The decisive factors are how the gateway responds to an unsuccessful installation and whether an intended recovery mechanism works. Such tests must only be performed according to the manufacturer’s instructions and under suitable conditions.

Testing data processing is particularly important. Does the new firmware continue transmitting the same measured quantities with the intended units, data points and timestamps? Do existing configuration files remain valid? Are previously stored measurement data correctly transferred?

The test environment should also demonstrate whether the previous firmware can actually be reinstalled. Merely having an archive containing the earlier firmware does not constitute reliable evidence of a working recovery procedure.

If testing is unsuccessful, the production rollout is postponed until the cause has been understood and the procedure adjusted.

7. Back Up Configuration, Measurement Data and Access Information

Before a firmware update, the information required to restore the previous functionality must be backed up. This includes considerably more than the firmware file itself.

A typical gateway contains fieldbus configurations, address assignments, measuring range parameters, units, MQTT topics, alarm thresholds, network settings and various communication certificates.

Systems with local data processing may also contain scripts, filtering rules, conversion formulas or customised applications. These settings must remain available after updating or be recoverable from an approved backup.

Local measurement data buffers must also be considered. If firmware installation modifies storage areas or reinitialises a storage medium, data that has not been backed up may be lost. The backup scope must therefore explicitly distinguish between configuration data and measurement data.

Backing up sensitive access information requires an appropriate security concept. Private keys, passwords and certificates must not be transferred unprotected into arbitrary backup files. Where the device does not support secure export, recovery must use the designated management or provisioning procedures.

Recoverability is also essential. A backup is only reliable once its readability and the necessary association with the device and firmware version have been verified.

Backup documentation should specify which data was saved, when the backup was created, which hardware revision it refers to and which permissions are required for restoration.

A simple copy of the device configuration does not automatically guarantee recovery of the complete operational state. Active measurement data, queues, certificates and external communication relationships may require additional measures.

8. Secure Firmware Authenticity and Update Transmission

A firmware update changes software that operates with extensive permissions on an industrial device. It is therefore essential to ensure that the firmware comes from a trusted source and has not been modified without detection during distribution.

Suitable update architectures use digitally signed firmware images and protected metadata, for example. The IETF standardisation approach described in RFC 9019 addresses the importance of authentication, integrity protection and secure recovery mechanisms, among other aspects.

A cryptographic hash can help compare a file with a known reference value. However, a hash alone does not establish that the source is trustworthy. Suitable authentication or signature verification is required for this purpose.

Compatibility must also be checked. A firmware file may be authentic but intended for a different hardware revision. The update procedure must therefore verify the device and version assignment.

For remote updates over a network, the permissions required to initiate updates and the communication paths must also be secured. Not every user authorised to read measurement data should automatically be allowed to install firmware on production gateways.

A suitable division of responsibilities distinguishes between approving an update, making it technically available and performing the actual installation within the designated maintenance window.

Automated update systems also require traceable logs. These must show which version was installed on which device and by whom or by which authorised process.

Secure firmware distribution and operational update approval complement each other. Even correctly signed firmware should not be rolled out to a production OT installation without control if its effects on measurement data processing have not yet been tested.

9. Prepare Reliable Rollback and Recovery Procedures

Rollback means returning to a previous, functioning software version in a controlled manner. This capability is particularly important for IIoT gateways if measurement data is no longer transmitted correctly after an update or a required function becomes unavailable.

However, automatic rollback is not a standard feature of every gateway. Its availability depends on the hardware, bootloader, memory layout and firmware architecture.

Certain devices use separate firmware partitions, such as an A/B system. In this arrangement, the new version can initially be prepared in another storage area. Following a successful start, the new partition is activated. Under appropriate conditions, it may be possible to revert if a fault is detected.

Other devices require a separate recovery environment, local service access or complete reinstallation. Some systems do not support directly downgrading to earlier firmware versions at all.

Security mechanisms may also prevent older, vulnerable versions from being reinstalled. Rollback must therefore not be assumed to be available without checking the manufacturer’s documented capabilities and limitations.

Changes to database or configuration structures are particularly critical. If new firmware migrates stored data into another format, the older firmware may no longer be able to read it.

A complete recovery plan must therefore consider software, configuration, data model and local measurement data separately. Successfully returning to the previous firmware does not automatically guarantee that the earlier data set has also been restored.

Before rollout, the permissible rollback versions, required backups, recovery time and responsible person should therefore be clearly defined.

A rollback is only complete when the gateway once again acquires and transmits the intended measurements and performs the required operational functions.

10. Distinguish Measurement Data Loss, Transmission Failure and Data Gaps

In IIoT operation, different types of data interruption are frequently grouped together under the term measurement data loss. However, a more precise distinction is necessary to assess the quality of subsequent analysis.

Acquisition failure: No new measurements are generated or recorded during a particular period. Without independent recording, the missing values cannot subsequently be reconstructed reliably.

Transmission failure: Measurements continue to be acquired but initially fail to reach the higher-level server. If they are reliably stored locally, subsequent transmission may be possible.

Time-related data error: The measurements are available but contain unsuitable or incorrect timestamps. Their order or temporal assignment may therefore appear incorrect later.

Semantic data error: After the update, a value is interpreted using another unit, scaling factor or data point assignment. Transmission works, but the physical meaning is no longer correct.

These types of errors have different consequences. A temporary cloud connection failure may not cause permanent data loss if buffering is available. A complete interruption of gateway acquisition, on the other hand, may create a genuine measurement gap.

Measurements transmitted retrospectively must therefore remain identifiable as historical measurements. The transmission time must not be confused with the original measurement time.

A continuous line in a dashboard is also insufficient evidence of uninterrupted acquisition. Some applications graphically connect existing data points or fill in missing values through interpolation. Such displays can conceal actual outages.

Reliable measurement data therefore requires information about acquisition status, timestamps and any existing data gaps.

11. Design Local Data Buffering and Store-and-Forward

Store-and-forward describes a method in which measurement data is initially stored locally during a transmission interruption and forwarded when the connection is restored.

For IIoT gateways, this function can make a significant contribution to data availability. However, it requires measurement acquisition and local storage to continue functioning during the interruption in question.

Volatile memory alone may be insufficient for a longer update interruption. During a restart, data may be lost if it exists only in RAM. Persistent storage may therefore be required to achieve the desired restart capability.

The local buffer should preserve the order of measurements, their timestamps and quality status. It must also handle incompletely transmitted data and repeated transmission attempts.

One important question concerns the behaviour when the storage capacity is exhausted. Some systems overwrite older data, others discard newer values or stop acquiring data. The appropriate strategy depends on the measurement task.

Overwriting the oldest measurements may be acceptable for a simple live trend display, for example. However, this behaviour may be impermissible for traceable quality records.

Stress on the storage medium must also be considered. Frequent write operations can affect the service life of unsuitable flash storage. Storage technology, write strategy and the required data integrity are therefore part of the system design.

The limitation of this method is also decisive: Store-and-forward protects against interrupted forwarding. If the gateway does not acquire new measurements during a firmware restart, its own buffer cannot generate those missing values afterwards.

To maintain uninterrupted acquisition during a complete gateway outage, independent recording in a PLC, separate data logger or suitably designed upstream component may be required. Whether historical retrieval is supported depends on the devices and protocols used.

Further information about local buffering and cloud connection failures is available in the ICS article Cloud Connection Failure: Planning Reliable Edge Alarms and Local Data Buffering.

12. Calculation Example: Buffer Capacity and Recovery Time

The required local buffer storage depends on the number of measurement points, acquisition interval, average record size and maximum interruption period to be covered.

The following relationship can be used for a simplified estimate:

S = N · f · B · t

Here, S is the required storage capacity in bytes, N is the number of measurement points, f is the acquisition rate per measurement point in 1/s, B is the average record size in bytes and t is the interruption duration in seconds.

One example considers a gateway with 120 measurement points. Each point is recorded every two seconds. Including the metadata assumed for this example, an average of 220 bytes is stored per record. A transmission interruption lasting three hours must be buffered completely.

The acquisition rate is therefore 60 records per second. Over three hours, 648,000 records are generated.

S = 120 · 0.5 s-1 · 220 bytes · 10,800 s
S = 142,560,000 bytes ≈ 142.6 MB

The following table summarises the example:

Example of Sizing a Local Measurement Data Buffer
Parameter Assumed Value Meaning
Number of measurement points 120 Separately acquired data points
Acquisition interval 2 s One measurement per data point every two seconds
Record size 220 bytes Assumed average storage requirement per record
Interruption duration 3 h Maximum planned transmission interruption in the example
Number of records 648,000 Records generated during the interruption
Calculated storage requirement approx. 142.6 MB Excluding additional storage and management reserves
With an illustrative 30 % reserve approx. 185.3 MB Additional planning allowance, not a universally applicable requirement

For actual sizing, the file system, database indexes, management information, possible retransmissions and other local storage tasks must also be considered. The storage used must remain available during the update or be reliably preserved.

Retransmission also requires time. Assume that after the connection is restored, the system can process a total of 100 records per second. At the same time, 60 new records per second continue to be generated. This leaves a net capacity of only 40 records per second for clearing the backlog.

Under these idealised conditions, clearing 648,000 buffered records would take approximately 4.5 hours. In practice, protocol overhead, server load and network connectivity may affect recovery time.

Sufficient storage capacity alone therefore does not guarantee a rapid return to current data. Transmission capacity after the interruption must also be sufficient to handle ongoing acquisition and the existing backlog.

The calculation explicitly considers only measurements that were actually generated and stored during the interruption. A complete acquisition failure during the gateway restart is not covered.

13. Preserve Timestamps and Clock Synchronisation

For an industrial measurement data history, the time reference is just as important as the measured value itself. A pressure reading of 6.2 bar can only be clearly assigned to a process condition if its acquisition time is known.

Several timestamps may occur in IIoT systems: the time of actual measurement acquisition, the time the gateway receives the measurement and the later time the value is received by a database or cloud platform.

These times are not necessarily identical, particularly during and after a firmware update. If measurements are buffered locally for several hours, they will reach the central database considerably later.

For historical analysis, the original acquisition timestamp should therefore be preserved. The transmission or storage timestamp may be documented additionally but must not silently replace the original measurement time.

A suitable data model distinguishes, for example, between timestamp_source, timestamp_received and timestamp_ingested. The exact names are project-specific, but the meaning of each field must be unambiguous.

OPC UA, for example, distinguishes between SourceTimestamp and ServerTimestamp. The OPC UA specification describes SourceTimestamp as the time reference originating from the data source and requires it to be passed on unchanged where such a timestamp is available.

After restarting, the gateway requires a reliable system time. If its clock is synchronised using NTP or another suitable time service, the behaviour during restart must be considered. Measurements acquired before successful time synchronisation may have an unreliable time reference.

Systems without a sufficiently backed-up real-time clock are particularly critical because they may initially start with an incorrect time after a power failure. Such values must not silently be marked as having valid timestamps.

It is therefore necessary to define how invalid or uncertain timestamps are handled. Suitable quality indicators can prevent such data from being used unchecked in subsequent calculations or reports.

Changes to time zones or daylight saving time must also not shift historical measurement data without detection. For industrial data models spanning multiple systems, a consistent time base, such as UTC with clearly documented presentation, is particularly useful.

14. Preserve Data Models, Units and Scaling

A firmware update can change not only communication functions but also the interpretation of measurement data. Changes to scaling factors, register assignments, units or data types are particularly problematic.

Consider a pressure sensor with an analogue signal of 4 … 20 mA. Before the update, the input is processed using a measuring range of 0 … 10 bar. Following an incorrect configuration migration, the gateway suddenly uses 0 … 16 bar. The same electrical input voltage or input current therefore produces different calculated pressure values.

Digital register processing may also be affected. For example, an updated Modbus driver may interpret data types, byte order or register addresses differently. Communication status may nevertheless continue to appear error-free.

Important measurement points should therefore be compared under the same known input conditions before and after the update. Where appropriate and safe for the measurement task, this includes representative values near the lower, middle and upper parts of the respective measuring range.

A particularly reliable approach uses versioned data models. If the unit, scaling or physical meaning of a data point changes, the system records from when each configuration was valid.

MQTT topics or data point identifiers should likewise not be renamed without control during a firmware update. Otherwise, historians, alarm rules or dashboards may receive data under an unexpected name or fail to associate it with the correct point.

The firmware version itself can be documented as additional metadata. However, it does not replace independent versioning of the measurement data model if the meaning of data can change independently of the firmware.

The complementary ICS technical article Versioning Units and Scaling in Measurement Data explains how to preserve the physical meaning of industrial measurements over time.

15. Handle MQTT, Delivery Acknowledgements and Duplicate Messages

MQTT is frequently used to transmit industrial measurement data between edge gateways, brokers and higher-level applications. Messages can be transmitted using different Quality of Service levels.

With QoS 0, a message is transmitted without the acknowledgement and retry procedure used by higher QoS levels. Message loss is possible. QoS 1 provides at least one delivery on the respective protocol connection, meaning duplicates may occur. QoS 2 uses a more extensive procedure to achieve exactly-once delivery within the respective MQTT protocol relationship.

However, these properties do not automatically mean that every measurement has been stored exactly once and permanently in a downstream historian database. A successful MQTT acknowledgement may indicate that the message was accepted by the relevant communication partner, while further processing takes place in another application.

For store-and-forward operation, it must therefore be defined when a message is considered successfully transmitted and under which conditions it may be deleted from the local buffer.

With QoS 1, messages may be repeated after connection interruptions. A robust data model should therefore enable unambiguous identification of measurements. Stable device identifiers, measurement point identifiers, timestamps and sequence numbers can be used for this purpose.

Based on these identifiers, downstream processing can recognise previously stored records and avoid duplicate entries. The combination needed to establish uniqueness depends on the acquisition method and application.

Another important consideration is message order. After reconnection, older buffered messages and new live data may arrive interleaved. A database must therefore not automatically interpret the reception time as the physical sequence of measurement events.

MQTT’s retained message function must also be distinguished from a measurement data history. Retained messages provide the last stored message content for a topic but do not replace a complete historical record.

The appropriate MQTT configuration therefore depends on data rate, availability requirements, permissible duplicates and the requirements of the complete measurement data chain.

16. Maintain Local Alarms and Safe Plant Operation

In IIoT applications, particular care must be taken to distinguish between an informative notification and a safety-related protective function.

A cloud dashboard may display limit violations and send notifications to maintenance personnel, for example. However, such a notification is not automatically suitable for initiating a safety-related shutdown.

If necessary alarm functions are executed directly within the IIoT gateway, they may be temporarily unavailable during a complete firmware restart. Local data buffering does not prevent this functional interruption.

For important process protection functions, it must therefore be determined whether independent evaluation in a suitable PLC, separate limit monitoring device or appropriately designed protection system is necessary.

The required independence follows from the hazard and functional assessment of the installation. An ordinary IIoT gateway with an alarm function is not automatically an approved functional safety system.

Before the update, it must be established which alarms must remain available during the maintenance window. If certain functions are temporarily suspended, an approved operational procedure and, where necessary, suitable alternative measures are required.

After changing the firmware, it must also be verified that all alarm thresholds, hysteresis values, delays and status indications have been transferred correctly.

The identification of historical alarms is also important. A limit violation transmitted retrospectively must not automatically be interpreted as a new current alarm if it occurred during an earlier buffering period.

A reliable architecture must therefore preserve the original event time, the measurement point status and the intended alarm processing unambiguously.

17. Verify Fieldbus, Sensor Connections and Communication After Restart

After a firmware update, checking the gateway’s cloud connection alone is insufficient. Communication with the field level is equally important.

For Modbus RTU connections, relevant parameters include serial interface settings, device addresses, register assignments and actual polling of connected instruments.

For example, a gateway may be accessible through its web interface after an update while individual Modbus devices fail to provide valid measurements because an interface configuration has changed.

For OPC UA, endpoint settings, security policies, certificates and the data point identifiers used must remain compatible with the existing installation.

HART or IO-Link connections may also require consideration of the relevant interface components and driver versions.

It is equally important to distinguish between successfully establishing a communication channel and receiving valid measurements. A successful TCP connection does not prove that the required process data is being read with the correct physical meaning.

A meaningful functional test therefore includes selected measurement points from all relevant groups of connected devices. Test values are compared with the intended references and configuration settings.

Reconnection after a network failure should also be verified. Following the firmware update, the gateway must be able to resume communication with the intended partners without requiring uncontrolled manual modifications.

Installations using changing network addresses, cellular routers or VPN connections require additional attention. In these cases, an apparent firmware fault may actually result from incorrect name resolution, certificates or network parameters.

18. Roll Out Firmware Updates to Multiple Gateways in Stages

Larger IIoT installations frequently include several gateways at different machines, production lines or sites. Updating every device simultaneously can increase the risk of a major operational interruption.

A staged rollout reduces this risk by initially updating a small, representative group of devices. Further groups are updated only after successful testing.

An appropriate sequence may begin with a test installation, continue with individual production gateways of relatively low operational criticality and finally extend to additional plant areas.

Grouping should not be based solely on location. Hardware revision, initial firmware version, interfaces used and operational importance are also relevant.

If a fault is detected, the rollout can be stopped before the same change affects further gateways.

Stable device identifiers and traceable version information are important for device and fleet management. After rollout, it must be possible to establish clearly which devices were successfully updated, which still use the previous firmware and which experienced an error.

Delayed status reports must also be considered. A gateway may be unreachable during a cellular connection interruption and come back online later. Its actual firmware and operational status must not be inferred solely from the last update command sent.

An installation reported as successful should therefore also be confirmed by an appropriate health and functional test.

For critical systems, a central facility to stop or suspend further updates is useful. Automated rollouts require clearly defined approval and abort conditions.

19. Perform a Controlled Firmware Update Procedure

A reproducible firmware update process combines technical testing, operational approval and traceable documentation. Implementation follows the instructions of the relevant gateway manufacturer and the approved OT change management procedure.

A suitable procedure may include the following steps:

  1. Assess the reason for updating: Clearly document the security vulnerability, bug fix or new function.
  2. Identify devices and dependencies: Record the hardware revision, initial firmware, measurement points, interfaces and affected plant functions.
  3. Verify firmware approval: Check the source, signature, compatibility, release notes and any required intermediate versions.
  4. Approve the operational maintenance window: Coordinate availability, necessary protective functions, recovery reserve and responsible personnel.
  5. Back up configuration and data: Secure existing settings and relevant local measurement data according to the recovery plan.
  6. Record the initial condition: Document connectivity, data quality, timestamps, buffer status and important reference measurements.
  7. Prepare the update for deployment: Use a suitable secure transmission path; if appropriate, transfer the firmware before the actual maintenance window.
  8. Install the firmware: Follow the manufacturer’s procedure and document status messages.
  9. Check the restart: Verify firmware version, system time, available services and device condition.
  10. Test field communication: Read representative measurements from the relevant sensor and control interfaces.
  11. Verify the measurement data path: Check the data model, units, scaling, timestamps, transmission status and local buffering.
  12. Test alarm and fault functions: Verify the intended operational notifications and required protective functions using the approved procedure.
  13. Make the rollback decision: If acceptance criteria are not met within the specified period, initiate recovery.
  14. Document completion: Record the final firmware version, test results, known limitations and approval for plant operation.

The order of testing is particularly important. First, the gateway must be technically operational again. Communication with the field level is then checked. Only after this can the complete measurement data transmission, including its semantic and temporal characteristics, be evaluated.

A successfully established cloud connection must not automatically eliminate the need for further tests. Complete functional verification also requires valid measurements and correct processing of connected data points.

All previously defined acceptance criteria should be fulfilled before the installation is released for production. Tests that have failed or were not performed must be explicitly documented.

20. Verify Measurement Data Quality and Operational Readiness After Updating

Successful installation of a firmware version is initially a technical status. Before an industrial IIoT measurement point can be released for operation, it must also be demonstrated that the intended measurement task is being performed correctly again.

An important test criterion is the freshness of the measurements. The data must not merely come from an old display value or an existing cache. The gateway and higher-level platform must process new measurements according to the intended acquisition interval.

It must also be checked whether measurement data from before the update remains available. Where buffering is used, it is necessary to establish whether unsent records have been fully and correctly transmitted afterwards.

The original timestamps are relevant for historical analysis. A database in which measurements several hours old suddenly carry the current time may contain all the numerical values but does not represent a correct time history.

Data quality must also be assessed. Invalid, stale or unavailable measurements should be marked according to the data model used. A missing measurement must not silently appear as a valid last known value.

Another test criterion is the order of the measurements. After a restart, buffered and new messages may arrive at different times. The downstream application must be able to assign them correctly.

Finally, measurement point identifiers, units, scaling, alarm parameters and any stored configuration versions must be compared with the approved initial state.

Complete verification therefore covers both device availability and the completeness, freshness and physical correctness of the measurement data.

Suitable performance indicators can be derived for operational quality assurance. Examples include time to restore availability, number of lost measurements, maximum buffer occupancy, duration of retrospective transmission and number of detected data point deviations.

The limits for these indicators depend on the application. They must be derived from the system requirements and established before the update.

21. Systematically Diagnose Typical Update Errors

After a firmware update, errors may occur that initially resemble ordinary network problems. A systematic diagnosis therefore separates firmware status, field communication, data processing and transmission.

Typical Errors After Firmware Updates on IIoT Gateways
Observation Possible Cause Suitable Check
Gateway is accessible, but individual measurements are missing Driver problem, changed register assignment or failed field connection Check the field interface, data point configuration and raw values
New data appears, but historical data is missing Buffer was deleted, was not stored persistently or has not yet been fully transmitted Check the local queue, database status and retransmission
Measurements appear with incorrect times System clock is not synchronised or reception time is stored as measurement time Check source timestamps, gateway clock and database time assignment
Individual measurements suddenly have different magnitudes Scaling, unit, data type or byte order has changed Compare reference values and the versioned data point configuration
Measurements are stored twice after the update Repeated transmission or missing duplicate detection Check MQTT QoS, sequence numbers and historian processing
MQTT connection is no longer established Certificate, authentication, TLS setting or broker configuration is incompatible Check connection logs, certificates, system time and access configuration
Gateway repeatedly restarts after updating Faulty firmware, configuration problem or storage/power supply disturbance Check boot logs and device status; initiate recovery if necessary
Dashboard shows normal values although a field connection has failed Old values continue to be displayed without suitable quality indicators Check measurement age, quality status and fault processing
New firmware operates, but rollback is not possible Downgrade restriction, missing recovery image or incompatible data structure Check manufacturer approval and the previously documented recovery procedure
Alarms are triggered again after historical values are retransmitted Past events are processed as current live notifications Investigate event timestamps, alarm status and replay processing

The table contains typical causes and useful diagnostic approaches. An observed fault pattern may have several technical causes. Another firmware installation or factory reset should therefore not be initiated immediately.

A factory reset can modify or delete existing configurations or local data. It is therefore only appropriate when the recovery scope has been clarified and the procedure approved.

A traceable diagnosis begins with the actual device status. Field communication, local data processing, the time base and external transmission are then checked in sequence.

22. Suitable IIoT Solutions from ICS Schneider

22.1 ICS IIoT Solutions: Edge Gateways, Data Models and Industrial Integration

The IIoT solutions from ICS Schneider connect industrial sensors, transmitters and controllers with higher-level IT and OT systems. The architectures described use technologies such as Modbus RTU, HART, IO-Link, OPC UA, Ethernet and MQTT/HTTPS.

ICS supports the selection and integration of suitable edge gateways, definition of data models, register mapping and secure transmission of measured values. Local data processing, alarm functions and buffering can also form part of an appropriate solution concept.

For the maintenance and firmware strategy, it is essential that the specific gateway hardware selected actually supports the necessary functions. These may include signed updates, persistent data buffering, secure configuration storage, defined recovery and suitable management capabilities.

These properties must be verified for the specific device. They must not be inferred merely from the general designation IIoT gateway or edge gateway.

22.2 IIoT Pressure Monitoring: Secure Processing of Pressure Sensor Data

The IIoT pressure monitoring solutions from ICS include the integration of pressure sensors and differential pressure transmitters into industrial data architectures.

One example of suitable field instrumentation is the IDCT 531i pressure sensor with RS-485/Modbus RTU. The sensor provides digital pressure values for processing by appropriately equipped master devices.

When updating the firmware of the higher-level gateway, particular attention must be paid to Modbus addressing, register interpretation, polling intervals and the documented scaling of measured values.

The pressure sensor itself does not replace gateway data buffering. If recording must continue during a complete failure of central data acquisition, the architecture requires an additional suitable acquisition or storage function.

22.3 WIKA NETRIS1: Wireless Transmission of Measurement Data

The WIKA NETRIS1 wireless unit enables suitable measuring instruments to be integrated wirelessly into IIoT applications. Depending on the configuration, supported technologies include LoRaWAN, mioty and Bluetooth.

For example, it can transmit measured values from suitable sensors with standard signals, making it of interest for remote monitoring of industrial installations.

However, the NETRIS1 is a wireless unit at the field or transmission level and must not be treated as equivalent to a freely programmable industrial edge gateway. Update, buffering and recovery functions must be verified separately for the respective system component.

In a complete IIoT architecture, it must also be clear which component generates the authoritative measurement timestamp and how data is handled during interruptions to downstream communication.

22.4 WIKA NETRIS3: IIoT Wireless Technology for Suitable Hazardous-Area Applications

The WIKA NETRIS3 wireless unit enables measurement data from suitable WIKA instruments to be transmitted via LoRaWAN. The series is designed for corresponding applications in hazardous areas.

It is used, for example, for wireless remote monitoring of industrial pressure, temperature or level measurement points.

Here too, the different levels of the IIoT architecture must be distinguished. Updating a central gateway or network server is not automatically equivalent to updating the firmware of the wireless field units.

Planning must therefore consider the relevant manufacturer approvals, operating conditions and communication dependencies separately. Equipment installed in hazardous areas is also subject to the applicable approval and maintenance requirements.

ICS Schneider offers further integration options in the areas of IIoT Temperature Monitoring and IIoT Level Monitoring.

For a complete IIoT solution, gateway management, update strategy, local data storage and future recovery requirements should be considered during the project design phase, alongside sensing technology and communication.

23. Conclusion: Treat Firmware Updates as Controlled Changes to the Measurement Data Chain

Firmware updates are part of the secure and reliable long-term operation of industrial IIoT gateways. They can address security vulnerabilities, correct errors and introduce new functions. At the same time, they can temporarily or permanently change measurement acquisition, data processing, communication and alarm operation.

A successful update must therefore not be judged solely by the installed firmware version or restoration of network connectivity. The decisive factor is whether the complete measurement data chain continues performing its intended functions after the update.

The most important prerequisites are an adequately planned maintenance window, a verified data backup, a recovery method actually supported by the device and a clear distinction between measurement acquisition and measurement transmission.

Store-and-forward can bridge an interruption of the higher-level connection. However, it cannot prevent a data gap if no new measurements are acquired during a complete gateway restart and no independent recording system exists.

Timestamps, units, scaling and data point identifiers must also remain traceable across the firmware change. A completely transmitted data set is only valuable for quality assurance if its physical meaning and temporal assignment are correct.

Assess the need for updating → Identify devices and dependencies → Define the maintenance window → Test firmware and recovery → Back up configuration and measurement data → Verify data buffering and alarm availability → Install the update in a controlled manner → Test field and IT communication → Verify measurement data quality → Document the rollout

The most important practical principle is therefore: A firmware update is only fully completed when the IIoT gateway is not merely accessible again but continues to acquire, timestamp, process and transmit industrial measurement data correctly.

24. Frequently Asked Questions About Firmware Updates for IIoT Gateways

24.1 Why Must Firmware Updates for IIoT Gateways Be Planned?

A firmware update can affect measurement acquisition, local processing, data transmission and alarm functions. A planned procedure ensures that maintenance time, data backup, recovery and functional testing are aligned with operational requirements.

24.2 Can a Firmware Update Cause Measurement Data Loss?

Yes. Data may be lost if the gateway does not acquire measurements during installation or local data is not stored persistently. However, interrupted transmission does not necessarily result in permanent data loss if the measurements are reliably buffered.

24.3 Is Local Store-and-Forward Sufficient to Prevent Data Loss?

Not in every case. Store-and-forward primarily protects against interruptions to data forwarding while measurement acquisition and local storage continue operating. During a complete gateway failure, data gaps may still occur without independent upstream recording.

24.4 How Large Should the Local Measurement Data Buffer Be?

The storage requirement depends on the number of data points, acquisition rate, average record size and intended interruption duration. Storage management, safety margins, write behaviour and retransmission capacity must also be considered.

24.5 What Does Rollback Mean for an IIoT Gateway?

Rollback is the controlled return to a previous functioning firmware version. Whether it can be performed automatically or manually depends on the specific device. Configuration and local data formats must also be compatible with the restored software.

24.6 Does Every Industrial Gateway Support Automatic Rollback?

No. Some systems use A/B firmware partitions or recovery images. Others require local service access or do not support directly returning to an earlier version. The actual capabilities must be checked against the manufacturer’s documentation before updating.

24.7 What Should Be Included in a Firmware Update Maintenance Window?

The maintenance window includes preparation, backup, installation, restart, communication and measurement verification, and an appropriate recovery reserve. Operational approval and a clearly defined abort point are also required.

24.8 Can Firmware Be Updated While Production Is Running?

This depends on the specific architecture and affected functions. If the update interrupts measurement acquisition or alarm operation, the consequences must be assessed and operational approval obtained. Necessary protective functions require suitable independent measures or an appropriately secured plant condition.

24.9 Why Are Timestamps Sometimes Incorrect After an Update?

Possible causes include an unsynchronised system clock, changed time settings or incorrect assignment of measurement and transmission times. Time synchronisation and the handling of historical measurements must therefore be checked after every relevant update.

24.10 Can Measurements Be Transmitted Retrospectively Without Changing Their Timestamps?

Yes, provided that the original measurement times were correctly recorded and stored together with the values, and the downstream system processes them accordingly. The later reception time should remain distinguishable from the original measurement time.

24.11 Why Do Duplicate MQTT Measurements Occur After a Firmware Update?

With certain transmission methods, such as MQTT QoS 1, messages may be delivered more than once. Without suitable identification and duplicate detection, this can produce duplicate historical records. Stable data point identifiers, timestamps or sequence numbers can help establish unambiguous identification.

24.12 Can a Firmware Update Change the Scaling of Measurements?

Yes, if configurations are migrated, register assignments change or data types are interpreted differently. Relevant measurement points should therefore be compared before and after the update using known reference values and the approved data point configuration.

24.13 Why Is Checking Only the Network Connection Insufficient After an Update?

A successful network connection does not automatically confirm correct sensor polling, scaling, timestamping or data transmission to the historian. Complete approval requires representative measurements and verification of the relevant data processing.

24.14 What Role Does Local Alarm Processing Play During a Firmware Update?

If alarms are calculated directly in the gateway, they may become unavailable during a restart. Necessary protective functions must therefore remain independently available based on the hazard assessment or be secured through an approved alternative procedure.

24.15 Should All IIoT Gateways Be Updated Simultaneously?

For larger installations, a staged rollout is often preferable. A representative group of devices is updated and tested first. Additional gateways follow only after successful assessment. This allows potential faults to be detected early and their impact limited.

24.16 What Security Requirements Apply to Firmware Updates?

Important requirements include the use of trusted firmware, integrity and authenticity verification, controlled permissions, suitable transmission paths and traceable update logs. Operational risks and required approval processes must also be considered for industrial OT systems.

24.17 When Is a Firmware Update Truly Complete?

When the gateway is running the intended firmware after updating and all required functions have been successfully verified. These include field communication, measurement data processing, timestamps, transmission, buffering and, where applicable, alarm and fault functions. A successful installation message alone is insufficient.

24.18 What Information Does ICS Schneider Need for Planning?

The required information includes the gateway manufacturer and model, hardware and firmware versions, field interfaces used, number of measurement points, acquisition intervals and connected IT and OT systems. Requirements concerning data availability, permissible interruption, local buffering, alarm functions, remote access and cybersecurity are also important. The update strategy additionally requires information about backup, rollback, maintenance windows, recovery times and any necessary testing and documentation requirements.

Diese Website benutzt Cookies. Wenn du die Website weiter nutzt, gehen wir von deinem Einverständnis aus.