Table of Contents
- Building a Critical Infrastructure Maintenance Plan
- SOCI Act Compliance Requirements for Critical Assets
- Preventive Maintenance Strategies for Data Centres
- The Case for Vendor-Agnostic Infrastructure Maintenance
- Integrating Predictive Maintenance and AI Monitoring
- Securing the Converged Threat Landscape: IT, OT, and Physical Security
- Budgeting for Resilience and Adapting to Climate Risks
- Frequently Asked Questions
Last Updated: September 13, 2026
Building a Critical Infrastructure Maintenance Plan
Most maintenance failures aren't technical. They're structural: no plan, no ownership, no budget line. Best practices for critical infrastructure maintenance start with a written, funded, reviewed plan that assigns every asset an owner, a service interval and an escalation path. This guide from Treske Pty Limited covers how to build one that survives audits, outages and budget season.
Critical infrastructure is any asset, system or network whose failure would disrupt essential services: power, cooling, water, communications, health, transport or finance. In a data centre, that means the UPS, switchgear, generators, cooling plant, rack enclosures and the monitoring layer tying them together.
The stakes are practical: a single unplanned outage cascades through clinical systems, payment platforms and emergency services within minutes. Regulation has tightened too, the Security of Critical Infrastructure Act 2018 (SOCI Act) imposes positive obligations on named asset classes, including data storage and processing (legislation.gov.au).
The Foundation: Asset Inventory and Risk Assessment
You cannot maintain what you haven't catalogued. A working asset inventory records, for every item: make, model, serial number, install date, expected service life, firmware or configuration baseline, physical location, and the business function it supports.
From that inventory, build a risk register: score each asset on likelihood and consequence of failure, then rank. The top of that list gets the most attention, the tightest monitoring and the fastest spares availability.
A practical sequence:
- Walk the floor. Physically verify every asset against procurement records and as-built drawings.
- Tag and photograph. Unique asset ID, location code and a dated photo for each item.
- Map dependencies. Note which assets feed which loads, including downstream systems.
- Score and rank. Use a simple 1-5 scale for likelihood and consequence.
- Assign an owner. One named person accountable for each asset's lifecycle.
A common mistake is treating the inventory as a one-off project. It's a living document. Every change, swap or decommission should update it within days, not quarters.
SOCI Act Compliance Requirements for Critical Assets
The SOCI Act governs how critical infrastructure owners and operators manage risk. It requires entities in named sectors to maintain a critical infrastructure risk management program, report certain cyber incidents, and register ownership and operational information with the relevant regulator.
In practice, compliance obligations vary by sector and asset class. Data storage and processing, electricity, gas, water, maritime, ports, hospitals and communications each carry different requirements, and the rules are updated periodically. Rather than rely on secondhand summaries, check the current requirements directly with the Department of Home Affairs critical infrastructure guidance and the Australian Cyber Security Centre's guidance for critical infrastructure.
What good compliance looks like on the floor:
- A documented risk management program reviewed at least annually
- Named accountable officers with defined reporting lines
- Incident response procedures with tested escalation paths
- Evidence trails for maintenance, testing and configuration changes
- Regular internal audits, not just annual external ones
Preventive Maintenance Strategies for Data Centres
Preventive maintenance is scheduled, condition-based servicing performed before failure occurs: inspection, testing and replacement of the components most likely to degrade, UPS batteries, cooling filters, fans, capacitors, electrical connections and monitoring sensors.


The core schedule most facilities work to:
| Asset | Typical Interval | Key Action |
|---|---|---|
| UPS batteries | Quarterly test, replace on measured degradation | Load test, impedance check, replace weak strings |
| Cooling filters and coils | Monthly visual, quarterly clean | Replace filters, clean coils, verify airflow |
| Standby generators | Monthly no-load, annual load bank | Fuel test, run under load, service engine |
| Switchgear and breakers | Annual | Thermal scan, torque check, insulation test |
| Monitoring sensors | Quarterly | Calibrate, verify alerting thresholds |
Two things separate a preventive program that works from one that just ticks boxes: maintenance must be scheduled around load, not convenience, and every visit should produce a written report with measurements, not just a signature.
Treske's preventative maintenance service is built around that principle: scheduled servicing of power and cooling assets with documented results, so you have both the uptime and the evidence trail.
The schedule is the easy part. The hard part is protecting the maintenance window when operations wants it back.
The Case for Vendor-Agnostic Infrastructure Maintenance
A vendor-agnostic maintenance provider isn't tied to a single manufacturer's equipment, software or spares channel. They maintain what you have, across mixed brands and generations, rather than pushing a replacement.
This matters most in facilities that grew by acquisition or upgrade. A typical mid-sized data centre might run UPS units from two manufacturers, cooling from a third, and monitoring software from a fourth, a single-vendor contract covers only part of that estate.

The practical benefits:
- One point of accountability across a mixed estate
- Independent advice on repair versus replace, without a sales incentive
- Longer asset life for legacy equipment that still performs
- Faster fault isolation when the fault spans multiple vendors' systems
The trade-off is real: a vendor-agnostic provider won't have the manufacturer's deepest engineering bench for every product line, and in-warranty work still routes through the OEM. For older assets out of warranty, though, agnostic maintenance is usually the difference between a repair and an unplanned capital purchase.
Integrating Predictive Maintenance and AI Monitoring
Predictive maintenance uses real-time data and analytics to flag degradation before failure. Instead of servicing on a fixed calendar, you service when the asset's own measurements say it's due. The practical detail, how data flows, how thresholds get set, where the approach fails, determines whether it works in your facility. Applying these diagnostic principles to energy storage systems requires a nuanced understanding of lithium battery maintenance to ensure that sensor readings accurately reflect the true state of health of your power infrastructure.
The enabling layer is monitoring software. Platforms like APC / Schneider EcoStruxure IT, Eaton Intelligent Power Manager, Sunbird DCIM and Ekkosense Datacenter Optimization collect data from UPS units, power distribution, cooling and environmental sensors, then surface anomalies through dashboards and alerts. Several use AI and machine learning to identify patterns that a human reviewing logs would miss.

How the data pipeline actually works
A predictive program has four layers, and weakness in any one of them breaks the whole chain:
- Sensors and telemetry. UPS units expose battery impedance, string voltage, load percentage and internal temperature; cooling plant exposes supply and return temperatures, humidity, fan speed and valve position; power distribution exposes current per phase, power factor and breaker status. If a device doesn't expose a metric over a protocol your platform can read (SNMP, Modbus, BACnet or a vendor API), that metric doesn't exist for analytics.
- Collection and normalisation. The platform polls each device on an interval, commonly every one to five minutes for power and cooling, longer for slow-moving metrics like battery impedance, and normalises units and timestamps so one vendor's temperature reading is comparable to another's.
- Baselining. The platform learns what 'normal' looks like for each asset under each operating condition, a UPS at 40% load in a 22°C room differs from the same unit at 80% load in a 28°C room. Baselining typically takes weeks to months of clean data before alerts become trustworthy.
- Anomaly detection and alerting. Once a baseline exists, the platform flags deviations, a slow upward drift in battery impedance, a cooling unit working harder to hold the same setpoint, a fan drawing more current than its peers. Machine learning helps where relationships are non-linear or patterns only appear across many assets.
What changes in practice
- Battery replacement moves from calendar to condition. Impedance trends show which strings are degrading, so you replace two, not twenty, the single largest cost saving most facilities realise, since battery strings are the most over-replaced asset in a data centre.
- Cooling optimisation becomes continuous. Thermal mapping identifies hot spots and overcooling simultaneously, so you can raise setpoints where there's headroom without risking inlet temperatures.
- Faults get caught in hours, not weeks. A rising temperature trend in one rack triggers an alert long before a threshold breach.
- Maintenance windows get targeted. Instead of servicing every unit on the same schedule, you service the ones the data flags and leave the rest alone.
The trade-offs nobody puts in the brochure
Predictive maintenance is not free of failure modes. Three matter most:
- Alert fatigue. A poorly tuned platform generates more alerts than a team can action, and the team starts ignoring them. Tuning thresholds and suppressing known-benign patterns is ongoing work, not a one-off setup task.
- Sensor gaps create blind spots. With one temperature sensor per aisle, no analytics platform will tell you what's happening in rack 14. Sensor density is the real investment, and it usually costs more than the software licence.
- Garbage in, garbage out. A sensor that has drifted out of calibration produces confident, wrong alerts. Calibration schedules matter as much as the analytics.
Monitoring software is only as good as the sensor coverage underneath it. Start with sensor density and data quality, then layer analytics on top, the reverse order produces dashboards nobody trusts.
Securing the Converged Threat Landscape: IT, OT, and Physical Security
Operational technology and information technology have converged: the same network that carries email now carries UPS telemetry and building management commands, collapsing the boundary between a cyber incident and a physical one.
A practical security posture for critical infrastructure rests on four foundations:
- Network segmentation. Keep OT traffic off the corporate network. Use dedicated VLANs, firewalls and, where possible, one-way data flows from OT to IT.
- Access control. Role-based access, multi-factor authentication for remote monitoring platforms, and physical access logs for every equipment room.
- System hardening. Change default credentials, disable unused services, apply firmware updates on a defined schedule, and maintain a configuration baseline.
- Perimeter and physical security. Camera coverage, door access control and visitor logging at every entry point to the plant.
Data encryption matters at both ends: in transit for monitoring traffic, and at rest for the logs and configuration records supporting compliance auditing.
Convergence also changes incident response: an escalation path covering only IT staff stalls when the incident is in the cooling plant. Build a single escalation matrix naming both IT and facilities contacts, with defined response times and a tested communication tree.
Budgeting for Resilience and Adapting to Climate Risks
Resilience costs money, and the budget conversation is where most maintenance programs quietly die. Win it by presenting maintenance as risk reduction with a number attached, not a cost centre. Most technical guides assume the budget exists; it usually doesn't, and the framework below closes the gap.
A four-step framework for the business case
Step 1: Quantify the cost of unplanned downtime. This is the largest number in the model and the one executives respond to. Build it from components you can defend:
- Revenue or service delivery lost per hour of outage
- Staff time diverted to incident response and recovery
- Customer or contractual penalties triggered by service level breaches
- Recovery cost, emergency callouts, expedited parts, overtime
- Secondary effects, regulatory reporting obligations, remediation, reputational damage
Where you can't get a hard number, present a range and state your assumptions. A defensible range beats false precision.
Step 2: Quantify the avoided cost of the maintenance program. This is the counterfactual, what the program prevents. Two components carry most of the weight:
- Deferred capital replacement. Every year of extended asset life that preventive maintenance buys is a year of avoided capital spend. If a UPS battery string costs a known amount to replace and condition-based servicing extends its life measurably, the annualised saving is straightforward arithmetic.
- Energy and efficiency gains. Cooling optimisation and condition-based battery replacement reduce operating cost, offsetting part of the maintenance spend. Continuous thermal monitoring typically identifies overcooling that can be corrected without risking inlet temperatures.
Step 3: Present the ratio, not the cost. Executives approve ratios and risk reductions, not line items. Frame it as: 'For every dollar of maintenance spend, we avoid X dollars of expected downtime cost and defer Y dollars of capital.' Even a conservative ratio beats a list of service activities.
Step 4: Tie it to obligations you can't opt out of. The Security of Critical Infrastructure Act 2018 (SOCI Act) imposes positive obligations on named asset classes, including data storage and processing. A documented, funded maintenance program with an evidence trail helps demonstrate a risk management program, reframing maintenance from discretionary spend to compliance necessity.
Climate adaptation belongs in the same budget
Higher ambient temperatures, more frequent extreme heat events and increased flood risk raise the load on cooling systems and the probability of utility disruption. These aren't hypothetical, they're already showing up in design conditions and insurance assessments. Practical adaptations include:
- Reviewing cooling capacity against projected design temperatures, not historical ones
- Raising or relocating equipment above flood-prone levels
- Adding on-site generation and fuel storage sized for longer outages
- Reviewing supply chain exposure for spares and consumables, particularly where a single supplier or a single port of entry carries the risk
- Reviewing the maintenance schedule itself against higher expected operating hours during heat events
Each has a cost and a risk-reduction value. Present them as a ranked list with cost and risk reduction side by side, so the board can choose the portfolio rather than approve or reject a single number.
Legacy system integration: the phased path
For legacy system integration, the budget question is different again. Replacing everything at once is rarely affordable, so the pragmatic path is to layer monitoring and maintenance across the existing estate, retire assets at end of life, and phase in new equipment as capital allows. A vendor-agnostic partner makes that work, servicing old and new side by side.
The budget model for a phased approach needs one extra column: the cost of maintaining the legacy asset another year versus accelerating its replacement. Often the maintenance cost is lower than the annualised capital cost of replacement, buying time to plan the upgrade properly rather than reactively.
Treske's UPS battery maintenance and replacement service is one example of that: condition-based servicing that extends the life of what you already own before you commit to replacement.
The strongest budget argument is not 'we need maintenance'. It's 'here is the expected cost of not doing it, here is the cost of doing it, and here is the ratio'. Bring the ratio to the board and the conversation changes.
Frequently Asked Questions
What are the key components of a critical infrastructure maintenance plan?
A comprehensive plan starts with a detailed asset inventory, categorising every piece of operational technology and IT equipment. It should include a risk assessment to prioritise assets based on their impact on business continuity. The plan must then define preventive maintenance schedules, incident response procedures, and clear performance metrics. Crucially, it needs to address regulatory requirements, such as those under the SOCI Act, and outline a strategy for managing both physical and cyber threats.
How does the Security of Critical Infrastructure Act 2018 impact maintenance requirements?
The SOCI Act 2018 requires owners and operators of critical infrastructure to manage both physical and cyber security risks. For maintenance, this means your activities must align with the Act's risk management program requirements. You need to document how maintenance tasks contribute to the resilience of critical assets. This includes ensuring that all work, from routine servicing to emergency repairs, follows strict access control and security protocols to prevent unauthorised interference.
What is the difference between preventive and predictive maintenance for critical facilities?
Preventive maintenance is schedule-based, where tasks are performed at regular intervals, like quarterly UPS battery tests. Predictive maintenance is condition-based, using real-time data from sensors and monitoring software to determine when a component is likely to fail. For example, instead of replacing a part every year, you replace it only when performance data indicates degradation. Predictive strategies, supported by tools like DCIM software, can reduce unnecessary maintenance and prevent unexpected downtime.
How can vendor-agnostic strategies improve infrastructure longevity?
A vendor-agnostic approach means you are not locked into a single manufacturer's ecosystem for maintenance and upgrades. This allows you to select the best-suited components and services for each specific need, regardless of brand. It also prevents the risk of a single vendor's supply chain issues or support limitations affecting your entire facility. By diversifying your maintenance partners and equipment, you build a more resilient and adaptable infrastructure.
What role does cybersecurity play in physical infrastructure maintenance?
Cybersecurity is now inseparable from physical maintenance because operational technology (OT) like power management systems and cooling units are increasingly networked. A vulnerability in a maintenance laptop or a compromised remote access point can be an entry point for a cyber-attack. Best practices require strict network segmentation, access control for all maintenance personnel, and ensuring all connected devices have the latest security patches and configurations.