Series: Lessons From The Field: 01
In a signalling equipment room, a small failure can look ordinary. A card comes out of a rack. The fault is found. The affected functions move to a restrictive state. The system has done the safe thing.
The engineer’s problem is not that hardware failed. Hardware fails. The problem begins when recovery depends on a module that is no longer supported, a programmable device that is no longer made, an engineering tool that will not run on a supported workstation, and a supplier that no longer holds the design knowledge.
At that point, the failed card is a safety assurance issue, a service availability issue and a cyber recovery issue in the same cabinet. Long-life rail assets keep running long after the support system around them has aged away.
The rail example is not theoretical
Network Rail, the public infrastructure manager for the UK mainline railway, in its Control Period 7 digital signalling plan, says a significant proportion of signalling assets are reaching life expiry and warns of an “undeliverable bow wave from the 2030s” if conventional renewal continues as before.
The Office of Rail and Road, Britain’s independent rail economic and safety regulator, adds a harder lesson. Its independent review found no evidence that Network Rail satisfied the requirements of an obsolescence management policy, no obsolescence management plans in place, and no single accountable owner for implementation.
Even in mature rail organisations, obsolescence sits between functions. Asset teams see spares and renewals. Safety teams see assurance and change control. Cyber teams see unsupported software, weak recovery paths and old interfaces. Finance sees another expensive renewal request. Each view is valid. None is complete.
The lesson
Obsolescence is not age. It is the slow disappearance of credible choices.
The asset still runs. The safety case still stands. The service still moves. But every unsupported tool, untested spare and missing recovery route leaves the railway with fewer safe ways back. Old systems become risky when confidence outlives evidence.
That is the deepest mismatch in rail OT. A signalling location case, control system, depot system or train fleet may be expected to operate for decades. Its operating systems, firmware branches, engineering laptops, cryptographic libraries and network equipment may become unsupported much earlier. When support ends, the operator has not been breached. But the trusted route to remediate defects has narrowed. Recovery evidence matters more, and residual risk must be visible to the people who own safety, service and funding.
Safe is not the same as running
A common mistake in cyber conversations about rail is to treat “fail safe” as if it means “nothing bad can happen.” Rail engineers know better.
Fail safe means an expected failure maintains or places equipment in a safe state. It does not promise that every fault is detected, that every degraded condition is operationally tolerable, or that every safe state preserves train service. A restrictive state protects passengers while surrendering availability. From the passenger, operator and maintainer’s point of view, “safe and stopped” is still a service problem.
The cyber lesson is similar. A firewall, jump host, endpoint agent, patch or monitoring probe that is fine on an office network is not automatically fine in signalling, control or rolling stock environments. Every defensive measure has to be judged against safety, availability, maintainability and recovery.
The budget decision is also a cyber decision
Obsolescence persists because deferral often looks rational in isolation. A renewal is expensive. Possession time is scarce. The current system is still running. The safety case still exists. The programme has other priorities.
Sometimes deferral is the right call. Replacing old assets too quickly can create new hazards and assurance burdens. But the decision has to be scored honestly. Extending the life of a rail OT asset is also a decision about supportability, recovery confidence, supplier dependence, skills retention and future assurance. If the obsolescence consequence is counted only in the asset plan and not in the cyber risk register, the railway is fooling itself.
The most dangerous cases are rarely the ones where everyone sees the risk and chooses to carry it. They are the ones where no one realised a choice had already been made. No one decided to lose the toolchain; it just stayed on an old laptop until the laptop failed. No one decided to make recovery impossible; the backups were simply never restored on a bench.
What good looks like
The answer is not to replace everything old. That would be unaffordable and, in some cases, operationally reckless. The answer is to manage obsolescence as a targeted programme where safety, availability and cyber recovery are assessed together.
Start with one verified lifecycle and cyber inventory for the assets that matter: hardware, firmware, operating systems, engineering tools, cryptographic mechanisms, licences, spares, backups, remote access, suppliers and competent people. Treat “unknown” as a risk state, not a blank cell.
Score supportability risk, not just age. The urgent asset is where criticality, exposure, failure likelihood, lack of spares, weak competence, poor monitoring, supplier fragility and unproven recovery meet.
Choose a treatment route deliberately: sustain with evidence, extend support by contract, isolate with engineered controls, monitor with triggers, migrate, replace or retire. “Do nothing” is acceptable only when it is recorded, owned, reviewed and accepted.
Prove recovery before you need it. Restore the backup. Rebuild the workstation. Reflash the controller. Validate the licences. Prove the spare is what the label says. A cyber incident, a failed board and a mistaken configuration change all become less frightening when the recovery route has already been rehearsed.
Procurement must also preserve future choices. Long-life rail contracts should require lifecycle transparency, vulnerability handling, security update routes, compatibility evidence, spare strategy and controlled access to critical configuration or design data.
The Monday test
The first useful test is not a six-month assessment. It is four questions.
For the most safety-critical assets, do you know the end-of-support date for every operating system, firmware branch, engineering tool and cryptographic mechanism, not just the hardware warranty?
If a stubborn legacy controller failed tonight, do you hold a spare you could actually trust, and could you prove it?
Have you restored the critical configurations, rebuilt the engineering workstation and reflashed the controller on a bench, using only the tools, licences and credentials the real maintenance team would have during an incident?
Does one named person own the lifecycle risk across spares, safety assurance, cyber security and funding, or is it split between departments with no single decision point?
If those answers are uncomfortable, that is the finding. The railway may not yet have a technology failure, but it has a management gap. That is cheaper to fix than a service-affecting incident.
Lesson, again
The engineer holding the failed card should never be the first person to discover that the design data is gone, the spare cannot be trusted, the support route has closed, and recovery has never been tested. Those facts are knowable years in advance. They belong in a register, on a roadmap and in a funded decision.
Obsolescence is not age. It is the slow disappearance of credible choices. Obsolescence has already arrived. The question is whether the railway is reading the countdown while there is still time to choose.