Network Resilience
The architecture behind a network that stays up.
The architecture behind a network that stays up.
By Rob Schrage, Director of Sales Engineering
Early in my career at Reuters, I worked on a foreign exchange trading platform called Dealing 2000. We were nominated for a quality award. Black-tie event. I was probably 24, and we were sure we were going to win. We didn’t. The award went to a Reuters journalist who had been shot while covering the U.S. invasion of Panama. I remember thinking, “Now I know what importance is.” That stayed with me. Somewhere at the other end of the technology is a person who needs it to work: a clinician waiting for an image, a dispatcher trying to understand a caller, a journalist trying to get a story out from a dangerous place. That’s why this particular topic, network brownouts, is especially important to me.
A clean outage is obvious. A circuit drops, monitoring lights up, a ticket opens and everyone knows what happened. The RFO gets logged and the team gets moving.
A network brownout is different. The service is still technically available, but the person relying on it may already be losing the ability to do what they need to do.
A network brownout is partial degradation. Latency climbs, packet loss builds, sessions drop or retransmit, but the condition may never cross the threshold that monitoring treats as an outage. Internally, everything can still look green while the person on the other end cannot reliably complete the task.
Monitoring can tell you the circuit is up. It cannot, by itself, tell you whether a radiology image arrived when it was needed, whether a 911 caller could be understood or whether a dispatcher got enough information to send help.
What degradation affects depends on the task and what happens when that task can’t be completed.
Healthcare
Healthcare is where this becomes impossible to dismiss as a performance metric. A delayed image or clinical record may mean someone is waiting for information before a decision can be made.
Imagine someone sitting in a physician’s office waiting to hear what an image means for a possible cancer diagnosis. The image is moving and the connection is up, but suppose retransmissions add five minutes. On a dashboard, that may look like degraded performance. To that person, it is five more minutes of not knowing.
Public safety / 911
Public safety is the same principle with even less margin for error. A person dialing 911 needs to reach the right emergency communications center, be understood clearly and get enough information through for help to be sent.
I have designed radio networks, and I do not think about a garbled transmission as a voice-quality issue. In an IP-based public safety environment, packet loss, jitter or partial degradation can interfere with voice quality, call handling and the applications dispatchers depend on. The person on the other end may be reporting a fire, a crash or a medical emergency, and the words that do not get through can change what happens next.
The same is true between an EMS crew and a hospital. If critical information is delayed or misunderstood, a clinician may be making a decision with incomplete information. Whether the links are technically available can be beside the point.
After a network brownout, the immediate job is to restore service. The team finds the stressed component, logs it in the RFO and makes a fix: add capacity, tune an alert, restart a service or change a threshold.
Sometimes that is exactly the right fix. And sometimes the risk is assuming the first visible symptom was the whole cause.
If the underlying issue is a shared dependency, an untested failover condition or monitoring that can see component status but not the user task, fixing the symptom does not prove the condition is gone.
Closing the ticket is progress. The harder question is whether the same failure could happen again the next time conditions degrade.
Redundancy and testing have to be judged against the critical task itself. For example, a second circuit can still share a dependency; failover can exist on paper and still not switch the way you expect under a partial failure; or monitoring can stay green while the user task degrades. You do not want the first real test of any of those safeguards to happen during an emergency.
That is what the infographic below examines: three common resilience moves and the gaps they can leave behind.
Download the Infographic: Redundancy vs. Resilience