Network Intelligence

Why uptime alone is not an SLA strategy

A practical framework for measuring service quality beyond a single availability percentage.

The Problem with a Single Number

99.9% uptime. Sounds impressive. Your vendor sends a monthly report with that number highlighted in green, and everyone moves on. But your operations team remembers the Tuesday afternoon when 23 sites went dark for 11 minutes. Your logistics client missed a critical dispatch window. The SLA report said: compliant.

Uptime percentage hides three things that actually matter:

1. When did it fail?

A link that drops at 3 AM on Sunday costs you almost nothing. The same link dropping at 9 AM Monday during peak order processing could cost you lakhs. Uptime percentage treats both identically.

2. How often did it flap?

A link that goes UP-DOWN-UP-DOWN 80 times in a day but never stays down for 9 minutes shows as 100% uptime in most systems. Your applications experienced 80 reconnection events. Your users experienced 80 disruptions. Your uptime report said: healthy.

3. Which vendor is responsible?

When you have 38 vendors across 154 links, aggregate uptime tells you nothing about accountability. One vendor could be responsible for 70% of your downtime while the overall number stays green.

A Practical SLA Framework

Instead of a single uptime percentage, measure four things:

Availability — Traditional uptime, segmented by business hours vs off-hours. A breach during business hours should carry 3x weight.

Stability — Flap count per link per day. A link with more than 15 flaps in 24 hours is unstable even if it never crosses your outage threshold. Chronic instability — links that flap on 5 or more days out of 30 — is a separate category entirely.

Recovery — Mean Time to Restore (MTTR). Two vendors can have identical availability numbers but one restores in 18 minutes and the other takes 4 hours. That difference is your exposure.

Accountability — Per-vendor scoring that weights severity, frequency, and recovery time together. This is the number you bring to vendor review meetings.

What Chronic Flappers Actually Cost You

In a real ISP deployment monitoring 154 WAN links, we found that 18 links (roughly 12%) qualified as chronic flappers — flapping on 5 or more days with 75 or more total flap events over 30 days. None of these links showed as "down" in uptime reports.

But each flap event represents a TCP session reset, an application reconnect attempt, a potential transaction failure, and a support ticket waiting to happen. Across 18 chronic links generating around 80 flaps per day each, that is 1,440 disruption events per day that uptime reports call healthy.

Building Your SLA Scorecard

A defensible SLA scorecard has three layers:

Layer 1 — Link Health Score: Per link, per month — Availability percentage plus Stability Index (inverse of normalized flap count) plus MTTR score. Weighted average gives you a single link health number.

Layer 2 — Vendor Score: Aggregate link health scores by vendor, weighted by number of links and criticality. This gives you a ranked vendor table — who is dragging your network down and by how much.

Layer 3 — Business Impact Score: Weight outages and flaps by time-of-day and link criticality. A logistics company's order management links should carry 5x the weight of a staff internet link.

The Operational Shift

The goal is not to produce a more complicated report. The goal is to shift from reactive firefighting to predictive intervention. When your monitoring system flags a link as a repeat offender before it becomes a critical failure, you have a conversation with the vendor backed by evidence — not after a major incident, but before one.

When your system identifies links that are both chronic flappers and repeat offenders, those links go on a watchlist. You schedule proactive maintenance. You brief your on-call team. You do not wait for the 2 AM escalation.

Conclusion

Uptime percentage is what vendors report. SLA intelligence is what you need to run a network. The difference is not complexity — it is visibility. Know when, know how often, know which vendor, and know the business impact. That is a strategy.

Request demo