Why Vendor Reviews Fail
Most carrier review meetings follow a predictable pattern. The ISP or enterprise brings a list of complaints. The carrier brings a PDF showing 99.7% uptime. Both sides argue about whose numbers are correct. Nothing changes. The meeting happens again next quarter.
The problem is not that vendors are dishonest — most are not. The problem is that both sides are measuring different things. Uptime percentage, measured differently, can produce genuinely different numbers from the same network events. Without a shared, evidence-based scorecard, vendor reviews produce friction instead of improvement.
What a Defensible Scorecard Measures
A defensible carrier scorecard measures four dimensions that vendors cannot easily dispute because they are derived from your own monitoring data, not theirs:
Availability Score — Your measured uptime percentage, calculated from your monitoring system's outage records. Not the vendor's reported uptime — yours. Segmented by business hours and off-hours, because a 3 AM failure and a 10 AM failure are not equivalent.
Stability Score — Flap frequency normalized per link per day. A vendor with five links averaging 40 flaps per day is less stable than one with 20 links averaging 8 flaps per day, even if their availability numbers are identical.
Recovery Score — Mean Time to Restore across all outages in the review period. Vendors who restore in 22 minutes should be scored differently from vendors who take 3 hours, even if both had the same number of outages.
Trend Score — Is performance improving or degrading? A vendor whose MTTR has dropped from 90 minutes to 25 minutes over 6 months deserves credit for that trajectory, even if their absolute numbers are not best-in-class yet.
The Problem Score Formula
Combining these four dimensions into a single vendor problem score allows you to rank all your carriers on one axis. A practical weighting: Availability 40%, Stability 25%, Recovery 25%, Trend 10%.
In a real deployment with 38 vendors across 154 links, this scoring reliably separates carriers into three tiers. The top tier (roughly 30% of vendors) delivers consistent performance and requires only routine review. The middle tier needs monitoring but not immediate escalation. The bottom tier — typically 3 to 5 vendors — accounts for a disproportionate share of total network incidents and requires active performance improvement plans.
Making It Defensible
The scorecard is only as defensible as the data behind it. Three practices make it hard to dispute:
Consistent methodology — Use the same outage definition (duration threshold, packet loss threshold) for every vendor, every month. If you change methodology, document it and restate prior periods.
Raw event access — Every score should be traceable to specific outage events with timestamps. When a vendor disputes their score, you can show them the exact 47 events that contributed to it, with start time, end time, and packet loss percentage.
Symmetric disclosure — Share the scorecard methodology with vendors before the first review. Surprise scorecards produce defensiveness. Shared methodology produces accountability.
Using the Scorecard Operationally
A vendor scorecard used only in quarterly reviews is a reporting tool. A vendor scorecard used operationally is a management tool. The difference is frequency and action threshold.
Operationally, vendor scores should update monthly. Any vendor whose score drops more than 15 points month-over-month should trigger an automatic review request — not waiting for the quarterly meeting. Any vendor whose worst link qualifies as both a chronic flapper and a repeat offender should receive a formal performance improvement notice.
Conclusion
Vendor accountability starts with measurement. Not the vendor's measurement — yours. A scorecard built on your monitoring data, applied consistently, shared transparently, and acted on promptly transforms vendor reviews from quarterly arguments into continuous improvement cycles. That is what defensible looks like.