Design Degraded Modes Before the Next Incident
Most reliability plans focus on keeping every capability available. That is an admirable goal, but it is not a complete operating strategy. Complex systems will eventually lose a dependency, exhaust a resource, or encounter demand they cannot serve normally. When that happens, teams need more than alerts and runbooks. They need an agreed way for the service to become less capable without becoming completely unavailable. Degradation Is a Product Decision A degraded mode defines which capabilities can be limited, delayed, simplified, or disabled to protect the service’s most important outcomes. An online service might pause recommendations while preserving core transactions. An internal developer platform might delay nonessential analytics while continuing deployments. A data pipeline might prioritize current operational events over historical backfills. Engineering cannot make these tradeoffs alone during an incident. Product, operations, security, and business leaders need to deci...