Graceful Degradation Under Server Overload


GitHub just had another major outage, the latest in a string of incidents this month, and the pattern behind most of them is the same: more load hits the system than it was built to absorb, and things that were fine on a normal day stop being fine. Retries pile onto services that are already struggling, and autoscaling doesn’t kick in fast enough or in the right place to save it.

It happens to platforms with some of the best infrastructure teams around, so it’s worth actually working through. Let’s see what we can do about it.

The First Version

Accept a request, do the work, send back a response. No limits, no concept of capacity. Fine under normal load, which was never the problem.

What Happens When Traffic Spikes?

Nothing in our server says no. Every request gets accepted, and if we can’t keep up, they just pile up in memory waiting for a worker to free up. Memory grows, latency grows with it, and eventually we either run out of memory and crash, or everything gets so slow that every request times out anyway. Nobody gets a clean error. Everybody, including users who showed up before the spike, gets a slow eventual failure.

Bounding the Queue

So we cap concurrency with a limiter — only N requests processed at once, the rest wait in a bounded queue, and once that’s full too, we return a 503 immediately instead of accepting work we know we can’t reach. Memory now has a ceiling, and some users get a fast honest error instead of everyone getting a slow one.

Picking the Limit

Little’s Law gives us a starting point: the number of requests in flight is throughput times latency. A service handling 500 requests a second that spends 40 milliseconds on each one needs about 20 in flight to keep up. Measure both under a load test and we have a number to work from.

The catch is that the number rots. We derived it from a latency we measured on a good day, and the moment a dependency slows down, that same limit admits far more work than we can actually hold — which is exactly the situation it was meant to protect us from. Adaptive limits solve this by not being a constant at all. We watch how long requests are sitting in the queue, compare that to the lowest we’ve seen recently, and shrink the limit when it starts climbing.

Shedding on the Way Out

The cap doesn’t fix everything. A request that’s been queued for eight seconds gets pulled off, worked on, and answered — except the client gave up after five and already closed the connection. We just paid CPU and a database round trip for nobody. Worse, doing that stale work delays the request behind it, making it more likely to also go stale by the time we reach it.

Fix: attach a deadline to everything the moment it queues. When a worker is free, the first check is whether the deadline already passed — if so, drop it without doing the real work. Do this as early as possible, ideally before parsing the body or hitting auth, so we’re not paying most of the request’s cost on work we’re about to throw away anyway.

Passing the Deadline Down

A deadline that stops at our process only solves half of it. If we call another service with 40 milliseconds left on the clock, it has no idea, and it will happily start work we’ve already decided we can’t wait for. So we send the remaining budget along with the call. Now the service on the other end can run the same check we do — is there enough time left for this to be useful? — and refuse immediately if there isn’t.

Each timeout has to fit inside whatever is left of the budget at the moment we make the call. And if we intend to retry, the retry has to fit in there too, or we’re spending load on a response that arrives after the caller already gave up.

Not All Requests Are Equal

Shedding evenly means checkout and a background analytics ping get treated the same. Split traffic into a few priority classes and shed from the bottom one first, so when we do have to say no, it’s not a random draw across everyone hitting us at that moment.

Saying no is one tool in out disposal. We can also return a partial response, skip the “you might also like” section or return a stale cache.

When the Problem Isn’t Us

Our server can behave perfectly and still go down because the database is slow tonight. Without a timeout on that call, every worker touching it sits there holding a connection, and once enough workers are stuck, requests that don’t even touch the database start queuing behind ones that do.

Timeouts and Circuit Breakers

Every outbound call gets its own timeout, well under our own request deadline. On top of that, a circuit breaker watches the failure rate and trips open once it crosses a threshold — we stop calling the dependency entirely, fail fast, and periodically test with a small trickle of real traffic to see if it’s recovered.

What Is a Circuit Breaker?

We wrap the call — a database query, a request to another service — in a counter that tracks how many of the recent ones failed. Once that share crosses a threshold we picked, we take it as an outage downstream and the circuit breaks. We stop making the call and fail immediately instead.

Our workers stop waiting around for an answer we can already guess. The database gets breathing room, which is usually what it needs to recover. Every so often we let a single real request through, and if it comes back healthy we go back to calling normally.

This is the part that actually sank GitHub. Clients kept retrying against a struggling service, the retries added more load than the original traffic did, and the service never got the room to recover. A circuit breaker without sane retry and backoff behavior on the client side just builds a bigger storm.

The Retry Storm

Retries are decided by our callers, and their defaults are usually generous.

Retry Amplification

Say the browser retries three times, the gateway it calls retries three times, and the service behind that gateway retries three times as well. One user action, one failure at the bottom, and the service that was already struggling sees up to 27 attempts. Layers don’t add up, they multiply, and every team along the way made a choice that looked reasonable on its own.

So retries belong at one layer — the one that knows whether the operation is safe to repeat — and everything above it should fail fast and pass the error up.

Retrying at a single layer still isn’t enough if every client retries at the same instant. They all failed at the same moment, so if they all wait two seconds, they all come back together and the recovery attempt is just the next spike. Exponential backoff spaces out the attempts of one client. Jitter, a random offset added to every wait, is what keeps separate clients from lining up with each other. It’s the half that gets skipped, and it’s the half that matters when thousands of clients are failing at once.

Backoff spreads retries over time but doesn’t reduce how many of them there are. When everything is failing, everyone is still retrying, just more politely. A retry budget puts a ceiling on it: we allow retries only up to some share of normal traffic, ten percent is a common starting point, and once we’re past that we stop retrying and fail immediately. Amplification is then bounded no matter how bad the outage gets.

None of this applies to work that isn’t safe to repeat. Retrying a read costs us a little extra load. Retrying a checkout that timed out after the charge already went through costs a customer twice. Either the operation carries an idempotency key, so the second attempt is recognized as the same one, or we don’t retry it at all.

And there’s our side of the contract. When we shed, we should say so. A 503 with a Retry-After header tells the client we turned it away deliberately and roughly when to come back. Without that, a shed request is indistinguishable from a broken one, and the sensible response to a broken one is to try again right now — which is precisely what we can’t afford.

Why the Outage Outlives the Spike

Traffic returns to normal and the system stays down. If the spike is what broke it, why doesn’t it recover when the spike is over?

Because by then the load isn’t coming from users anymore. Queued work, clients retrying, caches that emptied while we were failing and now pass every request straight through to the database — the broken state generates enough load to keep itself broken. Normal traffic is no longer survivable, because what we’re actually carrying is normal traffic plus everything the outage created. This is metastable failure: a system that had comfortable headroom at this exact traffic level an hour ago, and can’t handle it now.

Which means we can’t recover by waiting. Returning to normal load returns us to the level we’re already failing at. We have to go below it — shed hard, let the queues drain, give the dependencies room to catch up, and then bring traffic back gradually instead of opening the gates at once.

The instinct to restart everything is worth resisting for the same reason. It clears the bad state, but it also empties every cache, and the first wave of traffic after the restart goes straight through to the database that was the problem to begin with.

Autoscaling Won’t Save Us

The obvious objection to all of this is that we should just add servers. Usually we can, and usually it helps less than we expect.

Autoscaling has to be triggered by something, and that something is usually CPU. But a worker blocked on a slow database is barely using any CPU — it’s sitting idle holding a connection. The incident we would most like to scale our way out of is the one where our scaling metric looks perfectly healthy. Queue depth, latency, or rejection rate would catch it. CPU won’t.

Even when the trigger does fire, we’re on the wrong timescale. A new instance has to be provisioned, booted, connected, and warmed before it can take real traffic, and that is minutes. We’re making shedding decisions in milliseconds. Most of what was going to fail has already failed by the time the capacity shows up.

And when the bottleneck is downstream, scaling actively hurts. More instances means more concurrent connections pointed at the database that was the reason we were overloaded in the first place. We’ve scaled the part that was coping and added pressure to the part that wasn’t.

Summary

An unbounded server degrades for everyone the moment it’s overloaded. Capping concurrency contains that, but a capped queue still wastes work on requests nobody’s waiting for anymore unless we shed by deadline, not just by depth. A well-behaved server can still go down from a slow dependency it forgot to time out, or retries that turn a hiccup into a storm. Autoscaling saves us on the timescale of minutes, and only if it’s watching the right signal — shedding is what keeps us alive on the timescale of seconds until it arrives.

Graceful degradation was never about handling more load. It’s about making sure that when we can’t, the failure is small and honest instead of a multi-hour outage with a status page full of “investigating.”