A Server Response Crashed Thousands of iOS Apps at Launch. Your SDK Is a Remote Control.
What the 28 September Firebase incident says about code you ship but do not control.
What happened on 28 September
At about 5:41 p.m. Pacific on Monday 28 September 2026, thousands of iPhone apps began crashing in a third-party SDK crash that needed no release to cause. The common factor was the Firebase iOS SDK. According to reporting from 9to5Google and others, Google had sent the SDK an incorrectly formatted payload, and apps using Firebase Analytics fell over when they tried to launch.
Google pushed a server-side fix at about 7:52 p.m. Pacific, roughly two hours and ten minutes later. No app update was needed. But caching on devices kept some apps crashing for up to four hours after the fix, which is how a two-hour server incident became the "two to six hours" that The Pragmatic Engineer newsletter described. Developers reported spikes such as 2,700 crashes in 25 minutes.
The part that stung developers as much as the crash was the silence. As the Pragmatic Engineer Pulse noted, the status page was not updated and no postmortem followed, from a company with a strong incident-management reputation. We are working from public reports here. The precise failure inside the SDK has not been published.
A third-party SDK crash needs no release on your side
Most mobile teams think of crashes as something they introduce: a release goes out, crash-free sessions drop, someone rolls back. This incident had no release. Nothing in the app binary changed. A response from a vendor server changed, and the binary treated it as input it could not survive.
“If a library fetches instructions from a server at startup, the server operator can change your app's behaviour without a release.”
That is a different threat model. Your release process, staged rollouts and App Store review do nothing, because no new code shipped. Rollback does nothing either. You cannot roll back a payload you do not host.
Any SDK that pulls remote configuration at launch has this property. Analytics, crash reporters, A/B testing tools, feature-flag clients, ad SDKs and push providers all do. The question to ask of each is simple: if its server sends garbage, what happens to my process?
Why caching stretched the damage
The timing is worth studying. The server fix landed about two hours in, yet the incident lasted up to six. SDKs cache remote responses so they can start fast and work offline. That is sensible design, but it means a bad response is stored on the device and replayed on every launch until the cache expires or refreshes.
For a crash on launch, this is the worst combination. The app crashes while reading the bad cached value, so it never gets far enough to fetch the corrected one. The user taps the icon, sees it close, taps again, and the loop continues until the cache ages out. Server-side recovery and device-side recovery run on different clocks, and only the first is in the vendor's hands.
Guardrail one: get optional SDKs off the launch path
Analytics has no business value during the first hundred milliseconds of a launch. The user cannot see it, and nothing the app does depends on it. Yet many apps initialise it inside the app delegate, before the first screen, because the vendor quick-start guide says to.
Move optional SDKs to run after the first screen is interactive, on a short delay or after a first user action. You lose a little early-session telemetry. You gain the property that a vendor failure cannot prevent your own UI from appearing. This does not protect against a crash that occurs later on a background thread, because a crashing thread still kills the process. It removes the launch-time loop, which is the failure that locks users out entirely.
Guardrail two: a crash-loop breaker
If the app has crashed on the last three launches, something is wrong, and the safest response is to start without optional dependencies. The idea is old, borrowed from how operating systems offer safe mode. A minimal sketch in Swift:
enum LaunchGuard {
private static let key = "unstableLaunchCount"
private static var defaults: UserDefaults { .standard }
// True after three launches that never reached "stable".
static var skipOptionalSDKs: Bool {
defaults.integer(forKey: key) >= 3
}
// Call first thing in application(_:didFinishLaunchingWithOptions:).
static func launchBegan() {
defaults.set(defaults.integer(forKey: key) + 1, forKey: key)
// If we are still alive 10 seconds in, the launch was fine.
DispatchQueue.main.asyncAfter(deadline: .now() + 10) {
defaults.set(0, forKey: key)
}
}
}
// In the app delegate:
// LaunchGuard.launchBegan()
// if !LaunchGuard.skipOptionalSDKs { startAnalytics() }Two cautions. First, test that the counter write survives a hard crash on your minimum supported iOS version, because the whole mechanism rests on it. Second, decide what "optional" means in writing. Analytics and ad SDKs are optional. Your auth client is not. A breaker that skips the wrong thing turns a crash loop into a broken app that opens.
Guardrail three: a kill switch that does not depend on the vendor
Once you detect trouble, you want to disable an SDK for all users within minutes, without a release. That requires a remote flag, and the flag must come from somewhere other than the vendor you are switching off. A kill switch served by the same Firebase backend that is failing is a kill switch that fails together with the thing it controls.
Host the flag on infrastructure you operate or on a different provider: a small JSON file on your own CDN, fetched with a short timeout and a safe default. Cache it, but give the cache a short lifetime and a default of "SDK enabled" only if the last fetch succeeded recently. Remember the caching lesson from this incident: your flag cache can also hold a bad value, so keep its parsing defensive.
| Dependency type | What the server controls | Containment |
|---|---|---|
| Analytics / crash reporting | Sampling and config payloads | Deferred init, crash-loop breaker |
| Feature flags / A/B tests | Which code paths run | Last-known-good cache, local default values |
| Ad and attribution SDKs | Creative and config | Deferred init, kill switch on your own host |
| Push and messaging | Token and topic config | Init after first screen, failure isolation |
| Auth and payments | Critical session state | Not optional: vendor evaluation and fallback plan |
Build your own signal, because the status page may not come
Developers in this incident learned about the problem from their own crash dashboards and from each other, not from the vendor status page. That is likely to repeat. Status pages are written by humans who are also fighting the incident.
Two cheap signals help. Alert on crash-free-session rate dropping by a fixed margin within a window, not just on absolute count, so a launch crash across an unchanged release gets noticed. And tag crashes by the top stack frame's module, so a spike in a single vendor's frames is visible within minutes. If two dozen engineers across unrelated companies report the same stack at once, the cause is rarely in any one of their apps.
- Alert on crash-free-session drop with no release in flight.
- Group crashes by vendor module, not only by exception type.
- Keep a one-page runbook per optional SDK: how to disable it, who can flip the flag, how long the cache lives.
What to do this week
List every SDK that initialises before your first screen. For each, write down what its server can change, whether the app survives a malformed response, and how you would turn it off without a release. Most teams will find one or two entries with no good answer. Those are the ones to fix before the next vendor incident, which will arrive without a release, a warning or a postmortem to learn from.
Frequently asked questions
Related reading
Idempotency Keys Passed Code Review in 9 of 12 Payment Systems. All Nine Still Duplicated Writes.
A 2026 audit found idempotency logic that passed code review and unit tests still let duplicate writes through in nine of twelve payment systems. Six failure modes, and the fix for each.
Clerk’s Failover Didn’t Fire Because Postgres Was “Technically Still Online.” That’s a Design Category, Not a Fluke.
On February 19, 2026, Clerk’s session failover didn’t trigger because Postgres was “technically still online” — degraded enough that queries returning 200 took minutes instead of milliseconds.
Four AI Coding Agent Exploits Landed in Two Weeks. The Sandbox Boundary Failed in All Four.
Three separate AI coding agent compromises disclosed inside two weeks share a root cause that has nothing to do with tricking the model, and everything to do with what happens after it decides to act.