Field note
Hotfix or rollback after a crash spike
When crash volume jumps overnight, the first decision is recovery path. Here is a calm way to choose.
Confirm the spike is real and tied to the new version. Compare absolute crash counts and crash-free session rate against the previous stable build on a similar traffic day.
Identify whether one signature dominates. A single new family argues for a targeted hotfix if the fix is known and small. Many unrelated signatures after a large refactor lean toward rollback while you regroup.
Check store review lag and staged rollout percentage. A partial rollout that can pause is different from a full release already saturating your user base.
Write down the recovery choice and the metric that will reopen the gate. Teams under pressure forget why they rolled back; a short note keeps the next attempt disciplined.
After recovery, schedule a short post-incident triage so the same signature does not return in the following release. The triage brief becomes the bridge between crisis and the next assessment.