Load testing and change control are only part of peak readiness. On Black Friday, the systems you hardened in September are still operated by people, and some of the most consequential decisions of the year can land on whoever is holding the pager at 2am.
Their quality depends less on the code than on the arrangements around it. Is the person on call rested? Do they know what they are allowed to do without asking? Do the people on the call know who is in charge? And has anyone planned for what happens to the team once the traffic drops?
In September I argued that peak season is an engineering deadline, not a marketing one: load-test the revenue path, control change, rehearse failure. This is the companion piece. Black Friday 2026 is Friday 27 November, with Cyber Monday on 30 November, and October is when many teams set rotas and agree any freeze dates. It is also the right time to agree the people side.
Why the 2am decision is the one to design for
Design for a tired, stressed responder rather than hoping for a calm one. Google’s SRE book notes in its chapter on being on-call that stress can push engineers away from deliberate reasoning and towards quick, habitual reactions, and that the pressure of an outage can prompt incorrect choices. What it lists to make on-call less daunting is clear escalation paths, well-defined incident-management procedures and a blameless postmortem culture.
All three can be decided in advance, by leaders, in daylight. In my experience, the hardest part of a peak incident is often not the fix itself but the hesitation around it: who decides, who needs to know, and whether it is all right to act.
On-call rotas people can actually sustain
A rota that looks complete on a spreadsheet can still be fragile.
Keep shifts to a human length. For rotations that handle one or more pages a day, the SRE workbook’s on-call chapter recommends limiting shifts to 12 hours, noting that tired people make mistakes and that 24 hours of on-call without reprieve is not sustainable. It suggests splitting a week between a day person and a night person; for a four-day trading weekend, plan that split explicitly rather than leaving it to whoever volunteers.
Name a backup, and protect rest between shifts. A named backup means the primary knows who to contact; PagerDuty’s guidance on being on call is clear that nobody is expected to fix every issue alone. PagerDuty also suggests the backup shift follows the primary so context carries over. Adapt that sequence to your team’s workload: backup availability is a commitment, not automatically protected rest, so a tiring primary shift should not roll straight into another backup shift without considering recovery. Plan overnight cover alongside daytime responsibilities, and reassign daytime work for anyone handling incidents overnight so they can recover.
Assume someone will be ill. As the workbook puts it, “No one can promise on Monday that they won’t have the flu on Thursday.” Agree a swap policy and a named reserve well before peak, not on the Thursday.
Reduce pager noise before peak, not during it. Frequent low-priority alerts cause fatigue, and the SRE book warns that serious alerts can then be treated with less attention than they deserve. October is a sensible time to demote pages that rarely lead to action.
Spread the hard nights. Putting the most experienced engineers on every peak night concentrates both knowledge and exhaustion; pairing them with less experienced colleagues builds capability instead.
Check the employment basics. Agree compensation or time off in lieu in advance; Google describes offering either. Paying someone for on-call does not settle the rest question, though. In the UK, GOV.UK guidance on rest breaks sets out statutory rights including 11 hours’ rest between working days. How on-call interacts with working time can depend on the arrangement and the contract, so involve HR rather than guessing.
Written-down authority to roll back or switch off
At 2am, the question that slows a response is often not “what is wrong?” but “am I allowed to do this?”. If a rollback needs three senior approvals, mitigation may stall while customer impact continues.
Where it is safe, prefer rolling back to rushing a fix forward. The SRE workbook favours “roll back, fix, and roll forward”, noting that even a quick fix needs time to be tested, built and rolled out. It also notes that if a feature flag can disable one feature without affecting anything else, the decision to use it becomes much simpler.
Pete Hodgson’s article on feature toggles describes long-lived “kill switches” that let operators gracefully degrade non-vital functionality under unusually high load, and recommends a human-readable description for each toggle, which matters when the person flipping it did not build it.
The technical switch is only half of it. The other half is written-down authority, agreed with the business before peak: which switches exist and what customers see instead, which roles (not names) may use them, under what conditions, and who must be told afterwards.
For tested mitigations within agreed limits, responders should be able to act and notify without waiting for fresh approval. Actions outside those limits need a named, reachable decision-maker. Switches with commercial consequences, such as pausing a promotion, need the business to agree them knowingly, which is far easier in October than on the night.
Take a hypothetical example. A recommendation service is delaying product pages. The on-call engineer has authority to disable the widget using a tested flag, check that product selection and add-to-basket still work, and notify trading. Changing payment routing requires a separate escalation. With both boundaries written down, the engineer does not need to wake anyone to act on the first, and knows who to call for the second.
Clear incident roles, before anyone needs them
My earlier piece suggested a single incident commander for peak nights. A name on a rota is a start; it works when the people on the call understand that role and the others around it.
Separate coordination from fixing. The SRE book’s chapter on managing incidents opens with an unmanaged incident: the on-call engineer is too absorbed to communicate, executives demand updates, and a colleague’s uncoordinated change makes things worse. It then describes distinct roles, based on the Incident Command System:
- Incident commander holds the overall picture and assigns work.
- Operations applies the fixes, and should be the only group changing the system.
- Communication issues regular updates to responders and stakeholders.
- Planning handles the longer-running work, including handoffs.
PagerDuty’s role definitions are similar, adding a deputy, a scribe, subject matter experts and separate customer and internal liaisons. These are responsibilities, not necessarily separate people permanently on duty. A small response can start with one or two people combining roles, then separate coordination and communications from technical work as the incident grows.
PagerDuty’s incident commander training is clear that the commander coordinates and delegates rather than fixing, and needs no deep technical knowledge. It also gives the commander the final say on the call regardless of day-to-day seniority; a senior colleague who disagrees should take command formally rather than override. Agree its limits in advance: the commander coordinates the response within the agreed incident process and the organisation’s delegated authority; decisions such as changing payment routing still follow their own escalation.
In ecommerce, the communication role often carries more weight than engineering teams expect. Customer service, trading, marketing and fulfilment need one place to learn what is happening and what to tell customers, run by someone who is not also fixing the problem.
Finally, agree when to declare an incident (the SRE book suggests criteria such as customer visibility or needing a second team), how command is explicitly handed over at a shift change, and who is trained to step in.
Recovery is part of the plan
Peak does not end when Cyber Monday traffic subsides. For many retailers, December trading carries on, and the people who covered the long weekend are often expected to pick up the roadmap the following week. Treat recovery as a line in the plan, not a kindness offered if things go well: agree time off in lieu before the event, avoid major launches straight after peak, and tell stakeholders in advance that delivery will be slower that week.
Budget for follow-up work too. The SRE book reports that, in Google’s experience, handling an incident, including analysis, remediation and the postmortem, takes about six hours on average. That is one organisation’s figure, not a staffing benchmark, but a difficult night could create days of work. Capture the timeline while memories are fresh: PagerDuty’s postmortem process schedules the review meeting within three calendar days for its most severe incidents. But do not expect the same exhausted people to start on the fixes immediately.
A checklist to agree with the business before peak
Before peak, I would want each of these agreed and written down:
- Rota: sustainable shift lengths, with 12 hours as the proposed upper limit for this peak plan, informed by Google’s guidance rather than treated as a guarantee; a named primary and backup for every shift; protected rest between shifts; and cover for illness.
- Swaps and escalation: a documented swap policy and an escalation path that reaches a named decision-maker at any hour.
- Pager hygiene: non-actionable alerts removed or demoted before peak.
- Kill switches and rollbacks: for each, a plain description, customer impact, authorised roles, conditions for use and who is told; tested access and permissions; known limitations, including changes that cannot safely be reversed; how to verify the mitigation worked; and who authorises restoration, and under what conditions.
- Commercial switches: business sign-off now, rather than at 2am, for any switch with commercial impact.
- Incident roles: who covers incident commander, operations and communications on each peak shift, whether combined or separate, plus who steps in if one is unavailable.
- Agencies and critical suppliers: confirmed coverage hours and time zones, named contacts and a fallback if the first contact does not respond. Do not assume your own rota gives you out-of-hours access to payment, hosting or platform suppliers.
- Stakeholder channel: one place where trading, marketing and customer service get updates, with an agreed cadence.
- Declaration and handover: agreed criteria for declaring an incident and a script for handing over command.
- Compensation and rest: time off in lieu or pay agreed in advance and checked with HR, kept separate from statutory rest.
- Recovery: protected time after peak, a reduced-delivery week communicated to stakeholders, and a date for the postmortem.
Plan for the people holding the pager
A deploy freeze, where you use one, protects the systems from change. It does not protect the people from fatigue, ambiguity or second-guessing. Those protections come from decisions leaders make in October: who is on, what they may do, who is in charge, and how they recover.
Get those right, and the person holding the pager at 2am has a better chance of making a calm, well-supported call. That is part of the job too, and one more way leaders stay close to the craft.
Currency note: Black Friday and Cyber Monday 2026 dates, Google SRE guidance, PagerDuty incident response documentation, the feature toggles article and GOV.UK rest-break guidance were checked against public sources as of 6 October 2026. Nothing here is legal or HR advice; check current guidance for your own arrangements.
Sources and further reading
- Google SRE book — Being On-Call
- Google SRE book — Managing Incidents
- Google SRE workbook — On-Call
- PagerDuty Incident Response — Being On-Call
- PagerDuty Incident Response — Different Roles
- PagerDuty Incident Response — Incident Commander training
- PagerDuty Incident Response — Postmortem process
- Pete Hodgson on martinfowler.com — Feature Toggles (aka Feature Flags)
- GOV.UK — Rest breaks at work
- Peak season is an engineering deadline, not a marketing one
- Staying close to the craft