A release runs into trouble late in the evening. The engineer who understands the affected service stays online, identifies a workaround, and gets the application running again. The next morning, the team is relieved and the engineer returns to a full calendar of meetings. The incident is over, but the work needed to recover from it has not appeared anywhere in the schedule.
Occasional emergencies are part of operating software. The difficulty comes when this arrangement becomes routine, and the same people repeatedly absorb the gap between what the organisation promises and what its working hours can support. Their effort keeps delivery moving while making the underlying capacity problem less visible.
Fatigue Belongs in the Engineering Plan
The UK's Health and Safety Executive describes fatigue as a risk influenced by factors including workload, working hours, and opportunities for rest. Its guidance links fatigue with difficulties in attention, information processing, and decision-making. These are relevant to engineering work, although they do not establish that a particular number of hours will produce a particular software defect.
Software quality has many causes. Requirements, experience, review practices, tooling, and the design of the system all matter. Fatigue adds a condition under which those resources have to be used. Treating it seriously means considering whether people have a reasonable opportunity to do the work well, alongside checking whether the necessary technical safeguards exist.
What Extra Hours Can Conceal
Overtime can make an overloaded plan look achievable. A team completes the release, but only because testing moved into the evening and someone spent the weekend resolving integration problems. If those hours disappear from the review, the next estimate starts from an inaccurate account of what delivery required.
There may be deferred costs as well. Documentation, investigation of an intermittent fault, or replacement of a temporary workaround can be postponed to meet the immediate deadline. These decisions are sometimes justified, but their consequences need to remain visible. A successful release does not automatically settle the work left behind it.
The useful response is to record the trade-off and make an explicit decision about scope, timing, or capacity. Asking people to manage their time better will not resolve a schedule that consistently requires more work than the available hours allow.
Why Stopping Can Be Difficult
An engineer may stay online because they believe handing over would put the service at risk. They may be the only person familiar with a component, or the only one with the necessary access. In that situation, an instruction to get some rest leaves the operational dependency unresolved.
A viable handover requires someone able to receive it and enough information for that person to continue. The record should explain the current impact, the actions already taken, the evidence collected, and any changes that need to be reversed if the situation worsens. It should also make uncertainty explicit, so the next responder can distinguish a confirmed finding from a working theory.
The incoming person needs time to review that account and ask questions. Access, runbooks, and a clear escalation route should be in place before the handover becomes urgent. These arrangements are easier to build during ordinary work than in the middle of an extended incident.
Recurring dependence on one individual is therefore something to investigate. Pairing, documentation, and broader operational knowledge can make time off more practical while also reducing the risk of an unexpected absence. The organisation benefits from being able to continue without requiring the same person to remain available indefinitely.
Making Recovery a Scheduling Decision
After an overnight incident, the following day's plan needs review. Meetings may need to move, work may need to be reassigned, and a deadline may need to change. Leaving every commitment in place while encouraging someone to take it easy sends conflicting instructions.
A practical approach can include:
- Arrange operational cover before assigning on-call responsibilities.
- Review the next day's workload after substantial out-of-hours work.
- Make it clear how someone can report that they are too fatigued to continue safely.
- Track repeated overnight interruptions and address the services or staffing arrangements behind them.
These practices need to fit the work and the people doing it. There is no universal working week that guarantees sound architecture, and a shorter schedule can still be unreasonable if the same demands are compressed into it. The HSE's guidance on screen-work breaks similarly emphasises the nature and mix of demands rather than a single break pattern for every job.
Keeping Technical Safeguards in Place
Rest does not replace testing or review. A well-rested engineer can misunderstand a requirement or make an incorrect assumption. Equally, an automated test suite cannot establish that someone has considered every consequence of a rushed operational change. Reliable work needs several forms of protection.
During an extended incident, keeping changes small and recording what has been tried can make the situation easier to follow. Where practical, another qualified person should review consequential actions. Known rollback procedures and clear stopping points help the team avoid expanding the scope of an urgent fix simply because it has already invested hours in the investigation.
The aim is to reduce the amount that any one person has to remember or judge alone. Those practices are useful under ordinary conditions and become particularly valuable when the team is working under pressure.
What the Organisation Rewards
A team learns from the behaviour that receives recognition. Thanking someone for an emergency response is reasonable. Repeatedly celebrating their willingness to miss evenings while ignoring the work that would prevent those emergencies establishes a different expectation.
Preventive work deserves visibility too: a clearer runbook, a quieter on-call shift, a removed dependency, or a handover that lets a colleague leave on time. These improvements rarely have the drama of a late-night recovery, but they change the conditions under which the next incident will be handled.
Rest becomes part of engineering when the organisation plans for it and absorbs the consequences of making it possible. That means honest estimates, workable coverage, and adjustments after exceptional demands. A team should be able to deliver reliable software without making persistent exhaustion a condition of belonging to it.