opsmtrs.com
Pagerly
Directly manage and resolve operational incidents from Slack, streamlining the response process and improving efficiency.
#oncall is an active hashtag on Bluesky. In the last 30 days, 8 people shared 37 posts with it — around 1 a day. Activity is up 14% versus the previous week, peaking on Jul 16 with 4 posts.
Tags most often used together with #oncall.
opsmtrs.com
Pagerly
Directly manage and resolve operational incidents from Slack, streamlining the response process and improving efficiency.
opsmtrs.com
Pagerly
Directly manage and resolve operational incidents from Slack, streamlining the response process and improving efficiency.
opsmtrs.com
Pagerly
Directly manage and resolve operational incidents from Slack, streamlining the response process and improving efficiency.
opsmtrs.com
Pagerly
Directly manage and resolve operational incidents from Slack, streamlining the response process and improving efficiency.
opsmtrs.com
Pagerly
Directly manage and resolve operational incidents from Slack, streamlining the response process and improving efficiency.
opsmtrs.com
Pagerly
Directly manage and resolve operational incidents from Slack, streamlining the response process and improving efficiency.
opsmtrs.com
Pagerly
Directly manage and resolve operational incidents from Slack, streamlining the response process and improving efficiency.
greatcircle.com
Respecting fatigue isn’t coddling
Is it coddling when an on-call engineer takes the next morning off to recover after handling a production incident at 3 a.m., or is it a smart company managing a reliability risk? Here’s what that night actually looks like. The engineer gets paged at 3 a.m., then spends two hours diagnosing the problem, coordinating with fellow responders, and restoring service. By 5 a.m., the incident is resolved and they get back to bed, but it takes them a while to settle down and get back to sleep. Four hours later, they’re at standup. That afternoon, they’re in a planning meeting. That night, they’re still primary on the pager. This is the default at most companies. Nobody made a deliberate decision that it should work this way; it’s just what happens when there’s no explicit policy for post-incident recovery. And it carries more risk than most leaders realize. ## Incident response is more fatiguing than regular work Responding to an incident isn’t like a normal day of developing features and chasing bug reports. The cognitive demands are qualitatively different: rapid context-switching under time pressure, high-stakes decisions with incomplete information, coordinating across multiple people and systems, all while knowing that users are affected and stakeholders are watching. And there’s a physiological dimension that regular engineering work rarely triggers: adrenaline. Incident response activates the body’s stress response in a way that writing code or reviewing a design doesn’t. That heightened state feels productive in the moment, but it depletes reserves fast, and the crash afterward is steeper than the apparent effort would justify. This, incidentally, is one of the reasons that training and practice matter so much. Responders who’ve rehearsed the process and trust the framework around them experience a less intense stress response when real incidents hit. Turning incident response into a “routine emergency” doesn’t just improve efficiency; it reduces the physiological toll. An engineer who’s been actively responding for a few hours isn’t just tired in the way that a long day makes you tired. They’re measurably less effective at exactly the skills incident response demands: integrating new information, evaluating competing hypotheses, making decisions under ambiguity, and recognizing when a current approach isn’t working. ## The degradation is predictable Responder fatigue follows a recognizable pattern. As it sets in, people stop processing new information as effectively. They agonize over decisions they’d normally make quickly, or they stop making decisions altogether. They develop tunnel vision, fixating on the one theory they’re already pursuing instead of stepping back to consider alternatives. They fall into ruts, essentially pursuing “Plan A, again, with more feeling this time” instead of asking whether Plan A is still the right plan. They get less creative, more rigid, and more prone to mistakes. Fire departments study this, because it’s exactly the scenario firefighters face: interrupted sleep from overnight emergency calls, then back on duty the next day. The research consistently shows that the kind of fragmented, insufficient sleep they get around overnight calls degrades next-day cognitive performance to levels comparable to having had a couple of drinks. That’s why a growing number of fire departments have reconsidered their traditional 48-hour shifts; the performance degradation on day two is bad enough that departments are restructuring around it. ## Self-reporting isn’t enough The insidious part is that fatigue undermines exactly the capacity you need to recognize it. A fatigued responder genuinely believes they’re performing normally. That’s simply how fatigue works. “I’m fine, I can keep going” isn’t evidence of fitness. It’s one of the symptoms. This is the critical organizational point. If a company leaves fatigue management to individual judgment (“take it easy if you need to”), it’s built a system that depends on impaired people accurately assessing their own impairment. Aviation learned this the hard way. The FAA doesn’t ask pilots whether they feel too tired to fly. It sets hard limits on duty time and required rest periods, because decades of accident investigation proved that self-assessment under fatigue is unreliable. Pilots who’d been awake for 20 hours consistently reported feeling capable. The data said otherwise. If you’ve ever had a flight delayed while the airline sought a new crew because the original crew had “timed out,” you’ve seen these rules in action. The tech industry hasn’t had its equivalent reckoning yet, but the same cognitive science applies. An engineer who handled a two-hour incident at 3 a.m. and says they’re fine at 9 a.m. may well believe it. That doesn’t mean they’re right, and building your next day around that assumption is a gamble most companies don’t realize they’re taking. ## What active fatigue management looks like Companies that take responder fatigue seriously don’t rely on individual heroism or self-assessment. They build a few specific practices into their incident management capability. _Explicit rest expectations._ Not “take it easy if you need to,” but clear guidelines: an engineer who responds to a significant incident overnight is expected to start late or take the morning off, depending on duration and severity. The default is rest; working the next morning is the exception that requires a conscious choice, not the other way around. _The incident commander (IC) monitors for fatigue._ During extended incidents, it’s the IC’s responsibility to watch for fatigue signals in responders: slowed decision-making, tunnel vision, repeated questions, irritability, loss of situational awareness. This is the same responsibility a fire officer has for monitoring crew fatigue on a fireground. A fatigued responder who stays on the line isn’t being dedicated; they’re becoming a risk to the response, their teammates, and themselves. On incidents that stretch beyond a few hours, this includes planning responder reliefs early rather than waiting for someone to admit they’re spent. _Promote the backup._ After a significant overnight incident, consider moving the backup on-call engineer to primary for the next 12 to 24 hours. The person who spent two hours at 3 a.m. restoring service is not the person you want as your first line of defense if something else breaks that afternoon. _Rethink shift length._ Most teams seem to default to week-long on-call shifts with several weeks between shifts, but there’s a strong case for shorter, more frequent shifts. The same logic driving fire departments away from 48-hour shifts applies: shorter shifts mean less accumulated fatigue per shift, even if each person’s total on-call hours per quarter are similar. Shift design is a fatigue management decision, whether your company treats it as one or not. Some incident management platforms are starting to build this awareness into their tooling. incident.io, for example, detects overnight pages and proactively asks the responder the next day whether they’d like someone to cover their next shift. That’s the right instinct: making fatigue management a system-level concern rather than leaving it to the judgment of the person who’s least equipped to assess it. ## It’s a reliability decision Most companies aren’t actively choosing to ignore a fatigue problem. Rather, they have a fatigue problem that they haven’t noticed yet, because nobody has framed it as an operational risk. When a leader says “we trust our engineers to manage their own energy,” what they’re actually saying is: we have no organizational mechanism for ensuring that the people responding to our next incident are cognitively fit to do so. Respecting fatigue isn’t coddling. It’s protecting the quality of everything your engineers do the next day, including the next incident response. I’m writing **Incident Management for DevOps and SRE**, a practitioner’s book for incident responders, incident commanders, and people building incident management programs. It distills what I’ve learned leading incident management at Google and Slack and consulting for companies worldwide. Sign up at im4ds.com to hear when it’s available. If your company needs help with incident management, my consulting practice is all about engineering better incident management. ### Related Posts * Heroic saves are near misses * How often should your engineers be on call? * Routine Emergencies
opsmtrs.com
Pagerly
Directly manage and resolve operational incidents from Slack, streamlining the response process and improving efficiency.
Posts are pulled live from Bluesky and cached briefly. Posts with content labels are hidden.