What to Do When You Mess Up Prod (Incident Response Guide)
Hey, we are back. Episode three, right? Episode three. I mean, I think it's three. I think three sounds right. I can count to three. I think this is our third episode of the second season of databased. And I've got this feeling, James. I just think we're building strength as we go. I just think they're going to get better and better. Really? Is this saying the previous ones were imperfect? Uh, I agree. I am ever hopeful about improvement. Let's put it that way. Watch this terrible segue. Ready? You know what also improves over time, Jamie, at an organization? What? Ability to respond to incidents and be a reliable organization. The ability to to do your best when things don't go uh 100% well, I guess, is yeah. How to be on call, how to resolve incidents, how to build a reliable company. Uh this is stuff that is not well documented. Um a lot of people don't think about it and there's not a lot of good information about it on the internet to be honest. Very true. Yeah. I think that it's almost like anyone who makes anything whether you choose to have an on call rotation or not. You do. You discover, oh, if I'm going to be responsible for the thing I made, ultimately it sometimes as I evolve it, something will go wrong. and then I'm going to need to know what to do when something I changed has a unforeseen consequence. So, uh, as people that have done a lot of that kind of work in the past, we're hoping to share some wisdom on how to strive to do the best possible job of of that that kind of stuff. Yeah. So, we'll make a strong statement. If you have a product that people care about, if you have a product that matters, you will run into incidents, things will happen. There's no amount of testing, no amount of being careful that will prevent or avoid something bad happening. And so, a big part of building a a trustworthy, reliable company is getting really, really good at responding when bad things happen. And that's going to be the theme for this episode. It's like how to be on call, how to build an on call process, how to build an incident response process, how to how to be an operational company. I think we can start off with what I think is the least important part of this of this this episode which is the process. It's the part that everyone as in in almost all things this the parts that people think about are the least important parts right so the the least important part I would say is the the roles and the and the nouns and the process things like sev levels IMOS tlocks um status pages postmortems etc. Yep. Yep. Everybody focuses so quickly on the what to do, right? As opposed to the understanding about what matters and why. And we'll dive a lot more into that, but it is useful to just level set with the the basic roles and processes first. But I agree this is not this is the part everyone talks about, but might in many respects be the most straightforward part. Okay, glossery of terms. Let's go. Okay, you might hear someone talk about a sev. So at larger companies um when an incident happens there's something called a sev level ass assigned to it and a sev just I think is short for severity right so hey we have a sev 2 which means there's a severity 2 incident happening or we have a sev0 zero a severity zero incident happens and so the point of this is to kick off a process commensurate with how important this issue is and so typically within an organization sev you know they sometimes people have se threes and SE fours, but you know I you know I won't get out of bed for a se four. You know life life starts at se two or so you know I think it's fair to say if if that your sev level does not trigger any kind of particular action then it really doesn't actually exist to your point about like the point of a sev level designation is not to make you feel more or less bad about what you did. It's to trigger an active response of a different degree. Right? I mean that's the the point of these labels. Yeah. So I would say a SEV 2 is in practice the lowest the the least important incident where people have to really care. SEV 2 is like wow something pretty bad happened. Maybe one of our services is is is having some downtime or some unavailability. Um there's a a bug that people are complaining about. That's a SE two. You really need to kick off a process with an owner. And we'll get to ownership roles in a second. Sev one is like serious bad problem. You know, there's some downtime. That's a sev one. Um there's too much grade inflation in these levels, right? The you can only go to sev zero from sev one, right? So sev zero is reserved for this is a really really bad situation. And so if you have a SE zero and you don't loop in the entire executive team and maybe let the board of directors know, right, it might not be a SEZero, right? So So we, you and I have both been in in many SEZ zeros before, Jamie. Um, one of our jobs at Dropbox was to to like lead response for SEZO, but typically I would think of a sez zero as major major outage. Um, you need to get everyone in a room, no one's going home. stay over the weekend, like do whatever it takes to bring um the company back online. Um SEZ 0 should almost never happen. And SE one is is in practice the highest level you should ever see where you know SE one is kicking off a process of something really bad's happening. Let's mobilize a team. Yep. I mean, one way to put it is and obviously at a huge company like let's say Google, most of us aren't Google, right? But like I'm sure it's it's a much more nuanced story there, right? For most of us in the kind of places we work, I think a rule of thumb would be like a sev two is like someone should look into this and a sev one is often like hey this the team should stop and look into this and sev zero is like the company needs to look into this and so in a se zero if it's not the kind of instant where someone says well I can't help because I'm busy with this like a sez zero should be so company threatening that immediately everyone in the company is available if needed to to resolve it right so it's that level of threat I've seen situations where people have been at the airport about to go on a trip and they've been told leave the airport and come back like that. That's that that's the SE zero right there. Um, now SEV levels, you may not need them if you're small enough. Sev levels are only there as a means to codify a degree of how much you should care. So it sometimes it's easy to tell someone, hey, it's a SE one, and they and they know, oh my god, I need to really care right now. So it's just a it's just a placeholder, right? Yep. If I was gonna say if you're a group of to your point, James, if you're a group of like four people just starting a company, if something goes wrong, don't waste any time staring at each other and trying to agree on a sev level. Just fix it. Yeah. So, don't there's no point for any ceremony you don't need yet. But but eventually, these things will be valuable to to sort of trigger everyone to help as much as they should. Yeah. Now, there's there's there's two important things that come out, three things that come out of a SEV process, right? One is they invol involve assigning these two titles, these two roles called an IMO and a T- lock. One is they typically update a status page and one is they kick off a process with support and communications and they're the three things that kind of typically come out of a SEV. And I guess I should define what IMOK and T-lock are. What is an IMOK and a T-lock, James? Yeah. Yeah. So a T-lock stands for tech lead on call and IMO stands for incident manager on call. Um, and these at a at a larger organization are two different people for very good reasons. Okay, so let's say there's an incident. Let's say your site is down, completely down. Um, there's two very important things that have to happen. One, you have to fix the problem. And that's generally a technical response. And what you'll find, we'll talk about more of this later on. It's really easy for chaos to emerge because no one knows who's who's responsible for fixing it. It might take 10 people to restore service and someone has to be like tech being a tech lead like being a tech lead for a team. Someone's the tech lead for the incident and their focus is on fixing the problem. Now what you'll find uh with larger issues if that person is also simultaneously handling communication they're talking to the CEO who's freaking out. They're they're speaking to the support team asking about hey what can I tell our customers? then um miscommunications happen and the tech lead gets distracted etc. So typically the two jobs is I'm tilock you are in charge of making sure this problem's fixed. IMO you are in charge of messaging this and communicating this to the rest of the organization and to the customers. I have a I have a very managementoriented slightly hottake version of these two roles, right? Which is in an in a first of all in a sev 2 there normally isn't an iMok because there often isn't enough communication to make it very material. But if you need one, use one. But for big incidences, the most important person on the incident is the T-lock. And the IMOK's job is to protect their time and attention. Because like if you can imagine the higher the stakes get the more stakeholders are freaking out that all want to say constantly are we almost up again or is it almost solved like what's going on or does they have stress because their customers let's say are upset and they bring that stress into the technical team trying to fix the thing and they go people are pissed you guys have to [ __ ] hurry up or whatever right so I'm not saying companies are ever like that but just imagine they were hypothetically companies are like that in many respects the IMOS job is to be maybe the single person that occasionally steals little bits of time from the T-lock just to stay in stay in resonant of where we're at and just manages absorbing and communicating kind of all that anxiety with everybody else. So the IMOK's job is to protect the T-lock's time because as you all know like finding out like what's wrong with the technical system is a deep thinking kind of flow thing and if you're being interrupted constantly by people stressed out and asking you questions you're just not going to have enough concentration to stay on the task and figure out what's wrong. Absolutely. And so at drop at at Convex we are we are the executives. We're we're the bad annoying ones in an incident. So, cuz we can be freaking out and stuff. Now, both of us are engineers and I like to think that we're pretty reasonable, right? But, um, but there's a point um at Convex where someone will tell me to go away and I have to kind of listen to them. And there was certainly a point at Dropbox where either you or I would tell the seale leadership, do not talk to the engineers, you are not welcome in this room, you know, go away. This is a place for calm problem solving and I'm going to come with you and we're going to go for a walk and we're going to discuss the issue and then I'm going to come back in and discuss it with the with the T-lock. So the IMO kind of plays both sides. The IMO has to be technical enough to know how to interrupt the T-lock and know how to interrupt the team without being an annoying manager and then go out and handle the communication for the rest of the organization. Yep. We'll talk more about communications a little bit later on because it's a it's a more interesting topic. Um if you've ever but if you ever find because I think that because the T-lock is the most important um role that's the one that emerges organically. So again whether you knew you had a T- lock or not you did and it was probably the person everyone gathered around their keyboard if you were in person or whatever that was your de facto T-lock. But if you've ever found out that you were in an incident high enough that it was super stressful because people kept asking is it when is it coming back online and that's that need you were feeling is you needed an you need an IMO in your process to protect the technical leader that's trying to own solcting the information and solving the problem. Absolutely. And I'm Jamie, I'm I'm already getting bored of this section because this is this is the this is just the policy stuff, right? The last thing there's a status page and you got to update the status page. That's this is the mechanics, right? But what we really want to talk about and this is just kind of our ethos, Jamie and I, we care about the why in all that stuff. So, we want to orient this this podcast episode about like how to think about the art of maybe communication in an incident or how to think about the the the leadership of how to resolve an an incident. Um, so I think as we go more it'll become more clear about how to be a good IMO, how to be a good T- lock and how to do stuff like update a status page. Definitely. But yeah, I think maybe one of the sections we should talk a little bit more about is like you know when once you're in an incident like what do you do right like okay so how to actually resolve a problem right I tend to think that there is five stages of incident response um so the first is detection how do we know a problem happened right and we'll get to that in a little a little bit but um either either you know you got paged or someone internally noticed something or maybe a customer complained something gets surfaced and then there is an incident right and then at that point there is kind of this these four steps that normally happen in order and this is not a rule but in our experience I find that like triage stopping the bleeding like you know amelioration uh debugging and remediation in that order. Y um so what happens with an incident is you're going to be doing the rest of your job and then something happens. You're not waiting. You're not there watching ready and waiting, right? You're getting interrupted at an inconvenient time with an incident you didn't seem see coming most of the time, right? It's a stressful time. And so it's not possible and you should not attempt in most circumstances to jump straight into reading source code, fixing the underlying problem, right? It's really about this kind of journey of discovery, stopping the impact, removing the urgency, and then fixing the underlying problem. So triage, figure out what broke. Triage, figure out where what's coming from. stop. What can we do to make this problem go away? There's a lot of interesting things that we can talk about here. Debugging. What is the actual real problem here? Maybe there's actual a bug. Maybe there's a line of code that's wrong. Maybe some system sub-optimal. And then later on remediation. How do we actually prevent this from happening again? How do we actually fix the underlying problem? Yeah. No, I think these are great steps. And I I think a lot of incidences end up much longer than they need to be because people try to skip steps actually. It's sort of the go slow to go fast kind of thing. Like if you skip some of these steps, you'll burn a bunch of time and then you'll have to double back to the beginning anyway. I think a good example is like triage is such a great stage. It's easy to jump over because one of the most important things you have to figure out pretty early on is like do we even have the right people in the room that even know what might be affected here? Right? what system is this? Who might have touched it recently? Who knows stuff about it? Right. So, um just yeah, even that is is is this are the right set of people and the right set of information even present to be able to move on to anything else? Yeah. Now the the triage is a very difficult challenge and I think it's an area where engineers have to maybe check their ego a little bit because there's a temptation to rely on heroics and like intellectual strength in terms of like figuring out if I click seven dashboards deep um there's history list length is growing on my SQL and blah blah blah blah blah right um I think in in practice it's very hard to do this at 4:00 a.m. in the morning where you just got woken up and and you're nervous and there like it's a stressful situation. So typically this means building really really simple dashboards that that that as best you can point to the source of a problem. You know which service is throwing errors? Has there been an increase in load? you know, having as few dashboards as you can, like maybe six graphs on a page as best you can point to, oh, a specific customer is doing something weird or our network went down or maybe who knows what, database is getting slow. um really investing in that triage process because the most difficult and most stressful part of an outage from my perspective at least is not fixing the problem. It's figuring out what the problem is. It's it's like you've been woken up. It's an emergency. Your application's down. Something went bad and figuring out the source of it. Once you know what happened, then you have something to work with. It's a very stressful, confusing time until you figure that out. And so, a lot of the investment I think people should put into is is triage assistance, figuring out how to resolve this. Now, I'd be a little bit careful at relying on AI for this one because I'm sure there's a lot of great companies out there that do assist in triage, right? Yeah. But at the end of the day, you got to fix this problem. If this if if something bad happened, you got to fix it. You can't blame the model at that point. You you have to be able to find the problem. You can't Also, another thing about AI that is a is a mild weakness in most circumstances, but it's perhaps a more serious weakness in this circumstances. AI is always unusually confident in its assertions, right? And so often when you have humans doing triage, you know, if there especially if there's a few people involved, which trust me, if it's a SE zero, eventually you will not be alone. So maybe it's a whole another topic about when to sort of bring others in so you're not alone. But um is uh you know when people communicate things they see they'll say this level's a little high on this and really they're making all those things like resources available to T- lock and the group to be like oh maybe I should look into this other related thing because like you know Jill just said oh levels were elevated in this one graph over here right so but what Jill doesn't say is got it that's it this service was you know is you know there's this all this subtlety about like surf surfacing a little more information or a suspicion that ends up adding up to the kind of triage process to even discover do we understand what's going wrong and AI has a tendency to just want to solve your problem so badly it just declares an outcome too confidently and if you if you agree with it too quickly once again you'll waste time because you didn't finish triaging but you try to move on past that step yeah so two more things I have to say about triage one it is not helpful while you're doing triage to have someone being noisy about how to fix the problem, right? Um, hey, I think we could build a new system that did this differently. There's a temptation, especially engineers, we're problem solvers, right? To jump to like weigh down the process, things that we could be doing. Um, it's not helpful in that moment. You're trying to figure out what happened. Um, the other thing that's not helpful, by the way, is blaming someone. We can talk more about this, but like it is so so critical to have a zero blame culture in a in in a in an operational situation because if you have I'm not yes for cultural reasons but not just to be nice. If if there is a concept of blame or someone's going to get fired for an incident, then fear comes into incident response and defensiveness comes into incident response and people are inclined to not want to proactively call out a problem. They don't want to make themselves look bad. Right? If if there's any incentive against focusing everyone's energies on proactively identifying problems and fixing them, they will make you worse at responding to problems%. Yeah. Sorry. Go ahead. Yeah. So, typically blame is is useless. I I don't think I've ever and I've been in so many incidents, right? That was part big part of my job. I don't think I've ever seen someone ever get disciplined for causing an incident. Uh it would be a disaster and especi incident that would be a complete disaster. Like you just need everyone positive as weird as that sounds and excited about solving the problem. like it needs to be a zone of ultimate safety to to find together what the issue is without worrying about oh what if that ends up being that the root cause was something I did right so I actually think putting the IM mock hat on because being in management for a lot more last 10 years I've been IMO quite a lot it's one of the roles that IMOK can kind of play that's valuable is let's just say that when we're doing triaging let's say the way we ended up discovering about the problem is the customers complained before an alarm fired right which once in a while does happen and so like every business should probably have the goal that like our the systems tell us there's a problem before the before the customers do it might be both but it's not just the customers but as an IMO what I might do is just write that note down for later and do not bring that up right now it's a perfect example of something where later on when you're reviewing the process you say okay what do we all do about this later but anything anything like that that comes in early like why did a customer hear about this before we got paged like none of that energy in the room it's just all constructive toward how do we get everyone back online? What does everyone know? What ideas does everyone have? So, I agree with you completely. Blame has to be 100% out of the room if you want to get your system back online as fast as possible. Which by the way is why seale execs are not welcome, sales team are not welcome. Um, unless you have a culture where there's a super high trust, but like no like why did this happen should happen should be in this phase. Um, and they should never honestly like if something bad happen, it's a failure of process. It's a failure of systems. Um, right. It's not a failure of people. I I've seen a couple of times actual like sabotage happen. Not at a company I've worked at, but other companies I've advised. Okay. Sure. That's one thing. But if people are well-intentioned, no one gets in trouble for an outage, right? Yeah. what might be true is that like senior management gets in trouble for the outage because you did not invest into process enough or whatever but yeah there is accountability but it's more of a systems accountability not hey the engineer that accidentally typed the wrong command like that is never ever the issue so absolutely and we've and we've seen this happen so it's not a hypothetical this has seen this happened so many times the other thing I'm going to say is it's unsurprisingly right because I always go on about simplicity that's all I talk about Right? But if you have a complex system, it's very hard to triage. Right? Like if you have a very simple system with very simple algorithms and a very hierarchical structure and not tight coupling between services, it's it's normally quite easy to figure out what's failing. And when you have a sophisticated system, a complex system, it can become extremely hard to triage because all you know something's gone wrong. But who the hell knows where it's gone wrong? And this is something folks got to watch out for with AI um codegen, right? It is possible to build something so complex that you can't understand it. Um and then wow, you're in trouble when an incident happens. So of course I'm going to shill convex and say, well, if you're using convex, at least you have good building blocks and and and you're fine because most of the honestly most of the incidents happen deeper in the stack. Mostly it's on your infrastructure that that that these big incidents happen. Um, but designing for simplicity so so so important because when you build something you got to think to yourself, not do I understand the system, it's does someone else understand the system at 4:00 a.m. in the morning when they've just been woken up. Yep. Enough to fix. So triage. Great. Done. Right. Stop the bleeding. Maybe there's a maybe there's a piffy one-word phrase for this, but I think this is an important individual step distinct from debugging and fixing, right? Because often time you can just make a problem go away, right? Let's say there's a system that's corrupting data. I don't know, right? Turn it off, right? Or maybe your databases are overloaded. Um can we shut down some background load? Maybe you're doing a migration. Stop the migration. Right? If is there a um I don't know some service that's optional, um just change it to return yes all the time, whatever it is. But like often times it can happen in parallel with debugging. So obviously nothing can happen until triage happens because you don't know what to fix and then someone can start debugging right but really there should be a a person thinking about how do I stop the impact first because if you can stop the impact everyone takes a deep breath and then and then the criticality has been eliminated and then you can go and resolve the problem. One of the um very dangerous things that can happen in an incident is it becomes so stressful and high pressure people make it worse. People go and do a high-risk thing. I'm going to go change some code in prod and I'm going to fix this bug and guess what? You accidentally delete all your data, whatever. So if you can reduce or eliminate the critical dependency, everything changes. Take time off out of the equation if you can. Yeah. Yeah. And I think the the the version of that that was the example you alluded to briefly is a great illustration of that, right? Which is if it's a load related thing and you got some garbage collection job that runs in the background and cleans up old data or whatever, like one of the first things you can do is just pause that, right? Because like if that reduces your error rates and your customers are mostly back online again, now you can take a couple hours to figure out what's wrong and fix it the right way instead of feeling like we've somehow have to pull off a miracle in 15 minutes with limited information. So I I agree that this stop the bleeding thing. It's good to see if there's some quick mitigation that's not a full remediation for your issue. And if there is, it will give your team time to fix things come in a way that's composed and thoughtful without having to do some incredibly high risk move in production. Yeah. Now, I've seen situations where people are a little bit clever about this. Don't know how I feel about it, where they sand intentionally, like they maybe put a big file on all their discs. This is a bit of an old school thing to do, right? you put a big file there and then you've run out of disc space, you delete the file and then bam, you're you you're back or you have some background load. In general, I would not endorse like wasting resources that um you could have had paging thresholds on, etc. But I will say in any large system, you probably have a constant level of background test traffic. You're probably running some some benchmarking, some remediation scripts. you tend to have this like background workloads and those are a really good candidate to shut down. And by the way, they need to have a script ready to go. It can't be like, hey, let's go figure out how to stop whatever the table compaction job. Oh, I better push a new binary to do so. Uh it's really hard under duress to think through the implications of whether it's safe to turn something off. And so I would encourage folks when they're building systems to really segment them into like critical must be running services and they're live site services and then non-life site services that can get turned off. And there should be a script or a button that you just run and it turns all that stuff off. All right. If you have a background job doing analytics, great. Turn it off. Right. It doesn't have to run right now. Um but it should be just a single a single line to do so. Yep. Yeah. Um, at Convex, we've got all kinds of tools where the team just knows and we've got playbooks and all that fun stuff on call that like there's a quick command you can run that just immediately pauses all that background stuff just takes some load off the system to just again buy us the time and clarity to be able to identify and fix the real problem. Yeah. So after you've triaged and you've ideally stopped the bleeding, although it's not always possible, there's debugging. There's actually figuring out what's wrong. And unfortunately, there's no shortcuts here, right? Yeah. This is just regular old engineering. But to the point you made before, James, I think it's a very good one. If you've done a really good job at triage and and then if you've relieved a little pressure with stop the bleeding, the good news about debugging is it it it turns the problem into a regular old engineering problem you do all the time, right? So it's like how do you do debugging? Like yeah, you you do debugging a lot like you debug anything at that point, right? So if you've narrowed it in enough when triaging and debugging gets conflated is when debugging is the really hard version of debugging, right? Because you're both like trying to solve the problem but also not sure what the problem is at the same time. Yeah. So I mean debugging can be incredibly hard. Simple systems are easy to debug. Um, having engineers deeply understand how systems work helps, right? Having experts on your team who know how the code is structured helps. Um, a big part of the T-lock's job is to make sure someone's going to fix it. They are accountable. By the way, the most important thing about the the T-lock role is that they are accountable. It's it's they are in charge of making sure it gets fixed or finding someone else to take over the job. that other person needs to say the words, okay, I'm the T-lock now. It should never be ambiguous who's in charge of an incident because otherwise you'll see this like everyone's kind of doing stuff in parallel. No one quite knows what's going on. A big part of managing a complex incident is kind of shephering parallel efforts of a lot of engineers who are reporting status. I'm going to investigate what's happening in the logs. I'm going to go see look at the source code for blah. I'm going to look for exploits online. Um, and with a large incident, the T-Lock might not actually be doing any technical work themselves. They might just be coordinating the efforts of a bunch of engineers to figure out the problem and yeah, and collecting all that information, getting it all in one place and starting to do kind of a first pass almost like brainstorming on the any connections between anything people are seeing, right? Like helping to facilitate that conversation. Um, one kind of random tip, it's going to sound like an obvious one, but you know, it's a classic for a reason, is it's almost always something you just changed, right? Like after you do triage, if it wasn't the the environment or AWS is down or DNS isn't working or like or a new exploit in npm, right? Once you know after you've derisked some of those environmental things if it does seem like it might be something about our code that boy if you don't look really quick like you might want to focus a little early on what all has changed recently that went out into production because like like I said I mean in a pretty sh like not shocking way you know 80% of the time it's probably something that changed recently. So, um, I would say in our experience when there really is a thing that wasn't just, oh, we've run low on servers or someone's traffic got much higher, it wasn't kind of like a contextual thing, but it was something about our system, there's a very good chance it was a change that went out in the last week or two. So, um, now we're going to get to the power of zero a little bit later on, Jamie. So, uh, so we'll definitely re revisit this. Um, and then after debugging comes remediation and and and there's not much to say about that other than be so so careful you don't make it worse, right? Debug and remediation are just there's no rules for this stuff. It's hard and a lot of the work is done beforehand to make sure that these things are easier to do, right? There there are backups that have been tested there. There uh, you know, you can turn off non-critical load. um you maybe done some disaster recovery training to to test what happens when you restore from a backup etc. So debug and remediation are quite technical jobs. Um and a lot of the work on responding to an incident is is getting the conditions right so that that's easier. The communication is happening really well, the coordination's happening really well. Um Yep. And um I think that some of the preparatory work there because just like on triage and everything else, it's like the more you prepare the easier things go is there are some decisions you should work out probably as a team before you get into that spot which is stuff like how do you get around CI or whatever like and and is that a thing you can do or you're willing to do, right? So let's just say oh you've discovered the issue but it takes 90 minutes to run the test suite or three hours to run the test suite. Right? you may not want to wait that long to get the fix into production. Um, have you all agreed on a process to not use CI because you don't want to argue about this stuff when you suddenly have to make the decision you're down, right? So, even just stuff like how does your code management work in this play, right? Do you cherrypick onto the production branch or whatever? Or do you pull your whole like code base forward and then you might pull out other like like less tested problems, right? So the more you can think through ahead of time how are we going to deal with like an emergency patch of production the the less stress you'll have in trying to get your remediation out the door quickly and confidently. Yep. And look in any complianceoriented situation. You know Convex has you know all the compliance controls but there are exemptions for doing things in certain circumstances. There are exemptions for being able to push code through CI for example when there is incident and then documenting it appropriately. Think about this ahead of time. You don't want to have an argument with your CISO while while your company's down. This has to be done ahead of time. Um, and all the all the break glass stuff, what to do when um when your VPN doesn't work and who has access to to directly connect to the root account on AWS or whatever. Um, there's just a lot to be done there and it's hard to dis hard to distill it all. And so really the answer is the practice, right? I I guess eventually a company just has a bunch of incidents, right? And the team gets good at it and there will be folks on your teams who are good at incident response and maybe they're less good at some other like engineering planning stuff, but they're bringing real value to the organization through virtue of their incredible debugging skills. They're really valuable people to have, right? Don't um overutilize them. Don't burn them out, right? Don't make them the dirty work people. But you know really work on developing talent and practicing through simulated incidents what to do when stuff really goes down because um that's that's a stressful place. Now on the communication side we talked about the Tlock being the coordinator. We'll talk about customer communication later on. One other thing I want to add is like the value of live communication, right? Like like This is a this is not a time for communicating via email and it may not even be a time for communicating via Slack. This is probably a time to get on Zoom. It it very likely could be a time to get everyone into a conference room. It might be time to have a war room. It might be time to say, "Hey, yo, you can't go home right now." I mean, you know, within reason. Um because I think there is a lot of value to live communication within the within the response team. You're saying within the response team. I agree. Yeah, that's not that's not the whole company. And by the way, people need to get out of the room if they're not relevant. I I would say a huge percent of the time, I might even go as far as to say this is a bold claim. 50% of the time or more, the person that comes with up with the key piece of information to like unlock a really tricky incident is a kind of junior person on the team without as much context. Like a lot of times they're they're making fewer assumptions. They're looking in that one log. they go, you know what, folks, right around that time we had the issue, this one number spiked. And so, one of the benefits of the sort of like t like whether it's virtual or in person, but like the war room kind of scenario where everyone kind of has a live mic is it's just the freest environment for someone to just mention an aside of a thing they notice that suddenly the T-lock locks onto and goes, "Oh, hang on. Wait, what happened at 9:37?" Like so um I just I think it's very difficult to capture that kind of serendipity asynchronously over text. Like this is yeah this is so true. It's it's almost always like a quiet engineer who's a bit afraid to make a statement because they might be wrong and they say like, "Oh, it's it's weird that there was no log events between this time and that time." Right? And the job of the T-lock is like, "Wait a second. What's going on there? Hold on a minute." And also, frankly, encouraging the environment where someone feels comfortable sharing that. Yeah. I you it was a good thing you say quiet. We some of the strongest debuggers and triagers we've ever worked with have been some of the quietest people on the team. And all they might say is, "Huh?" And you know, as soon as they say that, the T-lock should be like, "What did you see?" Right. And so said, "Huh?" That's the kind of person that doesn't say, "Huh, unless there's a reason that unless there's a real reason." Yeah. So, um, you get the best shot of getting access to all the information before anyone's qualified. It is right. It just might be interesting if you just kind of get everyone on a live mic or over a table together. That's the best idea. So, the culture of being in the room, right, the the response room, right? Um, we've done it a lot of times. I some my best memories in my career have been in these rooms. I don't know if I should say that. is maybe is a bit of a weird weird weird thing to say but like I I just the the the camaraderie that's come from like feeling together on a problem like just that I think an incident can can go in two directions like it is stressful and there is a degree of guilt people feel there's a degree of um shame and fear and all these really negative emotions that are not conducive to resolving a situation and so in any productive incident I've seen. Like if you saw the war room, you might be quite confused cuz it seems like a fun place to be. At least in the incidents I've been involved in, it's like there is a little bit of gallows humor. There's a little bit of joking. There's a little bit of cynicism and and like and and laughing about stuff and and just being part of it because like managing stress is so important, right? I think it's safe to safe to say this. There was an incident we were part of. I mean this would have been over 10 years ago. Massive massive outage at Dropbox. Dropbox was kind of down for two days. And I remember some engineer in the in the room said, you know, good thing my resume is not in Dropbox, right? Yep. I remember that. And it was funny. I mean, but now now look, if someone had that level of cynicism and wasn't actively solving the problem, get out. like you were not part of the team, but this was someone who was in it. This was someone who was actively resolving the problem. And there outages, I don't know, they can be fun. It's um they're serious. They're important, but they're a time to work together on a hard problem. And there's a little bit of a culture where you got to make him a little bit about that about like let's go team and like feeling pride in yourselves and and excitement to go do something cool. Uh cuz that's the energy that that leads to to greatness, you know, and I I've seen like true you and I have seen incredible greatness in these instance. people have done incredible heroic things to solve problems, you know, just written insane amounts of code in a short period of time and and you're just so proud to be part of that moment. Like, wow, this is it. Like, this is like major leagues engineering right here. But it needs to feel that way. It can't feel like, oh my god, you're screwed up. What are you going to do about it? The IMO got to keep that away, right? Allow the to be the engineers are being a little bit sarcastic. Great, good, let them be. This is not the time to be, oh, I feel so bad about the users. Yes, we do feel bad about the users, right? But this is the time to to lean into like, I'm the engineer. I'm I'm the one who's going to fix this problem. And it's very healthy. I think it's I think it's extraordinarily hard to try to perform under under duress. Um, which is look why you got to keep sometimes people like Jamie and I out of the room because we're taking it so personally. We're we're the ones feeling it and we're the ones who got to go out to the customers and we we feel it, you know, and so sometimes, you know, an engineers almost seem too chill in an incident. And as long as they're working really hard, that's fine, right? Like what's the point of them feeling bad about themselves? The point is that happens later. Let's worry about postmortems later on. Now's the time to shine. Yep. Yeah. There's there's definitely a time for, okay, how did we end up here? What are we going to do to avoid this next time? And then and then that's significant enough that you know, you you be ready to change your plans for the next two weeks. We're not going to work on that feature. It ends up that we were more exposed here than we knew. And so we've got action items. We're going to go do some stuff. But during the incident is not the time. during the instance needs to be a time of excitement and and freedom and engineers like taking a lot of pride in pulling off some heroics in order to get the customers back online and that being a very safe space to just do some great engineering work. Um this the systemic like let's analyze all this and let's let's you know prevent this class of things or reduce this class of things from impacting the that's a that's a later thing. Don't worry about that during the incident. just just get everybody working and and have the IMO protect the engineers from all that energy. So yeah, in some respects the the T-lock does kind of a little bit report to the IMO. The IMO's normally a more senior person, normally more managerial person, a little bit keeping an eye on the team. I guess both of us have been in incidents that have gone for multiple days straight. I know certainly I've been working 36 plus hours non-stop on an incident, which sounds um it sounds crazy. It can happen. Um, it shouldn't happen very often, right? But in those moments, sustaining a team is so important, right? Like, and sometimes you're like cycling people off. You might have your best debugger um, you know, 12 hours into a stint and you're like, "Hey, go sleep." Yeah. You might have told people go get some sleep. And they probably gonna argue with you because they're going to want to stick around and help. And that's that's great. Uh, you know, it's it's it's a it's a it's a it's a complex thing, but maintaining your people is is important, right? You have to have some some gas in the tank. Yep. Um, I think both of us are smiling in this section, Jamie, because I don't know, I've had such really fond memories. I don't know. It's like the I have so this is the most respect I've developed for people and it's like, oh my god, it's been so good for company culture. But I think it bonds teams. I'll put it that way. Right. So by the team, it shows how much pride everyone has in what they made because they're willing to work so hard to get it back online. And and I think they you learn from each other about how sort of resilient you are and resourceful you are. And normally once you get back to the normal course of building when the incident is over, I do think the team is a little closer and trusts each other more when they do this kind of work together. So for that reason, even though you don't obviously you don't want to have too many of these, although I will say something kind of a hot take and say if you don't have any incidences at all, that's a problem too because the next one you have, you're kind of out of practice. So you definitely want to minimize the number of serious in incidents you have. But the truth is learning how to work this way together for the technical teams normally create stronger teams that are that are closer, you know, that work trust each other more and all that kind of stuff. So, look, I've got swag made before for, you know, I I remember we had an incident code named striped moose a long time ago, and we all got socks, striped socks with a moose on them, you know, and we all wore them with pride cuz it was like that was I don't know. The team the team killed it. Team was awesome, you know. Um, all right, enough of that. Detection. How do you know something happened? Um, that's the trigger, right? Um, there's an art to it, isn't there? It's like there's an art and a science. Um there's also metrics and logs. They're quite different things. Uh and may maybe maybe the the the separation I'd make, the distinction I'd make is metrics are like time series, up to-ate um f uh graphs and line charts and stuff that don't have a lot of debugging information in them. They're about triage detection. So typically you have paging and alerts on on on metrics right and then there's logs stack traces your you know exceptions in sentry or whatever um going to you know data dog and querying a bunch of logs and that's for debugging right I tend to try to keep this quite separate I think because logs have too much information in them they're too they're too slow like they they're probably a minute late or maybe they're half hour late Um, and it's it's hard to like alert on logs. Typically, I'm a metrics kind of guy. I want a small set of very informative metrics that will just give me enough information to triage. Then the metrics have done their job and then we go to debugging mode after that. Yeah, it's it's very fair to say this has a lot of parallels with triage versus debugging in the incident, right? Like when you're designing your your telemetry, be pretty disciplined about what things mean something is wrong versus what is wrong. Because like it once again, if you conflate those things too much, you could end up making like for example, if you try to trigger too much of your like something is wrong based on what really specifically is going wrong, you're going to end up with an absolute ton of alerts that fire too often and they're too noisy and and then people start ignoring them and stuff. So, you definitely want a few really highlevel quite clear something is wrong alerts that like you you can bank on. And I think many of those things that are closer to what the user experiences than one particular zoomed-in system is how like having a I mean those matter too, but they're a little bit less the stuff that wakes you up in the middle of the night versus like some higher level really reliable signal that like the user is experiencing a problem right now. Yes. and and there's so much like monitoring is a massive topic. We can probably do another whole podcast on monitoring and how to how to make good metrics and design them. But there's there's a few things you touched on that I really agree with. One is like noise. If you have too much noise, if you if your encore channel has like has 50 alerts in there, they just become you become used to the noise. It's you just depprioritize it in your mind. Typically, don't interrupt someone with something unless it's important, right? because then you can take it seriously and your mind is fresh. Um the other is that end to end metrics matter way more than system metrics. The classic case, right, is if you're tracking if you're alerting on errors on your web server, right? Right. And then someone misconfigures DNS and you don't get any more requests ever again because they're all falling on the floor. Well, guess what? your error rate will be zero because you're getting zero requests and zero of them are failing, right? Y and so it's it's it's it's very easy to have a service that itself has no errors because I don't know it's taking so long to respond that every client request is timing out or it's not even maybe it's turned off like it could literally be turned off. Um we've seen that before like a system is off and therefore it's not paging because the the errors graph is zero. Yeah. Right. Um, and so it's really important as best you can to monitor availability from the client side. So if you're writing a service in a larger company, put most of the metrics in the client bindings to talk to your service. It's a bit hard to do it on a website, you know, but you can have third party services that that monitor your monitor your product. Yeah. Actually, another thing it makes me think of another like bold statement I'll make, but based on, you know, poking this stuff for the last couple of decades. I think no matter how many things you log, the evolutionary/emergent property that ends up coming out, there's only like three alerts the whole company actually runs on, right? It's like you you might have 27 instrumented, right? And like some of them also might fire when one of the big three fire, but really everyone de facto ends up knowing those big three are the ones that matter. And so like since that's going to happen anyway, you can probably save yourself some time by really learning to think that way ahead of time, right? Like what are the like bigger signals and like the other stuff we should still measure, but that stuff is again more about like speeding up our debugging and our triaging. It's not about alerting. alerting. It's often pretty obvious when things aren't right. Right. Like if things are not right, they're not right. And you often get a little too clever about trying to design some really detailed like trigger. And it's often something pretty intuitive is like the right thing to actually measure. So errors per second measured from the client is your error rate from the client is too high. Yeah. Someone should look at that. Yeah. Um, now this leads us to like this philosophy that you and I uh, let's say originators of, Jamie. Why not? Let's say we invented this philosophy. Yeah. The power of zero. Um, and neither of us know which one of us invented it. Who knows? One of us did. I don't remember. Was it you, Jamie? I think it might have been me, but Okay, we discovered right now. I was leaving the door open for it to have been me. I don't remember though. It might have been you. It might have been you. Let's just say it was Jamie. Um, the power of zero, Jamie, what is the power of what is the power of zero? What is your what is this philosophy? The power of zero is a formalization of some of the stuff we've been like just talking about here casually between the two of us, right? Which is like it's so important to distinguish between signal and noise. And that like because because of humans and human behavior, right? So like you want this very clear signal between when something is wrong and when something isn't wrong. And so the best way to illustrate when you get this maybe you don't have this quite worked out is you know every system every complex system will usually have some low level of errors that happen for reasons you haven't gotten around to yet or you haven't you know whatever and so that might mean that like normally you have 14 errors a second or something like that right and the problem with that is it I if as a culture everybody gets kind of used to like oh it's always around 14 errors a second don't worry about those right like one day that might creep up to like 17 errors a second and then you off to be like, "Oh, is that real or is that just those 14 errors we talked about before that don't matter, you know, and then and so you end up creating this debt where everyone has to remember or try to somehow peel back the layers and be like, did was this change meaningful?" And so your real goal should be to to have your most important topline metrics have a pretty st definition of going wrong where when nothing is going wrong the number is at zero and when something is going wrong the number is not zero and and then that way this very clear cultural alignment that we act when the number is more than zero and when the number is zero we're we're okay. Um y and so in that case of those errors if let's say there were like seven errors that like for various reasons were difficult to permanently fix or they were encoding errors because the client did something weird then you should in your systems you should categorize those specific errors that way so that you have a line that is zero which is the uncatategorized ones or the ones we didn't foresee or whatever and the second that is not zero then you're like okay there's something we should look at. So, so you don't want to let noisy numbers like cloud your clarity on whether anything is actually wrong that you should take action on right now. Yeah. So, this is you and I are both led very difficult, very expensive initiative to get errors to zero. So, my case was was driving all the file system errors of Dropbox to zero and your case was driving the desktop client errors to zero and um they were expensive initiatives like took time a lot of engineers. It took money actually to get these errors to zero, but it sped the company up so much because if you know on Tuesday there's no errors and on Wednesday you ship something maybe to staging and there is an error probably something happened just now, right? Uh and so you can you have a lot more confidence in your systems and with confidence you can move faster, right? you you can really easily pinpoint when an issue happened because there was no errors and then there was errors. It's a very binary transition zero to not zero. It is and this is not Sorry, go forgive me. I was saying the trade-off is to your point is the team wastess a lot of time proving that the new errors aren't material if you didn't have a definition of zero before. Right. So if the numbers wiggles a little and then it wiggles a little more, your team ends up burning a bunch of time like de-risking, did that movement mean anything or not? Right? So versus James' point that your team is just more productive if you've really carved out that definition zero. It's still zero. It's very clear. You just keep moving or you definitely look into it. Yeah. Now, if you if you're if you're building a database or a storage system, right, and you're doing it properly, a big big fraction of your workload is just background jobs, like walking over your data and checking that it's still okay, right? Why is this? So, like if you're modeling um the durability of a storage system, like how many nines of durability or or or estimating how reliable your database is, a big factor in that is how quickly you can detect data loss or corruption and fix it, right? And if you don't have any validator jobs running, you might have some corrupt data sitting in production that you don't know about for 6 months, right? And you've completely blown out your projections because now your time to recovery from incident is 6 months plus one day whoever takes it you to fix it. So it's really important for any state in your system to be validating it constantly to know that it's a zero errors because if you don't like the power of zero is meaningless if you discover a a file in production that's corrupt for some reason and you have no idea when that happened if it could have happened in the past year. So a lot of what like a good systems engineer will do is have all these background jobs checking the data all the time. So you have really pretty high confidence that yesterday everything was fine and today it's not. Well, guess what? Something just happened. Now this is not about testing and pro. We're not saying just push stuff out to production and see what happens. This all this this push you can have a push process where you push to a staging cluster which has a copy of your production data. There's a whole other episode on that too. Mhm. But there's this incredible productive and operational power in having zero errors because it pinpoints when bad things happen. Yep. And just like we talked about with the errors, if you had those few things that were corrupt because something happened in, you know, 2021, they're oh that's right. And some of those people left, but there was some kind of Okay, then quarantine those things. Mark them as in a known unknown state and then reset to zero. now the all the block well all the everything besides those right so that way if a new one ever shows up you go oh this is something new we did right this isn't that thing that happened back then so but find a way to get to zero um and you'll be able to keep moving with confidence without constantly double-checking all your noisy signals to know whether they're still okay yeah now we've talked a lot about being a T-Lock or the technical side of operations um I think we should talk about IMO a little bit this is stuff that I don't know I learned pretty late in a career. The kind of art of communicating an incident externally, right? Because if you have a product people care about, you have to tell them when something goes wrong. And there's a real nuance to it. Um, and you can um the difference between an incident, people are like, "Okay, they've got it, no problems, and oh my god, I'm gonna move off your platform." A lot of it comes down to communication. Yep. Yeah. I it's you you have to communicate because like the number one thing everybody does is they check your status page, they check your Twitter, and if it looks like your team isn't working on it, then they rightfully assume maybe you don't know or care about the fact that I'm affected right now, you know? So, um, but there I think there's, um, you know, there's there's definitely kind of an art about like what to communicate and when, like, and also to recognize even if you're the one doing the status page updates, maybe you're the iMac or whatever, that like you are feeling a lot of pressure. And I think that one way you can put it is you can end up feeling kind of like the sales team pressure, right? To like declare victory too early cuz you really just want to reassure the customers like you're not down or you're resolved sooner than you are. And I think one of the most important things you need to do with communication is communicate as early as you can that you're down. But also too like probably maybe a rule of thumb thumb is the first thing you should say is like we are having an incident and say it as soon as possible right like don't speculate what it is you don't know yet don't because if you if you overstate confidence then you have to change your answer later once again your customers doubt whether you are operating with competence right like so the most important thing is people at first are just owed the information that you are aware of the incident and you are absolutely working on it and you will keep them updated as as things develop. So don't get too specific too early is definitely one piece of guidance I would give. Yeah, I I think in an incident like these two groups, the technical team and the communications team have to adopt the different personas, right? The technical team that's not really the time to be empathetic, right? Like it's not for them to be thinking about like the negative impact on all the people because they need to be like calm and resolving the incident, right? But the IMO has to be very empathetic, right? You have to think about what's going through the mind of the customer. Now, typically what's happened is you have an outage, therefore your customers are having an outage. Now, depending on what product you have, there might be end users are having an outage and and maybe just knowing your site is down and they can come back later. Maybe that's enough information for them. Maybe, okay, cool, whatever. They're down, I'll come back tomorrow. Or maybe your customer has customers of their own, right? So what's going through their mind is I'm having an incident now. The customer is having an incident. It's it's it's our fault, right? Um they need to be told one, yes, there is an incident. Two, we are on top of it, right? Three, ideally some indication of how long it will take, though don't overpromise this. Right? Now, if the customer knows that you are having incident and you're on it, they can relax a little bit. They are still having an issue, but it's not really their fault anymore, right? So, maybe they've gone from freaking out and panicking. What do I need to do to resolve this problem? And now they've switched into, well, I'm just on this ride. I've got to wait for it to get fixed. That's still a tough situation for them, but it does really take a lot of pressure off, right? So there's a lot communication says a lot to your customers that how they should reason about things. So typically I like to think about it like you said Jamie first acknowledge it, claim it. It's our fault. It's our problem and state we're on it. We are actively working on it. You'd be amazed that just saying these two things how much that reduces support call burden. Because if you're not like telling people that's an issue, you're going to get all these messages from everyone. What's what's going on? What's happening? And if you say, "Hey, we're having incident and by the way, like AWS is down, right? Maybe that's the situation." They're going to calm down a lot, right? Cuz they know, okay, cool. It's probably going to get fixed. Right now, once you've owned the incident, they're going to want to know some information at some point about what's what happened, what what needed to go. And my claim here for most almost all customers, there are a few corner case customers is really close relationship. They they kind of nerdy a bit. They want to know details. Most customers need actionable information. So reaching out to them and saying um Galactis hit a file descriptor limit caused by the restarts in the you know proton service, right? That's not useful for them, right? or like giving them kind of like too much irrelevant information or posing feedback to them where it kind of maybe passively suggests they need to do something but it's not clear what they need to do something. It's really bad communication. What you've got to do is give them actionable information. Do they have to do something right now? Is there anything they can do right now? And if not, like ideally when will it happen? When will it get resolved? So there's a lot of like keeping people up to date. Yes, we know this is still an issue. Our engineers are restoring from a backup. We'll expect to see resolution blah at this time. The the nuance in this communication is night and day difference into how customers experience an outage, right? Yeah. No, I agree. I mean, I think another like another way to put it is let's let's even say especially as you get deeper into an incident like customers will want you you know like In many technical products and probably other kinds of products too, customers do effectively ask a question basically like what happened, right? You get asked something like what happened and so probably the wrong answer is as James said, well the proton service hit a you know SEG fault at the same time that like goblygook, right? So, but uh but um understanding how they perceive your service and giving them a little bit information in kind of the altitude uh about you, what happened that is a part of how they perceive your service is valuable. And one of the biggest reasons for that isn't because they're going to be like, "Oh, got it. Did you check that the Frober was, you know, re recalibrated or whatever." It's because in part of what you're demonstrating to them is you continue to have a mastery of your system to be able to tell what went wrong and and also too it gives them some accountability that they're like well didn't the proton service have an outage last month too like are you all not investing in preventing right the things that have happened before right because I think one thing you might say about any service that you use is that like any sufficiently complicated service you use all of us know this e even as AWS customers right AWS has in inst incidences every now and then. But what you do expect is you do expect companies to always be working to prevent problems they know can occur from occurring in the future so that at least the frontier is moving a little bit on what goes wrong next, right? So I mean you can imagine if it was this is the fourth month in a row something's gone wrong with the proton service, you would start to lose a little confidence in your in your vendor that like are you all ever going to figure out this proton service, right? So, so I think owing that kind of clarity on like what part of your system like was it, you know, keeps some accountability going with you that like and here's what we're doing about that in a kind of customer level language and and gives them the kind of accountability information they need to be like, okay, got it. Well, this is a different problem than last time and it's a scale thing this time and last time it was a a bug in this one thing. So um so I think but but going back to the thing we said before like don't try to give that level of information too early because it is communicate don't overcommunicate and like almost like like James talks about this like for like cone of strategy or whatever you talk about you almost want like a a rever you know a reverse cone of information or whatever right so early on's like something is happening you're not crazy we are on it updates coming right and then like once you get a little bit more zoomed into like yep there was a data database issue, the team is on it, we'll get back to you soon, right? And don't worry about like at 9:04 a database incident blah blah blah. You can write up your postmortem a little later, but providing a little bit more information, but I would say also too as an IMO um communicate like I would I would suggest always communicate what your engineering team thought they knew 20 minutes ago, right? Because the truth is whatever your conclusions your engineering team, your T- lock and your your technical team made just now, there's a chance those might not might but might not be quite right because they're still iterating on them. So like always keep things a bit scaled back of the level of confidence that the engineering team has. Um and uh and yeah and then that's I think that's the way to kind of manage the ongoing communication as the incident moves forward. Yeah. Let let me let me selfidentify a problem where I time where I messed this up. You know, as you know, young engineer James, I think um the worst thing you can do with communication is to be too optimistic, right? If you say, "Hey, there's a problem and we're working on it blah blah blah blah blah." Great. If you say things are fixed, and they're not fixed, that's the worst. That's when that's when your customers will get mad. like if if that's could they'll be like wait a second I don't trust this team anymore. We've dropped vendors before for this reason. Like there's repercussions for this kind of stuff. We've had vendors who have done this to us and we've stopped using them. Um now there's a time we had a very large outage at Dropbox and someone on sales or someone reached out DM'd me and said, "Hey James, is is things fixed?" And I and I said, "Oh yeah, f file systems running." Right? Meaning the system I was working on is running. Now what I didn't mention is like and also image previews weren't working completely unrelated system I wasn't involved in because I wasn't speaking on behalf of the company I was this is speaking on behalf of young James the infra engineer right that made it to customers who were like wait a second Dropbox is not working right now it's not fixed yeah I don't know what a lot of very serious problems like it was a big problem uh that was like one of the worst things that happened in that incident Um, do not be overconfident. Yeah. Do not promise things that aren't true. Be very, very careful about saying you back up. And you could use shortcuts. You can say, "We believe service has been resolved, but we're monitoring the system for regressions." Right? You're leaving the door open for where we're keeping do not market incidents closed until it's closed because that that's when you really lose faith. An incident is an opportunity for your customers to evaluate you as a team. to decide is are these the people I want to bet on and how you act and respond really does does um does demonstrate whether you're worthy of trust. Yeah. So there's a way to put it is as the person managing communication let's say the IMO or just the person who's volunteering to do this. Yeah. I I would say try to be in a little bit of a time travel machine about 20 minutes behind both in terms of how specific you are but also how confident you are it's resolved. Right. So when the nge team is like looks up right that's doesn't mean immediately go and say we're up right like one of the things I do is I will clarify with the nge team like are you all comfortable with the all clear but I'll often wait 10 minutes after they say like we're up right you can always go change the time stamps later so that you don't look down for the extra 10 minutes but but you know don't don't ever the second you hear something in the war room overhear it immediately post that to your um to your status page for example or to your customers. I would say give it a little bit of a delay both in terms of how specific you're being because you'll always be a little more general and in terms of how confident you are that things are resolved. Um and uh and if you do that you're you're less likely to over project. I think if you imagine if you're the customer and somebody says like we've definitely figured out it's definitely a traffic problem and then 8 minutes later like nope it ends up it's definitely the databases right like you're going to be like you all don't know what's going on right so the right answer that early is our team is looking into it we'll as as we learn more we'll keep you up to date y we got it we're on it now the last the last phase of an incident is the learning phase the postmortem phase um again a phase that people take really literally um and and miss the point sometimes. So I feel I see it a lot happen at companies. They treat writing postmortems like they treat writing unit tests. They just go through the motion of running writing these things that they don't actually believe in and they don't think much value comes out of. Right now what this looks like is someone writes this big old long postmortem with all these timestamps. You know, at 5, 17 and 12 seconds, so and so said blah. And at 14 seconds, something else happened, right? Almost no one ever looks at this stuff, right? And when you get to be a big company, maybe having some loose timeline maybe is helpful to see maybe historically how long you're taking to respond. I never almost ever read this stuff when I was, you know, overseeing large incidents at previous companies. Um, and you can always just go back to Slack or whatever and piece things together. A postmortem should be about reflecting on what happened so it doesn't happen again, right? It's not a punitive measure. It's not like, "Hey team, you screwed up. Now you got to spend six hours writing a postmortem, right? In fact, that's sometimes the worst thing because the team that's having a lot of issues probably should be focusing their attention on fixing the problems underneath, not wasting their time every day writing postmodms. I have seen this happen where teams are struggling. They're getting paged all the time and then they're spending half their time writing postmodms as if someone's going to read these things. Right? So the point of a postmortm is identifying what happened, why it happened. How are we going to get better next time and the how we're going to get better is the point of the postmortem, right? If there's nothing to say about it, then I don't know. Is it important at all? Maybe for compliance reasons, I guess, but but really this is like writing down this thing happened, describing it a little bit so we can come back and reference like are these happening very frequently and then coming up with action items that we really do plan on following up on. Yep. The other thing I would say about the action item formation is that if you really want to move the frontier over time of where your issues are because again you never have no issues but you really do as you mature the system you because your your scale is going up you're adding features and so you always need to be kind of moving the frontier to be delivering excellent service. Um the mistake sometimes is to be too literal about the incident that went down, right? So for example, this hypothetical incident we just talked about, the proton service did whatever. You could be like the action item is fix the proton service so it doesn't do that in the future, right? But I think that you don't want to do five W's because then you'll get too abstracted like why are we softwareing it all, right? But you might want to ask why like twice, right? Be like well why did this happen? Well, because the proton service didn't like ran out of memory. like, okay, well, let's prevent Proton service from doing that. It might be more valuable to do one more why. Like, well, why haven't we fixed that yet? It's like, oh, we've been so busy shipping roadmap items because customers have really needed these things that Annie did point out. We've had these resource problems that we haven't had time to get back to recently. We probably got those across a few different services because then you'll be like, maybe let's look at the Electron service, too, because it might have the same problem. So, make sure you you you zoom out at least one more time than reacting to just the literal thing that went wrong because there's probably in like it's not like y'all were doing nothing, right? Most teams they it wasn't like they weren't doing anything. They were doing something, right? So, they were adding value to the business in some way, but clearly not in a way that would have prevented this. So, it's worth asking why at least kind of twice because then you'll be like, well, why haven't we invested in these things yet? Oh, we need a bigger team. Oh, we don't have the rel reliability expertise or we've been so focused on shipping because we this conference is coming up and so we've put off these reliability measures. And so if you do that, then you're going to have a better set of action items, right? Because you be like, well, let's clean up resource management for all these services otherwise we fix Proton this week and next week we have an incident on Electron because it'd have the same problem. Y it's an opportunity for reflection. It's a time for engineers who might be frustrated about things to have a a vehicle to write down what they think can be improved. It's it's you write it immediately after the incident or the next day or wherever cuz it's fresh in your mind. Then you have interesting reflections. It's not just a moment for documenting something that happened, right? It's hey, time for strategy. Something happened. What can we do about it? Now, the last thing I'll say about about action items is just to be careful about them. People get in brainstorming mode and they're like, "Well, this incident happened. Here's the 17 things we could have done to prevent it." Sure. Right. And they'll say like, "Oh, we could have a whole other cluster that we have a complete copy of production data and we do this thing and we we've designed a deterministic simulation testing framework and and you're not going to do it right." Like so if you want to use it as an opportunity for ideiation, okay, but like put them below a line like draw an H line in your dock and just say here's the things I guess one day we could do someday things. Yeah, I agree. You also want to avoid like like sees turning into Congress's equivalent of riers where people are like, I've always wanted to do this thing and now's my excuse to force it through. Right? And so someone's crazy science project even around reliability or whatever is sort of like I told you we should have done this. It's like no no no no that's still not the right answer. And so um do try to keep the action items focused on what really solves the problem and and still scoped appropriately for how much can your company really invest in this right now. But but there's a good chance that it the reason it went down isn't because you guys were sitting on your hands doing nothing. The reason it went down is because you haven't thought about or focused on or measured that impact. You've been focused on other things. And maybe this is a chance to double back and think more holistically about should we be spending more time on some of these areas of of reliability in the next few months or whatever. Yeah, the action items should result in action. If you go back and you're not doing your action items, then you're doing them wrong, right? They should actually go back and they should be about process. It's not about people. If your action item says Joe shouldn't be allowed near the whatever, um, that's not a real action item. It's about, hey, we could have a protection against typing this command wrong. If your action items say everybody stop making mistakes, that's not an action item because be better. The phrase be better should not show up anywhere in uh in the process because we're all humans. We all are fallible. Guess what? LMS are like us, too. They're also fallible. Yeah, you should you should change that to how can we make it safer to make mistakes, right? Like because people will make mistakes. So, all right. So, that's incident response, right? So, what what's what's the summary? There are roles, there are process, there's postmortems, there's a whole bunch of process. It's not as important as the spirit, right? The spirit is finding someone accountable to lead the incident response, the fixing the incident, finding someone accountable to communicate, right? Encourage a culture of constructive problem solving, not a culture of pumitive punitive blame. Try to go through these phases of like detection, triage, stopping the bleeding, debugging, and then remediate. And try not to jump too far ahead because you can kind of bamboozle yourself by just getting too distracted by fixing problems before you even know what's wrong. Um, and zero blame culture. U now ultimately like you said Jamie there is blame sometimes you know if you're the head of infrastructure at a company you are responsible for for shephering the the team towards reliability but in terms of the individual engineers you know it's it's really about the the process they're building not the mistakes that were made. Yep 100%. All right. Hopefully you never have an incident. No actually you know what I hope that you do have an incident. I hope you have an incident because you'll grow from this. You'll learn, but I hope your incidents are all easy to resolve. Um, now if you don't want to deal with this stuff, you could always use convex. You're probably going to have incidents too as a convex customer to be honest. But what's what's the lesson for every episode of database? Keep things simple. The simple your systems are, the more composable your building blocks are, the easier time you're going to have fixing things when they go wrong. Yep. Agree completely. All right, that's it for us. Thanks everybody. Thanks everyone. Yeah.
If you've got a product people care about, you're going to have incidents. No amount of testing or carefulness stops something from eventually breaking in production. We've both spent a lot of our careers on the other side of that moment: on call, staring at a dashboard, trying to figure out why something we shipped is now on fire. This is episode three of the second season of Databased. The theme is how to be on call, how to build an incident response process, and how to become the kind of operationally mature company that handles it well. It's not documented well anywhere. Most of what's written is either too abstract to use at 4 a.m. or written for a company ten times our size.
The boring stuff first
We'll start with the least important part of incident response, which also happens to be the part everyone thinks about first: SEV levels, on-call roles, status pages, postmortems. It's worth leveling-set on the vocabulary, but this is the most straightforward part of the whole topic, so we'll move quickly.
At larger companies, an incident gets a SEV level, short for severity, and the level is what tells the rest of the company how much to care without anyone having to have that conversation live. Here's the rough breakdown:
SEV Level
What It Means
Response
SEV 0
Truly catastrophic situation
Get everyone in a room, nobody's going home, stay over the weekend, whatever it takes. It should almost never happen. If the executive team (and maybe the board) isn't looped in, it might not really be a SEV 0.
SEV 1
A serious problem
The kind of thing that mobilizes a team
SEV 2
Something pretty bad happened, like downtime, or a bug customers are actively complaining about
Needs an owner and a process
SEV 3 / SEV 4
Sometimes defined by organizations, but in practice life starts at SEV 2
If a severity level doesn't trigger any real action, it doesn't exist
A SEV isn't there to make anyone feel better or worse about what happened. It's there to trigger a response that matches how bad things actually are, which is why a severity level that doesn't trigger any real action might as well not exist.
One rule of thumb that scales down to smaller teams: SEV 2 means someone should look into this, SEV 1 means the team should stop and look into this, and SEV 0 means the company needs to look into this. If you're four people just starting a company and something breaks, don't waste time agreeing on a SEV level. Just fix it. There's no point performing ceremony you don't need yet. The value shows up later, once you're big enough that you need a signal to tell everyone how much to care.
Out of the SEV process come two roles that matter more than the level itself. There's the tech lead on call, which we'll call the T-lock, and the incident manager on call, the IMO for most of this conversation (sometimes shorthanded IMOC, worth flagging since both terms point at the same role). At a larger organization these are two different people, and for good reason. When something's down, two things have to happen at once. Someone has to fix it, and someone has to manage communication. Try to do both with one person and it breaks down. Talking to the CEO or to support pulls the fixer off the technical problem, and miscommunication follows. The T-lock owns making sure the problem gets fixed. The IMO owns messaging to the rest of the org and to customers, and protects the T-lock's time and attention. The higher the stakes, the more stakeholders start asking for updates and piling on stress. At Convex, we're often on the annoying end of that dynamic. We're both engineers and like to think we're reasonable, but there's a point in a real incident where someone has to tell us to go away, and we have to listen. The IMO has to be technical enough to know when and how to interrupt the T-lock without being an annoying manager about it, while fielding everything coming from outside the room.
The last piece of the mechanics is the status page, which you keep updated, and knowing who's on call in the first place, since it's surprisingly common to lose a few minutes of an incident just figuring that out. That's the process. What we want to spend the rest of this on is the why. How to think about the art of communication and leadership during an incident. That's the part that determines whether people trust you afterward.
The five stages of incident response
Once something breaks, there's a rough order of operations: detection, triage, stopping the bleeding, debugging, and remediation. Incidents are stressful because you're doing the rest of your job when something interrupts you at the worst possible time. The temptation is to jump straight into reading source code and fixing the underlying problem. That's usually the wrong move. Think of it as a journey of discovery. Stop the impact first, take away the urgency, then fix what's broken. Skipping steps to save time is the classic way an incident ends up taking way longer than it needed to, because you double back once you realize you fixed the wrong thing.
The five stages of incident response: detection, triage, stopping the bleeding, debugging, and remediation, in sequence
1. Detection
Detection is where an incident starts. Metrics and logs do different jobs here, and it's worth keeping them separate in your head. Metrics are time series. They're good for triage and detection, and they're the thing you put alerts on. Logs (stack traces, exceptions) are for debugging. They're high-volume, often delayed by a minute or more, and hard to alert on cleanly. That's part of why streaming logs to a dedicated destination matters once you outgrow the built-in dashboard logs view. The signals for "something is wrong" should be small in number and close to the user experience, not a zoomed-in system metric. A service can look perfectly healthy on an internal metric while being completely unreachable. We've seen systems that never paged because their error graph read zero, purely because requests weren't arriving at all. Monitoring availability from the client side, not just the server side, catches that blind spot.
Across a lot of systems and a lot of years, the pattern holds. A company ends up leaning on maybe three core alerts that really matter, even with far more metrics instrumented than that. Everything else is useful for triage and debugging, not for waking someone up. If your alert channel is firing constantly, the alerts become background noise and get ignored exactly when you need them.
This is the idea we think of as the power of zero. Complex systems tend to run with some low background level of errors, say fourteen a second, that everyone gets used to and stops really looking at. Then a jump to seventeen sparks a debate about whether it means anything. The fix is to design your topline metric so that when nothing is wrong, the number is genuinely zero. Any movement off zero is real signal. If there's a known, hard-to-fix source of noise, like encoding errors caused by client behavior you don't control, categorize it out so the main metric can sit at zero. We've both led expensive, multi-quarter efforts to drive a category of error to zero: filesystem errors in one case, desktop client errors in another. The payoff wasn't just fewer errors, it was speed. Once you know Tuesday was clean, a new error on Wednesday tells you exactly when something happened. Without that baseline, every fluctuation costs you time proving whether it's real.
The same idea extends to data integrity. If you're running a database or storage system, a meaningful fraction of your workload should be background validation jobs scanning for corruption. The power of zero is worthless if you only discover the corruption months after it happened. Good systems engineering here means having enough confidence that yesterday was fine and today isn't, so a new problem announces itself instead of hiding in a slowly rotting number. If there are a few known-bad, ancient items you're never going to clean up, quarantine them explicitly and reset the metric around them. That way new noise doesn't hide behind old noise you've already accepted.
2. Triage
Triage means figuring out what broke and where it's coming from. It's a genuinely hard skill, because it tempts engineers toward heroics. Clicking seven dashboards deep and hunting through logs at 4 a.m. doesn't work well while stressed and half-awake. That's why it's worth investing ahead of time in a small number of dashboards that point straight at the likely source of a problem: which service is throwing errors, has load increased, is one customer doing something unusual, did the network drop, is the database slow. Somewhere around six graphs is enough to spot the issue quickly. More than that and you're back to hunting.
Two things actively hurt triage. First, someone loudly proposing a fix ("we could build a new system") before anyone knows what's wrong. Engineers naturally jump to solutions, and that jump derails the process of figuring out what happened. Second, blame. A zero-blame culture is operationally necessary here, not a nicety. If there's any fear of being blamed or disciplined for causing an incident, people stop calling out what they're seeing, and anything that discourages someone from raising a hand makes the whole team worse at responding. Between the two of us, across a lot of incidents, we can't think of a single time someone was disciplined for causing one. If your SEV designation is doing its job, the incident room stays a zone of safety to find the issue. Any "why did this happen" conversation gets written down for later, not raised in the moment.
We're cautious about leaning on AI heavily during triage. Some tools genuinely help. At the end of the day, though, you have to find and fix the real problem yourself, and you can't blame a model for getting it wrong. AI also tends to be unusually confident in its assertions during triage, which is the opposite of what the process needs. Human triage often works because someone notices a metric that's a little high, which nudges someone else to go look at something adjacent, and those small observations compound. AI wants to resolve the ambiguity and declare an answer. Accept that too quickly and you've effectively skipped triage, moving on before you know what's wrong.
Simplicity matters here too, more than it gets credit for. Complex systems are hard to triage almost by definition, and it's possible to build something so complicated that no one can reason about it during an incident. A useful design question, well before anything breaks: can someone else understand this system at 4 a.m., freshly woken up, well enough to fix it?
3. Stop the bleeding
Stopping the bleeding is distinct from debugging, and it's worth treating as its own step. You can often make a problem's impact go away quickly without understanding the root cause yet. If a system is corrupting data, turn it off. If a database is overloaded, shut down background load or pause a migration. If a service is optional, have it return a safe fallback. This can happen in parallel with debugging. Once the impact stops, everyone can breathe, and the team can work the real problem without the pressure of live customer impact.
This is where preparation pays off directly. In any large system there's background load (test traffic, benchmarking jobs, remediation scripts) that's a good candidate to shut off first. The value is in having a script or a button ready ahead of time, rather than trying to figure out how to stop a table compaction job under duress. At Convex, we have tools and playbooks that let the team run a single command to pause all background work and take load off the system, buying time and clarity to find the real problem. Segmenting live-site-critical services from things you can safely pause is a design decision worth making before you're in an incident, not during one.
4. Debugging
Once you've triaged and stopped the bleeding, debugging is regular engineering again. If triage narrowed the scope, debugging is a familiar kind of problem to work. It only gets brutal when you try to solve it without knowing what "it" actually is. Simple systems are easier to debug, and the engineers who deeply understand the system's structure are the ones who move fastest here. Part of the T-lock's job is making sure someone actually fixes it, or handing off explicitly to someone who will. Ambiguity about who's in charge means everyone acts in parallel and nobody's coordinated.
After triage, if the cause isn't environmental, AWS, DNS, or a new exploit, look hard at what changed recently. In our experience, something that changed in the last week or two is the cause roughly 80% of the time.
5. Remediation
Remediation is the "how do we prevent this from happening again" step, and it's where you have to be careful not to make things worse. Both debugging and remediation are technical work with no real shortcuts. A lot of what makes them tractable happens beforehand: tested backups, the ability to pause non-critical load, disaster recovery drills that actually exercise restoring from backup. Decide ahead of time, as a team, how you'll handle the awkward cases, so you're not arguing about it live during an outage. The same goes for compliance. Build in an exemption path for incident circumstances, with proper documentation, instead of arguing with your CISO while the company is down. Plan what happens when the VPN doesn't work, and know who has root on your infrastructure. The answer to all of this is practice. Companies that go through incidents get better at responding to them, so it's worth deliberately building that muscle, including through simulated incidents, without over-relying on the same one or two people who happen to be great at it.
Live communication is the multiplier
During an incident, this isn't the moment for email, and it might not even be the moment for Slack. Get on a call or into a room. There's real value in live communication for the response team specifically, not the whole company, just the people who are relevant. Anyone who isn't needed in the room should leave. The person who cracks an incident is often a junior engineer with less context, who hasn't made the same assumptions everyone else has and notices a spike in a log nobody else is watching. That kind of observation surfaces on a live call in a way it doesn't in a Slack thread. If someone says something like "huh, that's weird," the T-lock's job is to immediately ask what they saw, because that offhand noise is often exactly where the answer is hiding.
The culture in the room matters as much as the process. Incidents can go two ways. They can become sources of guilt, shame, and fear, none of which help anyone resolve anything faster. Or they can become gallows humor, joking, and the camaraderie of solving a hard problem together under pressure. Some of our best career memories come from exactly these rooms. During a massive outage at Dropbox, one engineer joked that it was a good thing their resume wasn't stored in Dropbox, and it landed because that same engineer was heads-down grinding on the fix. As long as people are working hard, a little levity is fine, even useful. Save the judgment for the postmortem, not mid-incident.
None of this means unlimited endurance is the goal. We've both been in incidents that ran multiple days, including stretches of 36-plus hours with no real break. It can happen, but it shouldn't happen often, and sustaining the team matters. If your best debugger has been at it for twelve hours, tell them to sleep, even if they argue. You want gas left in the tank for the next one.
Talking to customers without lying to them
Once you have a product people rely on, you have to tell them when something's wrong, and there's real nuance in how. The difference between a customer thinking "they've got this" and a customer thinking "I'm moving off this platform" often comes down entirely to communication, not to how bad the incident actually was.
The instinct to declare victory early is strong, a lot like the pressure in sales, because reassuring customers feels good in the moment. Resist it. The first message should just be "we're having an incident," sent as early as possible, without speculating about the cause before you know it. Overstate your confidence and you'll have to walk it back later, which costs you more trust than the delay would have. Customers need to know that you know something's wrong, that you're working on it, and that you have some sense of how long it might take, without overpromising on that last part.
The technical team and the communications side should run on different personas during the incident. The technical side stays focused on solving the problem, not on managing feelings. The person handling communication has to be the empathetic one, thinking about what the customer, or the customer's customer, actually needs to hear. That's rarely technical detail. Telling someone "the service hit a file descriptor limit" doesn't help them. Telling them whether they need to take action right now, and when to expect resolution if not, does. Run about ten to twenty minutes behind the war room in both specificity and confidence. Early conclusions in the room often change, and you don't want to have told customers something you then have to retract.
One of us got burned by this directly. During a big outage at Dropbox, someone DM'd asking if things were fixed, and the answer was "my filesystem is running." True personally, but not true company-wide, and it read as an official all-clear it was never meant to be. Speaking for yourself during an incident can easily be heard as speaking for the company. The better move is hedged language that's still useful: "we believe the issue is resolved, and we're monitoring for regressions," rather than declaring it closed before you're sure. Incidents are one of the clearest moments a customer gets to judge whether they trust you, and overpromising during one is a fast way to lose a customer who'd otherwise have stuck around.
Postmortems, and where this leaves off
The last phase is the postmortem, and it's easy for this to become a checkbox exercise. The real point of a postmortem is reflection that prevents the same thing from happening again, not a punitive record. Making a team that's already struggling spend half its time writing up a detailed timeline instead of fixing things is its own kind of waste. A good postmortem answers three things: what happened, why it happened, and what changes as a result. If there's genuinely nothing meaningful to say, it doesn't need to be long.
Postmortems are the close of the loop, not a formality bolted onto the end of it. The bigger throughline across all five stages is the same one that opened the episode. The goal was never to prevent every incident, since that's not achievable. It's to get good enough at responding to one that customers come away trusting you more, not less.
All gas, no breakages
Convex is the reactive backend platform that keeps up with you and your agents. Database, functions, workflow, sync, search, file storage, and more. All TypeScript, zero glue.