Blog Post

T-SQL Tuesday #202 SQL Server Outage You’ll Never Forget: A Roundup

,

When I put together the invitation for T-SQL Tuesday #202, I wasn’t sure what kind of stories would come out of it.

T-SQL Tuesday

The topic was simple:

That one SQL Server outage you’ll never forget.

I expected stories about bad queries, failed deployments, storage problems, or maybe a database that decided to have a very bad day. What I got was much more interesting.

There were stories about ransomware, corrupted databases, deleted storage, power and cooling failures, an identity column reaching its limit, and even a floppy disk. A server Meltdown too.

Some of these outages lasted hours. Others lasted days or even weeks.

And what I really enjoyed was that most of these stories weren’t just about what went wrong. They were about what people learned afterward.

Here are the stories I was able to collect.

Deborah Melkin: Sometimes a Floppy Disk Is All It Takes

Deborah Melkin shared a story from earlier in her career involving a server reboot that went very wrong. A bootable floppy disk had been left in the server. It contained an fdisk /mbr command. The server was rebooted, the command ran, and suddenly they had a much bigger problem than they expected.

The team had to find another server, reinstall SQL Server, and restore the databases from backup. The lesson seems obvious now, but that’s one thing I like about outage stories. Things that seem obvious after the fact aren’t always obvious when you’re standing in the middle of the problem.

Sometimes the lesson is simply to understand what is actually sitting in or connected to the server before you reboot it. And, of course, make sure your backups are somewhere other than the server you’re trying to recover.

Read Deborah’s full story

Andy Yun: The Outages That Live in Infamy

Andy Yun shares two non-SQL Server outages before getting to his SQL Server story. The first was when a backhoe literally dug up his company’s T1 line, leaving them without Internet for most of the day.

The second happened at a software company that supported market traders, where executives chose not to have a backup Internet connection after moving the office to VoIP. When the Internet went down, the entire office, call center, and data center were effectively offline, including their phones.

Both stories reinforced the same lesson: single points of failure hurt.

His SQL Server outage happened while Andy and another DBA were at PASS Summit, when a SAN administrator accidentally deleted the production transaction-log LUN for an instance with several hundred databases and several terabytes of data. Fortunately, their backups were good, and Andy used sp_restoregene to quickly generate the restore commands.

They initially tried four parallel restores, but the SAN couldn’t handle the load, so they stopped and worked with the business to prioritize the databases. The full recovery took three or four days.

What Andy realized afterward was that while they had tested restores for CHECKDB, they had never tested a full-instance restore at that scale. The experience also made him realize how important it was to understand the storage and infrastructure that your disaster recovery plan depends on.

Read Andy’s full story

Aaron Bertrand: When an INT Runs Out

Aaron Bertrand’s story is a good reminder that capacity problems aren’t always about disk space, memory, or CPU. In this case, an identity column reached the maximum value for an int. The result was an outage at Stack Overflow. The immediate solution wasn’t to change the column to bigint.

That would have been much more difficult to do during a live production outage. Instead, the identity was reseeded into the negative range, buying the team years of additional capacity. Problem solved. Until it happened again. A later operation involving IDENTITY_INSERT caused the identity value to move toward the positive limit again. So the team had another outage. And they reseeded it again.

I liked this story because it shows how an emergency fix can create another problem if the underlying issue isn’t eventually addressed. It also reminded me that we tend to think about capacity in terms of infrastructure. But data types have limits too. Sometimes the thing running out of room isn’t the disk.

Read Aaron’s full story

Rob Farley: When Corruption Leaves You With Very Few Options

Rob Farley wrote about an outage involving a bad disk controller that corrupted hundreds of database pages and even affected some backup files. This was one of those situations where the normal recovery path wasn’t enough. Rob and the customer’s CTO started looking at the database table by table, using clustered indexes, nonclustered indexes, DBCC PAGE, older backups, and other sources of information to reconstruct what they could.

One of the interesting parts of the story was discovering that a nonclustered index could still contain information that was missing from the corrupted clustered index. That meant even damaged parts of the database could potentially be useful during recovery.

Eventually, they were able to rebuild the tables and indexes and get the system back online. Reading this made me think about how different troubleshooting becomes when you’re no longer trying to find the best query plan or fix a blocking problem. When you’re dealing with serious corruption, you’re looking for anything that can help you recover the data.

Read Rob’s full story

Vlad Drumea: Two Weeks of Ransomware Recovery

Vlad Drumea shared probably one of the biggest incidents in this collection. His organization was hit by Ryuk ransomware in 2020. More than 60 SQL Server instances across more than 30 VMs were involved. Recovery took more than two weeks. This wasn’t simply a matter of restoring a database.

The environment itself had to be treated as compromised. Vlad described rebuilding VMs, recreating SQL Server directories, restoring system databases, restoring user databases, dealing with reinfection, fixing broken LSN chains, and rebuilding a VM from scratch. There was also a lot of automation involved using PowerShell, T-SQL, and dbatools. What really stuck with me was the human side of this story.

Recovery involved 16-hour workdays for more than two weeks, and Vlad talks about the burnout that followed. We spend a lot of time talking about backups, DR, security, and automation. Those things matter. But there are also people sitting in front of those computers at 2 AM trying to get a business back online. That’s part of the story too.

Read Vlad’s full story

Jeff Taylor: SQL Server on Fire, Literally!

Jeff Taylor shared two stories involving infrastructure problems. The first started with a power outage in an office that had effectively become a small data center. The servers had battery backup. The air conditioning didn’t. As the room heated up, the team started using fans and eventually began shutting servers down to keep the hardware from being damaged.

The temperature eventually approached 120°F. That incident resulted in a much larger infrastructure redesign, including better cooling, battery backup for the cooling, a generator, and fire suppression.

Then there was another incident involving a Dell server where Jeff was replacing memory and a drive. After powering it back on, he saw a flash. Then smoke. The problem apparently involved an iSCSI cable that had been damaged during the earlier heat incident and eventually shorted when the equipment was moved.

It’s a good reminder that SQL Server doesn’t operate in a vacuum. Power, cooling, storage, networking, and the physical environment are all part of keeping a database available.

Read Jeff’s full story

Thomas Rushton: When the Server Room Gets Too Hot

The Lone DBA shared a story about a server room in an old Victorian mill building that overheated after the air conditioning failed during a hot summer weekend.

The servers were shut down, and after things cooled down, most of them came back online without any obvious problems. One server, however, kept crashing intermittently. They patched it, updated drivers, replaced the memory, HBAs, CPUs, and even the storage, but nothing fixed the problem.

Eventually, while replacing the motherboard, an engineer discovered that a daughterboard had partially melted during the overheating event, causing an intermittent short circuit. It’s a good reminder that a server room getting too hot isn’t just a temporary availability problem. It can cause physical hardware damage that may not show up until much later.

Read The Lone DBA’s full story

M G: A Comment That Deserved to Be Part of the Roundup

A person who goes by their initial M G don’t have a blog, but left a detailed comment on my invitation. I thought the story was too interesting to leave out.

The incident involved SQL Server 2016 and a SharePoint database with around 15 million rows and more than 850 GB of binary data. Three databases were repaired, but the fourth became the real problem. There were hundreds of suspect pages, a corrupted clustered index that was also the primary key, and a DBCC CHECKTABLE ... REPAIR operation that had been running for weeks and failing.

What caught my attention was the eventual workaround. Changing the database’s PAGE_VERIFY setting from CHECKSUM to OFF allowed the primary key to be dropped. That exposed another problem: SharePoint had accumulated roughly 90,000 duplicate elements. This is exactly the kind of kind of troubleshooting story that is difficult to forget because there isn’t necessarily a clean checklist that tells you what to do next. You investigate. You try something. You learn something new. Then you try again. And sometimes the solution comes from a place you weren’t expecting.

What I Took Away From These Stories

After reading through all of these, I noticed something. The actual cause of the outage was often not SQL Server itself. It was the environment around SQL Server.

A floppy disk. A SAN administrator deleting a LUN. An identity value reaching its limit. A bad disk controller. Ransomware. A lack of cooling. Corruption inside a SharePoint database.

These are very different problems, but they have something in common.

You don’t always know what the outage is going to look like until you’re already in it.

That’s probably why these stories are useful. You can study SQL Server performance. You can learn backup and restore. You can learn Availability Groups. You can learn PowerShell and dbatools. You can learn monitoring. But eventually, something unexpected is going to happen.

The best thing we can do is learn from people who have already been there.

And that’s what I really liked about this month’s T-SQL Tuesday. These weren’t polished success stories. They were stories about things going wrong.

And those are often the stories I remember the longest.

Thank you to everyone who took the time to participate in T-SQL Tuesday #202, whether you wrote a full post or shared your experience in the comments.

And thank you to Steve Jones for giving me the opportunity to host this month’s T-SQL Tuesday.

Until the next outage…

No, God forbids. It ould be yours, and hopefully the stories above give you the resolution route.

The post T-SQL Tuesday #202 SQL Server Outage You’ll Never Forget: A Roundup first appeared on SQL, Code, Coffee, Etc..

Original post (opens in new tab)
View comments in original post (opens in new tab)

Rate

You rated this post out of 5. Change rating

Share

Share

Rate

You rated this post out of 5. Change rating