Press "Enter" to skip to content

Category: T-SQL Tuesday

100 Hours of Work

Steve Jones talks about a tough week:

I was a relatively new hire, a former intern, at a large electrical utility in Virginia. I worked as a network admin at a nuclear power plant in Surrey, VA. I showed up at work on New Year’s Eve at 5:00pm. We were planning on deploying a new database server running SQL Server, along with a new application to track radiation exposure for workers. This was a mandated change to our tracking, which needed to go live at midnight. I was supposed to be a bystander, helping developers from our internal group implement the server and then take over administration for the future.

What could possibly go wrong?

Leave a Comment

The Server Room in the Attic

Thomas Rushton whips out a story:

Anyway. Big Victorian mill-type building. First floor, document management, printing, post rooms, etc, going up. Second floor – erm… document processing? Finance? Third floor – IT, call center, network room, board room, lawyer-types. Fourth floor, just underneath that slate roof, held a smaller office full of debt recovery specialists, and the main server room.

What could possibly go wrong?

Leave a Comment

A Pernicious Floppy Disk

Deb Melkin deals with disaster:

One morning, the help desk support person came over and said this one client was having performance issues. This was one of the cases where we would just reboot. It was early and there would be plenty of time for everything to be back up and running before the system was needed. So I gave the go ahead to reboot the server.

This did not go well. Read on for the rest of the story.

Leave a Comment

A Trio of Outages

Andy Yun has been a part of several outages over the years:

My very first job out of college, I worked at a small boutique internet consulting firm that did some hosting in our back office. We had a T1 line coming into our office, which was a space that was part of a strip mall actually. One summer morning, we’re at our desks doing our thing and “hey, I just lost Internet connection, anyone else?” 

Click through for the conclusion of that story, as well as two other times everything went sideways. This might also be a good time to check your backups and make sure they’re working fine.

Leave a Comment

Fixing a Problem from the Naval Yard

Andy Levy tells us a story:

A buddy of mine invited me to take a road trip and join him in Philadelphia, PA to take a tour of the USS New Jersey while she was in drydock for its required museum ship maintenance. Quite possibly a once in a lifetime opportunity – we wouldn’t be touring the inside of the ship, we would be walking around underneath her, on the floor of the drydock. I hit the road on Friday, and picked him up at the airport. We got dinner, unwound from our respective travels, and planned out the following day.

Click through to see what happened next.

Leave a Comment

To Make a DBA

Andy Brownsword tells his origin story:

I enjoy a good war story.

But no, this isn’t the time Azure blindsided me, when service accounts were locked, or when I left a transaction open and halted a production system (easily fixed by closing SSMS 😅). Mine is a pivotal personal moment.

Andy humbly leaves out the part where he was hand-to-hand kung fu fighting a series of ninjas while doing all of this, but you can just assume that in there.

Leave a Comment

A Tale of Time

Chad Callihan gives us a rundown:

I’ve written about a few outages and issues in the past and wanted to write about a different one for this post. One that came to mind was from awhile ago when a code deployment appeared to be successful on the surface, but in reality turned into a time consuming mess.

Click through for the story. This is why I strongly recommend always storing dates and times in UTC format in SQL Server, so using GETUTCDATE() and its brethren versus GETDATE() and co. You can always translate from UTC into the relevant time zone, you don’t have to worry about daylight savings time fouling things up twice a year, and you always have a known starting point, regardless of where your server is located. This ties in with DR scenarios: if your main server is on the US east coast and your DR site is on the US west coast and both are using GETDATE() to store times, you’ll see problems after a failover.

Leave a Comment

An Outage from Corruption

Rob Farley tells a story:

Naturally when there’s a failure, there are questions about High Availability solutions. If the customer has an Availability Group, maybe they’ve avoided an outage completely – the primary fails over to the secondary, and the race starts to get the original machine back in case another failure (on the new primary) happens. Those situations aren’t the ones that keep coming back to mind, but are still stressful for the temporary loss of the safety net.

The bigger issues are when there is no quick failover available. About thirty years ago, I was dealing with a SQL Server 6.0 environment, where the customer’s DBA was cycling tapes (yes, tapes) for differential backups, except that he’d reused the tape that the full backup was on. When that system went down, I had a quick and harsh introduction to salvaging files off the disks from a failed computer. Fun times.

Also, it sounds like Rob’s surgery from last month has gone well, so hopefully he remains on the road to recovery.

Leave a Comment

Recovering from Ransomware

Vlad Drumea tells a story:

This is about the outage caused by a ransomware attack in October 2020 on my then employer, Steelcase, and the 2+ weeks required to get everything back online.

You can read more about the two-week halt of global order management, manufacturing, and distribution, as well as the fact that forensic investigation found no evidence of exfiltration, in this BleepingComputer article.

Disclaimer: I mention and link some products and companies in this blog post. This isn’t a sponsored post, and none of these companies have had a say in what I’m writing here.

Click through for Vlad’s story.

Leave a Comment

Blind Spots and Troubleshooting a Cluster Problem

Alexander Arvidsson tells a story:

The patient was a misbehaving SQL Server 2014 running in a two-node cluster. It was your garden-variety cluster with a shared disk and a remote witness. In my experience, as long as the customer has enough know-how to maintain a cluster like this, it’s essentially bulletproof.

This specific cluster was not.

Click through for the story, as well as a reminder of why people in mission-critical situations tend to follow checklists and say the items aloud.

Leave a Comment