Press "Enter" to skip to content

Curated SQL Posts

Asynchronous Snapshot Replication from ActiveCluster Pods to Other Arrays

Anthony Nocentino digs into some very neat and very expensive things:

I’ve been rebuilding my three-site SQL Server demo lab, and I ran into something I’ve wanted for a long time. If you’ve ever designed a SQL Server environment on ActiveCluster, you know the pattern: two FlashArrays running a synchronously replicated pod for zero RPO between sites, and a third array somewhere else for a longer retention, disaster recovery copy. The problem was that you couldn’t get the data to that third array directly from the pod. Protection groups inside a stretched pod simply couldn’t have an array target.

That’s changed, and it’s been possible longer than a lot of us realize. You can create a protection group inside an ActiveCluster pod, add a third FlashArray as a target, and asynchronously replicate snapshots to it on a schedule. Your synchronously replicated data gets a third copy, and you don’t have to build a parallel set of non-pod volumes to make it happen.

In this post, I’m going to show you how to configure this end to end with the Pure Storage PowerShell SDK2, so you can automate it. Let’s go.

I mean, sure, you need multiple FlashArrays to do this. But who doesn’t have a few of those floating around?

Leave a Comment

The State of JSON Indexing in SQL Server 2025

Greg Low shares some thoughts:

SQL Server 2025 finally gives developers a native JSON data type and, with it, a purpose-built way to index JSON documents with the new CREATE JSON INDEX statement. Before this, indexing JSON meant exposing individual properties through computed columns and building standard indexes on top.

It’s a major step toward closing the gap with databases like PostgreSQL, long praised for its JSON and JSONB support. However, as a preview feature, JSON indexing comes with real constraints DBAs and developers should understand before adopting it.

This guide has everything you need to know about JSON indexing in SQL Server 2025: what it is, how it works, and current limitations.

Click through to learn more.

Leave a Comment

Playing Poker in T-SQL

Brent Ozar is a madman and I love it:

Wanna play some Texas Hold ‘Em style poker and make a pile of Query Bucks? Wanna watch other database people playing?

Click through to learn more about the game itself, but also the strategic choices Brent made and a bit of an after-action report on what it took to generate this code.

Also, shout out to Brad Shultz, whose blog I miss. It was his T-SQL Tuesday on the APPLY operator that really opened my eyes to how good that operator is.

Finally, I’m giving this post the most coveted tag in Curated SQL: Wacky Ideas. It’s rare that I get to use this one.

Leave a Comment

Charting Average Full Database Backup Durations with R

Thomas Williams has a script:

DBAs spend time dealing with SQL Server performance, capacity, monitoring, troubleshooting, provisioning, etcetera; I’ve previously mentioned that R is a powerful, open-source language with a great ecosystem of libraries for analysis and visualisation, so it’s no surprise that I think mixing SQL Server and R Markdown for reporting goes together like a Vegemite and cheese sandwich…lunch perfection!

Here’s a couple of real-world examples of how I’ve used R Markdown connected to SQL Server (for a recap of how to do this from a technical perspective, see my earlier blog post “Connecting to a SQL Server database from R Markdown”):

This is a neat approach to visualizing database backup times as a process control chart.

Leave a Comment

Introducing sp_CheckHealth

Jeff Iannucci announces a new stored procedure:

This tool will give you a fast, comprehensive picture of a SQL Server instance. It gathers the kind of information you would otherwise collect by clicking through a dozen dialogs and running a handful of scripts, and it flags potential issues so you can decide what to deal with first. The findings are organized into categories like Recoverability, Security, Availability, Integrity, Reliability, and Performance, and each one comes with details and an action step so you aren’t left guessing about what to do next.

Click through for the script and how you can use it.

Leave a Comment

Surrogate Keys in Fabric Data Warehouse

Louis Davidson does some more digging:

Creating the simplest of tables, the first thing I start thinking about is making sure that the data is going to be protected from the user. Users (including myself on my own projects when I have my user hat on) don’t notice when they start inserting poor quality data sometimes. Oops, I hit F5 twice, I wonder if that will affect my data? Without proper constraints, it probably will.

In this second entry, I want to cover a few things about handling surrogate values that will help you avoid some of the (in retrospect) kind of dumb expectations that I had. Fabric Data Warehouse T-SQL feels so much like SQL Server Relational T-SQL that some stuff like choosing a surrogate key makes me think I am missing something.

One of the things Louis mentions is how the values get inserted and how it looks like they’re in different ranges. This makes sense, as each distributed node likely has its own range of identity values, similar to the way merge replication would work with identity keys to prevent overlap.

Leave a Comment

Row vs Page Compression in Animated Form

Brent Ozar has a new animation:

What’s the difference between SQL Server’s row compression and page compression, and when does each one make sense?

  • Row compression turns every fixed-length datatype into a variable-length datatype, using as little space as possible to store it
  • Page compression does that, AND adds a dictionary of repeated data on the page, getting more compression at the cost of more CPU

Here’s my dirty little secret: I don’t think row-level compression makes sense all that often, simply because I’m not sure I’ve ever seen negative consequences to page level compression, even in a variety of scenarios in very busy environments. I’m sure that there are specific cases, but I just default to page level compression because of how well it works.

Leave a Comment

What’s New with the VSCode MSSQL Extension

Yo-Lei Chen shares some updates:

Writing and maintaining SQL is easier when you can eliminate repetitive steps and keep your scripts cleanly formatted. With the MSSQL extension for VS Code v1.45, we’re introducing the Public Preview of the SQL Formatter alongside the General Availability of Azure SQL Database Provisioning and Shortcuts Configuration. You can now apply consistent T-SQL formatting across your projects, create free tier cloud databases with automated post-deployment actions, and streamline frequently used commands and queries directly inside Visual Studio Code.

The SQL Formatter is potentially interesting, inasmuch as you’re able to control the settings yourself rather than relying on a pre-defined format.

Leave a Comment

Database Application Security and High Availability Checklist

Andreas Wolter has an update:

I have updated the SQL Server Database Application Security & High Availability Checklist and moved the current version to the Sarpedon Quality Lab website: View here

The checklist is written for two audiences:

Database application vendors who want their SQL Server-backed products to be easier to approve in enterprise environments.

DBAs, security administrators, and architects who need to evaluate whether a vendor application can be deployed securely.

Click through to see what’s new, and check out the link for the full checklist.

Leave a Comment

Large Language Models for Data Professionals

Eugene Meidinger has a primer:

It’s tempting to think that working with LLMs and AI agents doesn’t require understanding anything about how they work. These tools are often presented as autonomous, intelligent, and self-explanatory. They communicate through conversational text, making them feel natural and intuitive, and can often even seem like magic.

In practice, however, these intuitions about LLMs are often very wrong. LLMs are a strange technology. As we stack tools and agent interfaces (or harnesses) on top of them, more and more misunderstandings also stack up. Treating them as an easy button leads to layers of frustration, waste, and, in the worst-case scenario, mistakes.

This is a nice baseline to get someone started with understanding LLM output behavior. It’s not a how-to guide, but rather a “Here’s what’s going on” type of guide. Eugene brings less snark to the topic than I do, but that’s par for the course for anyone who remembers the good ol’ days of the SQL Data Partners podcast.

Leave a Comment