A few months after a clean migration, replication between the nodes in the global primary AG in a Distributed Availability Group (DAG) just stopped. There were no new deployments, config changes, or specific warning signs prior to this. Windows patching had occurred, the servers rebooted like they usually would, and the AG never came back together.
I know for myself, one of the most frustrating phrases that gives me the sense of impending doom is “nothing changed, why is this breaking?!”. That’s exactly what was going on here.
Read on to see how an innocent-looking configuration setting can cause issues down the road.