Something Stronger Than Coffee
As anyone in IT will tell you, when an outage strikes, and all hands are on deck, that is something that will wake you up faster than any cup of coffee. That was me. Waking up trying to get my bearings to take part in addressing an outage with my bed-head in full effect (Thank you Microsoft Teams for giving that little preview window to show what you look like before you turn on your camera).
So What Happened?
This was a DR scenario that in my book was executed well by everyone involved. Each team and team member knew their role, and handled a chaotic situation very well. To keep a long story short, there were some unforeseen issues that resulted in the HVAC for our primary data center to go down. Temperatures rose drastically and we were at risk of having our hardware malfunction and potentially break due to the rising heat. To get out in front of this, among the many tasks the groups had to do to adapt to this issue, part of that was having to failover all of our Availability Groups that run out of this data center to our secondary data center and shut down those VMs in the primary data center to try and let our hardware cool down.
How We Were Impacted
With the now secondary replicas shut down, we were left in a new predicament: One of our two replicas in our AGs were offline for an extended period of time. Even with our AGs running in async, this meant a couple things:
- Log growth.
- CDC and Replication would continue to get further and further behind.
Obviously another impact is a potential lack of availability running off one replica, but I’m going to focus on point number two because this is where I learned something.
CDC And Replication In An AG With Only One Replica Available
We have CDC turned on for many of the databases in our production environment. We also have a fair amount of replication as well. When you combine an AG + replication + CDC + one or more unavailable replicas in your AG you get an ever growing number of built up uncaptured transactions. Looking back, I could expect replication to have issues, but I was surprised initially that CDC had problems as well. BUT, once I remembered that CDC leverages replication log reader magic under the hood, it clicked for me. The reason it stalls is that the log reader won’t advance past log records that aren’t hardened on every secondary, so a failover can’t lose replicated data. Here’s what I saw in replication monitor that clued me in as to what was going on for CDC for a database that is both replicated and has CDC turned on:

I saw this in both the details of the log reader agent, and the distribution agent (publisher to distributor tab).
What This Means And What To Do
My example database had both replication and CDC turned on, but the same stall applies even if you only have CDC configured, no replication involved.
The only way to fix this is to either pull your database in question out of the AG so that it is no longer waiting on a non-responsive replica to harden transactions, or bring the replica back online so the database can catchup, and then CDC and replication can start to flow. As soon as I pulled one of the databases where we needed CDC / replication to flow from the AG, everything caught up while the other replica was still offline. Of course, there’s a tradeoff to this. You pull the database(s) out of the AG, you lose HA, and then have to re-seed that database later on. But in my case, a few critical databases where downstream processes were dependent on CDC running made that tradeoff worth it.
How To Know When To Make A Decision
So how do you know when you need to pull a database from an AG facing these same issues? Well, it depends. It depends on business impact, database size, workload, duration of the outage, and much more. For me, there were several things at play:
- Business impact for CDC being behind was significant and directly affected critical processing.
- Database size was under 300GB for the few databases I pulled from their respective AG, so having to re-seed these databases wasn’t a huge deal.
- Workload against these databases was moderate and was receiving many intraday critical transactions that needed to be CDC’d to downstream processes.
- The duration of the outage was unknown at this time. At this point we were several hours into it with no end in sight at the time, so it was safer to take action now than to wait.
All of these are business and team decisions that are unique to your situation and environment that’ll influence your specific steps.
Another Option In The Playbook
Something I discovered after the fact is TF 1448. If enabled on the primary replica, it’ll allow the log reader agent to progress and not wait for any ASYNCHRONOUS replicas to acknowledge a transaction. There are some major tradeoffs though. The risk here is that if a failover happens while an async secondary is behind, your new primary could be missing transactions that were already sent downstream, so your subscribers or CDC consumers end up with data that doesn’t actually exist anymore. In my case, the secondary was fully offline, not just lagging, so there wasn’t really a failover target to worry about. Rather than dive deep into this myself, Garry Bargsley has an amazing writeup of this. If I had known about this at the time I definitely would have brought this up to my team to consider.
Tools If This Happens To You
Here is a script you can use to quantify built up transactions (works for CDC or replication) in your databases:
USE masterGOSELECT * from sys.dm_os_performance_countersWHERE counter_name = 'Repl. Pending Xacts'GO
And here is a snippet to check on what your capture jobs are doing for CDC:
USE [DatabaseName]GOselect *from sys.dm_cdc_log_scan_sessions
What I Learned
CDC and replication have a major dependency on synchronization health when dealing with them in an AG. If one replica is unresponsive or down for the count, it could leave you in a world of hurt.

Leave a comment