Hanging hundreds of Postgres shutdowns with a simple CDC plugin
Thursday, September 10 at 14:45–15:10
Cloud Nine
Intermediate
Whenever any simple switchover needed to be done, it took hours. The customer was killing the primary postmaster out of desperation because it was simply too long. It was reproducing perfectly every time, in every single production and most non-production environments. Was postgres hanging? Was it slow to finish something? Could some application hang a primary shutdown like this?
We will discuss debugging on Kubernetes, following the thread and debugging logical replication protocols, along with some simple pg_walreceiver patches