The Primary Is Down. Are You Sure?
October 20–23
Ubicloud runs managed PostgreSQL in three environments: our own bare metal fleet, AWS and GCP. Same control plane, same Postgres, same replication, same promotion sequence. So failover should work the same everywhere, right? It does not.
A failover is only safe if the old primary is really fenced before a standby starts taking writes. Each environment hands you different tools for that, and takes others away. Owning the hypervisor lets you do things a cloud API never will. Cloud APIs offer tools that only work if you planned your network for them. And sometimes the right answer is to not fail over at all.
This talk is the story of one failover procedure for all three: what we had to change per environment, what we refused to change, and the incidents that taught us the difference. We will cover detection, standby selection, fencing, DNS cutover and the WAL nobody archived.
You will leave knowing which of your failover assumptions are really assumptions about your infrastructure.