On Aug 3rd, 2026, from `5:02 AM UTC` till `5:13 AM UTC`, a routine automated update process failed to complete, causing a major downtime of our EU Database cluster.
Monitoring system alerted immediately our team, that started a manual override of the process, taking it to a successful end.
While our recurrent automated DB update process was running on the EU database it encountered an unexpected condition, that prevented the process to complete, leaving the cluster in an unconsistent state.
Some database engine items were not in the expected condition for the update to complete correctly and that caused the update to fail.
Moments after the initial failure of the automated process, our monitoring system alerted our infrastructure team.
Logical replication feature was not in the proper state after some previous replication activity in the cluster.
The current update process couldn't detect this situation in advance to avoid starting at all.
A few minutes after the beginning of the issue, our operators were able to reset the inconsistent parameters and apply the manual update procedure.
We are introducing a more robust update process that, by design, will not be prone to similar issues and will consistently limit the overall risk of errors during to database updates.