All connections to the unhealthy instance were terminated, and error rates fully recovered.
We're seeing recovery across the board, with runners no longer failing to acquire jobs. Queue time has returned to normal across all jobs.
This page lists incidents by Namespace and known upstream issues.