Shaman outage resolved
Around 15:24 UTC today, I noticed calls to shaman.ceph.com were getting refused. Upon investigation, I observed the nginx service on shaman.ceph.com (the load balancer in front of 1.shaman and 2.shaman) was in failed state. Oct 01 06:37:37 shaman systemd[1]: Stopping nginx.service - A high performance web server and a reverse proxy server... Oct 01 06:37:37 shaman nginx[2274090]: 2025/10/01 06:37:37 [emerg] 2274090#2274090: host not found in upstream "1.shaman.ceph.com" in /etc/nginx/sites-enabled/01-shaman.conf:35 Oct 01 06:37:37 shaman nginx[2274090]: nginx: configuration file /etc/nginx/nginx.conf test failed I manually restarted the service at 15:16:48 UTC. Some immediate actions we’re taking to improve this story: * Improved alerting. We used to have an nginx service that would have caught this and sent alerts but gluster ate the VM’s disk. We have a Grafana instance but its alerts are too noisy. I’ve asked Adam to create a way to prioritize alerts for outages like this. * Put entries in /etc/hosts for the downstream shaman hosts Next year, I aim to launch another status portal like we had at status.sepia.ceph.com. Any branches pushed between ~06:00 and 15:16 UTC should be force pushed to retrigger a build. -- David Galloway Ceph Engineering Labs – Infrastructure Architect +1 989 295 0091 - Mobile david.galloway@ibm.com IBM
participants (1)
-
David Galloway