Around 15:24 UTC today, I noticed calls to shaman.ceph.com were getting refused.
Upon investigation, I observed the nginx service on shaman.ceph.com (the load balancer in front of 1.shaman and 2.shaman) was in failed state.
Oct 01 06:37:37 shaman systemd[1]: Stopping nginx.service - A high performance web server and a reverse proxy server...
Oct 01 06:37:37 shaman nginx[2274090]: 2025/10/01 06:37:37 [emerg] 2274090#2274090: host not found in upstream "1.shaman.ceph.com" in /etc/nginx/sites-enabled/01-shaman.conf:35
Oct 01 06:37:37 shaman nginx[2274090]: nginx: configuration file /etc/nginx/nginx.conf test failed
I manually restarted the service at 15:16:48 UTC.
Some immediate actions we’re taking to improve this story:
Next year, I aim to launch another status portal like we had at status.sepia.ceph.com.
Any branches pushed between ~06:00 and 15:16 UTC should be force pushed to retrigger a build.
--
David Galloway
Ceph Engineering Labs – Infrastructure Architect
+1 989 295 0091 - Mobile
david.galloway@ibm.com
IBM