steve jung wrote:
Hello Ceph community,
I am seeking advice regarding an RGW notification issue encountered after upgrading our cluster from Reef to Tentacle.
We heavily use the RGW PubSub (topic/notification) feature. We recently upgraded to Tentacle specifically to leverage new reliability options such as time_to_live, max_retries, and retry_sleep_duration to prevent data loss during transmission.
The Problem:
After the upgrade, when using existing topics created in the Reef version, we started seeing the following error: "Failed to reserve notification on queue: error: -22" Consequently, all PUT requests to the affected buckets are failing with an error. (s3 response include error message "<Code>InvalidArgument</Code> Steps Taken (Troubleshooting):
To resolve this without deleting our existing topics/notifications, we have attempted the following: - Restarted all RGW instances. - Redeployed the RGW gateways. - Performed rados clearomap on the pubsub..bucket.xxxx objects found in the logs. Despite these efforts, the PUT operations continue to fail with the same error. We are trying to find a fix that doesn't involve recreating the topics and notifications, as we are still in the process of debugging the root cause.
Questions:
- Has anyone encountered this -22 (EINVAL) error specifically during a Reef-to-Tentacle transition involving PubSub? - Are there any internal OMAP structure changes in Tentacle that might cause incompatibility with legacy Reef topics? - Any suggestions or ideas on how to clear this "queue reservation" error without losing the current notification configuration? Any insights or guidance would be greatly appreciated.
Best regards,
You can look up https://tracker.ceph.com/issues/74713 could be due to this, and not be related to upgrade. You can run this bash script and check for the counter value, if that value is around 128Million then you have hit the above issue and unfortunately there is no easy way to clear the issue, you will have to modify the notification rados object to update the counter value. else you can patch the PR attached to tracker, that will automatically reset the counter Run the below script on node that has access to cluster, it will print the counter value for each of notification rados object, ```` #!/bin/bash set -euo pipefail NS="notif" OFF=$((0x5e)) SUFFIX="rgw.log" tmp="/tmp/rados_obj.$$" trap 'sudo rm -f "$tmp" 2>/dev/null || true' EXIT # Find the (single) rgw.log pool POOL=$(sudo rados lspools 2>/dev/null | awk '{print $NF}' | grep -E "${SUFFIX}$" | head -n 1 || true) if [[ -z "${POOL}" ]]; then echo "ERROR: Could not find a pool ending with '${SUFFIX}' via: sudo rados lspools" >&2 echo "Try: sudo ceph osd lspools" >&2 exit 1 fi # Only RGW* objects in notif namespace sudo rados -p "$POOL" -N "$NS" ls 2>/dev/null \ | awk '$0 ~ /^RGW/' \ | while IFS= read -r obj; do [[ -z "$obj" ]] && continue if sudo rados -p "$POOL" -N "$NS" get "$obj" "$tmp" 2>/dev/null; then val=$(dd if="$tmp" bs=1 skip="$OFF" count=8 status=none | od -An -t u8 | tr -d ' ') printf '%s\t%s\n' "$obj" "$val" else printf '%s\t%s\n' "$obj" "ERROR_GET_FAILED" >&2 fi done ```