[RGW] "Failed to reserve notification on queue: error: -22" after upgrading Reef to Tentacle(20.2.0)
Hello Ceph community, I am seeking advice regarding an RGW notification issue encountered after upgrading our cluster from Reef to Tentacle. We heavily use the RGW PubSub (topic/notification) feature. We recently upgraded to Tentacle specifically to leverage new reliability options such as time_to_live, max_retries, and retry_sleep_duration to prevent data loss during transmission. The Problem: After the upgrade, when using existing topics created in the Reef version, we started seeing the following error: "Failed to reserve notification on queue: error: -22" Consequently, all PUT requests to the affected buckets are failing with an error. (s3 response include error message "<Code>InvalidArgument</Code> Steps Taken (Troubleshooting): To resolve this without deleting our existing topics/notifications, we have attempted the following: - Restarted all RGW instances. - Redeployed the RGW gateways. - Performed rados clearomap on the pubsub..bucket.xxxx objects found in the logs. Despite these efforts, the PUT operations continue to fail with the same error. We are trying to find a fix that doesn't involve recreating the topics and notifications, as we are still in the process of debugging the root cause. Questions: - Has anyone encountered this -22 (EINVAL) error specifically during a Reef-to-Tentacle transition involving PubSub? - Are there any internal OMAP structure changes in Tentacle that might cause incompatibility with legacy Reef topics? - Any suggestions or ideas on how to clear this "queue reservation" error without losing the current notification configuration? Any insights or guidance would be greatly appreciated. Best regards,
steve jung wrote:
Hello Ceph community,
I am seeking advice regarding an RGW notification issue encountered after upgrading our cluster from Reef to Tentacle.
We heavily use the RGW PubSub (topic/notification) feature. We recently upgraded to Tentacle specifically to leverage new reliability options such as time_to_live, max_retries, and retry_sleep_duration to prevent data loss during transmission.
The Problem:
After the upgrade, when using existing topics created in the Reef version, we started seeing the following error: "Failed to reserve notification on queue: error: -22" Consequently, all PUT requests to the affected buckets are failing with an error. (s3 response include error message "<Code>InvalidArgument</Code> Steps Taken (Troubleshooting):
To resolve this without deleting our existing topics/notifications, we have attempted the following: - Restarted all RGW instances. - Redeployed the RGW gateways. - Performed rados clearomap on the pubsub..bucket.xxxx objects found in the logs. Despite these efforts, the PUT operations continue to fail with the same error. We are trying to find a fix that doesn't involve recreating the topics and notifications, as we are still in the process of debugging the root cause.
Questions:
- Has anyone encountered this -22 (EINVAL) error specifically during a Reef-to-Tentacle transition involving PubSub? - Are there any internal OMAP structure changes in Tentacle that might cause incompatibility with legacy Reef topics? - Any suggestions or ideas on how to clear this "queue reservation" error without losing the current notification configuration? Any insights or guidance would be greatly appreciated.
Best regards,
You can look up https://tracker.ceph.com/issues/74713 could be due to this, and not be related to upgrade. You can run this bash script and check for the counter value, if that value is around 128Million then you have hit the above issue and unfortunately there is no easy way to clear the issue, you will have to modify the notification rados object to update the counter value. else you can patch the PR attached to tracker, that will automatically reset the counter Run the below script on node that has access to cluster, it will print the counter value for each of notification rados object, ```` #!/bin/bash set -euo pipefail NS="notif" OFF=$((0x5e)) SUFFIX="rgw.log" tmp="/tmp/rados_obj.$$" trap 'sudo rm -f "$tmp" 2>/dev/null || true' EXIT # Find the (single) rgw.log pool POOL=$(sudo rados lspools 2>/dev/null | awk '{print $NF}' | grep -E "${SUFFIX}$" | head -n 1 || true) if [[ -z "${POOL}" ]]; then echo "ERROR: Could not find a pool ending with '${SUFFIX}' via: sudo rados lspools" >&2 echo "Try: sudo ceph osd lspools" >&2 exit 1 fi # Only RGW* objects in notif namespace sudo rados -p "$POOL" -N "$NS" ls 2>/dev/null \ | awk '$0 ~ /^RGW/' \ | while IFS= read -r obj; do [[ -z "$obj" ]] && continue if sudo rados -p "$POOL" -N "$NS" get "$obj" "$tmp" 2>/dev/null; then val=$(dd if="$tmp" bs=1 skip="$OFF" count=8 status=none | od -An -t u8 | tr -d ' ') printf '%s\t%s\n' "$obj" "$val" else printf '%s\t%s\n' "$obj" "ERROR_GET_FAILED" >&2 fi done ```
This may be related to Tracker #74713 (https://tracker.ceph.com/issues/74713) rather than the upgrade itself. If so, a fix would require backporting the associated PR. Otherwise, the only workaround is to delete and recreate the topic. To confirm if your environment is affected, please run the attached bash script. If the output shows a counter value near 128 Million, it confirms the issue is indeed related to the tracker. ``` #!/bin/bash set -euo pipefail NS="notif" OFF=$((0x5e)) SUFFIX="rgw.log" tmp="/tmp/rados_obj.$$" trap 'sudo rm -f "$tmp" 2>/dev/null || true' EXIT # Find the (single) rgw.log pool POOL=$(sudo rados lspools 2>/dev/null | awk '{print $NF}' | grep -E "${SUFFIX}$" | head -n 1 || true) if [[ -z "${POOL}" ]]; then echo "ERROR: Could not find a pool ending with '${SUFFIX}' via: sudo rados lspools" >&2 echo "Try: sudo ceph osd lspools" >&2 exit 1 fi # Only RGW* objects in notif namespace sudo rados -p "$POOL" -N "$NS" ls 2>/dev/null \ | awk '$0 ~ /^RGW/' \ | while IFS= read -r obj; do [[ -z "$obj" ]] && continue if sudo rados -p "$POOL" -N "$NS" get "$obj" "$tmp" 2>/dev/null; then val=$(dd if="$tmp" bs=1 skip="$OFF" count=8 status=none | od -An -t u8 | tr -d ' ') printf '%s\t%s\n' "$obj" "$val" else printf '%s\t%s\n' "$obj" "ERROR_GET_FAILED" >&2 fi done ```
Thanks for your information and script. The script didn't work as expected in my environment, so I ran the following command instead. Here are the results. root@ljc-ceph-132:~# sudo rados -p suwon-1.rgw.log -N "notif" ls 2>/dev/null \ | awk '$0' \ | while IFS= read -r obj; do if sudo rados -p suwon-1.rgw.log -N "notif" get "$obj" "./rados_1" 2>/dev/null; then val=$(dd if="./rados_$obj" bs=1 skip="$((0x5e))" count=8 status=none | od -An -t u8 | tr -d ' ') printf '%s\t%s\n' "$obj" "$val" else printf '%s\t%s\n' "$obj" "ERROR_GET_FAILED" >&2 fi done (Results) :sdt-webhook-topic-new 0 queues_list_object :sdt-webhook-topic-130 0
participants (2)
-
kchheda3@bloomberg.net
-
steve jung