Frequest LARGE_OMAP_OBJECTS in cephfs metadata pool
Hello Team , I am getting frequent LARGE_OMAP_OBJECTS 1 large omap objects in one of my cephfs metadata pools , anyone can explain why would this pool getting into this state frequently and how could I prevent this in future ? # ceph health detail HEALTH_WARN 1 large omap objects LARGE_OMAP_OBJECTS 1 large omap objects 1 large objects found in pool 'cephfs01-metadata' Search the cluster log for 'Large omap object found' for more details. Thanks , Uday
On Mon, Feb 24, 2020 at 11:14 AM Uday Bhaskar jalagam <jalagam.ceph@gmail.com> wrote:
Hello Team ,
I am getting frequent LARGE_OMAP_OBJECTS 1 large omap objects in one of my cephfs metadata pools , anyone can explain why would this pool getting into this state frequently and how could I prevent this in future ?
# ceph health detail HEALTH_WARN 1 large omap objects LARGE_OMAP_OBJECTS 1 large omap objects 1 large objects found in pool 'cephfs01-metadata' Search the cluster log for 'Large omap object found' for more details.
When was the file system created? What version is running? Please also share `ceph fs dump`. -- Patrick Donnelly, Ph.D. He / Him / His Senior Software Engineer Red Hat Sunnyvale, CA GPG: 19F28A586F808C2402351B93C3301A3E258DD79D
Hello Patrick, File system created around 4 months back. Using ceph version 14.2.3 version. [root@knode25 /]# ceph fs dump dumped fsmap epoch 577 e577 enable_multiple, ever_enabled_multiple: 0,0 compat: compat={},rocompat={},incompat={1=base v0.20,2=client writeable ranges,3=default file layouts on dirs,4=dir inode in separate object,5=mds uses versioned encoding,6=dirfrag is stored in omap,8=no anchor table,9=file layout v2,10=snaprealm v2} legacy client fscid: 1 Filesystem 'cephfs01' (1) fs_name cephfs01 epoch 577 flags 32 created 2019-10-18 23:59:29.610249 modified 2020-02-22 03:13:09.425905 tableserver 0 root 0 session_timeout 60 session_autoclose 300 max_file_size 1099511627776 min_compat_client -1 (unspecified) last_failure 0 last_failure_osd_epoch 1608 compat compat={},rocompat={},incompat={1=base v0.20,2=client writeable ranges,3=default file layouts on dirs,4=dir inode in separate object,5=mds uses versioned encoding,6=dirfrag is stored in omap,8=no anchor table,9=file layout v2,10=snaprealm v2} max_mds 1 in 0 up {0=2981519} failed damaged stopped data_pools [2] metadata_pool 1 inline_data disabled balancer standby_count_wanted 1 2981519: [v2:10.131.16.30:6808/3209191719,v1:10.131.16.30:6809/3209191719] 'cephfs01-b' mds.0.572 up:active seq 22141 2998684: [v2:10.131.16.89:6832/54557615,v1:10.131.16.89:6833/54557615] 'cephfs01-a' mds.0.0 up:standby-replay seq 2 [root@knode25 /]# ceph fs status cephfs01 - 290 clients ======== +------+----------------+------------+---------------+-------+-------+ | Rank | State | MDS | Activity | dns | inos | +------+----------------+------------+---------------+-------+-------+ | 0 | active | cephfs01-b | Reqs: 333 /s | 2738k | 2735k | | 0-s | standby-replay | cephfs01-a | Evts: 795 /s | 1368k | 1363k | +------+----------------+------------+---------------+-------+-------+ +-------------------+----------+-------+-------+ | Pool | type | used | avail | +-------------------+----------+-------+-------+ | cephfs01-metadata | metadata | 2193M | 78.1T | | cephfs01-data0 | data | 753G | 78.1T | +-------------------+----------+-------+-------+
It's probably a recently fixed openfiletable bug. Please upgrade to v14.2.8 when it is released in the next week or so. On Mon, Feb 24, 2020 at 1:46 PM Uday Bhaskar jalagam <jalagam.ceph@gmail.com> wrote:
Hello Patrick,
File system created around 4 months back. Using ceph version 14.2.3 version.
[root@knode25 /]# ceph fs dump dumped fsmap epoch 577 e577 enable_multiple, ever_enabled_multiple: 0,0 compat: compat={},rocompat={},incompat={1=base v0.20,2=client writeable ranges,3=default file layouts on dirs,4=dir inode in separate object,5=mds uses versioned encoding,6=dirfrag is stored in omap,8=no anchor table,9=file layout v2,10=snaprealm v2} legacy client fscid: 1
Filesystem 'cephfs01' (1) fs_name cephfs01 epoch 577 flags 32 created 2019-10-18 23:59:29.610249 modified 2020-02-22 03:13:09.425905 tableserver 0 root 0 session_timeout 60 session_autoclose 300 max_file_size 1099511627776 min_compat_client -1 (unspecified) last_failure 0 last_failure_osd_epoch 1608 compat compat={},rocompat={},incompat={1=base v0.20,2=client writeable ranges,3=default file layouts on dirs,4=dir inode in separate object,5=mds uses versioned encoding,6=dirfrag is stored in omap,8=no anchor table,9=file layout v2,10=snaprealm v2} max_mds 1 in 0 up {0=2981519} failed damaged stopped data_pools [2] metadata_pool 1 inline_data disabled balancer standby_count_wanted 1 2981519: [v2:10.131.16.30:6808/3209191719,v1:10.131.16.30:6809/3209191719] 'cephfs01-b' mds.0.572 up:active seq 22141 2998684: [v2:10.131.16.89:6832/54557615,v1:10.131.16.89:6833/54557615] 'cephfs01-a' mds.0.0 up:standby-replay seq 2
[root@knode25 /]# ceph fs status cephfs01 - 290 clients ======== +------+----------------+------------+---------------+-------+-------+ | Rank | State | MDS | Activity | dns | inos | +------+----------------+------------+---------------+-------+-------+ | 0 | active | cephfs01-b | Reqs: 333 /s | 2738k | 2735k | | 0-s | standby-replay | cephfs01-a | Evts: 795 /s | 1368k | 1363k | +------+----------------+------------+---------------+-------+-------+ +-------------------+----------+-------+-------+ | Pool | type | used | avail | +-------------------+----------+-------+-------+ | cephfs01-metadata | metadata | 2193M | 78.1T | | cephfs01-data0 | data | 753G | 78.1T | +-------------------+----------+-------+-------+ _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- Patrick Donnelly, Ph.D. He / Him / His Senior Software Engineer Red Hat Sunnyvale, CA GPG: 19F28A586F808C2402351B93C3301A3E258DD79D
Thanks Patrick, is this the bug you are referring to https://tracker.ceph.com/issues/42515 ? We also see performance issues mainly on metadata operations like finding file stats operations , however mds perf dump shows no sign of any latencies . could this bug cause any performance issues ? here is the perf dump metrics . https://pastebin.com/178anAe1 do you see any clue in this that could cause slow down in such operations ? our metadara pool has around 1.7 GB of data I gave mds cache 3 GB , I am not sure where to check how much used in the 3 GB or what is hit and miss count/ration in cache . We have huge cluster , there is definitely not enough IO that could saturate actual disk capacity so it is definitely MDS , not sure what to check here to pin point the issues. Could you point me where I can start to go deep in troubleshooting this ? Thanks, Uday.
On Mon, Feb 24, 2020 at 2:28 PM Uday Bhaskar jalagam <jalagam.ceph@gmail.com> wrote:
Thanks Patrick,
is this the bug you are referring to https://tracker.ceph.com/issues/42515 ?
Yes
We also see performance issues mainly on metadata operations like finding file stats operations , however mds perf dump shows no sign of any latencies . could this bug cause any performance issues ?
Unlikely.
do you see any clue in this that could cause slow down in such operations ? our metadara pool has around 1.7 GB of data I gave mds cache 3 GB,
3GB cache is probably too small for your cluster. How many users? Your perf dump indicates it probably is 8GB and not 3GB.
I am not sure where to check how much used in the 3 GB or what is hit and miss count/ration in cache .
"mds_co_bytes": 8160499164, in your perf dump. You can also look at inodes added/removed (to identify churn): "mds_mem": { "ino": 2740340, "ino+": 19461742, "ino-": 16721402, -- Patrick Donnelly, Ph.D. He / Him / His Senior Software Engineer Red Hat Sunnyvale, CA GPG: 19F28A586F808C2402351B93C3301A3E258DD79D
participants (2)
-
Patrick Donnelly
-
Uday Bhaskar jalagam