Ceph not showing full capacity
Hi, I have created a test Ceph cluster with Ceph Octopus using cephadm. Cluster total RAW disk capacity is 262 TB but it's allowing to use of only 132TB. I have not set quota for any of the pool. what could be the issue? Output from :- ceph -s cluster: id: f8bc7682-0d11-11eb-a332-0cc47a5ec98a health: HEALTH_WARN clock skew detected on mon.strg-node3, mon.strg-node2 2 backfillfull osd(s) 4 pool(s) backfillfull 1 pools have too few placement groups services: mon: 3 daemons, quorum strg-node1,strg-node3,strg-node2 (age 7m) mgr: strg-node3.jtacbn(active, since 7m), standbys: strg-node1.gtlvyv mds: cephfs-strg:1 {0=cephfs-strg.strg-node1.lhmeea=up:active} 1 up:standby osd: 48 osds: 48 up (since 7m), 48 in (since 5d) task status: scrub status: mds.cephfs-strg.strg-node1.lhmeea: idle data: pools: 4 pools, 289 pgs objects: 17.29M objects, 66 TiB usage: 132 TiB used, 130 TiB / 262 TiB avail pgs: 288 active+clean 1 active+clean+scrubbing+deep mounted volume shows node1:/ 67T 66T 910G 99% /mnt/cephfs
Can you post your crush map? Perhaps some OSDs are in the wrong place. On Sat, Oct 24, 2020 at 8:51 AM Amudhan P <amudhan83@gmail.com> wrote:
Hi,
I have created a test Ceph cluster with Ceph Octopus using cephadm.
Cluster total RAW disk capacity is 262 TB but it's allowing to use of only 132TB. I have not set quota for any of the pool. what could be the issue?
Output from :- ceph -s cluster: id: f8bc7682-0d11-11eb-a332-0cc47a5ec98a health: HEALTH_WARN clock skew detected on mon.strg-node3, mon.strg-node2 2 backfillfull osd(s) 4 pool(s) backfillfull 1 pools have too few placement groups
services: mon: 3 daemons, quorum strg-node1,strg-node3,strg-node2 (age 7m) mgr: strg-node3.jtacbn(active, since 7m), standbys: strg-node1.gtlvyv mds: cephfs-strg:1 {0=cephfs-strg.strg-node1.lhmeea=up:active} 1 up:standby osd: 48 osds: 48 up (since 7m), 48 in (since 5d)
task status: scrub status: mds.cephfs-strg.strg-node1.lhmeea: idle
data: pools: 4 pools, 289 pgs objects: 17.29M objects, 66 TiB usage: 132 TiB used, 130 TiB / 262 TiB avail pgs: 288 active+clean 1 active+clean+scrubbing+deep
mounted volume shows node1:/ 67T 66T 910G 99% /mnt/cephfs _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hi Nathan, Attached crushmap output. let me know if you find any thing odd. On Sat, Oct 24, 2020 at 6:47 PM Nathan Fish <lordcirth@gmail.com> wrote:
Can you post your crush map? Perhaps some OSDs are in the wrong place.
On Sat, Oct 24, 2020 at 8:51 AM Amudhan P <amudhan83@gmail.com> wrote:
Hi,
I have created a test Ceph cluster with Ceph Octopus using cephadm.
Cluster total RAW disk capacity is 262 TB but it's allowing to use of
only
132TB. I have not set quota for any of the pool. what could be the issue?
Output from :- ceph -s cluster: id: f8bc7682-0d11-11eb-a332-0cc47a5ec98a health: HEALTH_WARN clock skew detected on mon.strg-node3, mon.strg-node2 2 backfillfull osd(s) 4 pool(s) backfillfull 1 pools have too few placement groups
services: mon: 3 daemons, quorum strg-node1,strg-node3,strg-node2 (age 7m) mgr: strg-node3.jtacbn(active, since 7m), standbys: strg-node1.gtlvyv mds: cephfs-strg:1 {0=cephfs-strg.strg-node1.lhmeea=up:active} 1 up:standby osd: 48 osds: 48 up (since 7m), 48 in (since 5d)
task status: scrub status: mds.cephfs-strg.strg-node1.lhmeea: idle
data: pools: 4 pools, 289 pgs objects: 17.29M objects, 66 TiB usage: 132 TiB used, 130 TiB / 262 TiB avail pgs: 288 active+clean 1 active+clean+scrubbing+deep
mounted volume shows node1:/ 67T 66T 910G 99% /mnt/cephfs _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
On 2020-10-24 14:53, Amudhan P wrote:
Hi,
I have created a test Ceph cluster with Ceph Octopus using cephadm.
Cluster total RAW disk capacity is 262 TB but it's allowing to use of only 132TB. I have not set quota for any of the pool. what could be the issue?
Unbalance? What does ceph osd df show? How large is the standard deviation? Gr. Stefan
Yes, There is a unbalance in PG's assigned to OSD's. `ceph osd df` output snip ID CLASS WEIGHT REWEIGHT SIZE RAW USE DATA OMAP META AVAIL %USE VAR PGS STATUS 0 hdd 5.45799 1.00000 5.5 TiB 3.6 TiB 3.6 TiB 9.7 MiB 4.6 GiB 1.9 TiB 65.94 1.31 13 up 1 hdd 5.45799 1.00000 5.5 TiB 1.0 TiB 1.0 TiB 4.4 MiB 1.3 GiB 4.4 TiB 18.87 0.38 9 up 2 hdd 5.45799 1.00000 5.5 TiB 1.5 TiB 1.5 TiB 4.0 MiB 1.9 GiB 3.9 TiB 28.30 0.56 10 up 3 hdd 5.45799 1.00000 5.5 TiB 2.1 TiB 2.1 TiB 7.7 MiB 2.7 GiB 3.4 TiB 37.70 0.75 12 up 4 hdd 5.45799 1.00000 5.5 TiB 4.1 TiB 4.1 TiB 5.8 MiB 5.2 GiB 1.3 TiB 75.27 1.50 20 up 5 hdd 5.45799 1.00000 5.5 TiB 5.1 TiB 5.1 TiB 5.9 MiB 6.7 GiB 317 GiB 94.32 1.88 18 up 6 hdd 5.45799 1.00000 5.5 TiB 1.5 TiB 1.5 TiB 5.2 MiB 2.0 GiB 3.9 TiB 28.32 0.56 9 up MIN/MAX VAR: 0.19/1.88 STDDEV: 22.13 On Sun, Oct 25, 2020 at 12:08 AM Stefan Kooman <stefan@bit.nl> wrote:
On 2020-10-24 14:53, Amudhan P wrote:
Hi,
I have created a test Ceph cluster with Ceph Octopus using cephadm.
Cluster total RAW disk capacity is 262 TB but it's allowing to use of only 132TB. I have not set quota for any of the pool. what could be the issue?
Unbalance? What does ceph osd df show? How large is the standard deviation?
Gr. Stefan _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
On 2020-10-25 05:33, Amudhan P wrote:
Yes, There is a unbalance in PG's assigned to OSD's. `ceph osd df` output snip ID CLASS WEIGHT REWEIGHT SIZE RAW USE DATA OMAP META AVAIL %USE VAR PGS STATUS 0 hdd 5.45799 1.00000 5.5 TiB 3.6 TiB 3.6 TiB 9.7 MiB 4.6 GiB 1.9 TiB 65.94 1.31 13 up 1 hdd 5.45799 1.00000 5.5 TiB 1.0 TiB 1.0 TiB 4.4 MiB 1.3 GiB 4.4 TiB 18.87 0.38 9 up 2 hdd 5.45799 1.00000 5.5 TiB 1.5 TiB 1.5 TiB 4.0 MiB 1.9 GiB 3.9 TiB 28.30 0.56 10 up 3 hdd 5.45799 1.00000 5.5 TiB 2.1 TiB 2.1 TiB 7.7 MiB 2.7 GiB 3.4 TiB 37.70 0.75 12 up 4 hdd 5.45799 1.00000 5.5 TiB 4.1 TiB 4.1 TiB 5.8 MiB 5.2 GiB 1.3 TiB 75.27 1.50 20 up 5 hdd 5.45799 1.00000 5.5 TiB 5.1 TiB 5.1 TiB 5.9 MiB 6.7 GiB 317 GiB 94.32 1.88 18 up 6 hdd 5.45799 1.00000 5.5 TiB 1.5 TiB 1.5 TiB 5.2 MiB 2.0 GiB 3.9 TiB 28.32 0.56 9 up
MIN/MAX VAR: 0.19/1.88 STDDEV: 22.13
ceph balancer mode upmap ceph balancer on The balancer should start balancing and this should result in way more space available. Good to know that ceph df is based on the disk that is most full. There is all sorts of tuning available for the balancer, although I can't find it in the documentation. Ceph docu better project is working on that. See [1] for information. You can look up the python code to see what variables you can tune: /usr/share/ceph/mgr/balancer/module.py ceph config set mgr/balancer/begin_weekday 1 ceph config set mgr/balancer/end_weekday 5 ceph config set mgr mgr/balancer/begin_time 1000 ceph config set mgr mgr/balancer/end_time 1700 ^^ to restrict the balancer running only on weekdays (monday to friday) from 10:00 - 17:00 h. Gr. Stefan [1]: https://docs.ceph.com/en/latest/rados/operations/balancer/#balancer
Hi Stefan, I have started balancer but what I don't understand is there are enough free space in other disks. Why it's not showing those in available space? How to reclaim the free space? On Sun 25 Oct, 2020, 2:27 PM Stefan Kooman, <stefan@bit.nl> wrote:
On 2020-10-25 05:33, Amudhan P wrote:
Yes, There is a unbalance in PG's assigned to OSD's. `ceph osd df` output snip ID CLASS WEIGHT REWEIGHT SIZE RAW USE DATA OMAP META AVAIL %USE VAR PGS STATUS 0 hdd 5.45799 1.00000 5.5 TiB 3.6 TiB 3.6 TiB 9.7 MiB 4.6 GiB 1.9 TiB 65.94 1.31 13 up 1 hdd 5.45799 1.00000 5.5 TiB 1.0 TiB 1.0 TiB 4.4 MiB 1.3 GiB 4.4 TiB 18.87 0.38 9 up 2 hdd 5.45799 1.00000 5.5 TiB 1.5 TiB 1.5 TiB 4.0 MiB 1.9 GiB 3.9 TiB 28.30 0.56 10 up 3 hdd 5.45799 1.00000 5.5 TiB 2.1 TiB 2.1 TiB 7.7 MiB 2.7 GiB 3.4 TiB 37.70 0.75 12 up 4 hdd 5.45799 1.00000 5.5 TiB 4.1 TiB 4.1 TiB 5.8 MiB 5.2 GiB 1.3 TiB 75.27 1.50 20 up 5 hdd 5.45799 1.00000 5.5 TiB 5.1 TiB 5.1 TiB 5.9 MiB 6.7 GiB 317 GiB 94.32 1.88 18 up 6 hdd 5.45799 1.00000 5.5 TiB 1.5 TiB 1.5 TiB 5.2 MiB 2.0 GiB 3.9 TiB 28.32 0.56 9 up
MIN/MAX VAR: 0.19/1.88 STDDEV: 22.13
ceph balancer mode upmap ceph balancer on
The balancer should start balancing and this should result in way more space available. Good to know that ceph df is based on the disk that is most full.
There is all sorts of tuning available for the balancer, although I can't find it in the documentation. Ceph docu better project is working on that. See [1] for information. You can look up the python code to see what variables you can tune: /usr/share/ceph/mgr/balancer/module.py
ceph config set mgr/balancer/begin_weekday 1 ceph config set mgr/balancer/end_weekday 5 ceph config set mgr mgr/balancer/begin_time 1000 ceph config set mgr mgr/balancer/end_time 1700
^^ to restrict the balancer running only on weekdays (monday to friday) from 10:00 - 17:00 h.
Gr. Stefan
[1]: https://docs.ceph.com/en/latest/rados/operations/balancer/#balancer
Hi, In ceph, when you create an object, it cannot go any OSD as it fits. An object is mapped to a placement group using a hash algorithm. Then placement groups are mapped to OSDs. See [1] for details. So, if any of your OSD goes full, write operations cannot be guaranteed success. Once you correct the unbalance, you should see more available space. Also, you only have 289 placement groups, which I think is too few for your 48 OSDs [2]. If you have more placement groups, the unbalance issue will be far less severe. [1]: https://docs.ceph.com/en/latest/architecture/#mapping-pgs-to-osds [2]: https://docs.ceph.com/en/latest/rados/operations/placement-groups/
在 2020年10月25日,18:24,Amudhan P <amudhan83@gmail.com> 写道:
Hi Stefan,
I have started balancer but what I don't understand is there are enough free space in other disks.
Why it's not showing those in available space? How to reclaim the free space?
On Sun 25 Oct, 2020, 2:27 PM Stefan Kooman, <stefan@bit.nl> wrote:
On 2020-10-25 05:33, Amudhan P wrote: Yes, There is a unbalance in PG's assigned to OSD's. `ceph osd df` output snip ID CLASS WEIGHT REWEIGHT SIZE RAW USE DATA OMAP META AVAIL %USE VAR PGS STATUS 0 hdd 5.45799 1.00000 5.5 TiB 3.6 TiB 3.6 TiB 9.7 MiB 4.6 GiB 1.9 TiB 65.94 1.31 13 up 1 hdd 5.45799 1.00000 5.5 TiB 1.0 TiB 1.0 TiB 4.4 MiB 1.3 GiB 4.4 TiB 18.87 0.38 9 up 2 hdd 5.45799 1.00000 5.5 TiB 1.5 TiB 1.5 TiB 4.0 MiB 1.9 GiB 3.9 TiB 28.30 0.56 10 up 3 hdd 5.45799 1.00000 5.5 TiB 2.1 TiB 2.1 TiB 7.7 MiB 2.7 GiB 3.4 TiB 37.70 0.75 12 up 4 hdd 5.45799 1.00000 5.5 TiB 4.1 TiB 4.1 TiB 5.8 MiB 5.2 GiB 1.3 TiB 75.27 1.50 20 up 5 hdd 5.45799 1.00000 5.5 TiB 5.1 TiB 5.1 TiB 5.9 MiB 6.7 GiB 317 GiB 94.32 1.88 18 up 6 hdd 5.45799 1.00000 5.5 TiB 1.5 TiB 1.5 TiB 5.2 MiB 2.0 GiB 3.9 TiB 28.32 0.56 9 up MIN/MAX VAR: 0.19/1.88 STDDEV: 22.13 ceph balancer mode upmap ceph balancer on The balancer should start balancing and this should result in way more space available. Good to know that ceph df is based on the disk that is most full. There is all sorts of tuning available for the balancer, although I can't find it in the documentation. Ceph docu better project is working on that. See [1] for information. You can look up the python code to see what variables you can tune: /usr/share/ceph/mgr/balancer/module.py ceph config set mgr/balancer/begin_weekday 1 ceph config set mgr/balancer/end_weekday 5 ceph config set mgr mgr/balancer/begin_time 1000 ceph config set mgr mgr/balancer/end_time 1700 ^^ to restrict the balancer running only on weekdays (monday to friday) from 10:00 - 17:00 h. Gr. Stefan [1]: https://eur05.safelinks.protection.outlook.com/?url=https%3A%2F%2Fdocs.ceph....
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hi, For my quick understanding How PG's are responsible for allowing space allocation to a pool? My understanding that PG's basically helps in object placement when the number of PG's for a OSD's is high there is a high possibility that PG gets lot more data than other PG's. At this situation, we can use the balance between OSD's. But, I can't understand the logic of how does it restrict space to a pool? On Sun, Oct 25, 2020 at 5:55 PM 胡 玮文 <huww98@outlook.com> wrote:
Hi,
In ceph, when you create an object, it cannot go any OSD as it fits. An object is mapped to a placement group using a hash algorithm. Then placement groups are mapped to OSDs. See [1] for details. So, if any of your OSD goes full, write operations cannot be guaranteed success. Once you correct the unbalance, you should see more available space.
Also, you only have 289 placement groups, which I think is too few for your 48 OSDs [2]. If you have more placement groups, the unbalance issue will be far less severe.
[1]: https://docs.ceph.com/en/latest/architecture/#mapping-pgs-to-osds [2]: https://docs.ceph.com/en/latest/rados/operations/placement-groups/
在 2020年10月25日,18:24,Amudhan P <amudhan83@gmail.com> 写道:
Hi Stefan,
I have started balancer but what I don't understand is there are enough free space in other disks.
Why it's not showing those in available space? How to reclaim the free space?
On Sun 25 Oct, 2020, 2:27 PM Stefan Kooman, <stefan@bit.nl> wrote:
On 2020-10-25 05:33, Amudhan P wrote: Yes, There is a unbalance in PG's assigned to OSD's. `ceph osd df` output snip ID CLASS WEIGHT REWEIGHT SIZE RAW USE DATA OMAP META AVAIL %USE VAR PGS STATUS 0 hdd 5.45799 1.00000 5.5 TiB 3.6 TiB 3.6 TiB 9.7 MiB 4.6 GiB 1.9 TiB 65.94 1.31 13 up 1 hdd 5.45799 1.00000 5.5 TiB 1.0 TiB 1.0 TiB 4.4 MiB 1.3 GiB 4.4 TiB 18.87 0.38 9 up 2 hdd 5.45799 1.00000 5.5 TiB 1.5 TiB 1.5 TiB 4.0 MiB 1.9 GiB 3.9 TiB 28.30 0.56 10 up 3 hdd 5.45799 1.00000 5.5 TiB 2.1 TiB 2.1 TiB 7.7 MiB 2.7 GiB 3.4 TiB 37.70 0.75 12 up 4 hdd 5.45799 1.00000 5.5 TiB 4.1 TiB 4.1 TiB 5.8 MiB 5.2 GiB 1.3 TiB 75.27 1.50 20 up 5 hdd 5.45799 1.00000 5.5 TiB 5.1 TiB 5.1 TiB 5.9 MiB 6.7 GiB 317 GiB 94.32 1.88 18 up 6 hdd 5.45799 1.00000 5.5 TiB 1.5 TiB 1.5 TiB 5.2 MiB 2.0 GiB 3.9 TiB 28.32 0.56 9 up MIN/MAX VAR: 0.19/1.88 STDDEV: 22.13 ceph balancer mode upmap ceph balancer on The balancer should start balancing and this should result in way more space available. Good to know that ceph df is based on the disk that is most full. There is all sorts of tuning available for the balancer, although I can't find it in the documentation. Ceph docu better project is working on that. See [1] for information. You can look up the python code to see what variables you can tune: /usr/share/ceph/mgr/balancer/module.py ceph config set mgr/balancer/begin_weekday 1 ceph config set mgr/balancer/end_weekday 5 ceph config set mgr mgr/balancer/begin_time 1000 ceph config set mgr mgr/balancer/end_time 1700 ^^ to restrict the balancer running only on weekdays (monday to friday) from 10:00 - 17:00 h. Gr. Stefan [1]: https://eur05.safelinks.protection.outlook.com/?url=https%3A%2F%2Fdocs.ceph....
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
On 2020-10-25 15:20, Amudhan P wrote:
Hi,
For my quick understanding How PG's are responsible for allowing space allocation to a pool?
My understanding that PG's basically helps in object placement when the number of PG's for a OSD's is high there is a high possibility that PG gets lot more data than other PG's. At this situation, we can use the balance between OSD's.
But, I can't understand the logic of how does it restrict space to a pool?
The space available in the cluster is available for all pools that match the CRUSH rule. OSDs contain PGs which contain objects. With many PGs per OSD (100-200 depending on disk type) the balancer can optimize for equal placement of PGs, optimize for available space, or a mix of both. The amount of space available is determined by the fullest OSD, as that OSD will be the first to run out of space for new writes. The more PGs the more evenly they can be distributed accross the cluster and the lower the standard deviation will be. Gr. Stefan
Den sön 25 okt. 2020 kl 15:18 skrev Amudhan P <amudhan83@gmail.com>:
Hi,
For my quick understanding How PG's are responsible for allowing space allocation to a pool?
An objects name will decide which PG (from the list of PGs in the pool) it will end up on, so if you have very few PGs, the hashed/pseudorandom placement will be unbalanced at times. As an example, if you have only 8 PGs and write 9 large objects, then at least one (but probably two or three) PGs will receive two or more of those 9, and some will receive none just on pure statistics. If you have 100 PGs, the chance of one getting two out of those nine objects is much smaller. Overall, with all pools accounted for, one should aim for something like 100 PGs per OSD, but you also need to count the replication factor for each pool so if you have replication = 3 and a pool gets 128 PGs, it will place 3*128 PGs out on various OSDs according to the crush rules. PGs don't have a size, but will grow as needed, and since the next object to be written can end up anywhere (depending on the hashed result) ceph df must always tell you the worst case when listing how much data this pool has "left". It will always be the OSD with least space left that limits the pool.
My understanding that PG's basically helps in object placement when the number of PG's for a OSD's is high there is a high possibility that PG gets lot more data than other PG's.
This statement seems incorrect to me.
At this situation, we can use the balance between OSD's. But, I can't understand the logic of how does it restrict space to a pool?
-- May the most significant bit of your life be positive.
Hi Jane, I agree with you and I was trying to say disk which has more PG will fill up quick. But, My question even though RAW disk space is 262 TB, pool 2 replica max storage is showing only 132 TB in the dashboard and when mounting the pool using cephfs it's showing 62 TB, I could understand that due to replica it's showing half of the space. why it's not showing the entire RAW disk space as available space? Number of PG per pool play any vital role in showing available space? On Mon, Oct 26, 2020 at 12:37 PM Janne Johansson <icepic.dz@gmail.com> wrote:
Den sön 25 okt. 2020 kl 15:18 skrev Amudhan P <amudhan83@gmail.com>:
Hi,
For my quick understanding How PG's are responsible for allowing space allocation to a pool?
An objects name will decide which PG (from the list of PGs in the pool) it will end up on, so if you have very few PGs, the hashed/pseudorandom placement will be unbalanced at times. As an example, if you have only 8 PGs and write 9 large objects, then at least one (but probably two or three) PGs will receive two or more of those 9, and some will receive none just on pure statistics. If you have 100 PGs, the chance of one getting two out of those nine objects is much smaller. Overall, with all pools accounted for, one should aim for something like 100 PGs per OSD, but you also need to count the replication factor for each pool so if you have replication = 3 and a pool gets 128 PGs, it will place 3*128 PGs out on various OSDs according to the crush rules.
PGs don't have a size, but will grow as needed, and since the next object to be written can end up anywhere (depending on the hashed result) ceph df must always tell you the worst case when listing how much data this pool has "left". It will always be the OSD with least space left that limits the pool.
My understanding that PG's basically helps in object placement when the number of PG's for a OSD's is high there is a high possibility that PG gets lot more data than other PG's.
This statement seems incorrect to me.
At this situation, we can use the balance between OSD's. But, I can't understand the logic of how does it restrict space to a pool?
-- May the most significant bit of your life be positive.
在 2020年10月26日,22:30,Amudhan P <amudhan83@gmail.com> 写道: Hi Jane, I agree with you and I was trying to say disk which has more PG will fill up quick. But, My question even though RAW disk space is 262 TB, pool 2 replica max storage is showing only 132 TB in the dashboard and when mounting the pool using cephfs it's showing 62 TB, I could understand that due to replica it's showing half of the space. Your first mail shows 67T (instead of 62) why it's not showing the entire RAW disk space as available space? Number of PG per pool play any vital role in showing available space? I might be wrong, but I think the size of mounted cephfs is calculated by “used + available”. It is not directly related to raw disk space. You have unbalance issue, so you have less available space as explained previously. So the total size is less than expected. Maybe you should try to correct the unbalance first, and see if the available space and size go up. Increase pg_num, run balancer, etc. On Mon, Oct 26, 2020 at 12:37 PM Janne Johansson <icepic.dz@gmail.com<mailto:icepic.dz@gmail.com>> wrote: Den sön 25 okt. 2020 kl 15:18 skrev Amudhan P <amudhan83@gmail.com<mailto:amudhan83@gmail.com>>: Hi, For my quick understanding How PG's are responsible for allowing space allocation to a pool? An objects name will decide which PG (from the list of PGs in the pool) it will end up on, so if you have very few PGs, the hashed/pseudorandom placement will be unbalanced at times. As an example, if you have only 8 PGs and write 9 large objects, then at least one (but probably two or three) PGs will receive two or more of those 9, and some will receive none just on pure statistics. If you have 100 PGs, the chance of one getting two out of those nine objects is much smaller. Overall, with all pools accounted for, one should aim for something like 100 PGs per OSD, but you also need to count the replication factor for each pool so if you have replication = 3 and a pool gets 128 PGs, it will place 3*128 PGs out on various OSDs according to the crush rules. PGs don't have a size, but will grow as needed, and since the next object to be written can end up anywhere (depending on the hashed result) ceph df must always tell you the worst case when listing how much data this pool has "left". It will always be the OSD with least space left that limits the pool. My understanding that PG's basically helps in object placement when the number of PG's for a OSD's is high there is a high possibility that PG gets lot more data than other PG's. This statement seems incorrect to me. At this situation, we can use the balance between OSD's. But, I can't understand the logic of how does it restrict space to a pool? -- May the most significant bit of your life be positive.
Hi,
Your first mail shows 67T (instead of 62)
I have just given an approximate number the first given number is the right number. I have deleted all pools and just created a fresh pool for test with PG num 128 and now it's showing a full size of 248TB. output from " ceph df " --- RAW STORAGE --- CLASS SIZE AVAIL USED RAW USED %RAW USED hdd 262 TiB 262 TiB 3.9 GiB 52 GiB 0.02 TOTAL 262 TiB 262 TiB 3.9 GiB 52 GiB 0.02 --- POOLS --- POOL ID STORED OBJECTS USED %USED MAX AVAIL pool3 8 0 B 0 0 B 0 124 TiB So, PG number is not an issue in showing less size. I am trying other options also to see what made this issue. On Mon, Oct 26, 2020 at 8:20 PM 胡 玮文 <huww98@outlook.com> wrote:
在 2020年10月26日,22:30,Amudhan P <amudhan83@gmail.com> 写道:
Hi Jane,
I agree with you and I was trying to say disk which has more PG will fill up quick.
But, My question even though RAW disk space is 262 TB, pool 2 replica max storage is showing only 132 TB in the dashboard and when mounting the pool using cephfs it's showing 62 TB, I could understand that due to replica it's showing half of the space.
Your first mail shows 67T (instead of 62)
why it's not showing the entire RAW disk space as available space? Number of PG per pool play any vital role in showing available space?
I might be wrong, but I think the size of mounted cephfs is calculated by “used + available”. It is not directly related to raw disk space. You have unbalance issue, so you have less available space as explained previously. So the total size is less than expected.
Maybe you should try to correct the unbalance first, and see if the available space and size go up. Increase pg_num, run balancer, etc.
On Mon, Oct 26, 2020 at 12:37 PM Janne Johansson <icepic.dz@gmail.com> wrote:
Den sön 25 okt. 2020 kl 15:18 skrev Amudhan P <amudhan83@gmail.com>:
Hi,
For my quick understanding How PG's are responsible for allowing space allocation to a pool?
An objects name will decide which PG (from the list of PGs in the pool) it will end up on, so if you have very few PGs, the hashed/pseudorandom placement will be unbalanced at times. As an example, if you have only 8 PGs and write 9 large objects, then at least one (but probably two or three) PGs will receive two or more of those 9, and some will receive none just on pure statistics. If you have 100 PGs, the chance of one getting two out of those nine objects is much smaller. Overall, with all pools accounted for, one should aim for something like 100 PGs per OSD, but you also need to count the replication factor for each pool so if you have replication = 3 and a pool gets 128 PGs, it will place 3*128 PGs out on various OSDs according to the crush rules.
PGs don't have a size, but will grow as needed, and since the next object to be written can end up anywhere (depending on the hashed result) ceph df must always tell you the worst case when listing how much data this pool has "left". It will always be the OSD with least space left that limits the pool.
My understanding that PG's basically helps in object placement when the number of PG's for a OSD's is high there is a high possibility that PG gets lot more data than other PG's.
This statement seems incorrect to me.
At this situation, we can use the balance between OSD's. But, I can't understand the logic of how does it restrict space to a pool?
-- May the most significant bit of your life be positive.
participants (5)
-
Amudhan P
-
Janne Johansson
-
Nathan Fish
-
Stefan Kooman
-
胡 玮文