Hi everyone, During CDM today Ilya pointed out that there is an open pull request that adds on-wire compression to msgr v2 here: https://github.com/ceph/ceph/pull/36517 Before we proceed there, though, we decided we should have a broader discussion about how on-wire compression should be implemented. The current pull request implements this purely in the msgr layer. Benefits include that it applies to all messages--not just OSD replication but also client/OSD traffic, inter-MDS traffic, and so on. Downsides include that replicated writes are compressed multiple times--once for each replica. One alternate approach might be: - expand the Message interface to allow set_data()/get_data() to accept/expose compressed data (e.g., data + compression_disposition). The message header could include a field indicating what codec was used. - OSD replication code could compress the data once and pass it to both messages for both replicas This would only capture the data portion of the message payload, but that is probably the only part we really are about. It would also require some special support for all the users that want to take advantage of it... probably the osd replication backend and Objecter to start. One could also imagine extending this to allow compressed data to pass all the way through to bluestore, although that brings in some additional concerns (bluestore has a max chunk size and some alignment considerations, for instance). Another possibility is integration compression into bufferlist. I'm not sure that represents a very compelling set of trade-offs, however. Other thoughts? sage
On Wed, Feb 3, 2021 at 11:57 AM Sage Weil <sage@newdream.net> wrote:
Hi everyone,
During CDM today Ilya pointed out that there is an open pull request that adds on-wire compression to msgr v2 here:
https://github.com/ceph/ceph/pull/36517
Before we proceed there, though, we decided we should have a broader discussion about how on-wire compression should be implemented.
The current pull request implements this purely in the msgr layer. Benefits include that it applies to all messages--not just OSD replication but also client/OSD traffic, inter-MDS traffic, and so on. Downsides include that replicated writes are compressed multiple times--once for each replica.
One alternate approach might be: - expand the Message interface to allow set_data()/get_data() to accept/expose compressed data (e.g., data + compression_disposition). The message header could include a field indicating what codec was used. - OSD replication code could compress the data once and pass it to both messages for both replicas This would only capture the data portion of the message payload, but that is probably the only part we really are about. It would also require some special support for all the users that want to take advantage of it... probably the osd replication backend and Objecter to start. One could also imagine extending this to allow compressed data to pass all the way through to bluestore, although that brings in some additional concerns (bluestore has a max chunk size and some alignment considerations, for instance).
Another possibility is integration compression into bufferlist. I'm not sure that represents a very compelling set of trade-offs, however.
Other thoughts?
These interfaces seem a lot more complicated and require a lot of management outside the messenger — besides simply dealing with buffers, we need to handle negotiating compression techniques and keeping them uniform across servers running different versions. That sounds like a lot of pain to me. If the main concern is re-compressing replicated OSD data, and given that we're using bufferlists and bufferptrs to share that data amongst the messages anyway, perhaps we should just do memoization on those data structures when we compres? -Greg
sage _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
we may need a method like CEPH_OSD_OP_FLAG_FADVISE_* to bypass msg2 / bluestore compress. For objstore, client maybe already compress the data(attribute: user.rgw.content_encoding). So need recompress on msg layer or bluestore layer. I very much agree with Sage’s second idea. It like https://lwn.net/Articles/837816/ (Encode I/O). This can largely reduce network and CPU.
One point - this PR was presented months ago and the design was discussed in the team, and with cooperation of several team members - changing it after it was implemented seems like a non-friendly process :-(. Secondly the design supports user hints which may suggest that the data should not be compressed. It was not implemented because of time shortage (this feature is part of a collaboration with the academy and was performed by an experienced grad student, but under some time limits). Implementing this hint can solve the problem of compressed data sent by RGW. Regards, Josh On Thu, Feb 4, 2021 at 2:57 AM majianpeng <jianpeng.ma@intel.com> wrote:
we may need a method like CEPH_OSD_OP_FLAG_FADVISE_* to bypass msg2 / bluestore compress. For objstore, client maybe already compress the data(attribute: user.rgw.content_encoding). So need recompress on msg layer or bluestore layer. I very much agree with Sage’s second idea. It like https://lwn.net/Articles/837816/ (Encode I/O). This can largely reduce network and CPU. _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
On Wed, Feb 3, 2021 at 2:17 PM Gregory Farnum <gfarnum@redhat.com> wrote:
These interfaces seem a lot more complicated and require a lot of management outside the messenger — besides simply dealing with buffers, we need to handle negotiating compression techniques and keeping them uniform across servers running different versions. That sounds like a lot of pain to me.
I think the negotiation can be simple if it is simply based on OSDMap flags and we assume that specific versions of Ceph have the same supported codecs. If the main concern is re-compressing replicated OSD data, and given
that we're using bufferlists and bufferptrs to share that data amongst the messages anyway, perhaps we should just do memoization on those data structures when we compres?
Yeah, I think this would be the main other approach we should consider. My concern is that the buffer(list) code is already super complicated and I worry about overgeneralizing this. Mostly we need bufferlists for encoding and moving things around in memory and there are relatively few cases where we have large data buffers that may or may not be compressed (this is the only one, currently). Having bufferlist cache crcs, for instance, is still something that I have regrets about in retrospect. Hmm, which makes me wonder: if we widen the Message set_data/get_data interface to include compression disposition, I wonder if it should also include a checksum. I think it could then also cover the one case where we still rely on the bufferlist crc caching. (FileStore writes to the journal were the other, but we are probably past worrying about that now.) sage
Another thought: if ensure that this is gated by a feature flag (or similar) that is easily removed later (e.g., a msgr2 feature instead of a ceph feature bit), we could go with this solution now, and if/when we have something we like better gracefully remove this support in a future release.
+1 for this, and it should be simple to implement since the actual compression is decided based on the msgr2 connection feature negotiation. If we have an alternative we can just say that future OSDs will not accept this compression during connection feature negotiation. One more point - we should be very careful not to compress the data in the buffer list too soon (and please excuse my ignorance about the exact data layering implementation in Ceph) - some feature have to look into the data (dedup...), so we need to make sure that we don't compress above the level that these features need the data. Regards, Josh On Thu, Feb 11, 2021 at 2:13 AM Sage Weil <sweil@redhat.com> wrote:
Another thought: if ensure that this is gated by a feature flag (or similar) that is easily removed later (e.g., a msgr2 feature instead of a ceph feature bit), we could go with this solution now, and if/when we have something we like better gracefully remove this support in a future release.
_______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
participants (5)
-
Gregory Farnum
-
Josh Salomon
-
majianpeng
-
Sage Weil
-
Sage Weil