Thank you folks for the feedback. Clearly this is a controversial matter. So be prepared, long reply ahead :'D
@Ilya Dryomov, thanks for the naming hints. I'd definitely go for "backport:no-conflicts | has-conflicts".
@Ilya, @Loïc, regarding the size labels, I totally agree: talking here about "complexity" was too ambitious. "Lines of Code" is a pretty simplistic metric to describe whether something might be more risky or not (
one-liners can be fatal too), but OTOH from a reviewer perspective that's a pretty decent a-priori indicator of the time/effort a PR review might require (especially when there were conflicts). Having this kind of a-priori hints may help us better decide where to allocate our efforts.
As an test: a backporter/a reviewer/QA needs to approach these 2 PRs labeled this way:
And now they get the following hints:
We might debate whether these flags could bias reviewers in favor/against those, but it wouldn't be that different from when you see this and click "merge PR":
Regarding the concern about the number of labels, as raised by Loïc and Yuri, I checked the last 25 PRs and the median is 3 labels/PR (3.4 on average), and that including a cephadm batch PR with 14 labels! :D
Maybe it's just me but I don't think that 3 labels is too much, as long as it's useful. That said there are things we could improve here:
- Remove less informative labels (e.g.: "pybind" appears in 40% of the PRs and there's no such team or component).
- Check that labels are really useful. E.g.: is "needs-review" label improving reviews on stale PRs? Is "needs-rebase" really required?
- Better use of colors or grouping tags. This is an example from kubernetes:

@Nathan Cutler, if we fear that automation might increase the number of regressions because it makes backport stuff easier, we should also stop using "ceph-backports" and "backport-create-issue" scripts, shouldn't we? Or, on the contrary, we might also add automation to block backports from non-bug trackers, which we are not doing now (BTW ML-based PR classification is nothing new and there are lots of interesting proposals
[1] [2]).
I'm skeptical of the value of labels, but I do think it would be useful to have
Jenkins jobs checking:
1. whether the commits being cherry-picked are really in master
2. whether the master commits cherry-picked cleanly
3. whether the backport PR contains the same number of commits as the master PR
I agree, although I think that's more the implementation detail. I also agree with you here: we should have a "commit sanity" check (e.g.: block merge commits, detect "Fixes" line, etc.), but in the meantime we may start with one and then switch to the other.
@Yuri Weinstein interesting feedback. The current labels were mostly meant to dispatch PRs to the proper component bucket, but we may give it another thought for mapping that to QA suites. I remember Neha & Josh mentioned this in relation to this effort to bring teuthology runs to PR checks.
As the number of topics seems to be growing, shall we schedule a quick meeting and review this?
Thanks and nice weekend!