Programming Languages Research Group: Git

author	Charlie Turner <charlie.turner@arm.com>
	Tue, 27 Oct 2015 17:59:03 +0000 (17:59 +0000)
committer	Charlie Turner <charlie.turner@arm.com>
	Tue, 27 Oct 2015 17:59:03 +0000 (17:59 +0000)
commit	63fe1641e4f427e1283cddab98f8bf1c3790f064
tree	11be318322c4efc6641dc4adab3c5d284e0cd3e8	tree \| snapshot
parent	751d6ddd6dfeb0e66663c678080d0a10383f427d	commit \| diff

[SLP] Be more aggressive about reduction width selection.

Summary:
This change could be way off-piste, I'm looking for any feedback on whether it's an acceptable approach.

It never seems to be a problem to gobble up as many reduction values as can be found, and then to attempt to reduce the resulting tree. Some of the workloads I'm looking at have been aggressively unrolled by hand, and by selecting reduction widths that are not constrained by a vector register size, it becomes possible to profitably vectorize. My test case shows such an unrolling which SLP was not vectorizing (on neither ARM nor X86) before this patch, but with it does vectorize.

I measure no significant compile time impact of this change when combined with D13949 and D14063. There are also no significant performance regressions on ARM/AArch64 in SPEC or LNT.

The more principled approach I thought of was to generate several candidate tree's and use the cost model to pick the cheapest one. That seemed like quite a big design change (the algorithms seem very much one-shot), and would likely be a costly thing for compile time. This seemed to do the job at very little cost, but I'm worried I've misunderstood something!

Reviewers: nadav, jmolloy

Subscribers: mssimpso, llvm-commits, aemerson

Differential Revision: http://reviews.llvm.org/D14116

git-svn-id: https://llvm.org/svn/llvm-project/llvm/trunk@251428 91177308-0d34-0410-b5e6-96231b3b80d8

lib/Transforms/Vectorize/SLPVectorizer.cpp		diff \| blob \| history
test/Transforms/SLPVectorizer/AArch64/horizontal.ll		diff \| blob \| history