Multi-head attention with 8 heads provides a clear improvement over single-head attention on the WMT 2014 English-to-German task27.
What the paper measured
25.8→26.4+0.6
BLEU · §3.2.2 / Table 3
Why partial
BLEU figure of 26.4 matches Table 3 verbatim. The cited paper supports the qualitative direction of the claim, but the +0.6 BLEU gain is smaller than what “clear improvement” typically denotes in this literature.
Source
“Multi-head attention allows the model to jointly attend to information from different representation subspaces.”




