From one piece Attention is all you need 3 beliefs, in the piece's order there
-
korrents.com
Self-attention can yield more interpretable models.Their words
As side benefit, self-attention could yield more interpretable models.
-
Their words
To the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution.
-
korrents.com
Sequence transduction can use a simple architecture based solely on attention, without recurrence or convolutions.Their words
We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.