The transformer architecture replaced recurrence with a mechanism that lets every token look at every other token. The key idea is simple: represent a token as a query, then compare it to keys from the sequence.
Attention(Q,K,V) = softmax(\frac{QK^T}{\sqrt{d_k}})V
The scaling term keeps gradients well-behaved as the key dimension grows. In practice, this makes it possible to train on long sequences in parallel.