Machine learning

Limitations of normalization in attention mechanism

T. Mudarisov, M. Burtsev, T. Petrova, R. State

Arxiv (2025)

We demonstrate that transformer attention can only discriminate well at shorter context lengths, losing clarity as input length increases.

31231223
31231223
31231223
31231223
31231223
31231223
31231223
31231223
31231223
31231223
31231223
31231223
31231223
31231223
31231223
31231223