Google's 2017 paper "Attention Is All You Need" introduced the Transformer architecture, whose attention mechanism revolutionized sequence modeling and became the cornerstone of large language models.
打开完整版