Positional Encodings, Compared Positional Encoding 종류별 비교
Attention by itself doesn’t know order. The ways to inject position:
- Sinusoidal: the original. No learned parameters; extrapolation was the hope, but in practice it’s weak.
- Learned: one embedding per position. Beyond the training length it simply can’t.
- RoPE: rotates Q and K by position. Relative position falls out of the dot product naturally; it’s the current LLM standard.
- ALiBi: a distance-proportional penalty on the scores. Simple to implement, strong at length generalization.
The fact that today’s long-context extensions (YaRN and friends) work by tweaking RoPE’s rotation frequencies tells you why RoPE became the standard.
attention 자체는 순서를 모른다. 위치 정보를 넣는 방식들:
- Sinusoidal — 원조. 학습 파라미터 없음, 외삽 기대했지만 실제론 약함.
- Learned — 위치별 임베딩 학습. 학습 길이 밖은 아예 불가.
- RoPE — Q·K를 위치에 따라 회전. 상대 위치가 내적에 자연히 반영, 현 LLM 표준.
- ALiBi — 점수에 거리 비례 패널티. 구현 단순, 길이 일반화 강함.
요즘 긴 컨텍스트 확장(YaRN 등)이 RoPE의 회전 주파수를 조작하는 방식인 것도, RoPE가 표준이 된 이유를 보여준다.