Positional Encodings, Compared Positional Encoding 종류별 비교

Attention by itself doesn’t know order. The ways to inject position:

  • Sinusoidal: the original. No learned parameters; extrapolation was the hope, but in practice it’s weak.
  • Learned: one embedding per position. Beyond the training length it simply can’t.
  • RoPE: rotates Q and K by position. Relative position falls out of the dot product naturally; it’s the current LLM standard.
  • ALiBi: a distance-proportional penalty on the scores. Simple to implement, strong at length generalization.

The fact that today’s long-context extensions (YaRN and friends) work by tweaking RoPE’s rotation frequencies tells you why RoPE became the standard.

attention 자체는 순서를 모른다. 위치 정보를 넣는 방식들:

  • Sinusoidal — 원조. 학습 파라미터 없음, 외삽 기대했지만 실제론 약함.
  • Learned — 위치별 임베딩 학습. 학습 길이 밖은 아예 불가.
  • RoPE — Q·K를 위치에 따라 회전. 상대 위치가 내적에 자연히 반영, 현 LLM 표준.
  • ALiBi — 점수에 거리 비례 패널티. 구현 단순, 길이 일반화 강함.

요즘 긴 컨텍스트 확장(YaRN 등)이 RoPE의 회전 주파수를 조작하는 방식인 것도, RoPE가 표준이 된 이유를 보여준다.

← home← 랜딩으로