← Playground← 놀이터

PRISM · Statistical-Mechanics PlaygroundPRISM · 통계역학 놀이터

Hopfield → Attention홉필드 → 어텐션
Associative memory, capacity ≈ 0.14N, and the link to softmax attention연상 기억, 용량 한계 ≈ 0.14N, 그리고 softmax 어텐션과의 연결

Left: a classic Hopfield network stores patterns as attractors. Corrupt one and its dynamics flow back to the stored memory — until you exceed the capacity P ≈ 0.14 N, where recall breaks down. Right: the modern continuous Hopfield update ξ ← X·softmax(β Xᵀξ) retrieves in a single step — and that is exactly softmax attention with the stored patterns as keys and values.왼쪽: 고전적인 Hopfield 네트워크는 패턴을 끌개(attractor)로 저장한다. 하나를 손상시키면 그 동역학이 저장된 기억으로 다시 흘러들어간다 — 용량 한계 P ≈ 0.14 N를 넘어서면 복원이 무너지기 전까지는. 오른쪽: modern continuous Hopfield 갱신식 ξ ← X·softmax(β Xᵀξ)은 단 한 번의 단계로 검색해낸다 — 그리고 이것은 저장된 패턴을 key와 value로 삼은 softmax 어텐션과 정확히 같다.

Classic Hopfield — attractors & capacity고전 Hopfield — 끌개와 용량 한계
Store · corrupt · recover저장 · 손상 · 복원 target → probe → state타깃 → 프로브 → 상태
Controls컨트롤
Stored patterns P저장된 패턴 수 P of 6 memories전체 6개 기억 중4
Corruption손상 정도 fraction of flipped pixels뒤집힌 픽셀의 비율0.20
Target pattern타깃 패턴 which memory to probe어떤 기억을 검사할지#1
overlap m (state·target)오버랩 m (상태·타깃)
async sweeps비동기 스윕 수
A corrupted probe flows downhill in energy to the nearest stored memory.손상된 프로브는 에너지의 내리막을 따라 가장 가까운 저장된 기억으로 흘러간다.
Capacity용량 한계 recall rate vs load α = P/N · random patterns부하 α = P/N에 따른 복원 성공률 · 무작위 패턴
α = 0load α = P / N부하 α = P / Nα = 0.30
Modern Hopfield = softmax attention모던 Hopfield = softmax 어텐션
Attention over memories기억에 대한 어텐션 w = softmax(β · ⟨query, patternμ⟩)w = softmax(β · ⟨쿼리, 패턴μ⟩)
One-step retrieval한 번의 단계 검색 query → Σ wμ · patternμ쿼리 → Σ wμ · 패턴μ
Controls컨트롤
Inverse temperature β역온도 β attention sharpness어텐션의 날카로움4.0
top weight wμ최대 가중치 wμ
effective memories 1/Σw²유효 기억 수 1/Σw²
Modern Hopfield모던 Hopfield  ξnew = X·softmax(β·Xᵀξ)
Attention어텐션      out = V·softmax(Kᵀq/√d)
Set keys = values = stored patterns X, and β = 1/√d — they are the same operation. Low β blends memories (attention averaging); high β retrieves one (sharp recall).key = value = 저장된 패턴 X로 두고 β = 1/√d로 잡으면 — 둘은 완전히 같은 연산이다. β가 작으면 기억들이 섞이고(어텐션 평균화), β가 크면 하나만 검색된다(날카로운 복원).

Hebbian attractorsHebb 학습 끌개

Patterns are stored in symmetric weights Wᵢⱼ = (1/N) Σμ ξᵢ^μ ξⱼ^μ. Each memory becomes a local minimum of the energy E = −½ Σ Wᵢⱼ sᵢsⱼ; asynchronous updates sᵢ ← sign(Σⱼ Wᵢⱼ sⱼ) flow downhill to the nearest one.패턴은 대칭 가중치 Wᵢⱼ = (1/N) Σμ ξᵢ^μ ξⱼ^μ에 저장된다. 각 기억은 에너지 E = −½ Σ Wᵢⱼ sᵢsⱼ의 국소 최솟값이 되며, 비동기 갱신 sᵢ ← sign(Σⱼ Wᵢⱼ sⱼ)은 내리막을 따라 가장 가까운 기억으로 흘러간다.

The 0.14N cliff0.14N 절벽

Cross-talk between stored patterns grows with the load α = P/N. Above the Amit–Gutfreund–Sompolinsky limit αc ≈ 0.138, spurious minima take over and recall collapses — the sharp drop in the capacity curve.저장된 패턴 사이의 간섭(cross-talk)은 부하 α = P/N가 커질수록 늘어난다. Amit–Gutfreund–Sompolinsky 한계 αc ≈ 0.138을 넘어서면 가짜 최솟값이 우세해지고 복원이 붕괴한다 — 용량 곡선에서 나타나는 급격한 낙하가 그것이다.

Modern Hopfield → Transformers모던 Hopfield → Transformer

Continuous-state Hopfield networks (Ramsauer et al., 2020) retrieve in one update with exponential capacity, and that update is the Transformer's attention. Raising β is lowering the attention temperature: from soft averaging to hard retrieval.연속 상태 Hopfield 네트워크(Ramsauer et al., 2020)는 지수적 용량을 가지고 단 한 번의 갱신으로 검색하며, 그 갱신은 곧 Transformer의 어텐션 바로 그것이다. β를 키우는 것은 어텐션 온도를 낮추는 것과 같다 — 부드러운 평균화에서 단단한 검색으로.