GitHub ↗
Decoder-only Transformer · educational디코더 전용 Transformer · 교육용

TinyGPT, in NumPy

A complete GPT, forward and backward, written with NumPy only. Every gradient was derived by hand, implemented, then checked against two independent references. 순전파와 역전파까지 NumPy만으로 작성한 완전한 GPT입니다. 모든 그래디언트를 손으로 유도하고, 구현한 뒤, 서로 독립적인 두 가지 기준으로 검증했습니다.

GPT math.pdf
Handwritten derivations of every forward and backward formula · 22 pages모든 순전파·역전파 공식을 손으로 유도한 노트 · 22쪽
GS = A ⊙ [ GA − (GA ⊙ A) 11⊤ ] Softmax backward, derived by hand in the notes (pp. 16–19) and used as-is in the code노트(16–19쪽)에서 손으로 유도해 코드에 그대로 쓴 softmax 역전파
77,952 parameters파라미터 numpy the only dependency유일한 의존성 36 training sentences개 학습 문장 22 pages of handwritten derivation쪽 분량의 손글씨 유도 56/56 + 147/147 PASS
Method방법

From paper to verified code종이에서 검증된 코드까지

1

Derive by hand손으로 유도

Forward and backward formulas were worked through step by step on paper.순전파와 역전파 공식을 종이에 단계별로 직접 풀었습니다. GPT math.pdf

2

Implement in NumPyNumPy로 구현

ChatGPT turned the derivations into NumPy code, with no autodiff framework in the model.ChatGPT를 사용해 유도한 수식을 NumPy 코드로 옮겼습니다. 모델에는 자동 미분 프레임워크가 없습니다.

3

Compare line by line한 줄씩 대조

The code was checked against the notes: tensor shapes, forward values, every intermediate gradient.코드를 노트와 대조했습니다: 텐서 차원, 순전파 값, 모든 중간 그래디언트.

4

Check gradients twice그래디언트 이중 검증

Central finite differences and TensorFlow GradientTape, independently.중앙 차분과 TensorFlow GradientTape로 각각 독립적으로 확인했습니다.

Architecture아키텍처

Forward goes down, gradients come back up순전파는 아래로, 그래디언트는 위로

GPT-style decoder with Post-LayerNorm and weight tying. Each backward step is annotated with the page of the notes it comes from.Post-LayerNorm과 가중치 공유를 사용하는 GPT 스타일 디코더입니다. 각 역전파 단계에 해당 유도가 있는 노트 쪽수를 표시했습니다.

TinyGPT architecture with the backward pass FORWARD BACKWARD token ids E[token] + P[pos] × 2 blocks Masked multi-head attention 4 heads · dh = 16 · causal mask Add & LayerNorm Feed-forward · GELU 64 → 128 → 64 Add & LayerNorm residual logits = X E⊤ softmax + cross-entropy weight tying · same E GL = (P − Y) / T p.6 · softmax + CE GE = GEemb + GELM p.7 · LM head, tied LayerNorm backward p.7–14 · compact form FFN · GELU′ · residual p.15 WO, AV, softmax, QK⊤ p.16–19 · attention scatter-add into E, P p.20 · embedding
Dashed magenta rail: gradient flow. Dashed blue: the embedding matrix E is reused as the output layer, so its gradient has two contributions.자홍색 점선: 그래디언트 흐름. 파란 점선: 임베딩 행렬 E를 출력층에서 다시 쓰므로 그래디언트가 두 경로에서 합쳐집니다.
Setting설정Value값
Vocabulary (word-level)어휘 (단어 단위)169
d_model / heads / d_ff / layers64 / 4 / 128 / 2
Max context최대 컨텍스트11 tokens
Parameters파라미터 수77,952
Optimizer옵티마이저Adam · lr 2e-3 · clip 1.0
Batch배치1 sentence, 36 updates per epoch문장 1개, 에폭당 36회 업데이트
Trained for학습량700 epochs
Each attention head각 어텐션 헤드 Q = XWQ,  K = XWK,  V = XWV A = softmax( QK⊤√dh + M ),  Oh = AV M is the causal mask that blocks future positions.M은 미래 위치를 가리는 causal mask입니다.
The notes유도 노트

22 pages, mapped22쪽 노트 지도

Where each part of the derivation lives in각 유도가 들어 있는 위치: GPT math.pdf

LayerNorm backward, compact formLayerNorm 역전파, 간결한 형태 GR = 1D ⊙ [ GR̂ − mean(GR̂) − R̂ ⊙ mean(GR̂ ⊙ R̂) ],   GR̂ = GY ⊙ γ
Verification검증

Two independent gradient checks서로 독립적인 두 가지 그래디언트 검증

Components first (GELU, LayerNorm, attention head, block), then every parameter of the full model. LayerNorm γ, β and the biases are moved away from 1 and 0 so that a missing γ cannot hide.먼저 구성 요소(GELU, LayerNorm, 어텐션 헤드, 블록)를, 그다음 전체 모델의 모든 파라미터를 확인합니다. γ 누락 같은 버그가 숨지 못하도록 LayerNorm γ, β와 편향을 1과 0에서 벗어나게 설정합니다.

Typical relative error per check, log scale PASS < 1e-6 WARN FAIL 1e-16 1e-14 1e-12 1e-10 1e-8 1e-6 1e-4 relative error (log scale) · lower is better상대 오차 (로그 스케일) · 왼쪽일수록 정확 ① Numerical gradient① 수치 그래디언트 GELU, LN, attention, blockGELU, LN, 어텐션, 블록 ~1e-11 Full model, 32 param groups전체 모델, 파라미터 그룹 32개 ~1e-7 ② TensorFlow autodiff② TensorFlow 자동 미분 LayerNorm, attention headLayerNorm, 어텐션 헤드 ~1e-16 GELU, block, full modelGELU, 블록, 전체 모델 ~1e-10
Typical relative error reported by each script. Thresholds: PASS below 1e-6, WARN up to 1e-4, FAIL above.각 스크립트가 보고하는 대표 상대 오차입니다. 판정 기준: 1e-6 미만 PASS, 1e-4까지 WARN, 그 이상 FAIL.
gradcheck.py
56 / 56
PASS

Compares against central differences (L(θ+h) − L(θ−h)) / 2h. Writes an HTML report with an analytic-vs-numerical scatter and an error-vs-h curve.중앙 차분 (L(θ+h) − L(θ−h)) / 2h와 비교합니다. 해석적/수치 그래디언트 산점도와 h에 따른 오차 곡선이 담긴 HTML 리포트를 만듭니다.

tf_gradcheck.py
147 / 147
PASS

Re-implements the same model independently with TensorFlow ops and compares against GradientTape.같은 모델을 TensorFlow 연산으로 독립적으로 다시 구현하고 GradientTape 결과와 비교합니다.

Why the full model stops near 1e-7: the 1e-12 inside log(p + 1e-12) shifts the loss slightly for tokens with tiny probability. Why GELU rows sit at 1e-10 in TensorFlow: tf.nn.gelu already differs by about 1e-10 in the forward pass. Writing the same GELU formula by hand in TF brings all 147 checks to 1e-15.전체 모델이 1e-7 근처에서 멈추는 이유: log(p + 1e-12) 안의 1e-12가 확률이 아주 작은 토큰에서 손실을 조금 바꾸기 때문입니다. TensorFlow에서 GELU 항목이 1e-10인 이유: tf.nn.gelu가 순전파에서 이미 약 1e-10 차이가 납니다. 같은 GELU 공식을 TF로 직접 쓰면 147개 모두 1e-15까지 일치합니다.

Quick start빠른 시작

Clone, then generate클론하고 바로 생성

Trained weights (model.npz) ship with the repo, so generation works right away.학습된 가중치(model.npz)가 저장소에 포함되어 있어 바로 생성할 수 있습니다.

$ git clone https://github.com/ColumnBack/tinygpt-numpy.git
$ cd tinygpt-numpy
$ pip install -r requirements.txt   # numpy only
$ python generate.py
Live demo라이브 데모 Type a word and watch it finish the sentence단어를 입력하고 문장이 완성되는 과정을 보세요 The same 77,952 weights, running in your browser, next to the 36 sentences it learned from.같은 77,952개 가중치를 브라우저에서 실행하고, 학습한 36개 문장과 나란히 보여줍니다.

generate.py

  • Words that start a sentence work best: the, my, we, she문장을 시작하는 단어가 가장 잘 동작합니다: the, my, we, she
  • Unknown words (e.g. hello) are reported모르는 단어(예: hello)는 알려 줍니다
  • Each output is tagged as a training sentence or a new combination출력마다 학습 문장인지 새로운 조합인지 표시합니다
python generate.py --temperature 1.0 --samples 5

train.py

  • Saves model.npz after every epoch, Adam state included매 에폭 model.npz에 저장 (Adam 상태 포함)
  • Ctrl+C saves and stops; re-running resumes exactlyCtrl+C로 저장 후 중지, 다시 실행하면 그대로 재개
  • Changing the sentences triggers a fresh start학습 문장을 바꾸면 처음부터 다시 학습
python train.py --epochs 1000
python train.py --retrain
Files파일 구성

What is in the repo저장소 구성

tinygpt.pyThe model: TinyGPT forward/backward, Adam, tokenizer, save/load모델: TinyGPT 순전파/역전파, Adam, 토크나이저, 저장/불러오기
train.pyCorpus of 36 sentences and the training loop, with resume36개 문장 코퍼스와 학습 루프 (재개 지원)
generate.pyType one or two words, get a sentence단어 한두 개를 입력하면 문장 생성
gradcheck.pyGradient check ①, numerical, with an HTML report그래디언트 검증 ①, 수치 미분, HTML 리포트 생성
tf_gradcheck.pyGradient check ②, TensorFlow GradientTape그래디언트 검증 ②, TensorFlow GradientTape
model.npzTrained weights, 700 epochs학습된 가중치, 700 에폭
GPT math.pdfHandwritten forward/backward derivations손글씨 순전파/역전파 유도 노트
Limitations한계
  • It mostly memorizes: 36 sentences, most words seen once. Loss plateaus near 0.42 because some positions have many valid next words (over 20 sentences can follow the).대부분 외웁니다: 문장 36개, 대부분의 단어가 한 번씩 등장. 다음 단어 정답이 여러 개인 위치가 있어 손실이 0.42 근처에서 멈춥니다 (the 다음에 올 수 있는 문장이 20개 이상).
  • Word-level tokens, so unseen words cannot be handled.단어 단위 토큰이라 처음 보는 단어는 처리할 수 없습니다.
  • No batching; one sentence at a time on the CPU.배치 없이 CPU에서 문장 하나씩 처리합니다.
  • Post-LayerNorm (GPT-2 and later use Pre-LayerNorm).Post-LayerNorm 구조입니다 (GPT-2 이후는 Pre-LayerNorm).
Requirements요구 사항
  • Python 3.10+ (tested on 3.14)(3.14에서 테스트)
  • NumPy (tested on 2.3)(2.3에서 테스트)
  • TensorFlow only for tf_gradcheck.py (Python 3.13 + TF 2.21)TensorFlow는 tf_gradcheck.py에만 필요 (Python 3.13 + TF 2.21)