A complete GPT, forward and backward, written with NumPy only. Every gradient was derived by hand, implemented, then checked against two independent references. 순전파와 역전파까지 NumPy만으로 작성한 완전한 GPT입니다. 모든 그래디언트를 손으로 유도하고, 구현한 뒤, 서로 독립적인 두 가지 기준으로 검증했습니다.
GPT math.pdfForward and backward formulas were worked through step by step on paper.순전파와 역전파 공식을 종이에 단계별로 직접 풀었습니다. GPT math.pdf
ChatGPT turned the derivations into NumPy code, with no autodiff framework in the model.ChatGPT를 사용해 유도한 수식을 NumPy 코드로 옮겼습니다. 모델에는 자동 미분 프레임워크가 없습니다.
The code was checked against the notes: tensor shapes, forward values, every intermediate gradient.코드를 노트와 대조했습니다: 텐서 차원, 순전파 값, 모든 중간 그래디언트.
Central finite differences and TensorFlow GradientTape, independently.중앙 차분과 TensorFlow GradientTape로 각각 독립적으로 확인했습니다.
GPT-style decoder with Post-LayerNorm and weight tying. Each backward step is annotated with the page of the notes it comes from.Post-LayerNorm과 가중치 공유를 사용하는 GPT 스타일 디코더입니다. 각 역전파 단계에 해당 유도가 있는 노트 쪽수를 표시했습니다.
| Setting설정 | Value값 |
|---|---|
| Vocabulary (word-level)어휘 (단어 단위) | 169 |
| d_model / heads / d_ff / layers | 64 / 4 / 128 / 2 |
| Max context최대 컨텍스트 | 11 tokens |
| Parameters파라미터 수 | 77,952 |
| Optimizer옵티마이저 | Adam · lr 2e-3 · clip 1.0 |
| Batch배치 | 1 sentence, 36 updates per epoch문장 1개, 에폭당 36회 업데이트 |
| Trained for학습량 | 700 epochs |
Where each part of the derivation lives in각 유도가 들어 있는 위치: GPT math.pdf
Components first (GELU, LayerNorm, attention head, block), then every parameter of the full model. LayerNorm γ, β and the biases are moved away from 1 and 0 so that a missing γ cannot hide.먼저 구성 요소(GELU, LayerNorm, 어텐션 헤드, 블록)를, 그다음 전체 모델의 모든 파라미터를 확인합니다. γ 누락 같은 버그가 숨지 못하도록 LayerNorm γ, β와 편향을 1과 0에서 벗어나게 설정합니다.
Compares against central differences (L(θ+h) − L(θ−h)) / 2h. Writes an HTML report with an analytic-vs-numerical scatter and an error-vs-h curve.중앙 차분 (L(θ+h) − L(θ−h)) / 2h와 비교합니다. 해석적/수치 그래디언트 산점도와 h에 따른 오차 곡선이 담긴 HTML 리포트를 만듭니다.
Re-implements the same model independently with TensorFlow ops and compares against GradientTape.같은 모델을 TensorFlow 연산으로 독립적으로 다시 구현하고 GradientTape 결과와 비교합니다.
Why the full model stops near 1e-7: the 1e-12 inside log(p + 1e-12) shifts the loss slightly for tokens with tiny probability. Why GELU rows sit at 1e-10 in TensorFlow: tf.nn.gelu already differs by about 1e-10 in the forward pass. Writing the same GELU formula by hand in TF brings all 147 checks to 1e-15.전체 모델이 1e-7 근처에서 멈추는 이유: log(p + 1e-12) 안의 1e-12가 확률이 아주 작은 토큰에서 손실을 조금 바꾸기 때문입니다. TensorFlow에서 GELU 항목이 1e-10인 이유: tf.nn.gelu가 순전파에서 이미 약 1e-10 차이가 납니다. 같은 GELU 공식을 TF로 직접 쓰면 147개 모두 1e-15까지 일치합니다.
Trained weights (model.npz) ship with the repo, so generation works right away.학습된 가중치(model.npz)가 저장소에 포함되어 있어 바로 생성할 수 있습니다.
$ git clone https://github.com/ColumnBack/tinygpt-numpy.git $ cd tinygpt-numpy $ pip install -r requirements.txt # numpy only $ python generate.py
the, my, we, she문장을 시작하는 단어가 가장 잘 동작합니다: the, my, we, shehello) are reported모르는 단어(예: hello)는 알려 줍니다python generate.py --temperature 1.0 --samples 5
model.npz after every epoch, Adam state included매 에폭 model.npz에 저장 (Adam 상태 포함)python train.py --epochs 1000 python train.py --retrain
| tinygpt.py | The model: TinyGPT forward/backward, Adam, tokenizer, save/load모델: TinyGPT 순전파/역전파, Adam, 토크나이저, 저장/불러오기 |
| train.py | Corpus of 36 sentences and the training loop, with resume36개 문장 코퍼스와 학습 루프 (재개 지원) |
| generate.py | Type one or two words, get a sentence단어 한두 개를 입력하면 문장 생성 |
| gradcheck.py | Gradient check ①, numerical, with an HTML report그래디언트 검증 ①, 수치 미분, HTML 리포트 생성 |
| tf_gradcheck.py | Gradient check ②, TensorFlow GradientTape그래디언트 검증 ②, TensorFlow GradientTape |
| model.npz | Trained weights, 700 epochs학습된 가중치, 700 에폭 |
| GPT math.pdf | Handwritten forward/backward derivations손글씨 순전파/역전파 유도 노트 |
the).대부분 외웁니다: 문장 36개, 대부분의 단어가 한 번씩 등장. 다음 단어 정답이 여러 개인 위치가 있어 손실이 0.42 근처에서 멈춥니다 (the 다음에 올 수 있는 문장이 20개 이상).tf_gradcheck.py (Python 3.13 + TF 2.21)TensorFlow는 tf_gradcheck.py에만 필요 (Python 3.13 + TF 2.21)