LLM from scratch, part 28 – training a base model from scratch on an RTX 3090
↗A dense, data-driven diary of training a GPT-2 small base model from scratch on a single RTX 3090, detailing data prep, tokenization, sequence handling, speed tests, checkpoints, validation, and comparisons to OpenAI weights, with conclusions about feasibility and gaps to achieve a true Chinchilla-optimal train.
Dec 9, 20251%