TurboGPT, a tiny byte-level GPT trainer in CUDA C++, shown on HN

TurboGPT was shared as a Show HN project. The submission title says it trains a 22KiB transformer in 13 seconds. The README text available for this account does not mention the 22KiB model size or the 13-second training time, so those two figures are the submitter's claim only.
The README describes the project in one line: tiny byte-level GPT training in CUDA C++, under the MIT licence.
Building it requires Windows, the Visual Studio 2022 C++ tools and CUDA 13.4. The build command is .\build.ps1 -CudaArch 86, where CudaArch is the GPU compute capability taken from NVIDIA's CUDA GPU list (86 is the value in the example).
Training is started with .\build\turbogpt.exe --data hn1g.txt --log-to runs/ctx4. The run stores its checkpoint at runs/ctx4/ctx4.pt, and that file contains the model, optimizer, scheduler and trainer state. A run can be resumed by passing --load CHECKPOINT.pt.
A file called runs/ctx4/report.json is derived from the log directory. The logs are TensorBoard-compatible: one report per batch, capped at 8Mi reports, and flushed with periodic or final checkpoints.
The one result the README gives is for hn1g: 2.52435 BPB after 1.5G training tokens. Tests are run with python tests\verify.py.
Key facts
- TurboGPT is described in its README as tiny byte-level GPT training in CUDA C++, released under the MIT licence.
- The stated build requirements are Windows, Visual Studio 2022 C++ tools and CUDA 13.4, built with
.\build.ps1 -CudaArch 86. - The README reports 2.52435 BPB on hn1g after 1.5G training tokens.
- Checkpoints hold model, optimizer, scheduler and trainer state and can be resumed with
--load CHECKPOINT.pt. - Logs are TensorBoard-compatible, with one report per batch, capped at 8Mi reports.
Why it matters
TurboGPT is a small, focused training program: a byte-level GPT trainer written directly in CUDA C++. The Show HN title claims a 22KiB transformer trained in 13 seconds, which would make it a very fast way to train a micro-model. The README text itself gives no timing, so that claim rests on the submission title. What the README does show is a complete workflow: build, train, checkpoint, resume, log and verify.
Who it affects
People who want to train very small language models on an NVIDIA GPU and are working on Windows, since that is the platform the README covers. It is also relevant to anyone who likes reading compact CUDA C++ training code, as the project is MIT licensed.
How to use it
Install Windows, the Visual Studio 2022 C++ tools and CUDA 13.4. Build with .\build.ps1 -CudaArch 86, replacing 86 with the compute capability of your own GPU from NVIDIA's CUDA GPU list. Train with .\build\turbogpt.exe --data hn1g.txt --log-to runs/ctx4. The checkpoint lands at runs/ctx4/ctx4.pt; resume it with --load CHECKPOINT.pt. Logs can be opened in TensorBoard. Run the tests with python tests\verify.py. The project is MIT licensed.
How solid is it
The source is a short README excerpt, and the one quality figure in it is 2.52435 BPB on hn1g after 1.5G training tokens. The source text does not mention the 22KiB model size or the 13-second training time given in the title. No hardware (GPU model) is named for any timing or result. The source does not define BPB (bits per byte) or compare 2.52435 BPB to any baseline. No comparison with other frameworks (PyTorch, llm.c and so on) or speedup figure is given.
Risks and caveats
No Linux support or build instructions are given; only Windows is mentioned. The build needs a specific, recent toolchain (CUDA 13.4 and Visual Studio 2022 C++ tools) and an NVIDIA GPU whose compute capability you must supply. The source does not say what hn1g contains beyond its name, or how large the dataset is. The number of parameters and the architecture details (layers, heads, context length) are not stated, so the result cannot be put in context from the README alone.
“Tiny byte-level GPT training in CUDA C++. MIT.”
— TurboGPT README