05 / FIELD NOTETRAINING PIPELINE / IN DEVELOPMENT
SMALL LANGUAGE MODEL
Learning from
first principles.
An educational decoder-only language model with 97.5 million parameters, built from random initialization. Architecture, a byte-level BPE tokenizer, and checkpointed training workflows are implemented; a useful trained model is the next challenge.
CURRENT STATE /Training smoke workflows have been exercised. A substantively pretrained, evaluated model is not yet presented as available.
CONCEPT DIAGRAM / EXPLORE THE STAGES
01 / 03ARCHITECTURE
Twelve decoder blocks with RoPE, RMSNorm, and SwiGLU; 97,536,768 trainable parameters in the default configuration.
01 / ARCHITECTURE
Twelve decoder blocks with RoPE, RMSNorm, and SwiGLU; 97,536,768 trainable parameters in the default configuration.
02 / PIPELINE
A byte-level BPE tokenizer, streamed data, mixed precision, checkpoint/resume, and instruction-tuning workflows establish the route to training.
03 / NEXT MILESTONE
Meaningful pretraining and evaluation will determine what the model can really do.
SELECT A STAGE · SWIPE THE DIAGRAMPB / 2026
IDEAS MOVE FORWARD WHEN THEIR STATE IS CLEAR.PRABHAS BANGARUGARI © 2026