Learning what an improving training curve can tell me
In team coursework, I explored language models from two directions: adapting existing models and training a small GPT-style model. The experiments included BART with SAMSum, GPT-2 work with Reddit TIFU, and a small-model experiment with SimpleBooks92.
Different experiments, different evidence
The work brought together training curves, validation measurements, and generated text. Those views answer different questions. A falling loss can show that optimization is progressing, while a generated example can reveal a limitation the aggregate number does not explain.
In the approximately 3.26-million-parameter GPT-style experiment, recorded validation perplexity fell from 94.99 to 58.07 over six epochs. That result belongs to that small-model experiment; it is not a score for the BART or GPT-2 work.
Interpreting progress carefully
The important challenge was deciding what the measurements supported. Lower perplexity indicates better predictive fit under that experiment’s tokenization and validation setup. It does not establish factual accuracy or the usefulness of a summary. Keeping the tasks and evidence separate made the conclusions more precise.
What I can show today
This page preserves the documented coursework summary. The runnable scripts and checkpoints have not yet been recovered in the portfolio workspace, so I cannot offer a reproducible live demonstration or claim a fresh evaluation here.
The work reinforced two habits I want to carry into later projects: inspect examples alongside metrics, and preserve the artifacts needed to reproduce a result. Recovering the notebooks, dataset splits, tokenization, and checkpoints is the next step in turning this historical overview into a verifiable experiment.