Public preview · Try Sentiment, Image Captioning and Housing. Explore Forge & Voice and WiseShield through their project stories.
IAIbrahim Aldulaimi

Training experiments

Language Models

What historical language-model experiments taught me about evaluating generated text.

Project overview · no live application

Project
Historical team coursework · BART, GPT-2, small GPT
Work documented
Fine-tuning, small-model training, learning curves, and interpretation of generated outputs.
Evidence & limits
Small-model validation perplexity: 94.99 → 58.07 over six epochs. Runnable artifacts are not yet recovered.

Learning what an improving training curve can tell me

In team coursework, I explored language models from two directions: adapting existing models and training a small GPT-style model. The experiments included BART with SAMSum, GPT-2 work with Reddit TIFU, and a small-model experiment with SimpleBooks92.

Different experiments, different evidence

The work brought together training curves, validation measurements, and generated text. Those views answer different questions. A falling loss can show that optimization is progressing, while a generated example can reveal a limitation the aggregate number does not explain.

In the approximately 3.26-million-parameter GPT-style experiment, recorded validation perplexity fell from 94.99 to 58.07 over six epochs. That result belongs to that small-model experiment; it is not a score for the BART or GPT-2 work.

Interpreting progress carefully

The important challenge was deciding what the measurements supported. Lower perplexity indicates better predictive fit under that experiment’s tokenization and validation setup. It does not establish factual accuracy or the usefulness of a summary. Keeping the tasks and evidence separate made the conclusions more precise.

What I can show today

This page preserves the documented coursework summary. The runnable scripts and checkpoints have not yet been recovered in the portfolio workspace, so I cannot offer a reproducible live demonstration or claim a fresh evaluation here.

The work reinforced two habits I want to carry into later projects: inspect examples alongside metrics, and preserve the artifacts needed to reproduce a result. Recovering the notebooks, dataset splits, tokenization, and checkpoints is the next step in turning this historical overview into a verifiable experiment.

How it works

Experiment scope

Historical team coursework covered BART fine-tuning on SAMSum, GPT-2 work with Reddit TIFU, and training a small GPT-style model on SimpleBooks92. These are different experiments with different tasks and should not be treated as one model.

What was examined

The work recorded training curves, validation measurements, and generated text to inspect how optimization changed behavior. The small GPT-style experiment was recorded as approximately 3.26 million parameters, with validation perplexity decreasing from 94.99 to 58.07 over six epochs.

How to interpret the result

Perplexity measures predictive fit under the experiment’s tokenization and validation setup; it does not by itself establish summarization quality, factual accuracy, or useful dialogue behavior. The reported change belongs to the small language-model experiment, not to BART or GPT-2.

Available evidence and limits

The project overview preserves the previously reviewed coursework summary. Runnable training scripts and checkpoints for these experiments have not yet been recovered in the portfolio workspace. Exact optimizer settings, hardware timings, and fresh evaluation results are therefore not asserted.

Next reproducibility step

Recover the notebooks and checkpoints, record dataset splits and tokenization, reproduce the evaluation, and add representative outputs with known limitations before exposing a live app.