When a fluent caption tells the wrong story
Image Captioning began as a way to connect two kinds of learning: recognizing visual features and turning them into words. Using Flickr8k image-caption pairs, I worked with an EfficientNet feature extractor and a Transformer to generate a short description of a photograph.
Rebuilding the whole path to a sentence
Making the saved work usable in a browser meant restoring more than the image encoder. The Transformer weights, vocabulary, start and end tokens, and pixel preprocessing all had to agree with the training setup. I connected those pieces and added checks that the expected checkpoint variables were restored.
A working pipeline could still be wrong
A later smoke test made that distinction concrete: a flower photograph received a caption about a dress and leaves. The system had generated a sentence successfully, but the sentence did not describe the image. Restoring the model correctly had not solved its semantic accuracy limits.
That changed the product workflow. A generated caption is presented as a draft the user can inspect and edit before export. Changing the image clears the old result, and late responses cannot attach a caption to the wrong photograph.
The result and the next experiment
The app now supports upload, preview, generation, correction, and export with the restored model. Its value as a portfolio project is showing that full path, including a visible place for human judgment.
I learned to evaluate correctness at both levels: whether the software runs as intended and whether the model’s words fit the image. A systematic held-out evaluation and comparison of decoding strategies remain next steps. I do not claim a new caption-accuracy score from the interface improvements.