Public preview · Try Sentiment, Image Captioning and Housing. Explore Forge & Voice and WiseShield through their project stories.
IAIbrahim Aldulaimi

Computer vision

Image Captioning

A restored image-to-text pipeline with a deliberate review step before export.

Open application ↗
Project
Academic computer vision + inference app · TensorFlow
My contribution
Connected image features, Transformer weights, vocabulary, upload validation, and editable caption export.
Result & scope
Generates reviewable caption drafts. Restoration is verified; semantic accuracy remains a limitation.

When a fluent caption tells the wrong story

Image Captioning began as a way to connect two kinds of learning: recognizing visual features and turning them into words. Using Flickr8k image-caption pairs, I worked with an EfficientNet feature extractor and a Transformer to generate a short description of a photograph.

Rebuilding the whole path to a sentence

Making the saved work usable in a browser meant restoring more than the image encoder. The Transformer weights, vocabulary, start and end tokens, and pixel preprocessing all had to agree with the training setup. I connected those pieces and added checks that the expected checkpoint variables were restored.

A working pipeline could still be wrong

A later smoke test made that distinction concrete: a flower photograph received a caption about a dress and leaves. The system had generated a sentence successfully, but the sentence did not describe the image. Restoring the model correctly had not solved its semantic accuracy limits.

That changed the product workflow. A generated caption is presented as a draft the user can inspect and edit before export. Changing the image clears the old result, and late responses cannot attach a caption to the wrong photograph.

The result and the next experiment

The app now supports upload, preview, generation, correction, and export with the restored model. Its value as a portfolio project is showing that full path, including a visible place for human judgment.

I learned to evaluate correctness at both levels: whether the software runs as intended and whether the model’s words fit the image. A systematic held-out evaluation and comparison of decoding strategies remain next steps. I do not claim a new caption-accuracy score from the interface improvements.

How it works

Problem and dataset

Generate a short description of a photograph using a model trained on Flickr8k image-caption pairs. The app lets a visitor upload an image, inspect the prediction, edit inaccuracies, and export a reviewed caption.

Model architecture

A saved EfficientNetB0 feature extractor supplies image features to a Transformer encoder and decoder. The implementation uses TensorFlow and Keras with custom Transformer layers and a saved vocabulary. The training work used frozen image features, learning-rate warmup, early stopping, and saved checkpoints.

How inference works

The service validates the uploaded image, converts it to RGB, and resizes it to 299 × 299 pixels while preserving the original training pixel range. It restores the encoder and decoder checkpoint, begins with a start token, then selects the highest-scoring next token repeatedly until the end token or a 24-step limit.

Application engineering

Pillow handles image decoding, FastAPI handles upload requests, and the browser shows a local preview. File-size and pixel-count limits bound inputs. Changing the image clears the previous result; a late response cannot attach a caption to a different image. A separate editable field makes human review explicit before text export.

Verification and limits

Checkpoint restoration and generated captions were verified. The model can omit objects, invent details, or describe unfamiliar scenes poorly. Greedy decoding produces one candidate rather than a calibrated set of alternatives. No BLEU, CIDEr, or other new held-out score is claimed here.

Next steps

Evaluate a fixed held-out image set, inspect systematic errors, and compare decoding strategies. The reviewed caption should remain visibly distinct from the raw model prediction.