MACHINE LEARNING · INTERACTIVE PROJECT
A picture.
A new perspective.
Upload a photo and explore what a trained vision-and-language model describes. Built with EfficientNetB0 features and a Transformer on Flickr8k.
01 / THE IMAGEJPG · PNG · WEBP
Start with a photograph.
Your photo stays in this local workspace.
Through the
model’s eyes.
EfficientNetB0 reads the image. A Transformer turns those features into words.
Your generated caption will appear here.
A LEARNING PROJECT
Trained on Flickr8k. Descriptions can miss details or invent them—compare the result with what you see.
Academic project using saved model artifacts. Predictions can be imperfect; training results and limitations are shown alongside the implementation.
Review before you use it
Check the description against the photograph. Edit missing or incorrect details before exporting it.