Public preview · Try Sentiment, Image Captioning and Housing. Explore Forge & Voice and WiseShield through their project stories.
IAIbrahim Aldulaimi

Text classification

Sentiment Studio

Comparing three movie-review classifiers while preserving their original training inputs.

Open application ↗
Project
Academic model training + inference app · TensorFlow, Keras
My contribution
CNN, BiLSTM, and FNet workflows, restored preprocessing, and a shared review comparison interface.
Result & scope
Three saved networks run on the same text. Historical scores use different splits, so no benchmark winner is claimed.

Three models, one review, and a vocabulary problem

I trained different networks to classify movie reviews because I wanted to see how their approaches compared. A CNN, a bidirectional LSTM, and FNet gave me three ways to work on the same task. Bringing them into one app made the differences easier to explore—and exposed how much a model depends on its inputs.

The hidden part of a saved model

The BiLSTM integration hinged on vocabulary IDs. A word’s number is only useful if it is the same number the network saw during training. A new word mapping can produce a perfectly valid numeric sequence while giving the trained network the wrong meanings.

I restored the original IMDB word indices, special-token offsets, and sequence length. The CNN and FNet also keep their own saved preprocessing. That made the inference path consistent with the training work instead of assuming the weights alone were enough.

Turning a comparison into something inspectable

The interface captures one review and runs all three models on that same text. It shows the individual labels and scores, including disagreement. If the review changes during a request, an old result cannot quietly appear as the answer to the new text.

What I can—and cannot—conclude

The saved coursework records include 83.23% FNet test accuracy and 86.11% best BiLSTM validation accuracy. Those are different evaluation splits, so I do not use them to declare a winner. The app also labels model scores separately from calibrated confidence.

This project changed how I think about deployment: preprocessing belongs with the model, and a useful comparison needs controlled evaluation. The next step is a shared held-out review set, with confusion matrices and examples of errors on sarcasm and mixed opinions.

How it works

Problem and dataset

Classify an IMDB movie review as positive or negative and compare different neural-network approaches on the same text. This is academic model-training work extended with an interactive inference interface.

Models and tools

TensorFlow and Keras load the saved convolutional and bidirectional LSTM classifiers. KerasHub supplies the WordPiece tokenizer used by the FNet model. FastAPI exposes the classifiers to the browser. The comparison action submits the same captured review to each model and reports agreement or disagreement.

How the input is prepared

The CNN uses its saved end-to-end preprocessing, including lowercasing and punctuation/HTML handling. The BiLSTM restores the original IMDB word-index mapping, special token offsets, and a 200-token sequence. FNet uses its saved vocabulary and a 512-token WordPiece input. Reusing training-time mappings prevents a new vocabulary from changing what token identifiers mean.

Training and evaluation

The coursework examined embeddings, training curves, and validation behavior. Previously recorded results include 83.23% FNet test accuracy and 86.11% best BiLSTM validation accuracy. These come from different evaluation splits and should not be interpreted as a controlled head-to-head benchmark.

What the output means

A threshold of 0.5 converts each model’s positive-class score into a label. The raw scores are not calibrated confidence estimates. Agreement between models is informative but does not establish correctness. Sarcasm, mixed opinions, long inputs, and text outside movie reviews can produce misleading results.

Next evaluation step

Run every model on the same held-out review set, record confusion matrices and failure examples, and measure calibration. No new accuracy improvement is claimed from the comparison interface.