Three models, one review, and a vocabulary problem
I trained different networks to classify movie reviews because I wanted to see how their approaches compared. A CNN, a bidirectional LSTM, and FNet gave me three ways to work on the same task. Bringing them into one app made the differences easier to explore—and exposed how much a model depends on its inputs.
The hidden part of a saved model
The BiLSTM integration hinged on vocabulary IDs. A word’s number is only useful if it is the same number the network saw during training. A new word mapping can produce a perfectly valid numeric sequence while giving the trained network the wrong meanings.
I restored the original IMDB word indices, special-token offsets, and sequence length. The CNN and FNet also keep their own saved preprocessing. That made the inference path consistent with the training work instead of assuming the weights alone were enough.
Turning a comparison into something inspectable
The interface captures one review and runs all three models on that same text. It shows the individual labels and scores, including disagreement. If the review changes during a request, an old result cannot quietly appear as the answer to the new text.
What I can—and cannot—conclude
The saved coursework records include 83.23% FNet test accuracy and 86.11% best BiLSTM validation accuracy. Those are different evaluation splits, so I do not use them to declare a winner. The app also labels model scores separately from calibrated confidence.
This project changed how I think about deployment: preprocessing belongs with the model, and a useful comparison needs controlled evaluation. The next step is a shared held-out review set, with confusion matrices and examples of errors on sarcasm and mixed opinions.