Public preview · Try Sentiment, Image Captioning and Housing. Explore Forge & Voice and WiseShield through their project stories.
IAIbrahim Aldulaimi

Regression

Housing Explorer

A reproducible regression experiment that connects model selection to an inspectable prediction.

Open application ↗
Project
Coursework extension · scikit-learn, California Housing
My contribution
Leakage-aware pipelines, cross-validation, held-out evaluation, and scenario comparison.
Measured result
R² 0.835 and approximately $46,540 RMSE on 4,128 held-out districts, in 1990-era dollars.

Giving a prediction an evidence trail

The housing work started with a familiar task: estimate a district’s median house value from census features. When I extended the earlier tutorial work, I wanted to be able to explain why one model was selected and what its prediction actually meant.

Starting with a fair comparison

I compared a mean baseline, linear regression, and histogram gradient boosting. The challenge was keeping evaluation honest. If preprocessing learns from validation data, or the test set repeatedly guides model selection, an impressive score can tell the wrong story.

Each candidate therefore has its own preprocessing pipeline, fitted inside the training folds. Five-fold cross-validation selects the model using only the training partition. A separate 20% holdout measures the selected model afterward.

Making the result visible

The saved run selected histogram gradient boosting. On 4,128 held-out districts, it achieved approximately $46,540 RMSE and 0.835 R². These are results from the reproducible extension, measured in 1990-era dollars—not metrics I attribute to the original coursework.

I built the explorer so a visitor can change district features, save scenarios, and compare predictions alongside the evaluation. Each saved scenario keeps its inputs, making it possible to inspect what changed rather than just collect different numbers.

Knowing where the result stops

The model describes historical California district medians, not the current price of an individual home. Scenario differences show how this model behaves; they do not prove a causal effect. Working through those distinctions taught me to make data context and evaluation part of the interface. A geographic holdout would be a useful next test of how well the model transfers between areas.

How it works

Problem and dataset

Estimate historical California district median house values from eight census features. The runnable extension uses the scikit-learn California Housing dataset from the 1990 census, with targets measured in units of $100,000.

Tools and pipeline

Python, NumPy, scikit-learn, and joblib provide training and persistence. Each candidate is wrapped in a pipeline containing a SimpleImputer, StandardScaler, and estimator. Candidates are a mean baseline, linear regression, and histogram gradient boosting.

How training and selection work

A reproducible random split reserves 20% of districts for testing with seed 42. Five shuffled cross-validation folds on the training partition compare mean RMSE. The candidate with the lowest cross-validation error is selected, and its fitted model is evaluated once on the held-out test partition. Fitting preprocessing inside each pipeline keeps fold validation data out of the preprocessing fit.

Measured results

The saved run selected histogram gradient boosting. On 4,128 held-out districts it achieved approximately $46,540 RMSE in 1990-era dollars and 0.835 R². The report stores candidate scores, the selected model, feature names, split sizes, seed, and training-set feature medians.

How the explorer works

The API validates eight numeric inputs and uses the saved pipeline to predict. Default inputs are training medians. Visitors can save up to five predictions with their input values and compare each against the first scenario. These differences describe model behavior, not causal effects.

Limitations and provenance

This is a new reproducible extension of earlier housing tutorial work, not a claim that the original coursework produced these exact results. The split is random across districts rather than geographic. Predictions describe historical district medians, not current individual home prices or financial advice.