Giving a prediction an evidence trail
The housing work started with a familiar task: estimate a district’s median house value from census features. When I extended the earlier tutorial work, I wanted to be able to explain why one model was selected and what its prediction actually meant.
Starting with a fair comparison
I compared a mean baseline, linear regression, and histogram gradient boosting. The challenge was keeping evaluation honest. If preprocessing learns from validation data, or the test set repeatedly guides model selection, an impressive score can tell the wrong story.
Each candidate therefore has its own preprocessing pipeline, fitted inside the training folds. Five-fold cross-validation selects the model using only the training partition. A separate 20% holdout measures the selected model afterward.
Making the result visible
The saved run selected histogram gradient boosting. On 4,128 held-out districts, it achieved approximately $46,540 RMSE and 0.835 R². These are results from the reproducible extension, measured in 1990-era dollars—not metrics I attribute to the original coursework.
I built the explorer so a visitor can change district features, save scenarios, and compare predictions alongside the evaluation. Each saved scenario keeps its inputs, making it possible to inspect what changed rather than just collect different numbers.
Knowing where the result stops
The model describes historical California district medians, not the current price of an individual home. Scenario differences show how this model behaves; they do not prove a causal effect. Working through those distinctions taught me to make data context and evaluation part of the interface. A geographic holdout would be a useful next test of how well the model transfers between areas.