Used BMW & Audi prices, predicted within 6.5%.
A LightGBM model trained on 29,988 real Turkish listings, served behind a FastAPI backend on Railway. Every figure below is 5-fold out-of-fold — not a lucky holdout — and read from the same data the analysis reports.
The model's mean error is 43% lower than the median price of the same model and year. Beyond model and year, the gap is closed mostly by mileage and damage.
How it's served
Two tracks. Offline, listings become a DuckDB file and a model bundle in object storage. Online, FastAPI loads that bundle into memory at boot — the raw 30K rows never leave the backend.
Bundle: {model, tfidf, cat_maps, feat_cols} — 6 categoricals + 8 numerics + free-text model & series → TF-IDF → TruncatedSVD (170 dims). Train and serve can’t drift apart: one pickle carries its own preprocessing.
Scores
Three variants share one leak-free 5-fold split; the median baseline runs on the same folds.
| Variant | MAPE | R² | MAE ₺ | MedAE ₺ | RMSE ₺ |
|---|---|---|---|---|---|
| LightGBM · TF-IDF+SVD · the model | 6.49% | 0.9745 | 109,776 | 75,320 | 176,225 |
| CatBoost · TF-IDF+SVD ★ | 6.44% | 0.9745 | 110,085 | 74,927 | 176,130 |
| CatBoost · native | 6.58% | 0.9739 | 112,925 | 77,660 | 178,312 |
| Comparable median · laddered | 11.20% | 0.9235 | 191,224 | 130,000 | 305,050 |
★ wins the MAPE-only rule; the two TF-IDF+SVD variants are practically tied, and the site and the service use LightGBM. Metrics are out-of-fold (leak-free); final models train on all data. Serving returns a point price plus a fixed ±6.6% band.
What the data actually says
Findings that change how you'd read a price, not just decorate the model.
Data & method
Real TR-registered used-car detail pages, four snapshots, one listing per ad. A third of what the scraper returns is the same ad seen again.
The 15,171 rows between are the same ad re-scraped across four snapshots — scrape residue, not data. Dedup runs on ad_id before the CV split, so the same ad (ad_id) can't straddle folds and inflate the score.