What’s it worth?
What’s it worth?
A listing asks ₺1.54M. Too much? The model+year median — same model, same year, look at the median — misses by ₺191K. The model misses by ₺110K: 42% better.
Looking at the model+year median misses by ₺191K; the model misses by ₺110K. The remaining ₺81K is closed by what lies beyond model+year: km, damage, engine. The model isn’t a lookup table.
This baseline prices every car (falls back to model, then global, when no comp exists): 97.48% at model+year (₺178K), 2.01% model-fallback (₺561K), 0.51% global (₺1,075K). Without a comp the baseline collapses; the model stays at ₺110K everywhere. Full breakdown → 03.
One consequence: age and km are separate but correlated axes, and a low-km old car is where they diverge — it paid the age penalty but not the km one, so it stays systematically underpriced. Money on the table the model sees, beyond just pricing → 02 · how km and age bite.
The reader cuts both ways. Over-estimating costs the buyer, under-estimating costs the seller; one model, two readers. So the output is a 90% conformal range, not one number. The range stops holding on cheap cars (Q1 under-covers the target); the rationale is in 03.
- Price rare and edge cars by hand; the model scatters there and the range stops holding.
- Widen the range on cheap cars; Q1 under-covers the target, so a point estimate misleads.
- Retrain monthly; the market moved +5.3% in 5 months and the model is time-blind.
- ·The sale price. The model predicts the asking price, not what the car sells for; the haggling margin sits inside the target.
- ·Anything about brand. The BMW–Audi gap is too small to act on → 02.
- ·Causation. These are controlled associations; it won’t say “repaint it and the price drops”.
- ·The real value of a damaged car. “Damaged” spans a scratch to a rebuilt wreck, yet they all price into one low cluster the model can’t tell apart — so its single number for a specific one isn’t reliable → 03.
- ·Trim and modifications. A loaded or modified car looks the same to the model as a base one of the same specs — it can’t see the extras; part of the spread among identical-spec cars → 03.
Is this data any good?
Is this data any good?
On scraped marketplace data, duplication isn’t housekeeping, it’s a leakage risk: the precondition for the number in 00. What was broken on arrival, and what’s left:
Scale
45,159 snapshots → 29,988 listings. The 15,171 rows between are the same ad re-scraped: scrape residue, not data. Source: real TR-registered used-car detail pages (BMW + Audi), 4 snapshots (18 Jan – 27 Jun 2026), latest snapshot per ad_id. Of 117 raw columns, 15 are unusable → 05.
Median asking price ₺1.54M, ranging ₺0.84M–₺3.42M (P10–P90). Price is right-skewed; why the model trains on log-price → 05 · Target and preprocessing.
Leakage
Dedup runs on ad_id, before the CV split. Evaluation is 5-fold out-of-fold: every listing is predicted exactly once, by a model that never saw it.
The real risk is the one ad_id can’t see: the same car entering as two separate listings. That is a separate check — content-based duplication: 137 rows (0.46%) with every distinguishing field identical, 209 (0.70%) under the loosest definition → 05 · Content-based duplication. So the real repeats that could straddle folds sit under 1%.
Missingness isn’t random
15 columns drop together; co-missing correlation 1.00. This isn’t “missing data”, it’s listings where catalog matching collapsed: standard models match, niche variants don’t, and all their specs go blank at once. One block, one cause, one decision → 05 · What I dropped, and why.
Redundancy
Series is a coarsened view of model — a derived column, not independent information. Theil’s U (directional dependence): U(series | model) = 1.00 (model fully determines series), but U(model | series) = 0.39 (series leaves model ambiguous). The asymmetry gives direction. Cramér’s V is symmetric and couldn’t show this; the direction is the finding. The full 8×8 matrix → 05 · Leakage and redundancy checks.
Model fully determines series; series leaves model ambiguous. Series is a coarsened view of model — it carries no separate information.
The raw feed’s “G” segment isn’t real: nearly all its rows are MPV bodies (Active / Gran Tourer). This is neither a gap nor a redundancy — it’s a corrupt source. So raw gb_segment was dropped and segment re-derived from series; the MPV signal kept in body_type. Counts and heatmap → 05 · Leakage and redundancy checks.
What the rest does
What’s left is the structural fields: age, km, engine, body. The listing text isn’t here → text analysis. How well they price → 03 · How much to trust the number.
How this market builds a price
Three kinds of car in this market
Unsupervised clustering splits the market into 3 groups; each name comes from the axis that actually separates it (age+km, then engine or damage) — no “premium/economy” value words. This is where you see which cars the model is pricing.
Which features carry the price
The hedonic regression gives each driver's *controlled* effect on price (all else equal) — R² 0.931, n 29,554. Coefficients carry bootstrap confidence intervals; every 95% CI EXCLUDES zero → each driver is reliably significant.
cc–HP correlation · by fuel
VIF · multicollinearity check
How km and age actually bite
The hedonic model says km and age dominate; these two curves show the shape of that dominance. The drop is steep in the early years then flattens. Age costs 7.1% a year, km 14.6% per 100k km — two separate but correlated axes. A low-km old car is where the two diverge: it has paid the age penalty but not the km one → systematically under-priced, that’s where the arbitrage is.
Should brand affect price? No.
Brand says nothing separate because it already lives inside `model` (and `series`): both fully determine it. I measured this directly, not with a distributional test — adding brand on top of series+model doesn’t move the error.
Numerics (km · age · damage · hp) are held constant across all three; only the identity column changes — which is why series+model already reaches the full model’s ₺110K. The last two bars are identical: adding brand to series+model moves MAE by ₺1 and MAPE by 0.00 pts. The gain comes from series+model, not brand.
Both model and series fully determine brand (U=1.00) → brand sits inside both, no separate information. The ablation says the same.
▸ Brand · distributional test + ablation table
How much to trust the number
Is the model worth it?
3 model variants compared with 5-fold OOF (leak-free); final models trained on all data. Winner LightGBM · TF-IDF+SVD — MAPE 6.49%, 42% better than the model+year median (with fallback).
▸ Model+year median — per-tier breakdown (fallback penalty)
Ladder: (model, year) median → (model) median across all years → global median. When a (model, year) cell is absent in train, the baseline drops a tier; each drop enlarges the error sharply — a car with no comp has a weak baseline to begin with.
Cars with identical specs (same model · year · km · hp · body) still list ₺77K apart — 5,767 rows, 2,577 groups. That is a floor: sellers price the same car differently (options and mods, urgency, haggling) and no model can go below it. The model sits at ₺110K, 1.42× the floor, so the entire remaining headroom is ₺33K. A hyperparameter search typically claims ~₺5K of that, which is why none was run.
Target: log1p(price) · 25 features
Prediction calibration & residuals
OOF (leak-free) predictions vs actual — R² 0.975. Residual% centers on zero (mean -0.45%, std 9.20%) → no systematic bias.
Where the model is weak
The model’s weakness points to the same place from three angles: error by price quartile (harder on cheap cars), per-model sample size vs error (rare models scatter) and the best/worst predictions (edge cases). All of it: rare · edge · cheap.
How wide a range should you quote?
The weakness above is why you quote a range at all. Actual coverage of the conformal intervals by price quartile; target 90%. The upper quartiles hit it; Q1 (cheap cars) falls below — which is where 00’s “widen the range on cheap cars” comes from.
How long you can trust it
When does this model go stale?
Two lines of evidence give the same call. Distribution drift: the period curves nearly overlap, PSI stays <0.05 (shape stable), KS grows slightly (statistical — the big-n trap). Temporal backtest: train on an earlier period and test only on the next period’s NEW listings (leak-free); the cumulative strategy is more stable. Verdict: the market LEVEL shifted +5.3% but the SHAPE held → monthly retraining suffices.
How it was built
The decisions I took to work with dirty scraped data.
Feature selection
From 117 raw columns down to 25. Nothing dropped arbitrarily: constants, redundant kb/gb twins, leakage/identity, block-missing, collinearity, granular damage (aggregated) and an empirical audit.
Why 15 columns were dropped
Missingness isn’t random — it’s block-shaped: the high-missing columns below (spec/catalog + insurance, ~28–84%) sit empty together in the same listings. Because “being missing” is systematic, reliable imputation is impossible and leakage is a risk → these columns were dropped. (The 16 kept features are <2% missing — not shown here.)
This is MISSING correlation (presence/absence co-moves) — NOT value correlation (~0.59). Spec columns come from catalog matching: standard models match, special variants don’t → all specs go blank together. Each spec carries different information, but their presence/absence is tied to one source.
Content-based duplication
A check beyond ad_id-dedup: listings whose ad_id differs but whose every distinguishing feature (price, mileage, age, model, damage, engine) is identical. The rate is low — confirming clean collection; some are genuine re-posts, some coincidental overlaps among common models.
Categorical dependence (Cramér’s V + Theil’s U)
Cramér’s V gives association strength (symmetric); Theil’s U its direction (asymmetric) — both are for CATEGORICAL features: `model` (text), brand, series, segment, body, drivetrain, transmission, fuel. `model` almost fully determines the rest (U≈1) but not vice-versa → `model` is excluded from the hedonic OLS (too high-cardinality; it’s used in the LightGBM via TF-IDF text instead). Note: damage counters and engine specs (hp/cc) are NUMERIC, so they’re not here.
The raw feed had a “G” segment, but it isn’t real: nearly all G rows are MPV bodies (Active/Gran Tourer). Segment is derived from the series, the MPV signal kept in body_type — the asymmetry above, made visible.
Numeric correlation (Pearson + Spearman)
Correlation among numeric features — the numeric counterpart to the categorical dependence above. Green = positive, red = negative. Pearson measures linear, Spearman monotonic association. |r|>0.5 pairs are flagged for collinearity (also checked via VIF, all <3).
KMeans + PCA — method
k=3 was chosen by silhouette and corroborated with the PCA scatter: 3 groups separate along age × mileage × power, the damaged cluster on PC3. What the clusters are (name, median, damage) → 02. The damage signal appearing independently across the hedonic model, PCA and KMeans is a robustness check.
Why log price
Raw price is right-skewed (skew 1.62); a log transform pulls it toward symmetry (0.28). I trained on log1p(price): under squared loss the extremes were swallowing the whole error budget. A modelling decision, not a market finding — which is why it lives here.