Text & NLPsc.

An analysis of 29,988 listing descriptions. Free text does its work where the structured fields can't see: places where the ad copy and the ad form don't line up, equipment that has no column at all, and triage of the listings the price model gets wrong.

Text says clean
163
but the form records damage
Data-gap queue
3,964
damage in text, empty counter
Equipment inferred
35.9%
Sunroof / panoramic roof
Triple anomaly
5
review candidate
Thesis — why “not price”

Adding the text signals to the price model does not increase the variance explained — age, mileage, engine and the damage counters already do that work, and the description mostly restates them in prose. Text pays off somewhere else: in what the structured fields never record. Modifications and conversions, a seller contradicting their own declaration, equipment with no column at all, and the listings worth asking “why is this priced like that?” — all of it lives only in the text.

00

Claims, extraction and anomaly

[01]

The text says clean, the form says otherwise

cross-source · ad copy ↔ ad form

The word “clean” in the ad copy is compared against the damage counters the seller filled in on the same listing. Those counters are the seller’s own declaration — so nothing here is concealed; the shop window and the form simply don’t line up. In most of these the seller means “tidily repaired, presents clean”. In a minority they claimed something specific that their own form contradicts. The three figures below separate the two.

Text clean · Form shows damage
163
median ₺1.40M
Text damage · Form shows clean
3,964
median ₺1.82M
Both say clean
4,784
median ₺2.30M
Both say damage
14,781
median ₺1.36M
No text claim
5,239

So how far from clean are those 163?

One bar couldn’t show this: it averaged away both the reassuring half (the median car has two painted panels and nothing replaced) and the half that isn’t reassuring (a specific claim, or a write-off record).

Median painted panels
2
mean 2.17
Median replaced panels
0
mean 0.42
No replaced panel
68.7%
paint only
Write-off record
5
of 163
How heavy the damage is
How many panels — painted vs replaced
What the ad actually claimed
// heavy = write-off/heavy-damage record OR 3+ replaced panels · moderate = 1–2 replaced or 3+ painted · light = ≤2 painted & 0 replaced. The green bar is a general word (“flawless”, “spotless”) — there, “tidily repaired, presents clean” is a fair reading. The amber bars are specific: the ad said “no paint” or “no changed parts” while the seller’s own form records one. The arms can overlap, so they sum above 163.
Priced like what it is

These 163 listings are -12.8% CHEAPER than the market — the damage is already in the price. Holding vehicle specs and actual damage fixed, what's left is +1.4%: within noise in the OLS ladder (p=0.186), while the LGBM robustness arm puts it at +1.4% with a CI barely off zero (0.1…2.9). A systematic con would be expected to earn a premium; this group doesn't. (An association, not causation.)

The same thing in the titles

2,856 listings have a “clean” title but structural damage. Raw gap -9.4%; controlling for damage + series the premium is +0.0% (≈0). Same result: a “clean” title earns no premium either.

Examples — text says clean, the form has a record

The “phrase in text” column holds the words the detector found; raw listing prose is not published. Paint/Replaced is the counter from the seller’s own form.

ModelPricePaint/ReplacedPhrase in text
318i Pure₺1.73M1 / 0flawlessunpaintedno damage record
520i M Sport₺4.15M1 / 0flawlessunpaintedno damage record
A3 Sportback 1.6 TDI Attraction₺1.26M2 / 0flawless
A6 Sedan 45 TFSI Quattro Design₺4.29M3 / 0flawless
320i ED Luxury Line Plus₺1.55M0 / 2flawlessunpainted
A4 Sedan 2.0 TDI₺0.95M5 / 0no changed partsunpainted
A6 Sedan 2.0 TDI₺1.85M0 / 1flawless
520i M Sport₺1.71M3 / 1original

The other direction — damage in the text, blank form

The reverse direction: in 3,964 listings the text mentions damage while the structural counter is zero. Here the text is the honest half — it’s the form the seller left blank. This is where text COMPLETES the structured data.

ModelPricePaint/ChangedPhrase in text
320d Standart₺1.20M0 / 0damage recordlocal paintchanged panel
520d Standart₺0.70M0 / 0damage record
M5₺3.50M0 / 0damage record
320i ED Luxury Line₺1.57M0 / 0damage record
520d Standart₺0.80M0 / 0damage record
A3 Sedan 1.6 TDI Ambition₺1.38M0 / 0damage recordchanged panellocal paint
318i Edition Luxury Line₺1.50M0 / 0changed panel
316i Modern Line₺1.16M0 / 0damage record

How coverage was raised — LangExtract · Gemini 3.1 Flash Lite

Listing texts were run once, offline, through LangExtract for structured extraction (extraction model Gemini 3.1 Flash Lite): 13,904 listings · 41,866 extractions (avg 3.0 per listing). Every extraction carries a character interval aligned to the text (match_exact) — the model only marks phrases that literally occur, no free generation.

ClassExtractionsAttributesExample
Damage19,800part · conditionsağ arka kapı boyalı” → parça: sağ arka kapı · durum: boyalı
Maintenance14,086part · conditionağır bakımlar yeni yapılmıştır” → parça: genel · durum: yeni yapıldı
Modification4,433part · conditionstage 2 yazılım” → parça: motor · durum: yazılımlı
Horsepower3,547power value270 beygir” → güç: 270
⚠ Its role is coverage ONLY

The LLM’s condition vocabulary (boyalı · tramer kayıtlı · lokal boyalı · değişmiş · hatasız / değişensiz) was distilled into the regex detectors and is used as ground truth in permanent regression tests. The LLM is NOT a price feature and does not run in production — the regex does; the LLM read the text once and we took its vocabulary. Because coverage is 13,867/29,988 = 46.2%, this is PARTIAL ground truth: the measured precision is an agreement rate, not a “regex error” rate.

// Circularity note: because the regex vocabulary was distilled from the LLM output, measuring the regex against those same LLM labels is partly circular (it inflates agreement). The coverage test should be read as a REGRESSION test (“did this change break coverage?”), not as independent proof of accuracy. That is why the LLM was deliberately not used to audit the examples above — those counters were read by hand.
// The naive match was 1,220; about 1,057 were dropped — maintenance work (“gearbox replaced”), local touch-up only, partial disclosures (“except for X”), and scoped phrasing like “flawless engine”. That leaves 163: MOST of the charitable readings were already applied. The counters are seller-declared, so this isn’t an accusation of concealment — it’s a way of telling the buyer to read the form too.
[02]

Information extraction → data completion

equipment absent from the structured schema

Equipment absent from structured fields (sunroof, heated seats, xenon…) is inferred from free text (regex + normalize) — for these, text is the ONLY source. Useful for search/filtering, carries no price claim.

Equipment mention rate (%)
// Regex + normalize, negation-safe (“no sunroof” doesn’t count). Meaning-similarity search (multilingual-e5-large-instruct embeddings) is built as a separate layer but isn’t in this report: live free-text search needs a server side, and it is deferred.
[03]

Anomaly queue

triple concordance · triage

Review candidates where two independent signals cross: a POSITIVE price residual (pricier than expected) ∧ text conversion/mod ∧ text-hp ≫ field-hp (triple concordance). An anomaly = a review candidate, NOT evidence.

resid ∧ text · 48resid ∧ hp · 14triple · 5
ModelAgeField HPText HPPriceResid
750i Long17413600₺5.30M+52.7%
M320343750₺2.90M+48.8%
640i15320650₺5.60M+41.1%
M314420630₺3.98M+15.8%
520d M Sport19188350₺1.13M+15.2%

Field-contradiction queue — examples (text ↔ structured)

ModelFieldTextStructured
A3 Sportback 1.6 TDI Sport LinefuelDieselPetrol
520d M SporttransmissionManualAutomatic
420i Gran Coupe Edition M Sportdrivetrain4WDRWD
A3 Sedan 35 TFSI DynamicbodySedanHatchback
420d M Sportengine size3.0L1995
116i M Sportyear20172012
335i Standarthp600306
116i M Sportmodelm6116i M Sport
// Triple = positive residual-outlier ∧ conversion verb ∧ hp-contradiction (cross-source). The queue holds 48 records; counts are computed every run (nothing hardcoded). Triage for human review — not an automatic decision or a population statistic.
[04]

Controlled coefficient table

hedonic log-OLS · HC3 + bootstrap

Each text signal's *controlled* price association (a single hedonic log-OLS, rich control set, simultaneous). % = exp(β)−1, 95% CI (HC3) + 1,000× bootstrap 95%. Grey = not significant (p≥0.05). Controlled ≠ raw: the raw premium misleads. n=29,562, R² 0.931.

Controlled % effect (± 95% CI)
SignalRaw %Ctrl %HC3 95%boot 95%pn
Premium audio+83.5%+5.5%4.8…6.14.8…6.1<0.0013,055
Mod suspension-9.2%+5.1%3.7…6.53.7…6.5<0.001418
Mod exhaust-13.0%+3.5%1.4…5.61.5…5.5<0.001285
Service history+18.4%+3.1%1.3…4.91.4…5.1<0.001226
Franchised service+53.2%+2.0%1.4…2.51.4…2.5<0.0013,248
Mod engine/tune-21.0%+1.9% ·-0.5…4.4-0.4…4.40.118211
Clean claim vs. declared damage-12.9%+1.9% ·-0.2…3.9-0.1…3.90.078161
Navigation+51.3%+1.5%1.0…1.91.0…2.0<0.0016,558
Heated seats+44.4%+1.5%1.0…2.01.0…2.0<0.0016,866
Panoramic roof+38.0%+0.9%0.5…1.30.5…1.3<0.00110,749
Driver assist+91.2%+0.8%0.3…1.30.2…1.30.0035,570
Mod wheels/body-10.8%+0.4% ·-1.2…2.1-1.3…2.10.615226
Warranty+13.5%+0.1% ·-0.4…0.6-0.4…0.60.7672,878
Leather seats+22.3%-0.4% ·-0.9…0.1-0.9…0.10.1054,752
// exp(β)−1, HC3 robust SE, all signals estimated simultaneously in one control set. boot≈HC3 → stable. When raw ≫ controlled the signal was a segment/model proxy. An ASSOCIATION — not causation/deployment; predictive power is a separate question (ΔR²~0). “·” = p≥0.05 (not significant).
[+]

Listing-language overview (NMF)

descriptive · not a price archetype

3 NMF themes — a seller REGISTER map, not a price archetype. The themes re-derive seller type at ~33% (Cramér’s V 0.339); they carry no new price information.

Topic 134.7%
elektrikli (electric)sensörü (sensor)sistemi (system)koltuklar (seats)ledkoltuk (seat)
Dealer · 83%
Topic 29.6%
oldukça (quite)orijinaldir (is original)herhangi (any)aracımda (on my car)uzun (long)satıyorum (I'm selling)
Private seller · 98%
Topic 355.1%
kredi (finance)parça (part)boya (paint)vardır (there is)yeni (new)bakımları (servicing)
Dealer · 68%
01

The surface of the text — who writes how, where the model errs

[05]

Seller register / communication

descriptive

Writing metrics per seller type: CAPS ratio, average emoji, description length, phone-drop rate. Descriptive — a register difference, not a quality claim.

SellerListingsCAPS %Emoji avgExcl avgDesc lenPhone %
Dealer20,09490.8%0.220.5078422.9%
Private seller9,72730.5%0.210.105281.4%
Franchised dealer16791.2%0.000.1091138.9%
[06]

Ad-title mining

hook distribution · title claims

Which “hook” appears in ad titles and how often + counts for “clean” titles and mismatched years. The title’s “clean” claim also carries a ~0 controlled premium.

spec (year / km / engine) · 79%equipment · 41%clean claim · 31%condition / praise · 8%urgency / promo · 1%
Avg words
8.5
Title says clean
2,856
Year mismatch
530
“Clean” ctrl
+2.1%
hatasız (flawless) 5,610sport 4,888boyasız (unpainted) 4,056tdı 3,821line 2,845dan (from) 2,577tfsı 2,016motors 1,888quattro 1,793sportback 1,776değişensiz (no changed parts) 1,738bakımlı (well-maintained) 1,669
[07]

Damage status (3 groups)

descriptive · the bare median misleads

Listing count + median price for three groups: damage-claim / clean-claim / no-mention. DESCRIPTIVE: the bare median comparison misleads (clean-claim cars are already younger/lower-km).

Damage-claim
₺1.43M
18,745 listings
Clean-claim
₺2.10M
6,004 listings
No claim
₺1.50M
5,239 listings
// Once vehicle specs + objective damage counters are controlled, the clean-claim premium is only +2.7% — “significant” due to large-n but small in practice. The raw gap (₺0.67M) misleads.
[08]

Residual signals (triage)

top-5% under-predict · robust-filtered

Where the price model errs most: text signals concentrated in the top-5% UNDER-predicted (pricier than expected). Not a mispricing CLAIM — triage of where the model errs.

Under-predict signals (concentration in the top-5% · lift)
// Lift = the signal’s density in the top-5% under-predicted / its overall rate (robust: n≥30, lift≥1.2, p<0.05). n_top5% = 1,500. Non-robust signals (e.g. “high horsepower”, lift≈1.0) are not published. Modified/conversion/M-RS models are pricier than the model expects — a review candidate, not evidence.
[09]

Cross-source field contradictions

text ↔ structured field · counts

How many listings have a value in the text that contradicts the structured field, by field. Most are seller errors or a swap/conversion signal. A review queue, not evidence.

Field-contradiction count (text ↔ structured, by field)
// Mileage was dropped from this list: a km figure in the text can’t be separated by pattern-matching into odometer vs service/purchase/swap/policy mileage, which made it an unreliable contradiction signal.