An analysis of 29,988 listing descriptions. Free text does its work where the structured fields can't see: places where the ad copy and the ad form don't line up, equipment that has no column at all, and triage of the listings the price model gets wrong.
Adding the text signals to the price model does not increase the variance explained — age, mileage, engine and the damage counters already do that work, and the description mostly restates them in prose. Text pays off somewhere else: in what the structured fields never record. Modifications and conversions, a seller contradicting their own declaration, equipment with no column at all, and the listings worth asking “why is this priced like that?” — all of it lives only in the text.
Claims, extraction and anomaly
The text says clean, the form says otherwise
The word “clean” in the ad copy is compared against the damage counters the seller filled in on the same listing. Those counters are the seller’s own declaration — so nothing here is concealed; the shop window and the form simply don’t line up. In most of these the seller means “tidily repaired, presents clean”. In a minority they claimed something specific that their own form contradicts. The three figures below separate the two.
So how far from clean are those 163?
One bar couldn’t show this: it averaged away both the reassuring half (the median car has two painted panels and nothing replaced) and the half that isn’t reassuring (a specific claim, or a write-off record).
These 163 listings are -12.8% CHEAPER than the market — the damage is already in the price. Holding vehicle specs and actual damage fixed, what's left is +1.4%: within noise in the OLS ladder (p=0.186), while the LGBM robustness arm puts it at +1.4% with a CI barely off zero (0.1…2.9). A systematic con would be expected to earn a premium; this group doesn't. (An association, not causation.)
2,856 listings have a “clean” title but structural damage. Raw gap -9.4%; controlling for damage + series the premium is +0.0% (≈0). Same result: a “clean” title earns no premium either.
Examples — text says clean, the form has a record
The “phrase in text” column holds the words the detector found; raw listing prose is not published. Paint/Replaced is the counter from the seller’s own form.
The other direction — damage in the text, blank form
The reverse direction: in 3,964 listings the text mentions damage while the structural counter is zero. Here the text is the honest half — it’s the form the seller left blank. This is where text COMPLETES the structured data.
How coverage was raised — LangExtract · Gemini 3.1 Flash Lite
Listing texts were run once, offline, through LangExtract for structured extraction (extraction model Gemini 3.1 Flash Lite): 13,904 listings · 41,866 extractions (avg 3.0 per listing). Every extraction carries a character interval aligned to the text (match_exact) — the model only marks phrases that literally occur, no free generation.
The LLM’s condition vocabulary (boyalı · tramer kayıtlı · lokal boyalı · değişmiş · hatasız / değişensiz) was distilled into the regex detectors and is used as ground truth in permanent regression tests. The LLM is NOT a price feature and does not run in production — the regex does; the LLM read the text once and we took its vocabulary. Because coverage is 13,867/29,988 = 46.2%, this is PARTIAL ground truth: the measured precision is an agreement rate, not a “regex error” rate.
Information extraction → data completion
Equipment absent from structured fields (sunroof, heated seats, xenon…) is inferred from free text (regex + normalize) — for these, text is the ONLY source. Useful for search/filtering, carries no price claim.
Anomaly queue
Review candidates where two independent signals cross: a POSITIVE price residual (pricier than expected) ∧ text conversion/mod ∧ text-hp ≫ field-hp (triple concordance). An anomaly = a review candidate, NOT evidence.
Field-contradiction queue — examples (text ↔ structured)
Controlled coefficient table
Each text signal's *controlled* price association (a single hedonic log-OLS, rich control set, simultaneous). % = exp(β)−1, 95% CI (HC3) + 1,000× bootstrap 95%. Grey = not significant (p≥0.05). Controlled ≠ raw: the raw premium misleads. n=29,562, R² 0.931.
Listing-language overview (NMF)
3 NMF themes — a seller REGISTER map, not a price archetype. The themes re-derive seller type at ~33% (Cramér’s V 0.339); they carry no new price information.
The surface of the text — who writes how, where the model errs
Seller register / communication
Writing metrics per seller type: CAPS ratio, average emoji, description length, phone-drop rate. Descriptive — a register difference, not a quality claim.
Ad-title mining
Which “hook” appears in ad titles and how often + counts for “clean” titles and mismatched years. The title’s “clean” claim also carries a ~0 controlled premium.
Damage status (3 groups)
Listing count + median price for three groups: damage-claim / clean-claim / no-mention. DESCRIPTIVE: the bare median comparison misleads (clean-claim cars are already younger/lower-km).
Residual signals (triage)
Where the price model errs most: text signals concentrated in the top-5% UNDER-predicted (pricier than expected). Not a mispricing CLAIM — triage of where the model errs.
Cross-source field contradictions
How many listings have a value in the text that contradicts the structured field, by field. Most are seller errors or a swap/conversion signal. A review queue, not evidence.