What is data leakage in machine learning?
Assesses fundamental understanding of Machine Learning conventions, runtime behavior, and memory/performance considerations.
Hiring managers look for precision, avoidance of ambiguous jargon, and ability to explain trade-offs under real production conditions.
Data leakage is when information unavailable at prediction time seeps into training, producing optimistic validation scores that collapse in production.
Common sources:
- Target leakage: a feature that is a proxy for the label, such as account_closed_date when predicting churn.
- Train-test contamination: fitting scalers, imputers, target encoders or SMOTE on the full dataset before splitting.
- Temporal leakage: using future data to predict the past, as with random splitting of time series.
- Group leakage: the same entity appears in both train and test.
Prevention:
- Split first, then fit all preprocessing on training folds only, ideally inside a pipeline.
- Use time-based splits for temporal problems.
- Audit features for whether they are known at prediction time.
- Use grouped splits when rows share an entity.
Pipeline([("scaler", StandardScaler()), ("clf", LogisticRegression())])
Leakage often shows as suspiciously high accuracy, so treat such results with suspicion.
Candidate Response Strategy & Interview Tips
- Start with a concise one-sentence summary: Deliver a direct, confident answer first before expanding into nuances.
- Demonstrate real-world trade-offs: Discuss where this approach excels and when you would avoid it in production systems.
- Discuss complexity & edge cases: Proactively explain time/space complexity or boundary conditions (null values, scale limits).
- Prepare for interviewer follow-ups: Technical hiring panels frequently probe deeper into concurrency, backward compatibility, or alternative libraries.