How do you handle missing data in an analysis?
Assesses fundamental understanding of Data Analysis & BI conventions, runtime behavior, and memory/performance considerations.
Hiring managers look for precision, avoidance of ambiguous jargon, and ability to explain trade-offs under real production conditions.
First understand why data is missing, because the mechanism determines the right approach.
- MCAR: missing completely at random, safe to drop rows with a small loss of power.
- MAR: missing at random given observed variables, so imputation using other features is reasonable.
- MNAR: missing not at random, where the missingness itself carries information, such as high earners refusing to report income. Simple imputation then biases results.
Options: drop rows or columns, mean or median imputation, model-based imputation such as MICE or k-NN, or flag missingness with an indicator variable and let the model use it. For time series, forward-fill carefully and never leak future values backward.
df["income"] = df["income"].fillna(df["income"].median())
df["income_missing"] = df["income"].isna().astype(int)
Always report how much data was missing and test sensitivity across methods.
Candidate Response Strategy & Interview Tips
- Start with a concise one-sentence summary: Deliver a direct, confident answer first before expanding into nuances.
- Demonstrate real-world trade-offs: Discuss where this approach excels and when you would avoid it in production systems.
- Discuss complexity & edge cases: Proactively explain time/space complexity or boundary conditions (null values, scale limits).
- Prepare for interviewer follow-ups: Technical hiring panels frequently probe deeper into concurrency, backward compatibility, or alternative libraries.