What is linear regression and what are its assumptions?
Assesses fundamental understanding of Data Science & Statistics conventions, runtime behavior, and memory/performance considerations.
Hiring managers look for precision, avoidance of ambiguous jargon, and ability to explain trade-offs under real production conditions.
Linear regression models a continuous target as a linear combination of predictors plus error: y = b0 + b1*x1 + ... + e. Coefficients are estimated by ordinary least squares, which minimizes the sum of squared residuals, or by maximum likelihood.
from sklearn.linear_model import LinearRegression
model = LinearRegression().fit(X, y)
Assumptions:
- Linearity between predictors and the mean of the target.
- Independence of errors, violated by time series or clustered data.
- Homoscedasticity: constant error variance.
- Normally distributed errors, mainly for inference and intervals.
- Little multicollinearity; check variance inflation factors.
- No influential outliers distorting the fit.
Violations affect inference more than prediction. Diagnose with residual plots, QQ plots and influence measures. Remedies include transformations, robust standard errors and adding interaction or polynomial terms.
Candidate Response Strategy & Interview Tips
- Start with a concise one-sentence summary: Deliver a direct, confident answer first before expanding into nuances.
- Demonstrate real-world trade-offs: Discuss where this approach excels and when you would avoid it in production systems.
- Discuss complexity & edge cases: Proactively explain time/space complexity or boundary conditions (null values, scale limits).
- Prepare for interviewer follow-ups: Technical hiring panels frequently probe deeper into concurrency, backward compatibility, or alternative libraries.