How do you handle late-arriving data in a pipeline?
Assesses fundamental understanding of Data Engineering conventions, runtime behavior, and memory/performance considerations.
Hiring managers look for precision, avoidance of ambiguous jargon, and ability to explain trade-offs under real production conditions.
Late data breaks the assumption that all events for a time window have arrived. Common strategies:
- Watermarks: define how long to wait for late events before closing a window. A ten-minute watermark holds windows open ten minutes past the latest event time. Events after that are dropped or routed separately.
- Allowed lateness: keep window state longer and update results when late events arrive, emitting revised aggregates downstream.
- Reprocessing: store raw immutable events in the lake, then rerun the affected partition to correct results. Idempotent writes and partitioning by event date make this safe.
- Reconciliation: combine a streaming near-real-time view with a batch correction that overwrites the same partitions.
The right choice balances latency, cost and correctness. Billing pipelines usually favor a batch correction layer, while dashboards may accept approximate streaming results.
Candidate Response Strategy & Interview Tips
- Start with a concise one-sentence summary: Deliver a direct, confident answer first before expanding into nuances.
- Demonstrate real-world trade-offs: Discuss where this approach excels and when you would avoid it in production systems.
- Discuss complexity & edge cases: Proactively explain time/space complexity or boundary conditions (null values, scale limits).
- Prepare for interviewer follow-ups: Technical hiring panels frequently probe deeper into concurrency, backward compatibility, or alternative libraries.