Reviewed by Aditya Kumar · Last reviewed 2026-03-24
Sliding window: window(timeColumn, windowDuration, slideDuration) where slideDuration < windowDuration creates overlapping windows. Example: df.withWatermark("event_time", "10 minutes").groupBy(window(col("event_time"), "1 hour", "10 minutes"), col("user_id")).count(). **Why...
This hard-level Spark/Big Data question appears frequently in data engineering interviews at companies like Fragma Data Systems, Swiggy. While less common, it tests deeper understanding that distinguishes strong candidates. Mastering the underlying concepts (spark, window) will help you answer variations of this question confidently.
This is a senior-level question that tests architectural thinking. Lead with the high-level design, then drill into specifics. Discuss trade-offs explicitly - there is rarely one correct answer. Show awareness of scale, fault tolerance, and operational complexity.
Sliding window: window(timeColumn, windowDuration, slideDuration) where slideDuration < windowDuration creates overlapping windows. Example: df.withWatermark("event_time", "10 minutes").groupBy(window(col("event_time"), "1 hour", "10 minutes"), col("user_id")).count(). Why watermark: Late data would grow state unbounded; watermark drops events older than (max_event_time - delay) and allows state cleanup. Scalability: State grows with (unique keys × windows); for high cardinality, consider approximate aggregations or truncate state. Output modes: Append only emits final results when watermark passes; Update emits partial results; Complete emits full state (use sparingly). Cost implication: Sliding windows with small slides (e.g., 1min slide, 1hr window) create many overlapping windows—state and compute scale with 1/slide. Architectural choice: Tumbling (slide=window) is cheapest; sliding trades cost for smoother curves. Best practice: Align watermark delay with late-arrival SLA; monitor state store size.
Red Flag: Using sliding windows without a watermark—state grows indefinitely. Pro-Move: 'We use 10-min watermark based on our 99th percentile event delay; we monitor late-data metrics to tune the watermark and avoid premature dropping.'
Some links below are affiliate links. If you buy through them we may earn a small commission at no extra cost to you — it helps keep DataEngPrep free.
According to DataEngPrep.tech, this is one of the most frequently asked Spark/Big Data interview questions, reported at 2 companies. DataEngPrep.tech maintains an editor-reviewed database of 1,863 data engineering interview questions across 7 categories.