Reviewed by Aditya Kumar · Last reviewed 2026-03-24
Broadcast join sends the small table to every executor; join happens locally without shuffle. **Trigger**: broadcast(df) hint or spark.sql.autoBroadcastJoinThreshold (default 10MB). **Why use it**: Avoids shuffling the large table—the dominant cost in sort-merge join. **When**:...
This medium-level Spark/Big Data question appears frequently in data engineering interviews at companies like Delivery Hero, Fragma Data Systems. While less common, it tests deeper understanding that distinguishes strong candidates. Mastering the underlying concepts (join, spark, sql) will help you answer variations of this question confidently.
Break this problem into components. Identify the core trade-offs involved, then walk the interviewer through your reasoning step by step. Demonstrate awareness of edge cases and production considerations - this is what separates good answers from great ones.
Broadcast join sends the small table to every executor; join happens locally without shuffle. Trigger: broadcast(df) hint or spark.sql.autoBroadcastJoinThreshold (default 10MB). Why use it: Avoids shuffling the large table—the dominant cost in sort-merge join. When: One table fits in executor memory (typically < 100MB after serialization). Scalability trade-off: Driver fetches and distributes; oversized broadcast causes driver OOM. Cost implication: Broadcast is free (no shuffle) for the large table; sort-merge shuffle scales with data size. Architectural logic: Fact–dimension joins (fact large, dimension small) are prime candidates. Anti-pattern: Broadcasting a 500MB table—exceeds memory, spills, or fails. Best practice: Set autoBroadcastJoinThreshold based on executor memory; use broadcast hint when optimizer chooses wrong side; monitor driver memory.
Red Flag: Broadcasting a table without knowing its size. Pro-Move: 'We broadcast our 5MB dimension table; we had to increase autoBroadcastJoinThreshold to 20MB when we added columns; we monitor driver memory in production.'
Some links below are affiliate links. If you buy through them we may earn a small commission at no extra cost to you — it helps keep DataEngPrep free.
According to DataEngPrep.tech, this is one of the most frequently asked Spark/Big Data interview questions, reported at 2 companies. DataEngPrep.tech maintains an editor-reviewed database of 1,863 data engineering interview questions across 7 categories.