Reviewed by Aditya Kumar · Last reviewed 2026-03-25
**Implementation**: Use `when`/`otherwise` with modulo—each row gets non-null in one column. ```python from pyspark.sql.functions import when, col, lit df = df.withColumn("even", when(col("num") % 2 == 0, col("num")).otherwise(lit(None))) .withColumn("odd", when(col("num")...
This medium-level Spark/Big Data question appears frequently in data engineering interviews at companies like KPMG. While less common, it tests deeper understanding that distinguishes strong candidates. Mastering the underlying concepts (partition, python, spark) will help you answer variations of this question confidently.
Break this problem into components. Identify the core trade-offs involved, then walk the interviewer through your reasoning step by step. Demonstrate awareness of edge cases and production considerations - this is what separates good answers from great ones. The expert answer includes a code example that demonstrates the implementation pattern.
Implementation: Use when/otherwise with modulo—each row gets non-null in one column.
from pyspark.sql.functions import when, col, lit
df = df.withColumn("even", when(col("num") % 2 == 0, col("num")).otherwise(lit(None)))
.withColumn("odd", when(col("num") % 2 != 0, col("num")).otherwise(lit(None)))
Why This Approach: Single pass; no shuffle; Catalyst optimizes the expressions. Alternative: filter + unionByName yields separate DFs but doubles scans.
Scalability Trade-offs: Wide table (2x columns) vs. narrow with union—choose based on downstream. For TB-scale, avoid UDFs; built-in modulo is codegen'd.
Cost Implications: No shuffle cost; minimal CPU. Partition by even/odd only if downstream aggregations benefit from locality.
Pro-Move: 'We use this pattern for routing odd/even keys to different sinks in a fan-out pipeline.' Red Flag: Using Python UDF for modulo—built-in is 10x faster.
Some links below are affiliate links. If you buy through them we may earn a small commission at no extra cost to you — it helps keep DataEngPrep free.
According to DataEngPrep.tech, this is one of the most frequently asked Spark/Big Data interview questions, reported at 1 company. DataEngPrep.tech maintains an editor-reviewed database of 1,863 data engineering interview questions across 7 categories.