Reviewed by Aditya Kumar · Last reviewed 2026-03-25
**Why it matters**: At scale, design choices directly impact reliability, latency, and cost. Wrong decisions compound across jobs and teams. Columnar formats (Parquet, ORC): store by column not row. Benefits: (1) Compression—similar data compresses well; (2) Predicate...
This hard-level Spark/Big Data question appears frequently in data engineering interviews at companies like Disney+ Hotstar. While less common, it tests deeper understanding that distinguishes strong candidates. Mastering the underlying concepts (optimization, partition) will help you answer variations of this question confidently.
This is a senior-level question that tests architectural thinking. Lead with the high-level design, then drill into specifics. Discuss trade-offs explicitly - there is rarely one correct answer. Show awareness of scale, fault tolerance, and operational complexity.
Why it matters: At scale, design choices directly impact reliability, latency, and cost. Wrong decisions compound across jobs and teams.
Columnar formats (Parquet, ORC): store by column not row. Benefits: (1) Compression—similar data compresses well; (2) Predicate pushdown—read only needed columns; (3) Better for analytics—aggregations, scans; (4) Schema embedded. Parquet is widely supported; ORC has better compression in Hive. Example: SELECT date, sum(amount) reads only those columns. Best practices: use for analytics workloads; choose appropriate block size; enable predicate pushdown; leverage nested types when useful.
Scalability trade-offs: Partition/parallelism limits; single points of failure; horizontal vs vertical scaling. Cost implications: Sizing, spot vs reserved, optimization ROI.
Red Flag: Row format for analytics. Pro-Move: 'Parquet/ORC; predicate pushdown; block size.'
Some links below are affiliate links. If you buy through them we may earn a small commission at no extra cost to you — it helps keep DataEngPrep free.
According to DataEngPrep.tech, this is one of the most frequently asked Spark/Big Data interview questions, reported at 1 company. DataEngPrep.tech maintains an editor-reviewed database of 1,863 data engineering interview questions across 7 categories.