Interview questions
Preparing for a data engineering interview at PWC? This page contains 15 real interview questions sourced from verified PWC interview experiences. Questions are sorted by frequency β the ones asked most often appear first.
PWC data engineering interviews typically focus on Spark/Big Data, and System Design/Architecture. The interview bar skews toward harder problems (11 hard vs. 1 easy), suggesting emphasis on depth and system-level thinking.
Use the difficulty filters above to focus your preparation. For each question, attempt your own answer first, then compare with our expert solution. You can also practice these questions in our AI Mock Interview Coach for real-time feedback.
Design a cost-aware resource strategy for a Databricks workload with spiky and batch jobs. Explain Dynamic Resource Allocation, when to disable it, and how min/max executors and spot instances affect cost and SLAs.
Explain how Adaptive Query Execution changes the economics of Spark tuning. What problems does it solve at runtime, and when might you still need manual intervention (e.g., salting, broadcast hints)?
Explain Delta Time Travel and the purpose of the vacuum command.
Explain the architecture of Spark, including the roles of driver, executors, DAGs, and SparkContext.
How do you handle bad data in Databricks?
How do you resolve merge conflicts in Databricks notebooks?
How do you use Spark UI to debug stages, tasks, and performance issues?
How does Optimize command improve query latency in Delta tables?
How does the driver program handle task scheduling?
How is Git version control implemented in Databricks?
How would you identify and resolve a shuffle spill in Spark UI?
What are the limitations of the REORG command with respect to large datasets?
What causes Out of Memory (OOM) issues in Databricks, and how do you resolve them?
Can Schema Evolution lead to data inconsistencies? If so, how do you manage them?
Differentiate between Schema Enforcement and Schema Evolution.
Type or paste your answer to any of these questions and our AI Coach scores it, highlights gaps, and rewrites it at FAANG quality. Free to try.
The Data Engineering Interview Answer Vault bundles 750+ reviewed answers into 7 focused PDF volumes β SQL, Spark, Python, System Design, Cloud, Behavioral, and Data Modeling. Study on any device, no subscription required.
800+ hands-on courses β Grokking System Design, Coding Patterns, and AI mock interviews for your DE loop.
Turn any topic or your own notes into an interactive, personalized course in 60 seconds.
The book that gets data engineers through system-design rounds. Essential reading.
Some links below are affiliate links. If you buy through them we may earn a small commission at no extra cost to you β it helps keep DataEngPrep free.