Questions tagged partition Β· medium
What is the difference between repartition and coalesce in Apache Spark?
Write an SQL query to find the second-highest salary from an employee table.
What is the difference between cache() and persist() in Spark? When would you use each?
What is the difference between groupByKey and reduceByKey in Spark?
What is the difference between narrow and wide transformations in Apache Spark? Explain with examples.
Demonstrate the difference between DENSE_RANK() and RANK()
Explain the differences between Data Warehouse, Data Lake, and Delta Lake
Explain the differences between Repartition and Coalesce. When would you use each?
What is the difference between partitioning and bucketing in Spark, and when would you use bucketing?
What strategies can you use to handle skewed data in Spark?
Describe a scenario where partitioning and bucketing would improve query performance.
Explain the types of triggers in ADF, including schedule, tumbling window, and event-based triggers.
How do you remove duplicate rows in BigQuery?
Explain the difference between Spark's map() and flatMap() transformations.
Tell me about a time when you faced a challenging situation at work and how you handled it.
What challenges did you face, and how did you tackle them?
What would you do if a pipeline failed and you couldn't find the reason?
How would you read data from a web API? What steps would you follow after reading the data?
Difference between ROW_NUMBER(), RANK(), and DENSE_RANK() with examples.
Explain SQL Window Functions with examples.