Real questions from top companies Β· easy
What is the difference between managed and external tables in Hive or Spark SQL?
What is the difference between map and flatMap in Spark transformations?
What is the purpose of the VACUUM command in Delta Lake?
What limitations do you face when using Delta Tables in a multi-cloud environment?
What metrics do you use to determine whether a Spark job is going well or not?
Which Spark version are you using in your project, and why did you choose it?
Why does Hive use Derby by default, and what alternatives are used in production?
Worked with UDFs - share examples
Write PySpark code to filter and count records.
Write PySpark code to filter records based on specific conditions and add a calculated column.
Write a PySpark code snippet to filter rows with a specific condition.
Write the Spark command to rename an existing column in a DataFrame.
Writing Excel sheets to Delta tables in Databricks
You are given 10 worker machines with 100 GB RAM and 25 CPU cores. How would you determine the number of executors and the size of each executor?
How do you handle exceptions in data ingestion?
Type or paste your answer to any of these questions and our AI Coach scores it, highlights gaps, and rewrites it at FAANG quality. Free to try.