Real questions from top companies in Spark/Big Data · easy
Why does Hive use Derby by default, and what alternatives are used in production?
Worked with UDFs - share examples
Write PySpark code to filter and count records.
Write PySpark code to filter records based on specific conditions and add a calculated column.
Write a PySpark code snippet to filter rows with a specific condition.
Write the Spark command to rename an existing column in a DataFrame.
Writing Excel sheets to Delta tables in Databricks
You are given 10 worker machines with 100 GB RAM and 25 CPU cores. How would you determine the number of executors and the size of each executor?
Type or paste your answer to any of these questions and our AI Coach scores it, highlights gaps, and rewrites it at FAANG quality. Free to try.