Real questions from top companies in Spark/Big Data Β· easy
What is the difference between Managed and External tables in Hive/Spark?
When would you architecturally choose Dataset[T] over DataFrame in a Scala Spark pipeline, and what are the scalability and portability trade-offs? Include type-safety benefits vs. operational constraints.
What is the difference between Managed and External Tables in Databricks?
Approaches to handling multiple tasks within a sprint?
Can you share a time when you had to shift focus due to urgent tasks?
Create a DataFrame with default column types
Explain your approach to monitoring and logging Spark jobs in AWS. What tools would you use to identify performance bottlenecks?
How do you compare the time investment and value of a task?
How do you handle bad data in Databricks?
Sqoop Incremental Import?
Sqoop command for importing multiple tables
Suppose you have a DAG that ingests data from multiple databases. How would you increase task parallelism in Airflow to improve performance without overloading the system?
Suppose you need to import 5 tables from an external RDBMS (like MySQL) into Hadoop HDFS. Write the Sqoop command
Task Dependencies in DAG
What are the advantages of using Delta Lake over Parquet?
What are the differences between %pip and %conda commands in Databricks?
What are the different modes in which you can submit Spark jobs? Explain each.
What are the performance considerations when using Auto Loader?
What are transient clusters in EMR, and when would you use them?
Write PySpark code to filter and count records.
Type or paste your answer to any of these questions and our AI Coach scores it, highlights gaps, and rewrites it at FAANG quality. Free to try.