Essential cookies keep authentication working. With your permission, we also use analytics cookies to understand and improve the product. Read our Privacy Policy

DataEngPrep.tech
QuestionsPracticeAI CoachDashboardPricingBlog
ProLogin

Interview Questions

Real questions from top companies in Spark/Big Data

700+ Easy450+ Medium650+ Hard
All CategoriesBehavioralSpark/Big DataSQLPython/CodingSystem Design/ArchitectureCloud/ToolsGeneral/Othereasymediumhard
161

Steps to link a Databricks notebook to an ADF pipeline

Spark/Big Datahardspark0.6 min read
Kaseya
β†’
162

Steps to mount storage in Databricks.

Spark/Big Datamediumsparkwindow0.5 min read
Chubb
β†’
163

Suppose you have a DAG that ingests data from multiple databases. How would you increase task parallelism in Airflow to improve performance without overloading the system?

Spark/Big Dataeasyairflowsql0.6 min read
Dunnhumby
β†’
164

Suppose you need to import 5 tables from an external RDBMS (like MySQL) into Hadoop HDFS. Write the Sqoop command

Spark/Big Dataeasyairflowsql0.6 min read
Meesho
β†’
165

Task Dependencies in DAG

Spark/Big Dataeasyairflow0.5 min read
Verizon
β†’
166

Trade-offs between batch processing (Spark) vs. real-time streams (Kafka)

Spark/Big Datahardpartitionspark0.7 min read
PayPal
β†’
167

Transformation vs. Action in PySpark?

Spark/Big Datamediumjoinpartitionspark0.6 min read
Comcast
β†’
168

Usage of UDFs?

Spark/Big Datahardoptimizationpythonsql0.6 min read
Citi
β†’
169

Walk through how you would debug the data ingestion process to identify slow stages.

Spark/Big Datahardpartitionspark0.6 min read
Swiggy
β†’
170

Walkthrough Spark's architecture, focusing on driver, executors, and DAGs

Spark/Big Datahardoptimizationpartitionspark2.5 min read
KPMG
β†’
171

What Hadoop command would you use to merge multiple files into one?

Spark/Big Datamediumpartitionspark0.5 min read
Infosys
β†’
172

What are Spark optimizations, and can you explain them?

Spark/Big Datahardjoinoptimizationpartition0.6 min read
Cognizant
β†’
173

What are the advantages of using Delta Lake over Parquet?

Spark/Big Dataeasy0.5 min read
Puma
β†’
174

What are the differences between %pip and %conda commands in Databricks?

Spark/Big Dataeasypython0.6 min read
TCS
β†’
175

What are the different modes in which you can submit Spark jobs? Explain each.

Spark/Big Dataeasyspark0.5 min read
Dunnhumby
β†’
176

What are the limitations of the REORG command with respect to large datasets?

Spark/Big Datamediumpartition0.5 min read
PWC
β†’
177

What are the performance considerations when using Auto Loader?

Spark/Big Dataeasy0.5 min read
TCS
β†’
178

What are transient clusters in EMR, and when would you use them?

Spark/Big Dataeasyetl0.5 min read
Persistent Systems
β†’
179

What causes Out of Memory (OOM) issues in Databricks, and how do you resolve them?

Spark/Big Datamediumpartitionspark0.5 min read
PWC
β†’
180

What is Predicate Pushdown and AQE with Example

Spark/Big Datahardjoinoptimizationpartition0.6 min read
Nagarro
β†’

Reading isn't practice. Get AI feedback on your answers.

Type or paste your answer to any of these questions and our AI Coach scores it, highlights gaps, and rewrites it at FAANG quality. Free to try.

Try AI Answer Coach β€” FreeStart a Mock Interview
Previous1...7891011Next
Categories
All QuestionsSQLSpark / Big DataPython / CodingSystem DesignCloud / ToolsBehavioral
By Company
AmazonGoogleDatabricksSnowflakeAWSAzureMicrosoftNetflixUberTCS
Interview Guides
All GuidesTop SQL QuestionsTop Spark QuestionsPySpark QuestionsTop Python QuestionsTop System DesignKafka QuestionsAirflow QuestionsSQL Window FunctionsETL QuestionsData Modeling
Products
AI Interview CoachAnswer AnalyzerSQL PlaygroundResume AnalyzerAnswer Vault PDFsPricing
Company
About & Editorial PolicyContact UsAI DisclosureDisclaimerTerms of ServicePrivacy Policy
Β© 2026 DataEngPrep.tech. All rights reserved.
AboutBlogContactDisclaimer