Essential cookies keep authentication working. With your permission, we also use analytics cookies to understand and improve the product. Read our Privacy Policy

DataEngPrep.tech
QuestionsPracticeAI CoachDashboardPricingBlog
ProLogin

Interview Questions

Real questions from top companies

700+ Easy450+ Medium650+ Hard
All CategoriesBehavioralSpark/Big DataSQLPython/CodingSystem Design/ArchitectureCloud/ToolsGeneral/Othereasymediumhard
741

Split a DataFrame such that even numbers appear in one column and odd numbers in another

Spark/Big Datamediumpartitionpythonspark0.5 min read
KPMG
β†’
742

Sqoop Incremental Import?

Spark/Big Dataeasysql0.6 min read
Altimetrik
β†’
743

Sqoop command for importing multiple tables

Spark/Big Dataeasyairflowsql0.5 min read
Meesho
β†’
744

Steps to link a Databricks notebook to an ADF pipeline

Spark/Big Datahardspark0.6 min read
Kaseya
β†’
745

Steps to mount storage in Databricks.

Spark/Big Datamediumsparkwindow0.5 min read
Chubb
β†’
746

Suppose you have a DAG that ingests data from multiple databases. How would you increase task parallelism in Airflow to improve performance without overloading the system?

Spark/Big Dataeasyairflowsql0.6 min read
Dunnhumby
β†’
747

Suppose you need to import 5 tables from an external RDBMS (like MySQL) into Hadoop HDFS. Write the Sqoop command

Spark/Big Dataeasyairflowsql0.6 min read
Meesho
β†’
748

Task Dependencies in DAG

Spark/Big Dataeasyairflow0.5 min read
Verizon
β†’
749

Trade-offs between batch processing (Spark) vs. real-time streams (Kafka)

Spark/Big Datahardpartitionspark0.7 min read
PayPal
β†’
750

Transformation vs. Action in PySpark?

Spark/Big Datamediumjoinpartitionspark0.6 min read
Comcast
β†’
751

Usage of UDFs?

Spark/Big Datahardoptimizationpythonsql0.6 min read
Citi
β†’
752

Walk through how you would debug the data ingestion process to identify slow stages.

Spark/Big Datahardpartitionspark0.6 min read
Swiggy
β†’
753

Walkthrough Spark's architecture, focusing on driver, executors, and DAGs

Spark/Big Datahardoptimizationpartitionspark2.5 min read
KPMG
β†’
754

What Hadoop command would you use to merge multiple files into one?

Spark/Big Datamediumpartitionspark0.5 min read
Infosys
β†’
755

What are Spark optimizations, and can you explain them?

Spark/Big Datahardjoinoptimizationpartition0.6 min read
Cognizant
β†’
756

What are the advantages of using Delta Lake over Parquet?

Spark/Big Dataeasy0.5 min read
Puma
β†’
757

What are the differences between %pip and %conda commands in Databricks?

Spark/Big Dataeasypython0.6 min read
TCS
β†’
758

What are the different modes in which you can submit Spark jobs? Explain each.

Spark/Big Dataeasyspark0.5 min read
Dunnhumby
β†’
759

What are the limitations of the REORG command with respect to large datasets?

Spark/Big Datamediumpartition0.5 min read
PWC
β†’
760

What are the performance considerations when using Auto Loader?

Spark/Big Dataeasy0.5 min read
TCS
β†’

Reading isn't practice. Get AI feedback on your answers.

Type or paste your answer to any of these questions and our AI Coach scores it, highlights gaps, and rewrites it at FAANG quality. Free to try.

Try AI Answer Coach β€” FreeStart a Mock Interview
Previous1...3637383940...47Next
Categories
All QuestionsSQLSpark / Big DataPython / CodingSystem DesignCloud / ToolsBehavioral
By Company
AmazonGoogleDatabricksSnowflakeAWSAzureMicrosoftNetflixUberTCS
Interview Guides
All GuidesTop SQL QuestionsTop Spark QuestionsPySpark QuestionsTop Python QuestionsTop System DesignKafka QuestionsAirflow QuestionsSQL Window FunctionsETL QuestionsData Modeling
Products
AI Interview CoachAnswer AnalyzerSQL PlaygroundResume AnalyzerAnswer Vault PDFsPricing
Company
About & Editorial PolicyContact UsAI DisclosureDisclaimerTerms of ServicePrivacy Policy
Β© 2026 DataEngPrep.tech. All rights reserved.
AboutBlogContactDisclaimer