Essential cookies keep authentication working. With your permission, we also use analytics cookies to understand and improve the product. Read our Privacy Policy

DataEngPrep.tech
QuestionsPracticeAI CoachDashboardPricingBlog
ProLogin

Interview Questions

Real questions from top companies Β· hard

700+ Easy450+ Medium650+ Hard
All CategoriesBehavioralSpark/Big DataSQLPython/CodingSystem Design/ArchitectureCloud/ToolsGeneral/Othereasymediumhard
141

Describe a challenging project where you optimized a complex ETL process.

SQLhardetljoinoptimization0.5 min read
Goldman Sachs
β†’
142

Describe a situation where you had to redesign a data model to meet changing business needs

SQLhardjoinoptimizationpartition3.6 min read
Kagina
β†’
143

Design a Custom API that can query a backend server and return customer data such as the number of orders placed by a user based on their user ID

SQLhardjoinoptimizationpartition3.6 min read
Meesho
β†’
144

Design a daily ETL pipeline to ingest API data into BigQuery.

SQLhardbigqueryetljoin3.6 min read
Google
β†’
145

Design a financial database system focusing on database models, schema design, partition keys, and query optimization techniques.

SQLhardjoinoptimizationpartition3.6 min read
Flipkart
β†’
146

Design a relational data model for a sales database, incorporating normalization techniques

SQLhardjoinoptimizationpartition3.6 min read
Morgan Stanley
β†’
147

Design a structure (data model) that allows efficient querying of movies based on multiple search criteria (title, genre, actor, director).

SQLhardjoinoptimizationpartition3.6 min read
Wayfair
β†’
148

Design the data model for an ETL pipeline that ingests data from a database and loads it into Snowflake

SQLhardetljoinoptimization3.6 min read
Meesho
β†’
149

Designing backend architecture for SQL Warehouse?

SQLhardjoinoptimizationpartition3.6 min read
Snowflake
β†’
150

Designing scalable data models - explain approach

SQLhardjoinoptimizationpartition3.6 min read
Lumiq
β†’
151

Difference Between Truncate/Delete and Union/Union All – Performance and Usage

SQLhardetlspark0.5 min read
Presidio
β†’
152

Explain BigQuery Architecture.

SQLhardbigqueryjoinoptimization3.6 min read
EY
β†’
153

Articulate the architectural decisions, scalability trade-offs, and cost implications of designing an AWS data platform. How would you justify glue vs. EMR, Redshift vs. Athena, and when would each choice become cost-prohibitive at scale?

SQLhardjoinoptimizationpartition3.6 min read
Freecharge
β†’
154

Explain the architectural rationale for using LeftAntiJoin vs. NOT IN vs. NOT EXISTS in a distributed context. When does LeftAntiJoin become a performance or scalability bottleneck, and how do broadcast vs. shuffle joins affect cost?

SQLhardjoinpartition0.6 min read
Infosys
β†’
155

Explain the architectural trade-offs when optimizing a query on 100M+ rows: indexing vs. partitioning vs. materialized views. When does each approach become cost-prohibitive or operationally burdensome, and how do you quantify impact?

SQLhardoptimizationpartitionwindow0.5 min read
Bristol Myers Squibb
β†’
156

Explain bloom filters in Spark: how they reduce I/O and when they introduce false positives that hurt performance. What are the scalability and cost implications of enabling dynamic partition pruning and bloom filter pushdown at petabyte scale?

SQLhardjoinoptimizationpartition0.5 min read
American Express
β†’
157

Design a star schema for retail analytics (e.g., Adidas). Explain the dimensional modeling choices, SCD strategy, and how you would scale this schema for global multi-currency, multi-region deployments. What are the refresh and storage cost implications?

SQLhardjoinoptimizationpartition3.6 min read
Adidas
β†’
158

Explain the Medallion Architecture (Bronze, Silver, Gold).

SQLhardjoinoptimizationpartition3.6 min read
Databricks
β†’
159

Given a CSV file with raw customer transactions, design an ETL pipeline that cleans data, aggregates total sales by region and product, and loads into target table

SQLhardetljoinoptimization3.6 min read
McKinsey
β†’
160

How can you automate data insertion into BigQuery using Python?

SQLhardairflowbigquerypython2 min read
Aarete
β†’

Reading isn't practice. Get AI feedback on your answers.

Type or paste your answer to any of these questions and our AI Coach scores it, highlights gaps, and rewrites it at FAANG quality. Free to try.

Try AI Answer Coach β€” FreeStart a Mock Interview
Previous1...678910...22Next
Categories
All QuestionsSQLSpark / Big DataPython / CodingSystem DesignCloud / ToolsBehavioral
By Company
AmazonGoogleDatabricksSnowflakeAWSAzureMicrosoftNetflixUberTCS
Interview Guides
All GuidesTop SQL QuestionsTop Spark QuestionsPySpark QuestionsTop Python QuestionsTop System DesignKafka QuestionsAirflow QuestionsSQL Window FunctionsETL QuestionsData Modeling
Products
AI Interview CoachAnswer AnalyzerSQL PlaygroundResume AnalyzerAnswer Vault PDFsPricing
Company
About & Editorial PolicyContact UsAI DisclosureDisclaimerTerms of ServicePrivacy Policy
Β© 2026 DataEngPrep.tech. All rights reserved.
AboutBlogContactDisclaimer