Essential cookies keep authentication working. With your permission, we also use analytics cookies to understand and improve the product. Read our Privacy Policy

DataEngPrep.tech
QuestionsPracticeAI CoachDashboardPricingBlog
ProLogin

Interview Questions

Real questions from top companies

700+ Easy450+ Medium650+ Hard
All CategoriesBehavioralSpark/Big DataSQLPython/CodingSystem Design/ArchitectureCloud/ToolsGeneral/Othereasymediumhard
521

Design a daily ETL pipeline to ingest API data into BigQuery.

SQLhardbigqueryetljoin3.6 min read
Google
β†’
522

Design a financial database system focusing on database models, schema design, partition keys, and query optimization techniques.

SQLhardjoinoptimizationpartition3.6 min read
Flipkart
β†’
523

Design a relational data model for a sales database, incorporating normalization techniques

SQLhardjoinoptimizationpartition3.6 min read
Morgan Stanley
β†’
524

Design a structure (data model) that allows efficient querying of movies based on multiple search criteria (title, genre, actor, director).

SQLhardjoinoptimizationpartition3.6 min read
Wayfair
β†’
525

Design the data model for an ETL pipeline that ingests data from a database and loads it into Snowflake

SQLhardetljoinoptimization3.6 min read
Meesho
β†’
526

Designing backend architecture for SQL Warehouse?

SQLhardjoinoptimizationpartition3.6 min read
Snowflake
β†’
527

Designing scalable data models - explain approach

SQLhardjoinoptimizationpartition3.6 min read
Lumiq
β†’
528

Difference Between Truncate/Delete and Union/Union All – Performance and Usage

SQLhardetlspark0.5 min read
Presidio
β†’
529

Explain BigQuery Architecture.

SQLhardbigqueryjoinoptimization3.6 min read
EY
β†’
530

Explain CTE vs Temp Table. What are the differences and use cases?

SQLmediumjoin0.5 min read
Fractal
β†’
531

Explain Data Modeling SCD Types (Type 1, 2, 3).

SQLmediumjoin0.5 min read
EY
β†’
532

Articulate the architectural decisions, scalability trade-offs, and cost implications of designing an AWS data platform. How would you justify glue vs. EMR, Redshift vs. Athena, and when would each choice become cost-prohibitive at scale?

SQLhardjoinoptimizationpartition3.6 min read
Freecharge
β†’
533

Explain the architectural rationale for using LeftAntiJoin vs. NOT IN vs. NOT EXISTS in a distributed context. When does LeftAntiJoin become a performance or scalability bottleneck, and how do broadcast vs. shuffle joins affect cost?

SQLhardjoinpartition0.6 min read
Infosys
β†’
534

Describe a cross-team data project where you had to align architectural boundaries, ownership, and SLAs. How did you handle conflicting priorities, technical debt, and the scalability of communication as the number of stakeholders grew?

SQLeasy0.5 min read
American Express
β†’
535

Walk through a production incident where data freshness or correctness was at risk. How did you balance immediate mitigation vs. root-cause remediation? What architectural changes would prevent recurrence, and what are the cost vs. reliability trade-offs?

SQLeasy0.5 min read
Adidas
β†’
536

Explain the architectural trade-offs when optimizing a query on 100M+ rows: indexing vs. partitioning vs. materialized views. When does each approach become cost-prohibitive or operationally burdensome, and how do you quantify impact?

SQLhardoptimizationpartitionwindow0.5 min read
Bristol Myers Squibb
β†’
537

Implement a recursive query for hierarchy (employee-manager). Explain the termination guarantees, depth limits, and when a recursive CTE becomes a scalability bottleneck. What alternatives exist for graph-scale hierarchies in Spark or a data lake?

SQLmediumjoinspark0.6 min read
American Express
β†’
538

Explain bloom filters in Spark: how they reduce I/O and when they introduce false positives that hurt performance. What are the scalability and cost implications of enabling dynamic partition pruning and bloom filter pushdown at petabyte scale?

SQLhardjoinoptimizationpartition0.5 min read
American Express
β†’
539

Design a star schema for retail analytics (e.g., Adidas). Explain the dimensional modeling choices, SCD strategy, and how you would scale this schema for global multi-currency, multi-region deployments. What are the refresh and storage cost implications?

SQLhardjoinoptimizationpartition3.6 min read
Adidas
β†’
540

Compare Glue partition discovery with Hive MSCK/ADD PARTITION. Explain the operational and cost implications of crawler-based vs. partition-projection approaches. When does partition projection become necessary, and what are its limitations?

SQLmediumpartition0.5 min read
Capco
β†’

Reading isn't practice. Get AI feedback on your answers.

Type or paste your answer to any of these questions and our AI Coach scores it, highlights gaps, and rewrites it at FAANG quality. Free to try.

Try AI Answer Coach β€” FreeStart a Mock Interview
Previous1...2526272829...47Next
Categories
All QuestionsSQLSpark / Big DataPython / CodingSystem DesignCloud / ToolsBehavioral
By Company
AmazonGoogleDatabricksSnowflakeAWSAzureMicrosoftNetflixUberTCS
Interview Guides
All GuidesTop SQL QuestionsTop Spark QuestionsPySpark QuestionsTop Python QuestionsTop System DesignKafka QuestionsAirflow QuestionsSQL Window FunctionsETL QuestionsData Modeling
Products
AI Interview CoachAnswer AnalyzerSQL PlaygroundResume AnalyzerAnswer Vault PDFsPricing
Company
About & Editorial PolicyContact UsAI DisclosureDisclaimerTerms of ServicePrivacy Policy
Β© 2026 DataEngPrep.tech. All rights reserved.
AboutBlogContactDisclaimer