Real questions from top companies in SQL
Implement a recursive query for hierarchy (employee-manager). Explain the termination guarantees, depth limits, and when a recursive CTE becomes a scalability bottleneck. What alternatives exist for graph-scale hierarchies in Spark or a data lake?
Explain bloom filters in Spark: how they reduce I/O and when they introduce false positives that hurt performance. What are the scalability and cost implications of enabling dynamic partition pruning and bloom filter pushdown at petabyte scale?
Design a star schema for retail analytics (e.g., Adidas). Explain the dimensional modeling choices, SCD strategy, and how you would scale this schema for global multi-currency, multi-region deployments. What are the refresh and storage cost implications?
Compare Glue partition discovery with Hive MSCK/ADD PARTITION. Explain the operational and cost implications of crawler-based vs. partition-projection approaches. When does partition projection become necessary, and what are its limitations?
Explain how partitioning and bucketing in Hive/Spark optimize queries. What are the trade-offs in bucket count, partition cardinality, and small-file problem? When does over-partitioning or over-bucketing become counterproductive?
Explain the Medallion Architecture (Bronze, Silver, Gold).
Explain the concept of window functions in SQL and provide an example
Explain the difference between Star and Snowflake schemas. When would you choose one over the other?
Explain the difference between partition count and query performance in Spark.
Given a CSV file with raw customer transactions, design an ETL pipeline that cleans data, aggregates total sales by region and product, and loads into target table
How do you design a scalable and fault-tolerant data warehouse on a cloud platform?
How would you design a data model for an e-commerce platform?
How would you handle data type changes for an existing column?
How would you handle duplicate or corrupted data in a batch ETL job?
How would you handle null values in a dataset, especially in a single column?
How would you handle nulls in a SQL join? Provide examples using COALESCE.
How would you identify duplicate records based on a composite key in SQL?
How would you optimize a SQL query for better performance when working with large datasets?
How would you optimize a query fetching sales data across multiple countries with billions of rows?
How would you optimize a query with multiple joins and subqueries?
Type or paste your answer to any of these questions and our AI Coach scores it, highlights gaps, and rewrites it at FAANG quality. Free to try.