Reviewed by Aditya Kumar · Last reviewed 2026-03-24
Glue: Crawler infers partitions from S3 path (e.g., s3://bucket/table/year=2024/month=01/); catalogs in Glue Catalog. Hive: MSCK REPAIR TABLE or ADD PARTITION; expects Hive-style paths. Why Glue: AWS-native; integrates with Athena, Redshift Spectrum; no cluster to run. Partition...
This medium-level SQL question appears frequently in data engineering interviews at companies like Capco. While less common, it tests deeper understanding that distinguishes strong candidates. Mastering the underlying concepts (partition) will help you answer variations of this question confidently.
Break this problem into components. Identify the core trade-offs involved, then walk the interviewer through your reasoning step by step. Demonstrate awareness of edge cases and production considerations - this is what separates good answers from great ones.
Glue: Crawler infers partitions from S3 path (e.g., s3://bucket/table/year=2024/month=01/); catalogs in Glue Catalog. Hive: MSCK REPAIR TABLE or ADD PARTITION; expects Hive-style paths. Why Glue: AWS-native; integrates with Athena, Redshift Spectrum; no cluster to run. Partition projection: Define partition schema (type, range); no crawler needed; queries resolve partitions from path at query time. Operational: Crawler runs on schedule, costs DPU time; projection has no crawler cost. Cost: High-cardinality partitions (e.g., by hour, user_id) make crawler expensive; projection scales. Limitations: Projection requires predictable path pattern; no validation of actual data; complex patterns need custom projection. Best practice: Use projection for time-series or bounded cardinality; crawler for ad-hoc or evolving schema.
Red Flag: Running crawler on every pipeline run—cost and latency pile up. Pro-Move: Use partition projection for time-partitioned tables; reserve crawler for schema discovery or low-frequency catalog updates.
Some links below are affiliate links. If you buy through them we may earn a small commission at no extra cost to you — it helps keep DataEngPrep free.
According to DataEngPrep.tech, this is one of the most frequently asked SQL interview questions, reported at 1 company. DataEngPrep.tech maintains an editor-reviewed database of 1,863 data engineering interview questions across 7 categories.