Interview questions
Preparing for a data engineering interview at Citi? This page contains 20 real interview questions sourced from verified Citi interview experiences. Questions are sorted by frequency β the ones asked most often appear first.
Citi data engineering interviews typically focus on Spark/Big Data, SQL, and General/Other. The interview bar skews toward harder problems (10 hard vs. 3 easy), suggesting emphasis on depth and system-level thinking.
Use the difficulty filters above to focus your preparation. For each question, attempt your own answer first, then compare with our expert solution. You can also practice these questions in our AI Mock Interview Coach for real-time feedback.
What is the difference between repartition and coalesce in Apache Spark?
What is the difference between SparkSession and SparkContext in Spark?
What is the difference between partitioning and bucketing in Spark, and when would you use bucketing?
What strategies can you use to handle skewed data in Spark?
What is the difference between Managed and External tables in Hive/Spark?
What is a window function? Explain with an example.
Explain the concept of checkpointing in Spark and why it is important.
Shell commands for renaming a file?
Shell: change permissions?
Shell: command to check processes running in the background?
Given 1TB of a file, how to check word count?
Teradata to Hadoop migration and handling data with SCD Type 2?
What is a Kafka topic, and how do you choose the number of partitions for it?
Explain the concept of consumer groups in Kafka. How do they affect message processing?
Explain the difference between TriggerDagRunOperator and ExternalTaskSensor in Airflow.
How would you design a Kafka-based pipeline for processing streaming data in real-time?
Usage of UDFs?
Describe an end-to-end data pipeline project you worked on, highlighting your role and the technologies used.
Describe how Kafka ensures data durability and fault tolerance.
Introduce your recent project, explaining its goal, architecture, tools, and technologies.
Type or paste your answer to any of these questions and our AI Coach scores it, highlights gaps, and rewrites it at FAANG quality. Free to try.
The Data Engineering Interview Answer Vault bundles 750+ reviewed answers into 7 focused PDF volumes β SQL, Spark, Python, System Design, Cloud, Behavioral, and Data Modeling. Study on any device, no subscription required.
800+ hands-on courses β Grokking System Design, Coding Patterns, and AI mock interviews for your DE loop.
Turn any topic or your own notes into an interactive, personalized course in 60 seconds.
The book that gets data engineers through system-design rounds. Essential reading.
Some links below are affiliate links. If you buy through them we may earn a small commission at no extra cost to you β it helps keep DataEngPrep free.