Interview questions
Preparing for a data engineering interview at American Express? This page contains 21 real interview questions sourced from verified American Express interview experiences. Questions are sorted by frequency β the ones asked most often appear first.
American Express data engineering interviews typically focus on Python/Coding, Spark/Big Data, and SQL. The interview bar skews toward harder problems (10 hard vs. 5 easy), suggesting emphasis on depth and system-level thinking.
Use the difficulty filters above to focus your preparation. For each question, attempt your own answer first, then compare with our expert solution. You can also practice these questions in our AI Mock Interview Coach for real-time feedback.
What is the difference between SparkSession and SparkContext in Spark?
Discuss the data size challenges in your previous projects. How did you optimize storage and processing?
What were the biggest infrastructure-level challenges you faced, and how did you resolve them?
Why do you want to join American Express?
What are your strengths, and how do they align with the Data Engineer role?
Create a Python program to demonstrate the use of set operations (union, intersection).
Describe Spark's memory management model. How do you handle heap memory overhead issues?
Explain the difference between mutable and immutable objects in Python.
Explain the differences between multiprocessing and multithreading.
Implement a Python function to count unique words from a file and write them to another file.
Write a decorator function to log the execution time of a function.
Describe a cross-team data project where you had to align architectural boundaries, ownership, and SLAs. How did you handle conflicting priorities, technical debt, and the scalability of communication as the number of stakeholders grew?
Implement a recursive query for hierarchy (employee-manager). Explain the termination guarantees, depth limits, and when a recursive CTE becomes a scalability bottleneck. What alternatives exist for graph-scale hierarchies in Spark or a data lake?
Explain bloom filters in Spark: how they reduce I/O and when they introduce false positives that hurt performance. What are the scalability and cost implications of enabling dynamic partition pruning and bloom filter pushdown at petabyte scale?
How would you optimize a query with multiple joins and subqueries?
Code a simple PySpark job to read a JSON file, filter records, and write output in Parquet format.
Explain a scenario-based question on Spark optimization and how you would troubleshoot performance issues.
Explain repartition vs. coalesce. Which one would you use to reduce shuffle operations?
How did you handle data ingestion and processing for large datasets?
Describe the architecture of an ETL pipeline you built in your previous project.
Type or paste your answer to any of these questions and our AI Coach scores it, highlights gaps, and rewrites it at FAANG quality. Free to try.
The Data Engineering Interview Answer Vault bundles 750+ reviewed answers into 7 focused PDF volumes β SQL, Spark, Python, System Design, Cloud, Behavioral, and Data Modeling. Study on any device, no subscription required.
800+ hands-on courses β Grokking System Design, Coding Patterns, and AI mock interviews for your DE loop.
Turn any topic or your own notes into an interactive, personalized course in 60 seconds.
The book that gets data engineers through system-design rounds. Essential reading.
Some links below are affiliate links. If you buy through them we may earn a small commission at no extra cost to you β it helps keep DataEngPrep free.