Reviewed by Aditya Kumar · Last reviewed 2026-08-08
To flatten a nested list into a single, non nested list, a common and elegant Pythonic approach involves recursion, often combined with a list comprehension and the sum() function. This method…
This easy-level Python/Coding question appears frequently in data engineering interviews at companies like Cognizant. While less common, it tests deeper understanding that distinguishes strong candidates.
Start by clearly defining the core concept being asked about. Interviewers want to see that you understand the fundamentals before diving into implementation details. Structure your answer with a definition, then explain the practical application with a concise example. The expert answer includes a code example that demonstrates the implementation pattern.
To flatten a nested list into a single, non-nested list, a common and elegant Pythonic approach involves recursion, often combined with a list comprehension and the sum() function. This method efficiently handles lists nested to arbitrary depths.
The core idea is to iterate through the input list. For each element, check if it's a list itself using isinstance(item, list). If it is, recursively call the flattening function on that sub-list. If it's not a list, treat it as a single element by wrapping it in a new list [item]. The sum() function, when used with an empty list [] as its starting value, effectively concatenates all the resulting sub-lists (from recursive calls) and individual elements (wrapped in lists) into a single flat list. This leverages Python's list concatenation behavior.
For arbitrarily nested lists, a recursive solution is often the most concise and readable:
def flatten(nested_list):
return sum([flatten(item) if isinstance(item, list) else [item] for item in nested_list], [])
While elegant, this recursive approach has trade-offs. Python has a default recursion depth limit (typically 1000), which can lead to a RecursionError for very deeply nested lists. For extremely large or deeply nested structures, an iterative approach (e.g., using a stack to manage pending items) or leveraging libraries like itertools.chain.from_iterable (for shallow nesting or when all sub-elements are iterable) might be more robust or memory-efficient. In data engineering, where memory and stack limits are critical for processing large datasets, these performance considerations are important.
Discuss how this problem relates to parsing complex JSON structures (e.g., using explode in Spark SQL or PySpark to flatten arrays of structs), denormalizing data, or handling semi-structured data in systems like Snowflake. Mention generator expressions for memory-efficient flattening without building intermediate lists, which is crucial for large datasets.
Pro-Move: Iterative with stack for deep nesting. Red Flag: Recursion overflow on deep input.
Some links below are affiliate links. If you buy through them we may earn a small commission at no extra cost to you — it helps keep DataEngPrep free.
According to DataEngPrep.tech, this is one of the most frequently asked Python/Coding interview questions, reported at 1 company. DataEngPrep.tech maintains an editor-reviewed database of 1,863 data engineering interview questions across 7 categories.