Reviewed by Aditya Kumar · Last reviewed 2026-03-24
In Databricks, you run one notebook from another using the %run magic command. This command executes the specified notebook within the same Spark session as the calling notebook, making its variables,…
This easy-level General/Other question appears frequently in data engineering interviews at companies like Hexaware. While less common, it tests deeper understanding that distinguishes strong candidates. Mastering the underlying concepts (python) will help you answer variations of this question confidently.
Start by clearly defining the core concept being asked about. Interviewers want to see that you understand the fundamentals before diving into implementation details. Structure your answer with a definition, then explain the practical application with a concise example. The expert answer includes a code example that demonstrates the implementation pattern.
In Databricks, you run one notebook from another using the %run magic command. This command executes the specified notebook within the same Spark session as the calling notebook, making its variables, functions, and DataFrames directly accessible.
%run is invoked, the target notebook's code is executed sequentially, effectively embedding its logic. Any objects (e.g., Python variables, PySpark DataFrames, temporary views) created or defined in the "callee" notebook become available in the "caller" notebook's scope. This is powerful for modularizing common setup steps, utility functions, or sequential ETL stages, promoting code reuse and organization. You can also pass parameters to the called notebook using the $param=value syntax, which are then accessible via dbutils.widgets.get("param") in the callee. To return a value from the callee, dbutils.notebook.exit("return_value") can be used, which terminates the callee and passes the string back to the caller.
%run is excellent for interactive development, breaking down complex tasks, and sharing common logic within a single job, it has important considerations. Over-reliance can lead to implicit dependencies and a less clear execution flow, making debugging and understanding data lineage challenging. The shared Spark session means that changes in one notebook (e.g., spark.conf settings, temporary views) affect the entire session. For robust, production-grade pipelines, it's generally recommended to encapsulate reusable logic into Python modules or libraries (e.g., .py files, wheel files) and install them into your cluster. These modules can then be imported and used across notebooks or jobs, offering better version control, testability, and explicit dependency management. Orchestrate these modules and notebooks using Databricks Jobs or external orchestrators like Airflow for scheduled execution and monitoring.
# Caller Notebook
%run ./my_utility_notebook $input_path="/data/raw" $output_table="processed_data"
# Now, variables/functions defined in my_utility_notebook are available here.
display(processed_df)
In the interview, also mention the implications of the shared Spark session context, such as resource management and potential side effects from shared state.
Pro-Move: 'Common setup in shared notebook; %run with env param. Production jobs use .py modules—notebooks for exploration only.'
Some links below are affiliate links. If you buy through them we may earn a small commission at no extra cost to you — it helps keep DataEngPrep free.
According to DataEngPrep.tech, this is one of the most frequently asked General/Other interview questions, reported at 1 company. DataEngPrep.tech maintains an editor-reviewed database of 1,863 data engineering interview questions across 7 categories.