The short answer: Spark interviews test whether you understand how Spark executes your code across a cluster: transformations vs actions and lazy evaluation, the driver and executors, how data is partitioned, what causes a shuffle and why it’s expensive, caching, join strategies and data skew, and the DataFrame/Spark SQL API most teams use today. Experienced roles add Structured Streaming and performance tuning scenarios (“This job is slow. What do you check?”). Below are the common questions with short answers and PySpark examples.
Core concepts
“What is an RDD?”
The Spark documentation describes a resilient distributed dataset (RDD) as “a collection of elements partitioned across the nodes of the cluster that can be operated on in parallel.” It’s Spark’s low-level abstraction; most new code uses DataFrames.
“What’s a DataFrame, and why prefer it over RDDs?”
In Spark’s words, “a DataFrame is a Dataset organized into named columns,” conceptually like a database table. Because Spark knows the structure of the data and the computation, Spark SQL can optimize DataFrame queries, which makes them usually faster and easier than equivalent RDD code.
“What’s the difference between a transformation and an action?”
Transformations (such as select, filter, map, join) create a new dataset; actions (such as count, collect, write, show) return a result or write output. The docs explain that “all transformations in Spark are lazy”: they’re only computed when an action needs a result.
“Why is lazy evaluation useful?”
Spark can see the whole plan before running it, combine steps, push filters down and avoid unnecessary work.
“What do the driver and executors do?”
The driver runs your main program, builds the execution plan and schedules tasks; executors run the tasks on worker nodes and hold cached data.
Execution and performance
“What’s the difference between narrow and wide transformations?”
Narrow transformations (filter, map, select) work within a partition. Wide transformations (groupBy, join, distinct, repartition) need data from many partitions, causing a shuffle: data is redistributed across the network and written to disk, which is expensive.
“What’s the difference between repartition and coalesce?”
repartition(n) does a full shuffle to create n evenly sized partitions (it can increase or decrease them); coalesce(n) reduces partitions without a full shuffle, which is cheaper but can leave uneven partitions.
“When should you cache or persist?”
When a DataFrame is reused by several actions and is expensive to recompute. Unpersist it when done, and choose a storage level that fits memory.
features = raw.filter("event_date >= '2026-01-01'").select("user_id", "event", "value")
features.cache()
daily = features.groupBy("event").count()
users = features.select("user_id").distinct().count()
“How does Spark choose a join strategy?”
For a small table, a broadcast hash join sends it to every executor and avoids shuffling the large table; otherwise Spark typically uses a sort-merge join, which shuffles both sides by the join key. You can hint a broadcast:
from pyspark.sql.functions import broadcast
orders.join(broadcast(countries), "country_code")
“What is data skew, and how do you fix it?”
A few keys hold most of the data, so a few tasks run much longer than the rest. Fixes: broadcast the smaller side, salt the hot keys (add a random suffix and aggregate in two steps), filter or handle the hot keys separately, or use adaptive query execution’s skew-join handling in recent Spark versions.
Spark SQL and data
- “How does Spark SQL optimize queries?” It builds a logical plan, applies rule-based optimizations (such as predicate pushdown and column pruning), chooses a physical plan, and generates code. Use
explain()to see the plan. - “Why use Parquet?” Columnar storage, compression and predicate pushdown mean Spark reads only the columns and row groups it needs.
- “What’s partition pruning?” If data is partitioned on disk (for example by date) and you filter on that column, Spark skips irrelevant partitions.
Streaming
“What is Structured Streaming?”
Spark’s streaming engine treats a live stream as a table that keeps growing, so you write the same DataFrame operations as for batch. Know triggers, output modes (append, update, complete), checkpointing for fault tolerance, and watermarks for handling late data in windowed aggregations.
Scenario questions
- “A job that used to take 10 minutes now takes 2 hours. What do you check?” The Spark UI: stages with long tasks (skew), large shuffle reads and writes, spills to disk, input size changes, the number of partitions, and whether a broadcast join became a sort-merge join as a table grew.
- “An executor runs out of memory.” Check partition sizes and skew, avoid
collect()on large data, tune memory settings and partitions, and avoid caching more than needed. - “You see thousands of tiny output files.” Reduce partitions before writing (
coalesce), or compact files later.
See also our data engineer interview questions, Kafka interview questions and Hadoop interview questions.
Frequently asked questions
What Spark questions are asked in interviews?
What is lazy evaluation in Spark?
What causes a shuffle in Spark?
How do you handle data skew in Spark?
Should I use RDDs or DataFrames?
Sources
- Apache Spark, RDD Programming Guide. Read on 1 October 2026.
- Apache Spark, Spark SQL, DataFrames and Datasets Guide. Read on 1 October 2026.