Huru
Technical Interviews

Spark Interview Questions and Answers (2026)

Apache Spark interview questions with short answers and PySpark code: RDDs vs DataFrames, transformations and actions, lazy evaluation, the driver and executors, narrow vs wide transformations and shuffles, partitioning, caching, joins and data skew, Spark SQL and the optimizer, Structured Streaming, and performance tuning scenarios.

By Amine Boukioud Updated 4 min read
The Huru mascot holding a laptop and pointing at a cluster of glowing server nodes processing a stream of colored cubes into organized stacks around a bright spark

The short answer: Spark interviews test whether you understand how Spark executes your code across a cluster: transformations vs actions and lazy evaluation, the driver and executors, how data is partitioned, what causes a shuffle and why it’s expensive, caching, join strategies and data skew, and the DataFrame/Spark SQL API most teams use today. Experienced roles add Structured Streaming and performance tuning scenarios (“This job is slow. What do you check?”). Below are the common questions with short answers and PySpark examples.

Core concepts

“What is an RDD?”

The Spark documentation describes a resilient distributed dataset (RDD) as “a collection of elements partitioned across the nodes of the cluster that can be operated on in parallel.” It’s Spark’s low-level abstraction; most new code uses DataFrames.

“What’s a DataFrame, and why prefer it over RDDs?”

In Spark’s words, “a DataFrame is a Dataset organized into named columns,” conceptually like a database table. Because Spark knows the structure of the data and the computation, Spark SQL can optimize DataFrame queries, which makes them usually faster and easier than equivalent RDD code.

“What’s the difference between a transformation and an action?”

Transformations (such as select, filter, map, join) create a new dataset; actions (such as count, collect, write, show) return a result or write output. The docs explain that “all transformations in Spark are lazy”: they’re only computed when an action needs a result.

“Why is lazy evaluation useful?”

Spark can see the whole plan before running it, combine steps, push filters down and avoid unnecessary work.

“What do the driver and executors do?”

The driver runs your main program, builds the execution plan and schedules tasks; executors run the tasks on worker nodes and hold cached data.

Execution and performance

“What’s the difference between narrow and wide transformations?”

Narrow transformations (filter, map, select) work within a partition. Wide transformations (groupBy, join, distinct, repartition) need data from many partitions, causing a shuffle: data is redistributed across the network and written to disk, which is expensive.

“What’s the difference between repartition and coalesce?”

repartition(n) does a full shuffle to create n evenly sized partitions (it can increase or decrease them); coalesce(n) reduces partitions without a full shuffle, which is cheaper but can leave uneven partitions.

“When should you cache or persist?”

When a DataFrame is reused by several actions and is expensive to recompute. Unpersist it when done, and choose a storage level that fits memory.

features = raw.filter("event_date >= '2026-01-01'").select("user_id", "event", "value")
features.cache()
daily = features.groupBy("event").count()
users = features.select("user_id").distinct().count()

“How does Spark choose a join strategy?”

For a small table, a broadcast hash join sends it to every executor and avoids shuffling the large table; otherwise Spark typically uses a sort-merge join, which shuffles both sides by the join key. You can hint a broadcast:

from pyspark.sql.functions import broadcast
orders.join(broadcast(countries), "country_code")

“What is data skew, and how do you fix it?”

A few keys hold most of the data, so a few tasks run much longer than the rest. Fixes: broadcast the smaller side, salt the hot keys (add a random suffix and aggregate in two steps), filter or handle the hot keys separately, or use adaptive query execution’s skew-join handling in recent Spark versions.

Spark SQL and data

  • “How does Spark SQL optimize queries?” It builds a logical plan, applies rule-based optimizations (such as predicate pushdown and column pruning), chooses a physical plan, and generates code. Use explain() to see the plan.
  • “Why use Parquet?” Columnar storage, compression and predicate pushdown mean Spark reads only the columns and row groups it needs.
  • “What’s partition pruning?” If data is partitioned on disk (for example by date) and you filter on that column, Spark skips irrelevant partitions.

Streaming

“What is Structured Streaming?”

Spark’s streaming engine treats a live stream as a table that keeps growing, so you write the same DataFrame operations as for batch. Know triggers, output modes (append, update, complete), checkpointing for fault tolerance, and watermarks for handling late data in windowed aggregations.

Scenario questions

  • “A job that used to take 10 minutes now takes 2 hours. What do you check?” The Spark UI: stages with long tasks (skew), large shuffle reads and writes, spills to disk, input size changes, the number of partitions, and whether a broadcast join became a sort-merge join as a table grew.
  • “An executor runs out of memory.” Check partition sizes and skew, avoid collect() on large data, tune memory settings and partitions, and avoid caching more than needed.
  • “You see thousands of tiny output files.” Reduce partitions before writing (coalesce), or compact files later.

See also our data engineer interview questions, Kafka interview questions and Hadoop interview questions.

Frequently asked questions

What Spark questions are asked in interviews?
RDDs vs DataFrames, transformations vs actions, lazy evaluation, driver and executors, narrow vs wide transformations and shuffles, partitioning, caching, join strategies, data skew, Spark SQL optimization, Structured Streaming, and performance tuning scenarios.
What is lazy evaluation in Spark?
Transformations are not computed immediately; Spark records them and only runs them when an action needs a result, which lets it optimize the whole plan.
What causes a shuffle in Spark?
Wide transformations such as groupBy, join, distinct and repartition, which need data from many partitions and redistribute it across the cluster.
How do you handle data skew in Spark?
Broadcast the smaller side of a join, salt hot keys and aggregate in two steps, handle hot keys separately, or use adaptive query execution's skew-join handling.
Should I use RDDs or DataFrames?
Usually DataFrames, because Spark SQL can optimize them. RDDs are for low-level control or unstructured data where DataFrames don't fit.

Sources

Practice this with Huru AI

Turn what you just read into a rehearsal. Unlimited AI mock interviews with instant, specific feedback.

Amine Boukioud

Amine Boukioud is Head of Growth at Huru and writes Huru’s company interview guides. Each one starts from the company’s own careers pages and what candidates report, and follows Huru’s editorial policy.

Published · Updated

How this article was made: researched and drafted with AI assistance, then written, fact-checked and signed by Amine Boukioud on October 1, 2026. Sources are linked where they are used. How we work: editorial policy.