spark

Solid

Use when building distributed data pipelines with Apache Spark. Covers partitioning, shuffles, skew, joins, caching, and reading the Spark UI to find why a job is slow.

Data & Documents 26 stars 3 forks Updated 3 weeks ago MIT

Install

View on GitHub

Quality Score: 83/100

Stars 20%
48
Recency 20%
90
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
80
License 10%
100
Description 5%
100

Skill Content

# Spark ## Purpose Write Spark jobs whose cost is understood. Almost all Spark performance problems are one of three things: too much shuffle, skewed partitions, or reading far more data than the query needs. ## When to Use - Building or reviewing a Spark pipeline. - A job that is slow, failing with out-of-memory errors, or has one straggling task. - Tuning partitioning and join strategy. - Reading the Spark UI to diagnose a stage. ## Capabilities - Partitioning strategy and repartitioning. - Shuffle minimization and broadcast joins. - Skew detection and mitigation. - Caching and persistence levels. - File-format and predicate-pushdown optimization. - Spark UI interpretation. ## Inputs - The job, its input data volume, and its physical plan. - The Spark UI: stage timings, task distribution, shuffle read/write. - Cluster resources. ## Outputs - A plan with fewer or smaller shuffles. - Balanced partitions with no straggling tasks. - Measured improvement in wall-clock time and cost. ## Workflow 1. **Read the plan first** — `df.explain(True)`. Every `Exchange` is a shuffle, and a shuffle writes to disk and crosses the network. It is the dominant cost. 2. **Prune early** — Select the columns and filter the rows you need before joining, not after. With Parquet, this pushes down to the file reader and never reads the data at all. 3. **Broadcast the small side** — A join where one side fits in memory (roughly under 100 MB) should be a broadcast join. That eliminates the s...

Details

Author
nimadorostkar
Repository
nimadorostkar/Claude-Skills-collection
Created
1 months ago
Last Updated
3 weeks ago
Language
Python
License
MIT

Similar Skills

Semantically similar based on skill content — not just same category