Try it
Add the skill to a bot, then ask your Chief of Staff:
“Use the Spark Engineer skill on this: [describe the job, or paste your notes].”
Apache Spark is the de facto standard for large-scale distributed data processing. This skill covers the internals, optimization techniques, and best practices needed to write Spark applications that are both correct and performant at terabyte-to-petabyte scale.
What it covers
- API Comparison: RDD vs DataFrame vs Dataset
- Partitioning Strategies
- Shuffle Optimization
- Broadcast Joins
- Caching and Persistence
- Spark SQL
- Structured Streaming
- UDFs: When and How