# Apache Spark

> Unified analytics engine for large-scale data

- Category: [Data Engineering](https://ailandscape.org/category/data-engineering) › Data Processing
- Homepage: https://spark.apache.org
- Repository: https://github.com/apache/spark
- Crunchbase: https://www.crunchbase.com/organization/apache
- Tags: batch, distributed, python, scala
- Added to the landscape: 2026-03-18

## Similar tools in Data Processing

- [Apache Flink](https://ailandscape.org/tool/apache-flink): Stateful stream processing framework
- [Apache Hadoop](https://ailandscape.org/tool/apache-hadoop): Distributed storage and processing framework
- [Apache Kafka](https://ailandscape.org/tool/apache-kafka): Distributed event streaming platform
- [Apache Pulsar](https://ailandscape.org/tool/apache-pulsar): Cloud-native distributed messaging and streaming
- [dbt](https://ailandscape.org/tool/dbt): Data transformation tool for analytics engineers
- [Redpanda](https://ailandscape.org/tool/redpanda): Kafka-compatible streaming platform without ZooKeeper

---

Part of [AI Landscape](https://ailandscape.org), an open map of the AI ecosystem. Web page: https://ailandscape.org/tool/apache-spark · Index for AI assistants: https://ailandscape.org/llms.txt
