Data Processing & Pipelines

6 tools

Apache Spark

Why I use it: The industry standard for large-scale data processing. MLlib for distributed ML training.

4.3
dbt

Why I use it: Transform data in your warehouse with SQL. Essential for preparing ML training data at scale.

4.2
Databricks

Why I use it: Managed Spark with lakehouse architecture. The premium data + ML platform with MLflow built in.

4.6
Airflow

Why I use it: Industry-standard workflow orchestration. Define ML data pipelines as Python DAGs.

4.4
Dagster

Why I use it: Modern Airflow alternative with software-defined assets and built-in testing. Better developer experience.

4.2
Prefect

Why I use it: The most Pythonic workflow orchestrator. Easy to start with, powerful at scale.

4.5