Data Engineers build and maintain the infrastructure and architecture for data generation, transformation, and storage. They create robust data pipelines that enable data scientists and analysts to do their work effectively.
Just some of the skills Data Engineering professionals use. One may not need to know all of them, but one technology from each area would make you well-rounded.
- Python
- Java
- Scala
- Go
- SQL / Relational (PostgreSQL, MySQL, SQL Server)
- NoSQL (MongoDB, Cassandra, DynamoDB)
- Data Warehouses (Snowflake, Redshift, BigQuery)
- Graph databases (Neo4j) - not critically important
- SQL (Advanced)
- Query optimization
- Index management
- Apache Spark (PySpark, Scala)
- Apache Hadoop
- Pandas / Dask
- Apache Flink
- Apache Airflow
- Luigi
- Prefect
- Dagster
- dbt (data build tool)
- Apache Kafka
- Apache Flink
- AWS Kinesis
- Google Pub/Sub
- Apache Storm
- AWS: S3, Lambda, DynamoDB, Kinesis, Batch, EMR, Glue, Redshift
- GCP: BigQuery, Dataflow, Pub/Sub, Cloud Storage, Dataproc
- Azure: Data Factory, Synapse Analytics, Event Hubs, Databricks
- Serverless frameworks: AWS CDK, Serverless Framework, Chalice
- Dimensional modeling
- Star schema / Snowflake schema
- Data vault
- Normalization / Denormalization
- Git/GitHub
- GitLab CI/CD
- Jenkins
- GitHub Actions
- Docker
- Kubernetes
- Docker Compose
- Great Expectations
- dbt tests
- Unit testing for data pipelines
- CloudWatch
- Datadog
- Prometheus & Grafana
- ELK Stack (Elasticsearch, Logstash, Kibana)
- The Data Engineering Cookbook - Free comprehensive guide
- Data Engineering with Python
- Data Engineering Nanodegree by Udacity
- The Complete SQL Bootcamp by Jose Portilla
- SQL for Data Analysis
- Database Design and Basic SQL
- Apache Spark with Python by Jose Portilla
- Big Data Specialization by UC San Diego
- Spark and Scala for Big Data and Machine Learning
- dbt Fundamentals - Free official course
- Analytics Engineering with dbt
- "Designing Data-Intensive Applications" by Martin Kleppmann
- "The Data Warehouse Toolkit" by Ralph Kimball
- "Fundamentals of Data Engineering" by Joe Reis and Matt Housley
- "Data Pipelines Pocket Reference" by James Densmore
- "Streaming Systems" by Tyler Akidau, Slava Chernyak, and Reuven Lax
- "Spark: The Definitive Guide" by Bill Chambers and Matei Zaharia
- DataExpert.io - Data engineering bootcamp
- Seattle Data Guy YouTube Channel
- Work on personal ETL/ELT projects with public datasets
- Contribute to open-source data tools (Airflow, dbt, etc.)
- Data Engineering Weekly Newsletter
- The Data Engineering Podcast
- Reddit r/dataengineering
- dbt Community
- Locally Optimistic
- Data Engineering on Medium
- SQL proficiency
- Understanding of ETL concepts
- Basic Python or Java
- Version control (Git)
- Cloud platform basics
- Design and build data pipelines
- Experience with orchestration tools (Airflow)
- Cloud platform expertise
- Data modeling skills
- Performance optimization
- Architecture design
- Mentoring junior engineers
- Complex data pipeline development
- Cost optimization
- Cross-team collaboration
- Technical leadership
- System architecture
- Strategic planning
- Team management
- Technology evaluation and adoption
- ETL Pipeline: Build an end-to-end ETL pipeline using Airflow
- Data Warehouse: Design and implement a data warehouse with dimensional modeling
- Streaming Pipeline: Create a real-time data processing pipeline with Kafka and Spark
- Cloud Data Lake: Build a data lake on AWS S3 or GCP Cloud Storage
- dbt Project: Transform raw data into analytics-ready datasets using dbt
- API Integration: Build pipelines to extract data from various APIs
- Data Quality Framework: Implement data quality checks and monitoring
- AWS: AWS Certified Data Analytics - Specialty
- GCP: Google Cloud Professional Data Engineer
- Azure: Azure Data Engineer Associate
- Databricks: Databricks Certified Data Engineer
- Snowflake: SnowPro Core Certification
- Apache: Databricks Apache Spark Developer Certification
- SQL queries (joins, aggregations, window functions)
- Data modeling (star schema, normalization)
- System design for data pipelines
- Big data technologies (Spark, Hadoop)
- Cloud services and architecture
- Data pipeline optimization
- LeetCode SQL problems
- HackerRank SQL challenges
- Python data structure problems
- Design a data warehouse
- Design a real-time analytics system
- Design an ETL pipeline for specific use case
- Stay Current: Follow new tools and technologies in the data engineering space
- Build Portfolio: Maintain GitHub with data engineering projects
- Network: Join data engineering communities and attend meetups
- Learn from Others: Study architecture of companies like Uber, Airbnb, Netflix
- Documentation: Practice writing clear technical documentation
- Cost Awareness: Understand cloud cost optimization techniques
- Security: Learn about data security and compliance (GDPR, HIPAA)