Skip to content
View diegobarbosaa's full-sized avatar

Block or report diegobarbosaa

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
diegobarbosaa/README.md

Hi there, I'm Diego Vanderlei πŸ‘‹

Data Engineer | Cloud Data Architecture | Data Governance & Modern Data Stack

πŸ“ Olinda, Brazil β€’ 🌐 Open to Worldwide & Canada Opportunities

LinkedIn ProtonMail Power BI Dashboard Juliane Bezerra ImΓ³veis


πŸš€ About Me

I am a Data Engineer with over 6 years of experience operating at the intersection of Cloud Architecture, Data Governance, and Scalable ETL/ELT Pipelines.

  • πŸ›οΈ Proven Institutional Impact: Led database modernization, architected AWS Medallion Lakehouse environments (Bronze, Silver, Gold), and built enterprise-wide analytical solutions for the FIEPE System (SENAI/SESI/IEL).
  • πŸ›‘οΈ Data Governance & Quality: Hands-on implementation of metadata cataloging, data lineage, and data quality frameworks using OpenMetadata.
  • βš™οΈ Modern Data Engineering: Experienced in distributed processing (Apache Spark / PySpark), open table formats (Apache Iceberg + Nessie), infrastructure as code (Terraform), and workflow orchestration (Apache Airflow 2 & 3).

πŸ› οΈ Tech Stack & Tooling

Domain Technologies & Frameworks
Languages & Core Python SQL PySpark T-SQL
Cloud & Lakehouse AWS Azure Apache Iceberg MinIO Databricks
Orchestration & DevOps Apache Airflow Docker Terraform GitHub Actions
Governance & Analytics OpenMetadata Power BI DataOps

🌟 Featured Repositories & Projects

🏒 Enterprise Data Integration Platform

End-to-end data platform on Microsoft Azure implementing Medallion Architecture (Bronze β†’ Silver β†’ Gold), Azure Functions, ADF, Service Bus, and Terraform IaC.

Azure β€’ ADF β€’ Python β€’ Terraform β€’ Zero Trust

🧊 Local Lakehouse (Iceberg + Nessie + MinIO)

Local modern Data Lakehouse environment using Apache Iceberg table format with ACID transactions, time travel, Project Nessie REST catalog, and MinIO S3 storage.

Apache Iceberg β€’ PyIceberg β€’ Nessie β€’ MinIO β€’ Docker

⚑ 1 Billion Row Challenge (Python)

High-performance parsing and aggregation of 1,000,000,000 temperature records (~14 GB) in Python using memory-efficient parallel processing techniques.

View Repository β†’

☁️ AWS Serverless Data Pipeline

Event-driven data pipeline leveraging AWS S3, SQS, Lambda, AWS Glue (ETL & Data Catalog), and querying with Amazon Athena.

View Repository β†’


πŸ“œ Certifications & Achievements

  • πŸ… Astronomer β€” DAG Authoring for Apache Airflow 3 (2026)
  • πŸ… Databricks β€” Databricks Fundamentals (2025)
  • πŸ… Astronomer β€” Apache Airflow 2 Fundamentals (2025)
  • πŸ… AWS Academy β€” Data Engineering (2024)
  • πŸ… Microsoft β€” Certified: Azure Data Fundamentals (DP-900) (2022)
  • πŸ… CertiProf β€” General Data Protection Regulation / LGPD (2021)

πŸ’¬ "Transforming data complexity into clear, scalable, and governed business value."

πŸ“« Let's connect on LinkedIn or via Email!

Pinned Loading

  1. One-Billion-Row-Challenge-Python One-Billion-Row-Challenge-Python Public

    High-performance processing and aggregation of 1 billion rows (~14GB) in Python using memory efficiency, chunking, and parallel computing techniques.

    Python 1

  2. de-dataops de-dataops Public

    Complete DataOps pipeline featuring Docker containerization, Python ETL, MySQL storage, and automated CI/CD workflows via GitHub Actions.

    Python

  3. enterprise-data-integration-platform enterprise-data-integration-platform Public

    Enterprise Data Platform on Azure: Medallion Architecture (Bronze/Silver/Gold), Azure Functions, ADF, Service Bus, Azure SQL, Terraform IaC & Zero Trust.

    Python

  4. graphFrames graphFrames Public

    Distributed data analysis and graph processing with Apache Spark, PySpark, Spark SQL, and Spark GraphFrames.

    Jupyter Notebook

  5. lakehouse-iceberg-minio lakehouse-iceberg-minio Public

    Local Modern Data Lakehouse with Apache Iceberg, Project Nessie REST catalog, MinIO S3 storage, PyIceberg & Docker Compose.

    Python

  6. pipeline_AWS pipeline_AWS Public

    Serverless data pipeline on AWS leveraging S3, SQS, Lambda, AWS Glue (ETL & Data Catalog), and analytical querying with Amazon Athena.