A Data & Software Engineer with a B.S. in Computer Science (UFOP), focused on building scalable ETL/ELT pipelines, robust backend systems, and deterministic AI-assisted workflows. With production experience across AWS and GCP environments, I bridge data engineering, machine learning, and DevOps to deliver reliable, high-impact data products.
- Languages: Python, SQL, TypeScript, Bash.
- Data Engineering: Apache Airflow, dbt, Apache Spark, DuckDB, MongoDB, Medallion Architecture.
- Cloud Platforms: AWS (Redshift, Lambda, S3, Glue, SQS), GCP (BigQuery, GCS).
- Software Engineering & AI: FastAPI, Next.js, Autonomous Agents, RAG Systems, Scikit-learn, NLP (SciBERT).
- DevOps & Infra: Terraform, Docker, GitHub Actions (CI/CD), Arch Linux (WSL).
- Data Engineer / DevOps @ Beegol: Automated CI/CD pipelines via GitHub Actions, architected data schema validations using Pandera, and maintained high-availability data ingestion infrastructure.
- Data Engineer @ Booster / Rabbot: Designed and orchestrated complex ETL/ELT pipelines using Airflow and dbt across AWS (Redshift) and GCP (BigQuery), integrating relational and NoSQL (MongoDB) sources.
- Previous Roles: Data Engineering Intern @ Media.Monks (GCP pipelines & Airflow) and Data Intern @ Enacom Group (Machine Learning PoCs & EDA).
- Pokémon AI Lakehouse & ML Pipeline: A comprehensive multi-layer lakehouse featuring automated DLT ingestion, PySpark processing, and scikit-learn model serving.
- Olist Lakehouse: Robust cloud infrastructure managed via Terraform (VPC, EC2, RDS, AWS DMS) paired with PySpark Iceberg merges for advanced data modeling.
- Financial Lake Serverless: A cloud-native, event-driven data lake built on AWS (Lambda, SQS, S3, Glue) and entirely provisioned via Terraform for real-time financial tracking.
- Brasileirão 2026 Pipeline: An end-to-end system combining custom Python web scraping, Apache Airflow orchestration, and dbt models for sports analytics.
- TalkDoc: A modular Retrieval-Augmented Generation (RAG) application demonstrating full-stack engineering, combining a Next.js frontend with a robust FastAPI backend for context-aware document interactions.
- AI-Assisted Development Infrastructure: A strictly parameterized Zsh router (
zsh-ai-workflow) and an architectural template designed to orchestrate autonomous AI agents and manage cross-agent context handoffs directly from the terminal.
- SLSDT: A published, scikit-learn compatible Python package for inducing oblique decision trees using Stochastic Local Search (LAHC).
- SciSci-BERT: Academic research applying SciBERT language models and HDBSCAN clustering for semantic space validation and collaboration network analysis in Brazilian scientific publications.
Feel free to explore my repositories or reach out for engineering collaborations. I am always open to discussing complex challenges across Data Engineering, AI/ML systems, and Software Architecture.
