Examples of Apache Flink® v2.1 applications showcasing the DataStream API, Table API in Java and Python, and Flink SQL, featuring AWS, GitHub, Terraform, Streamlit, and Apache Iceberg.
-
Updated
Jan 13, 2026 - Java
Examples of Apache Flink® v2.1 applications showcasing the DataStream API, Table API in Java and Python, and Flink SQL, featuring AWS, GitHub, Terraform, Streamlit, and Apache Iceberg.
Automation framework to catalog AWS data sources using Glue
Smart City Realtime Data Engineering Project
A CLI tool to back up and restore AWS Glue catalog resources such as Databases, Tables, and Connections as JSON files. Useful when you don't have AWS Backup or versioning enabled in your account.
This project repo 📺 offers a robust solution meticulously crafted to efficiently manage, process, and analyze YouTube video data leveraging the power of AWS services. Whether you're diving into structured statistics or exploring the nuances of trending key metrics, this pipeline is engineered to handle it all with finesse.
Tool to migrate Delta Lake tables to Apache Iceberg using AWS Glue and S3
It is a project build using ETL(Extract, Transform, Load) pipeline using Spotify API on AWS.
Quality_Movie_Data_Analysis
This project demonstrates how to use Terraform to enable Tableflow in Kafka to generate and store the Iceberg Table files in an AWS S3 bucket. Then, configure Snowflake to read the Iceberg Tables using AWS Glue Data Catalog and the AWS S3 bucket where Tableflow produces the Iceberg files.
Example using the Iceberg register_table command with AWS Glue and Glue Data Catalog
Creating an audit table for a DynamoDB table using CloudTrail, Kinesis Data Stream, Lambda, S3, Glue and Athena and CloudFormation
End-to-end AWS data analytics pipeline for product risk detection and customer dissatisfaction analysis.
Working with Glue Data Catalog and Running the Glue Crawler On Demand
Enterprise track: Step Functions/EventBridge + Glue + data quality on top of the v1 serverless ELT
End-to-end YouTube Data Engineering Pipeline built on AWS using Amazon S3, Lambda, Glue, Athena, and Python. Automated ETL transforms raw CSV/JSON data into Parquet for analytics, with interactive insights delivered through Power BI.
Unveiling job market trends with Scrapy and AWS
Engaging, interactive visualizations crafted with Streamlit, seamlessly powered by Apache Flink in batch mode to reveal deep insights from data.
End to end Superstore Data Analysis using AWS services
A serverless, medallion-architecture data pipeline on AWS that ingests YouTube trending video data, cleans and enriches it through Bronze → Silver → Gold layers, and produces analytics-ready tables for querying in Athena (and, downstream, BI tools like QuickSight).
End-to-end AWS Data Pipeline designed to ingest, cleanse, and transform daily trending YouTube video data using Medallion Architecture (Bronze, Silver, Gold). Built with AWS S3, PySpark Glue ETL, Lambda, and Athena for automated validation, serverless reporting, and high-performance SQL analytics querying.
To associate your repository with the aws-glue-data-catalog topic, visit your repo's landing page and select "manage topics."