Databricks framework to validate Data Quality of pySpark DataFrames and Tables
-
Updated
Sep 10, 2026 - Python
Databricks framework to validate Data Quality of pySpark DataFrames and Tables
Metadata-driven framework for Databricks Spark Declarative Pipelines. Config-driven, pattern based approach to batch & streaming across the medallion architecture. Deploys via Declarative Automation Bundles. Built for simplicity, extensibility, and alignment with the Databricks product roadmap.
Open-source community study guide for all six Databricks certifications (Data Engineer Associate / Professional, Data Analyst Associate, ML Associate / Professional, GenAI Engineer Associate). Aligned to the 2025-2026 official exam guides. Obsidian-flavoured Markdown; PRs welcome.
Medallion Architecture for Data Engineering projects
Medallion analytics for the Ceres Open Data Index on Databricks — Lakeflow Declarative Pipelines (Bronze → Silver → Gold), runs on Free Edition.
Databricks-native data trust pipeline — intake certification, drift gating, and control benchmarking in a single deployable product.
Bring your Claude Code skill unchanged and run it as a governed Databricks job. Publish once to a Unity Catalog volume, reuse from any job, chain skills into a pipeline (markdown in, branded PowerPoint out). No external API key. Runs on Free Edition.
Databricks SQL in Action — End-to-end medallion architecture lab using Unity Catalog, Volumes, Streaming Tables, Materialized Views, AI SQL functions, dashboards, lineage, and workflow orchestration.
Built an end-to-end retail lakehouse using the Medallion architecture with batch and streaming pipelines for scalable business analytics.
Hands-on Azure Databricks learning project — medallion pipeline, Lakeflow SDP, Delta Lake, Unity Catalog security, SQL analytics and Jobs. Built end-to-end on Azure with real executed notebooks.
Demo of Databricks Lakeflow Jobs Automation with StackQL and Databricks Asset Bundles
Sample Databricks Asset Bundle: hotel daily performance KPIs with Lakeflow SDP, UC Metric Views, AI/BI dashboard (Brickstar styled), and Genie NL→SQL
Built a metadata-driven Azure lakehouse that incrementally ingests Spotify-style SQL data with ADF, processes new Parquet files with Databricks Auto Loader, and publishes SCD-managed Delta facts and dimensions.
End-to-end NYC Taxi data engineering pipeline using Azure Databricks, Apache Spark, Delta Lake, ADLS Gen2, Terraform, and Medallion Architecture
Metadata-driven schemaless MongoDB Debezium CDC ingestion into a Unity Catalog medallion lakehouse on Databricks Lakeflow Declarative Pipelines, using the VARIANT data type
Tsuga Logs community connector for Databricks Lakeflow Connect — scheduled log ingestion into Delta tables you own
Built a real-time Azure lakehouse that streams synthetic ride-booking events from FastAPI through Event Hubs into Databricks, unifies them with historical data, and publishes SCD-managed facts and dimensions.
Built an OAuth-authenticated Spark streaming lakehouse that consumes NASA GCN Fermi gamma-ray burst notices from Kafka, parses classic-text messages, and publishes a governed analytical snowflake schema.
Event-driven lakehouse on Databricks with real CDC from PostgreSQL via Debezium — medallion architecture, Lakeflow Declarative Pipelines, SCD, CDF, and liquid clustering. Infra as code with Terraform.
To associate your repository with the lakeflow topic, visit your repo's landing page and select "manage topics."