Databricks framework to validate Data Quality of pySpark DataFrames and Tables
-
Updated
Jul 27, 2026 - Python
Databricks framework to validate Data Quality of pySpark DataFrames and Tables
Metadata-driven framework for Databricks Spark Declarative Pipelines. Config-driven, pattern based approach to batch & streaming across the medallion architecture. Deploys via Declarative Automation Bundles. Built for simplicity, extensibility, and alignment with the Databricks product roadmap.
Medallion Architecture for Data Engineering projects
Open-source community study guide for all six Databricks certifications (Data Engineer Associate / Professional, Data Analyst Associate, ML Associate / Professional, GenAI Engineer Associate). Aligned to the 2025-2026 official exam guides. Obsidian-flavoured Markdown; PRs welcome.
Medallion analytics for the Ceres Open Data Index on Databricks — Lakeflow Declarative Pipelines (Bronze → Silver → Gold), runs on Free Edition.
Databricks-native data trust pipeline — intake certification, drift gating, and control benchmarking in a single deployable product.
Databricks SQL in Action — End-to-end medallion architecture lab using Unity Catalog, Volumes, Streaming Tables, Materialized Views, AI SQL functions, dashboards, lineage, and workflow orchestration.
Built an end-to-end retail lakehouse using the Medallion architecture with batch and streaming pipelines for scalable business analytics.
Demo of Databricks Lakeflow Jobs Automation with StackQL and Databricks Asset Bundles
Bring your Claude Code skill unchanged and run it as a governed Databricks job. Publish once to a Unity Catalog volume, reuse from any job, chain skills into a pipeline (markdown in, branded PowerPoint out). No external API key. Runs on Free Edition.
Databricks Lakeflow Jobs Demo This repository demonstrates the core features of Databricks Lakeflow Jobs through practical examples. It is designed to help data engineers understand how to build, configure, and orchestrate production-ready workflows using Lakeflow Jobs.
Lakeflow Designer pipeline + AI/BI dashboard turning session, attendance, and feedback logs into governed session-performance analytics
Sample Databricks Asset Bundle: hotel daily performance KPIs with Lakeflow SDP, UC Metric Views, AI/BI dashboard (Brickstar styled), and Genie NL→SQL
Tsuga Logs community connector for Databricks Lakeflow Connect — scheduled log ingestion into Delta tables you own
Metadata-driven pipeline orchestration pattern on Databricks — dynamic DAG from a Delta control table, Jobs API v2.2, Unity Catalog, Lakeflow. Deployable via Databricks Asset Bundle.
End-to-end Azure Databricks Data Engineering Pipeline with Medallion Architecture, Delta Lake, Unity Catalog, and Lakeflow Jobs.
PayGuard Streaming Fraud Detection is an end-to-end real-time fraud monitoring project that simulates card transactions with Python and Kafka, processes them through Databricks Bronze/Silver/Gold pipelines, generates fraud and high-value transaction alerts, sends email notifications, and visualizes insights in a Lakeview dashboard.
Enterprise-pattern Databricks lakehouse — Unity Catalog, Auto Loader, Lakeflow Declarative Pipelines, native SCD2, one-bundle deploy.
Event-driven lakehouse on Databricks with real CDC from PostgreSQL via Debezium — medallion architecture, Lakeflow Declarative Pipelines, SCD, CDF, and liquid clustering. Infra as code with Terraform.
Add a description, image, and links to the lakeflow topic page so that developers can more easily learn about it.
To associate your repository with the lakeflow topic, visit your repo's landing page and select "manage topics."