pyiceberg
Here are 18 public repositories matching this topic...
Lakevision is a tool which provides insights into your Apache Iceberg based Data Lakehouse.
-
Updated
Apr 11, 2026 - Python
Sample code to collect Apache Iceberg metrics for table monitoring
-
Updated
Aug 18, 2024 - Python
A poc open framework to manage data ingestion into apache iceberg tables
-
Updated
Sep 11, 2024 - Python
Official companion repository for the "Unleash the power of Apache Iceberg on AWS" Book. Includes code samples, deployment guides, ETL examples, and best practices for implementing scalable Apache Iceberg lakehouses on AWS.
-
Updated
Nov 26, 2025 - Jupyter Notebook
Tansu schema-backed topics, instantly accessible as Apache Iceberg tables with pyiceberg
-
Updated
May 31, 2025 - Just
-
Updated
Jan 23, 2026 - Go
Model Context Protocol (MCP) server for Apache Polaris. Enables AI agents and LLMs to interact with Polaris Catalog Management, audit data governance and access control (RBAC) policies, and perform Iceberg table inspections using PyIceberg.
-
Updated
Jul 5, 2026 - Python
A low-cost online Apache Iceberg lab using Cloudflare R2 Data Catalog
-
Updated
Jul 11, 2026 - Python
Apache Iceberg with PyIceberg (no Spark/JVM): catalog setup, time travel, CDC, schema evolution, upserts
-
Updated
May 1, 2026 - Python
Apache Polaris + PyIceberg で Iceberg REST Catalog を動かして確かめた実験記録(2026年7月時点)
-
Updated
Jul 21, 2026 - Shell
provenance and immutability checks for iceberg lakehouses
-
Updated
May 21, 2026 - Python
Apache Iceberg feature store and data lakehouse on Backblaze B2 — versioned feature tables, schema evolution, snapshot rollback, and DuckDB time-travel queries via PyIceberg, with B2 as the warehouse (no data warehouse to run)
-
Updated
Jun 25, 2026 - TypeScript
Postgres CDC → Kafka → Apache Iceberg with exactly-once, out-of-order merge (source-LSN high-water-mark) and a source-to-lakehouse reconciler that quantifies drift in dollars. Ships a Spark Structured Streaming job, a docker-compose CDC stack, and Terraform.
-
Updated
Jul 8, 2026 - Python
Locally reproducible Apache Iceberg lakehouse with read-only FastAPI/DuckDB serving, near-real-time micro-batch ingestion, and a read-only NL agent over the API. Containerized with CI/CD, single-host go-live (Caddy auto-HTTPS), and OpenTelemetry→Datadog observability. Live synthetic demo: https://demo.agentic-data-platform.me/docs
-
Updated
Jul 11, 2026 - Python
Provides a pyiceberg.io.FileIO implementation that uses hdfs-native client.
-
Updated
Aug 27, 2025 - Python
Data Lakehouse implementations showcasing two approaches: one built with Open Source technologies and another with Databricks
-
Updated
Sep 13, 2025 - Jupyter Notebook
Improve this page
Add a description, image, and links to the pyiceberg topic page so that developers can more easily learn about it.
Add this topic to your repo
To associate your repository with the pyiceberg topic, visit your repo's landing page and select "manage topics."