Skip to content

Navigation Menu

Sign in
Appearance settings

Search code, repositories, users, issues, pull requests...

Provide feedback

We read every piece of feedback, and take your input very seriously.

Saved searches

Use saved searches to filter your results more quickly

Appearance settings
This repository was archived by the owner on Jul 26, 2026. It is now read-only.
Open more actions menu

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

4,555 Commits
4,555 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

⚠️ Archived — This repository is no longer maintained and will not receive updates, including security patches. It is preserved in read-only form for reference.

Overview Build Status Maven Central

Herd is big data governance for the cloud. The herd unified data catalog helps separate compute from storage in the cloud. Herd job orchestration manages your ETL and analytics processes while tracking all data in the catalog. Here is a quick summary of features:

  • Unified Data Catalog A centralized, auditable catalog for operational usage and data governance.
  • Track Lineage Capture data ancestry for regulatory, forensic, and analytical purposes
  • Manage Clusters Create and launch clusters; load data into clusters from catalog entries
  • Orchestrate Jobs Orchestrate clusters and catalog services to automate processing jobs

Find out more about herd features on our GitHub project page

Quick Start

The best way to start learning about herd is through these links. The demo installation process is quick and easy - you can have herd up and running in AWS in 10-15 minutes and start registering data immediately afterwards.

Get Involved

We are actively seeking organizations and individuals that are interested in adopting herd and contributing to the development effort. Find out more in the contributions section of our GitHub project page. If you have any questions or discussion topics, post them on GitHub Issues or email us at herd@finra.org.

License

Herd is licensed under Apache License 2.0

About

⚠️ Archived — This repository is no longer maintained and will not receive updates. Herd is a managed data lake for the cloud. The Herd unified data catalog helps separate storage from compute in the cloud. Manage petabytes of data and make it accessible for data processing and analytical purposes by any cloud compute platform.

Resources

Stars

Watchers

Forks

Releases

Packages

Used by

Contributors

Languages

Morty Proxy This is a proxified and sanitized view of the page, visit original site.