data-profiling

Here are 75 public repositories matching this topic...

open-metadata / OpenMetadata

OpenMetadata is a unified platform for discovery, observability, and governance powered by a central metadata repository, in-depth lineage, and seamless team collaboration.

Updated May 28, 2024
TypeScript

DataKitchen / data-observability-installer

Star

Installer for DataKitchen's Open Source Data Observability Products. Data breaks. Servers break. Your toolchain breaks. Ensure your team is the first to know and the first to solve with visibility across and down your data estate. Save time with simple, fast data quality test generation and execution. Trust your data, tools, and systems end to end.

Updated May 28, 2024
Python

great-expectations / great_expectations

Star

Always know what to expect from your data.

Updated May 28, 2024
Python

polyaxon / haupt

Star

Lineage metadata API, artifacts streams, sandbox, API, and spaces for Polyaxon

Updated May 28, 2024
Python

dqops / dqo

Star

Data Quality and Observability platform for the whole data lifecycle, from profiling new data sources to full automation with Data Observability. Configure data quality checks from the UI or in YAML files, let DQOps run the data quality checks daily to detect data quality issues.

monitoring data-quality-checks data-quality data-profiling data-ops data-quality-measurement data-quality-monitoring data-quality-report data-observability

Updated May 28, 2024
Java

opendatadiscovery / odd-platform

Star

First open-source data discovery and observability platform. We make a life for data practitioners easy so you can focus on your business.

Updated May 28, 2024
Java

sodadata / soda-core

Star

⚡ Data quality testing for the modern data stack (SQL, Spark, and Pandas) https://www.soda.io

Updated May 28, 2024
Python

Desbordante / desbordante-core

Star

Desbordante is a high-performance data profiler that is capable of discovering many different patterns in data using various algorithms. It also allows to run data cleaning scenarios using these algorithms. Desbordante has a console version and an easy-to-use web application.

data-science data-mining exploratory-data-analysis tabular-data feature-selection data-engineering feature-extraction data-analytics knowledge-discovery data-wrangling data-preprocessing feature-engineering spreadsheets data-exploration data-mining-algorithms data-cleaning data-profiling anomaly-detection data-cleansing correlations

Updated May 28, 2024
C++

cleanlab / cleanlab

Star

The standard data-centric AI package for data quality and machine learning with messy, real-world data and labels.

Updated May 28, 2024
Python

open-metadata / openmetadata-site

Star

Open Standard for Metadata. A Single place to Discover, Collaborate and Get your data right.

Updated May 23, 2024
CSS

ptypes-nlesc / data-profiling

Star

Data preparation and exploration scripts

tdd test-driven-development data-cleaning data-profiling tag-analysis

Updated May 24, 2024
Python

ydataai / ydata-profiling

Star

1 Line of code data quality profiling & exploratory data analysis for Pandas and Spark DataFrames.

Updated May 23, 2024
Python

cleanlab / cleanlab-studio

Star

Client interface for all things Cleanlab Studio

Updated May 23, 2024
Python

SJTU-Quant / awesome-ml-data-quality-papers

Star

Papers about training data quality management for ML models.

machine-learning data-management data-quality data-profiling data-debugging data-valuation data-centric-ai ai4db db4ai

Updated May 21, 2024

tsegall / fta

Star

Metadata/data identification Java library. Identifies Semantic Type information (e.g. Gender, Age, Color, Country,...). Extensive country/language support. Extensible via user-defined plugins. Comprehensive Profiling support.

java metadata date data-discovery data-profiling semantic-types semantic-typechecking semantic-type-detection data-profiler