Data Quality

Lighthouse

Data quality as an API. One call from any pipeline.

42
built-in check templates across seven quality dimensions
0-100
quality score for every set, against thresholds you tune
1
API call to add a quality gate to any pipeline

The Problem

Every data team has quality checks; the problem is where they live, and how they are applied.

Scattered across pipeline scripts, rewritten for every engine, and invisible once the job finishes, they fail silently until bad data reaches a submission, a dashboard, or a downstream system, and by then the cleanup costs far more than the check ever would have.

What Lighthouse Does

Lighthouse is an API-driven data quality validation engine. Your pipelines call it mid-run, from Airflow, Dagster, Prefect, or any orchestrator you already use, and it checks the dataset against your rule library before the data moves on. Every rule returns green, yellow, or red, every batch gets a 0-100 quality score, and your pipeline gets a clear answer: proceed or stop.

The rules are modular and yours. Start with 42 built-in check templates spanning seven quality dimensions, then add your own domain packs, T-MSIS submissions, claims validation – whatever your business runs on – as configuration rather than code. Define a rule once and Lighthouse translates it to each data source: PostgreSQL, Snowflake, or files in your data lake.

Because every result is stored, quality becomes something you manage instead of something you react to. Dashboards track health scores by program and subject area, trends show whether quality is improving or drifting, and critical failures can raise tickets automatically. Lighthouse runs containerized in your cloud or on-prem, standalone or paired with DataLoom™.

Credentials

Containerized: runs in your cloud or on-prem
Role-based access control and full audit logging
PHI/PII-conscious logging and telemetry
CMS T-MSIS/NCQA DQ aligned
Works standalone or with DataLoom™

How It Fits

Your Pipeline. Your Rules. Our Quality Check.

The whole architecture in one picture. Follow a batch of data through it.

ANY ORCHESTRATORONE API CALLSUBMIT · POLL · RESULTSAIRFLOWDAGSTERPREFECTRULE LIBRARY+ YOUR RULE PACKSDATABASESWAREHOUSESCLOUD FILESCHECKED WHERE THE DATA LIVESBATCH SCORE94/100THRESHOLDS YOU TUNEPROCEEDSTOPRETURNED TO YOUR PIPELINE,BEFORE BAD DATA MOVES ONHEALTH SCORE TRENDSCORES BY PROGRAM · OVER TIMETOP FAILING RULESAUTO-TICKETED1 · YOUR PIPELINE CALLS2 · LIGHTHOUSE CHECKS3 · THE GATE4 · MONITOR

Your orchestrator, Airflow, Dagster, Prefect, or anything that can call an API, submits a validation run mid-pipeline. Submit, poll, fetch results. No new platform to adopt.

Lighthouse runs your rule library against the data where it already lives: databases, warehouses, and cloud file storage. Table to table, across databases, file to table, and file structure checks.

Every rule lands green, yellow, or red against thresholds you tune. The batch gets a 0 to 100 quality score, and your pipeline gets a clear answer: proceed or stop.

Every run builds history: health scores, trends over time, and the rules failing most, with incidents filed automatically in your ticketing system.

How It Works

Key Capabilities

  • One call from any pipelineAirflow, Dagster, Prefect, Azure Data Factory, or cron: submit a validation run, poll progress, fetch results, gate the batch.
  • Checks data where it livesPostgreSQL, Snowflake, and Azure Data Lake files (CSV, Parquet, JSON) today, with a growing connector library.
  • 42 templates, seven dimensionsCompleteness, uniqueness, integrity, validity, accuracy, consistency, and timeliness: from null checks to cross-database reconciliation.
  • Bring your own rule packsT-MSIS, claims, Veterans programs, omic information, rural health: your standards become templates and rules, added without a code deploy.
  • Thresholds with judgmentGreen, yellow, and red bands you tune per rule, with severity levels that decide what blocks a pipeline and what just files a ticket.
  • Quality you can watchHealth scores, trends over time, top failing rules, and incidents raised automatically in Jira, ServiceNow, or Azure DevOps.

How We Compare

Checks Buried in Pipeline Code vs. Lighthouse

Checks Buried in Pipeline CodeLighthouse
Validation logic scattered across ETL scriptsOne rule library, managed as data, versioned and auditable
Rewritten for every engine and SQL dialectDefine a rule once; Lighthouse translates it to each source
Results vanish into job logsEvery run scored 0-100, stored, and trended over time
Pass or fail, nothing in betweenGreen, yellow, red thresholds, from zero-tolerance to flexible
New checks mean another deploymentRules and domain packs added as configuration, no redeploy

Why It Wins

The Differentiator

Data quality tools usually arrive embedded in an ETL platform you must absorb, or as a framework your engineers wire into every job by hand. Lighthouse is a standalone API-first validation engine, so the pipelines you already run gain superior quality control without re-platforming, and your rules live in one governed library instead of a hundred scripts. Use it on its own, or let it power data quality inside DataLoom™.

Works Well With

Complementary Products

Data Intelligence

DataLoom™

Seven data capabilities in one platform. Zero fragmentation.

Learn more
CMS Compliance

TorQE

T-MSIS Outcome Review and Quality Engine. Catch errors before they reach CMS.

Learn more
Master Data Management

MDM Workbench

One trusted identity for every patient, provider, and member across your systems: real-time, explainable entity resolution with stewardship, analytics, and governance built in.

Learn more
By designAPI-FirstAny OrchestratorRules as Data

Put a Quality Gate in Every Pipeline.

A working session with the engineers behind Lighthouse: bring one pipeline, leave with a plan.