Lighthouse
Data quality as an API. One call from any pipeline.
The Problem
Every data team has quality checks; the problem is where they live, and how they are applied.
Scattered across pipeline scripts, rewritten for every engine, and invisible once the job finishes, they fail silently until bad data reaches a submission, a dashboard, or a downstream system, and by then the cleanup costs far more than the check ever would have.
What Lighthouse Does
Lighthouse is an API-driven data quality validation engine. Your pipelines call it mid-run, from Airflow, Dagster, Prefect, or any orchestrator you already use, and it checks the dataset against your rule library before the data moves on. Every rule returns green, yellow, or red, every batch gets a 0-100 quality score, and your pipeline gets a clear answer: proceed or stop.
The rules are modular and yours. Start with 42 built-in check templates spanning seven quality dimensions, then add your own domain packs, T-MSIS submissions, claims validation – whatever your business runs on – as configuration rather than code. Define a rule once and Lighthouse translates it to each data source: PostgreSQL, Snowflake, or files in your data lake.
Because every result is stored, quality becomes something you manage instead of something you react to. Dashboards track health scores by program and subject area, trends show whether quality is improving or drifting, and critical failures can raise tickets automatically. Lighthouse runs containerized in your cloud or on-prem, standalone or paired with DataLoom™.
Credentials
| ✓ | Containerized: runs in your cloud or on-prem |
|---|---|
| ✓ | Role-based access control and full audit logging |
| ✓ | PHI/PII-conscious logging and telemetry |
| ✓ | CMS T-MSIS/NCQA DQ aligned |
| ✓ | Works standalone or with DataLoom™ |
How It Fits
Your Pipeline. Your Rules. Our Quality Check.
The whole architecture in one picture. Follow a batch of data through it.
Your orchestrator, Airflow, Dagster, Prefect, or anything that can call an API, submits a validation run mid-pipeline. Submit, poll, fetch results. No new platform to adopt.
Lighthouse runs your rule library against the data where it already lives: databases, warehouses, and cloud file storage. Table to table, across databases, file to table, and file structure checks.
Every rule lands green, yellow, or red against thresholds you tune. The batch gets a 0 to 100 quality score, and your pipeline gets a clear answer: proceed or stop.
Every run builds history: health scores, trends over time, and the rules failing most, with incidents filed automatically in your ticketing system.
How It Works
Key Capabilities
- One call from any pipelineAirflow, Dagster, Prefect, Azure Data Factory, or cron: submit a validation run, poll progress, fetch results, gate the batch.
- Checks data where it livesPostgreSQL, Snowflake, and Azure Data Lake files (CSV, Parquet, JSON) today, with a growing connector library.
- 42 templates, seven dimensionsCompleteness, uniqueness, integrity, validity, accuracy, consistency, and timeliness: from null checks to cross-database reconciliation.
- Bring your own rule packsT-MSIS, claims, Veterans programs, omic information, rural health: your standards become templates and rules, added without a code deploy.
- Thresholds with judgmentGreen, yellow, and red bands you tune per rule, with severity levels that decide what blocks a pipeline and what just files a ticket.
- Quality you can watchHealth scores, trends over time, top failing rules, and incidents raised automatically in Jira, ServiceNow, or Azure DevOps.
How We Compare
Checks Buried in Pipeline Code vs. Lighthouse
| Checks Buried in Pipeline Code | Lighthouse |
|---|---|
| Validation logic scattered across ETL scripts | One rule library, managed as data, versioned and auditable |
| Rewritten for every engine and SQL dialect | Define a rule once; Lighthouse translates it to each source |
| Results vanish into job logs | Every run scored 0-100, stored, and trended over time |
| Pass or fail, nothing in between | Green, yellow, red thresholds, from zero-tolerance to flexible |
| New checks mean another deployment | Rules and domain packs added as configuration, no redeploy |
Why It Wins
The Differentiator
Data quality tools usually arrive embedded in an ETL platform you must absorb, or as a framework your engineers wire into every job by hand. Lighthouse is a standalone API-first validation engine, so the pipelines you already run gain superior quality control without re-platforming, and your rules live in one governed library instead of a hundred scripts. Use it on its own, or let it power data quality inside DataLoom™.
Works Well With
Complementary Products
TorQE
T-MSIS Outcome Review and Quality Engine. Catch errors before they reach CMS.
Learn moreMDM Workbench
One trusted identity for every patient, provider, and member across your systems: real-time, explainable entity resolution with stewardship, analytics, and governance built in.
Learn morePut a Quality Gate in Every Pipeline.
A working session with the engineers behind Lighthouse: bring one pipeline, leave with a plan.
