An enterprise healthcare data-archival provider had a problem its own team described plainly: hundreds of millions of legacy records — structured rows and unstructured blob storage — and not enough visibility into what was actually in them. As a subcontractor to a nearshore engineering services prime, StandardData designed and built a proof-of-concept ingestion pipeline that answered that problem directly. A working POC was standing by the end of the kickoff week. The full end-to-end demo followed within about three weeks, and was presented to the client's board, well received.
This was a paid discovery engagement: prove, in a working artifact rather than a slide deck, that ingestion, schema mapping, validation, and data discovery for a legacy healthcare archive could be automated and made observable. The deliverable was real code, handed cleanly to the prime's engineering team for productization. The notes below are anonymized — the client and the prime are unnamed, and no commercial figures appear — but the technical story is StandardData's own.
The problem — "we don't know what's in this data"
The archive draws from many disparate source systems, the way legacy healthcare records always do. On the live showcase, the client framed the core risk in their own words: hundreds of millions of rows in one place, blob storage in another, and limited confidence about the contents or the ownership of either. That is not a tooling preference. For a regulated archive, not knowing what is in your data is an operational and compliance risk, and the team said so directly.
Legacy ingestion made it worse. Validation and schema mapping happened late — at the end of the pipeline, after records had already moved downstream. When something was missing or malformed, the cost of finding out came after the work, not before it. The mandate, set during scoping, was to invert that: catch data-quality problems at intake, and replace bespoke per-source mapping code with something that learns the shape of the data.
What we built
Front-loaded validation. We moved validation and schema mapping to the very front of the pipeline, so malformed or incomplete records are rejected — with a reason — before any downstream processing. On the showcase, the benefit was stated plainly: doing it up front automatically saves a lot of time, but it also tells you right off the bat if something is missing. That is the whole point of front-loading: the pipeline fails fast and fails legibly, instead of failing quietly three steps later.
AI-assisted schema mapping. Instead of mapping each source field to the target schema by hand, the pipeline scores source fields against the target schema using sentence-transformer vector embeddings and cosine similarity. High-confidence matches map automatically; low-confidence fields fall to a human-in-the-loop review threshold. The manual, field-by-field labor that dominates a legacy migration becomes review-and-approve instead of write-from-scratch.
Automated failure alerting. When an ingest fails, the system emails the exact missing required fields — not a generic error. We demoed it live: an alert that arrived at 8:41am listing the specific gaps, including a missing medical number, address, zip code, and enterprise medical number. The operator does not have to go hunting for the cause; the cause comes to them.
Elasticsearch observability. Ingested data lands in Elasticsearch as the data layer, which gives per-field statistics, distributions, missing-field detection, and anomaly detection across both structured and unstructured inputs. This is the part that answers the "we don't know what's in this data" problem head-on: instead of guessing, the team can see the shape of the archive. It also sets up what comes next — document OCR and a semantic-search layer over archived records.
- Validation-first ingestion: malformed or incomplete records are rejected with a reason at the front of the pipeline, not after downstream rework.
- Embeddings-based schema mapping: cosine-similarity scoring against the target schema, with a confidence threshold and human-in-the-loop review for the uncertain cases.
- Self-explaining failures: a failed ingest auto-emails the specific missing required fields, demoed live during the client showcase.
- Observability over heterogeneous data: Elasticsearch surfaces per-field statistics, distributions, and anomalies across structured and unstructured records.
Why this approach is different — four reasons
Validation at the front, not the end
Missing or invalid data is caught at intake, with a reason, before any downstream processing — instead of surfacing after the work is already done.
AI mapping replaces manual field work
Sentence-transformer embeddings and cosine similarity score each source field against the target schema. A confidence threshold routes only the uncertain cases to a human.
Failures that explain themselves
A failed ingest auto-emails the exact missing required fields. The operator gets the cause delivered to them — no log archaeology.
Observability over what you already hold
Elasticsearch surfaces per-field statistics, distributions, and anomalies across structured and unstructured data — directly answering "we don't know what's in this data."
From kickoff to the board
The schedule is the part we are most comfortable putting numbers to, because it comes straight from our own delivery record. A working ingestion POC was stood up during the kickoff week. A full working end-to-end demo notebook existed within about three weeks. The proof of concept was then presented to the client's board, and reported back as well received.
We are deliberate about the scope of those numbers. The demonstrated POC ran against synthetic ambulatory data to prove the approach end to end. We are not claiming a production migration here, and we have left out the throughput and data-quality percentages that the POC did not actually measure. What we are claiming is what the record supports: a working, observable, AI-assisted ingestion pipeline, stood up fast, and good enough to take in front of a board.
What the client said
The reaction we value most was the one on the live showcase, from the client's engineering lead — captured because the client recorded the session to share with their leadership.
It's clear you guys were listening. Very clear you guys are listening. I think it's fantastic. It's got my creative juices running.Client engineering lead, live POC showcase (recorded by the client to share with leadership), 31 January 2025.
The team also named where they expected the value to land. With a consistent way to discover and identify records, people would stop hunting and pecking to figure out who owns what data — which, in their words, was a real risk. That is a client-stated potential benefit, not a measured saving, and we frame it that way: the discoverability layer is built to remove that hunting, and the client saw the payoff before we did.
Why this travels beyond one archive
The pattern here — validation at the front, embeddings-based schema mapping, self-explaining failures, and observability over heterogeneous data — is not specific to one client. Any organization with many source systems, limited mapping resources, and a mix of structured and unstructured data has the same problem. Healthcare makes it acute because the data is regulated and the stakes of "we don't know what's in it" are high, but the architecture is general.
It also pairs with the rest of our data-engineering record. We have run Elasticsearch as a production data layer on other engagements — heterogeneous data, an observability layer that makes it legible, and a migration the customer can actually operate afterward. The healthcare work is that same pattern applied to a regulated archive.
If you are sitting on a legacy archive you can't fully see into — too many source systems, too much manual mapping, validation that only catches problems after the fact — we are happy to walk through how this maps to your environment. Honest scope, a working proof of concept, and numbers we can stand behind.
About this account. Anonymized at the client's and prime's level. Figures and quotes are drawn from StandardData's own delivery record for the engagement — the proof-of-concept repository, the recorded client showcase (31 January 2025), and our internal delivery log. Throughput and data-quality percentages that the proof of concept did not measure have been deliberately omitted.

