Data Engineer, Red Tape Index - Labrynth
ABOUT LABRYNTH
Labrynth accelerates progress by streamlining regulatory complexity. We build AI-powered platforms that navigate complex regulations, generate audit-level documentation, and provide certainty, not shortcuts. Our technology serves clients across heavily regulated industries including energy, compliance, and government regulations.
We operate as a forward-deployed engineering organization: small, high-velocity teams embedded directly with clients to rapidly discover needs and ship production-quality solutions.
ABOUT THE ROLE
We are hiring a Data Engineer to build the data platform behind our regulatory indices: acquiring fragmented public data, transforming it into clean, auditable datasets, and constructing the index methodology that turns it into published rankings.
This is a data platform role more than a pure pipeline or backend role. You will sit close to the raw sources and close to the math. The work spans three modes:
- Acquire: source data from fragmented and often hostile places, including government open-data portals, APIs, HTML, PDFs, legacy Excel formats, login-protected portals, and commercial sites behind anti-bot protection.
- Transform: normalize inconsistent jurisdictional data through bronze → silver → gold pipelines with idempotent ingestion, content hashing, and run-level lineage.
- Construct: turn clean data into transparent, auditable indices through winsorization, percentile ranks, weighting, composites, and sensitivity testing.
WHAT YOU'LL DO
- Ship scrapers and ingestion flows against messy, sometimes adversarial sources, using HTTP/2 clients, TLS-fingerprint evasion, and browser automation fallbacks, and keep them resilient as sources change
- Own Postgres schema design and migrations end to end across per-country and per-domain schemas
- Build and maintain medallion (bronze → silver → gold) transforms that are idempotent, content-hashed, and lineage-tracked
- Implement and defend index metho...