Senior Data Engineer - Data Flow

Revelio Labs
Revelio Labs

Software Engineering, Data Science

Remote

USD 150k-220k / year + Equity

Posted on Aug 10, 2026

About Revelio Labs:

Revelio Labs builds workforce intelligence products powered by large-scale, real-world labor market data. We collect, standardize, enrich, and serve data about companies, employees, job postings, skills, compensation, sentiment, and workforce movement so our customers can understand how organizations and labor markets change over time.

Our customers include investors, corporate strategists, HR teams, academics, and governments. They rely on Revelio because we turn fragmented, noisy, fast-changing raw data into trustworthy datasets, metrics, and products. That requires more than moving data from one place to another; it requires engineering systems that are correct, observable, resilient, explainable, and continuously improving.\

Overview:

The Data Flow team owns the systems that turn raw inputs into reliable, production-grade datasets used across Revelio’s products. The team sits at the center of Revelio’s data lifecycle and works closely with Data Acquisition, Infrastructure, Data Science, Economics, and client-facing teams. In essence, the Data Flow team is responsible for making sure high-volume, messy, constantly changing source data becomes consistent, trusted, and usable.

About the role:

We’re hiring a Senior Data Engineer to help design, build, operate, and improve Revelio’s core data flow systems.

This is not a narrow ETL role. You will own production data systems end to end: source ingestion, transformation logic, orchestration, infrastructure integration, data quality, stakeholder communication, and long-term system design. You will work on systems where correctness and timeliness matter, but where the environment is inherently messy. The right person will be excited by that ambiguity.

What you’ll do:

  • Own and operate production data flow pipelines end to end: ingestion, transformation, enrichment, validation, publication, and monitoring.
  • Build and improve reusable flow patterns for high-volume datasets such as profiles, companies, job postings, and related enrichments.
  • Debug production failures across multiple layers: source data, code, orchestration, storage, Kubernetes, data warehouses, models, and product expectations.
  • Investigate data correctness issues, including unexpected or anomalous metrics and client-facing discrepancies.
  • Design and implement guardrails: validation checks, lifecycle checks, monitors, alerts, data contracts, runbooks, and safer release patterns.
  • Improve pipeline reliability, performance, and cost efficiency by tuning resources, execution patterns, storage layouts, and orchestration logic.
  • Partner with Data Acquisition, Infrastructure, Data Science, Economics, and client-facing teams to understand upstream/downstream dependencies and communicate impact clearly.
  • Raise the engineering bar for the team by simplifying brittle systems and building patterns that other engineers can safely reuse.
  • Bring strong judgment to ambiguous situations: know when to patch, when to roll back, when to rerun, when to escalate, and when to redesign.

What we’re looking for:

  • 4+ years of professional experience in a production software or data engineering role.
  • Strong programming ability in Python, Go, and/or Rust.
  • Strong SQL skills and a proven ability to debug complex data quality issues.
  • Strong expertise in building or operating large production data pipelines.
  • Experience with incremental processing, idempotency, data versioning, schema evolution, and operational recovery.
  • Comfort working across the full data system stack: source inputs, transformation code, workflow orchestration, storage, containers, Kubernetes, data warehouses, and downstream serving layers.
  • High agency and ownership: when something breaks or looks wrong, you form hypotheses, gather evidence, reason about blast radius, pull in the right people, and drive toward resolution.
  • A systems mindset: you naturally ask yourself how a local change affects upstream producers, downstream consumers, operational load, future debugging, and team maintainability.

Bonus points:

  • Experience with Flyte, Argo Workflows, or similar workflow orchestration systems.
  • Experience with Snowflake, ClickHouse, or other analytical data platforms.
  • Experience with Terraform, AWS, and production infrastructure workflows.
  • Experience with large-scale web-derived datasets, job postings, workforce data, marketplace data, profiles, company data, or other noisy external data sources.
  • Experience integrating machine learning or NLP models into production data pipelines.
  • Experience with entity resolution (people, postings, or companies), taxonomy systems, or other data normalization problems.

How we work:

We value engineers who are intellectually honest, pragmatic, and have strong, independent judgement. We care about whether the data is right, whether the system can be operated safely, whether teammates can understand it later, and whether customers can trust what we say.

This role is a strong fit for someone who wants autonomy and responsibility, but not isolation. You will collaborate closely with engineers, data scientists, infrastructure owners, economists, client-facing teams, and product stakeholders. You will be expected to communicate clearly, make good tradeoffs, and own outcomes when the work is ambiguous and the answer is not obvious.

If you are excited by difficult data systems, messy real-world inputs, high-stakes ownership, and the chance to solve problems that are at the core of an entire company’s operations, we’d like to meet you.

Location:

Our offices are based in New York City, but the position can be done remotely.

Compensation:

The pay range for this position in New York City is $150,000 - $220,000 per year. The salary range for performing this role outside of New York City may differ. Base pay offered may vary depending on job-related knowledge, skills, and experience. Additionally, you may be eligible to participate in our company’s equity program, plus benefits, including medical, dental, vision, retirement, and other. The range above is for the expectations as laid out in the job description, however we are often open to a wide variety of profiles, and recognize that the person we hire may be more senior or have different experience than this job description as posted. If that ends up being the case, the updated salary range will be communicated to you as a candidate.

How to reach us:

Please email your resume to recruiting@reveliolabs.com as a PDF file. Please include your GitHub and highlight any projects that you’ve worked on that may be relevant.