Senior Data Engineer
About CloudX
At CloudX we’re building a new supply-side advertising platform for mobile publishers. We’re convinced that an AI-native product will significantly advance the state of the art in mobile advertising; we recently raised a $30M Series A in order to bring this dream to life. We have a long history of innovation in this space — our founding team previously built MoPub (sold to Twitter for $350M) and MAX (acquired by AppLovin). Our new platform combines real technical improvements like verifiably fair auctions with a truly AI-native product experience to give publishers unprecedented control over their ad monetization strategy.
About The Team
Our Engineering team is distributed and remote — spanning UTC-8 to UTC+6, with core working hours of roughly the US Eastern business day. We have a strong ownership culture and are heavily collaborative, relying primarily on asynchronous, written, communication for coordination. We ship daily and believe that fast CI and good test coverage is the best way to remain productive as we scale. We’re small and high trust; we optimize for rapid iteration and experimentation. Everyone has access to the latest AI tools, but rather than generating vibe-slop we use them pragmatically to build better products. We are lucky to work closely with our talented Product and Business teams to make sure we’re building the right things. It’s a true early-stage startup with lots of important work to go around.
What you'll do
We are looking for a Senior Data Engineerto build the platform that turns billions of auction events into numbers that our customers and our company can rely on. We have a basic system in place, but as we scale we're looking for someone to come in and take ownership of a modern streaming data platform. You'll work closely with the rest of the technical staff, as well as our analytics and business teams, to make it all work. Your key responsibilities will be:
- Ingestion and storage: own the path from edge event to queryable table — throughput, cost, schema evolution, partitioning, and retention, plus the instrumentation that tells you it’s healthy. The events must flow.
- Orchestration and transformation: choose our scheduling and transformation layer, stand it up, and then use it to build the derived datasets — fact tables, aggregates, materialized views — with tests, lineage, and backfills.
- Serving the product and the business: our customer-facing reporting is a product surface with latency and freshness expectations, not a batch job. You’ll partner closely with analytics, and build the platform that lets engineering, account management, and support answer their own questions.
- Incident response: alongside the rest of us, you’ll be on-call for the systems we build. We keep a green steady state by default and ruthlessly fix or silence issues as they appear. We run a follow-the-sun global rotation.
Who you are
We encourage you to apply if you meet these requirements:
- Hands-on expertise: you may have managed or led at points in your career, but you still code regularly and are interested in continuing to do so. You’ve been in charge of larger projects or initiatives and know how to work well with others.
- Strong written communication skills: you are used to writing about, speaking about, and generally communicating complex technical subject matter both to other engineers and to non-engineers. You’ll also be documenting what our data means, which is its own discipline.
- Early-stage mentality: you understand that success at a startup involves grit and determination. You have good taste when it comes to trading off speed vs. perfection. You know when to cut corners but aren’t afraid to advocate for rigor when you believe it’s necessary.
- Data platform experience: you’ve owned ingestion, storage, and scheduling for a system with real consumers, where a wrong or late number had consequences you had to answer for. You’ve built the freshness and reconciliation checks that caught those problems first.
- Large-scale event streams: you’ve run high-throughput streaming and queueing systems in production — Kafka, RedPanda, Kinesis, Flink, Spark Streaming, Pulsar, NATS, or others — and dealt with partitioning, consumer lag, duplicate delivery, backpressure, and replay after an outage.
- Columnar stores at scale: ClickHouse, Druid, BigQuery, Snowflake, or similar. We work primarily with Clickhouse so expertise there is particularly appreciated, but any experience with columnar stores is essential.
- Data modeling: you have opinions about grain, idempotent incremental loads, late-arriving data, and slowly-changing dimensions. You write the transforms yourself and don’t consider that someone else’s job.
- Tooling judgment: you’ve chosen an orchestrator or transformation framework for a small team, and can argue for a specific answer — Airflow, Dagster, Temporal, dbt, SQLMesh, or a few hundred lines of Go — on the merits rather than the popularity of the tool.
- AI forward: you are actively experimenting with or using AI as part of your software engineering practice. You don’t send vibe-coded slop to your teammates to review, but you use AI appropriately to achieve great results.
- High ownership: you care a lot about your work and when you ship a product, you make sure it continues to solve problems for the customer. You care a lot about the customer, the overall business, and are constantly trying to help achieve success — with or without code.
While not required, we’re particularly interested in candidates with:
- Adtech experience: you’ve worked in adtech, particularly mobile adtech, and understand the event taxonomy, the discrepancy problems, and why nobody’s numbers ever match on the first try. Equivalent experience from other high-volume, money-denominated domains — payments, marketplaces, exchanges — is a real substitute.
- Data governance instincts: advertising data carries real obligations around consent, regional restrictions, and retention. Experience building systems where those constraints are enforced by the pipeline rather than by a policy document is a strong positive.
- Stack experience: we run in AWS on EKS, our services are written in Golang, we put application state in Postgres, we store a lot of event data in Clickhouse, we use Datadog for observability. Familiarity with any or all of these, particularly Clickhouse at volume, is a strong positive. Everything above the storage layer is genuinely undecided, and we’d like to hear from people with strong opinions.
In general, we’re looking for people with grit, passion, and talent. If you’re not sure if this role is an exact fit, we encourage you to apply. Many members of our team have had interesting career paths and we relish the chance to work with extraordinary individuals.
Pay and Benefits
The annual US base salary range for this role, and other engineering roles, is $150,000 – $300,000. This salary range is broad in order to accommodate a wide range of candidates; the interview process will narrow it down based on a number of factors, including your experience, qualifications, and location.
We offer equity compensation and top-tier medical, dental, and vision benefits.
We also have a generous hardware budget for a computer, monitor, and other core equipment necessary to work effectively on a remote team.
We care about the quality of your work more than the specific hours you spend getting it done, and try to minimize the number of synchronous meetings in favor of greater flexibility. There is no in-office requirement.