InterviewHack.ai
Empezar gratis
Vacantes / Meetdavis

Founding Data Engineer

Meetdavis·Parissenior

En corto

  • →Construyes y gestionas el conjunto de datos desde cero para un modelo de IA que genera edificios como grafos geométricos.
  • →Eres responsable de todo el ciclo: datos reales (planos) y sintéticos, representación estandarizada, curación, entrenamiento a gran escala y pruebas de ablation
  • →El dataset es parte esencial del algoritmo, y tu trabajo define qué datos realmente mejoran el modelo.

No se requiere idioma específico, pero se menciona trabajo en entornos internacionales con

Postularme en la empresa ↗Compartir por WhatsApp

En ~1 minuto te damos: quién te entrevista, las preguntas probables con respuestas desde tu CV, y tu CV adaptado a esta vacante. Gratis, sin tarjeta.

🎧¿Llegás a la entrevista? Llevá el copiloto. Nuestra extensión escucha la entrevista en vivo y te muestra anclas de 3-4 palabras desde tu CV y tu preparación — mirás, conectás, hablás. Gratis. Ver la extensión →

¿Qué piden?

  • ✓Experiencia comprobada construyendo datasets desde cero para pre-entrenamiento masivo.
  • ✓Capacidad para arquitectar y ejecutar pipelines de datos a gran escala, desde lo real hasta lo sintético.
  • ✓Habilidades de ingeniería sólidas en Python con enfoque en código limpio, tipo estático y concurrencia.
  • ✓Experiencia directa con generación sintética de datos, curation, deduplicación y puntuación de calidad.
  • ✓Capacidad para diseñar experimentos de ablation y dar feedback al generador de datos.
  • ✓Enfoque centrado en datos como problema primario, no solo como entrada para modelos.

¿No cumplís todo? Es lo normal — tu dossier gratis te dice qué gaps tenés y cómo cubrirlos en la entrevista.

PythonTypeScriptPostgreSQLRedisDockerKubernetesAirflowMLflowWandbHugging Face

¿A quién escribirle en Meetdavis?

Tu dossier gratis identifica a las personas que te entrevistarían — con su background, qué valoran y cómo escribirles para destacar antes de aplicar.

TL;DR Davis is hiring a Founding Data Engineer to build the data behind our foundation model, trained from scratch to generate buildings as geometric graphs. You will own the data end to end, from raw floorplans and synthetic generation to a canonical representation, curation, large-scale pretraining and the ablations that tell us which data actually moves the model. Here the dataset is part of the algorithm. About Davis Davis is an AI-native real estate company accelerating early-stage development and architectural design. Today developers coordinate four to five fragmented stakeholders over weeks or months. Soon they will need only one: Davis. We turn every input that shapes a development decision into decision-ready outputs: investor-grade feasibility studies, investment analysis, and architect-certified designs, delivered in days. Every stage pairs our proprietary AI systems with expert review, so velocity never comes at the cost of reliability. We closed a $5.5M pre-seed co-led by Heartcore Capital and Balderton Capital , with Yellow, Evantic and Entrepreneur First, alongside angels from the founding teams of Spacemaker, Black Forest Labs, Hugging Face, Supabase, Cleo and Spore Bio. We already work with leading developers and expect to support hundreds of projects over the coming year, deepening our research, our hiring, and our coverage of the development process end to end. The Role You will own the data our foundation model learns from, a model we train from scratch to generate buildings as geometric graphs. Part of the corpus comes from real floorplans as images and PDFs that have to become clean, standardized graphs. A large part will be synthetic, procedurally generated building graphs, geometry and rendered floorplans, with controlled variation in style, scan noise, annotations and furniture, each kept with its ground-truth graph automatically. You will think about the whole loop, from raw and synthetic data to a canonical structured representation, curation and validation, the training dataset, large-scale pretraining, evaluation, data ablations, and back to improving the generator and the data mixture. You are senior enough to architect the data stack and set the data strategy, and hands-on enough to write the pipelines, run the experiments, train the models and do the ablations yourself. What you will be working on Corpus from raw sources. Turn real floorplans (images, PDFs, scans) into clean, standardized building graphs, with the geometry and semantics that make them trainable. Synthetic data generation. Explore strategies to expand the dataset with synthetic data. Canonical representation and curation. Define the standardized representation, then filter, deduplicate, quality-score and validate at scale, with versioning, provenance and lineage. Pretraining data and mixtures. Assemble the training datasets, design the mixture and the curriculum, and blend synthetic and real data for large-scale pretraining. Data ablations. Train models to learn which data actually helps, read the results, and feed them back into the generator and the mixture. Evaluation. Build the eval harness (datasets, metrics, regression tests, monitoring) that tracks data and model quality over time. What We Are Looking For You have built the dataset, not just trained on it. You have personally built or generated the data used for a large pretraining run, from raw or synthetic sources, rather than only training on a dataset someone handed you. Senior and deeply hands-on. At least 5 years of strong experience, senior enough to architect the data stack and set strategy, but still coding the pipelines, running the experiments and doing the ablations yourself. Data as a first-class problem. A track record where the data itself is the object: curation, filtering, deduplication, quality scoring, mixtures, synthetic generation. Strong engineering. Deep Python, clean and typed code, async and concurrency, distributed data pipelines, TDD culture.

Más empleos como este

Empleos remotos de Data Engineer

¿Buscando algo parecido?

Dejá tu email y te avisamos cuando salgan vacantes que coincidan con tu perfil.

No apliques sin prepararte

Investigamos quién te entrevista, adaptamos tu CV y te ensayamos en vivo — gratis la primera.

InterviewHack.ai

Preparate para la entrevista exacta: quién te entrevista, tu CV a medida y coach real.

Producto

VacantesEmpresas contratandoRevisar CV (ATS) gratis¿Cómo suena tu inglés?¿Te pagan bien?Reporte de sueldos LATAMCursos gratisBlogCV a medidaPráctica habladaEs gratis

Empleos remotos

ReactPythonFull-StackLATAMArgentinaMéxicoVer todas →

Preparate

Práctica habladaFrontendBackendAI EngineerPor empresaVendete con tu CV

Empresa

Buscás talentoAcerca deContactoPrivacidadTérminos

© 2026 InterviewHack.ai · Tu CV es tuyo. Nunca se usa para entrenar nada. · Un producto de IA-PTY

Vacantes similares activas

AI Engineer - Full time

Meetdavis · Paris

→

AI Researcher - Full time

Meetdavis · Paris

→

AI Research Intern

Meetdavis · Paris

→