InterviewHack.ai
Empezar gratis
Vacantes / blackforestlabs

Member of Technical Staff - VLM

blackforestlabslead

En corto

  • ▸Desarrollar modelos visión-lenguaje (VLM) de vanguardia integrados directamente en el sistema FLUX.
  • ▸Innovar en arquitecturas multimodales y mejorar la calidad generativa con enfoque en controlabilidad y escalabilidad.
  • ▸Rol líder técnico con foco en investigación avanzada, no en ajuste fino (SFT/LoRA) de modelos existentes.

Comfortable at the research/production boundary — you care whether the work ships and gene

Postularme en la empresa ↗Compartir por WhatsApp

En ~1 minuto te damos: quién te entrevista, las preguntas probables con respuestas desde tu CV, y tu CV adaptado a esta vacante. Gratis, sin tarjeta.

¿Qué piden?

  • ✓Experiencia en pre-entrenamiento o avance significativo de un VLM con despliegue en producción o lanzamiento público.
  • ✓Historial claro de investigación o resultados en producción sobre arquitecturas multimodales.
  • ✓Comprensión profunda de la interacción entre representaciones visuales y lingüísticas: alineación, tokenización, atención cruzada.
  • ✓Experiencia en entrenamiento distribuido a escala multi-nodo.
  • ✓Capacidad de operar en la frontera entre investigación y producción, con enfoque en que el trabajo se implemente y generalice.
  • ✓Experiencia con modelos generativos basados en difusión o flujo es un fuerte plus.

¿No cumplís todo? Es lo normal — tu dossier gratis te dice qué gaps tenés y cómo cubrirlos en la entrevista.

Stable DiffusionFLUXLatent DiffusionVision-Language ModelsDiffusion ModelsFlow-based Generative ModelsCross-modal AttentionDistributed TrainingMulti-node TrainingModel Deployment

¿A quién escribirle en blackforestlabs?

Tu dossier gratis identifica a las personas que te entrevistarían — con su background, qué valoran y cómo escribirles para destacar antes de aplicar.

About Black Forest Labs We're the team behind Latent Diffusion, Stable Diffusion, and FLUX — foundational technologies that changed how the world creates images and video. Our models power the tools used by millions of creators, developers, and businesses worldwide, and FLUX is among the most advanced generative systems in the world. Headquartered in Freiburg, Germany with a growing presence in San Francisco, we're scaling fast while staying true to what makes us different: research excellence, open science, and building technology that expands human creativity. Why This Role Vision-language models are becoming foundational to how people interact with generative AI — but most VLM research happens in isolation from the generation stack. At Black Forest Labs, we're integrating VLMs directly into FLUX in ways that make our models more powerful, more controllable, and more aligned with what creators actually want. This role is about pioneering that integration. You won't be applying off-the-shelf VLMs — you'll develop novel approaches, innovate on architectures, and answer questions that haven't been solved yet: how vision and language representations inform each other, how multimodal understanding improves generation quality, and how to make these capabilities deployable at scale without compromising what makes FLUX exceptional. This is a Staff / Senior IC role. We're looking for someone who has pretrained or significantly advanced a VLM, not just fine-tuned one. What You'll Work On • Lead development and training of state-of-the-art multimodal vision-language models within the FLUX stack — innovating on architectures, not just applying existing ones • Design fine-tuning strategies that adapt VLMs to specialized creative use cases (captioning, editing instructions, prompt enhancement) that general-purpose models can't handle • Research integrations between VLM/LLM capabilities and our diffusion and flow pipelines — finding creative ways to improve generation quality and controllability without computational bottlenecks • Evaluate emerging multimodal architectures, translating the best of recent research into practical improvements What We're Looking For • You've pretrained or significantly advanced a VLM (not just SFT'd or LoRA'd one) that was deployed in a production system or released publicly • Strong publication record or unambiguous production track record showing you push the frontier on multimodal architectures • Deep understanding of how vision and language representations interact: tokenization, alignment, grounding, cross-modal attention, and the failure modes of each • Experience with distributed training at multi-node scale • Comfortable at the research/production boundary — you care whether the work ships and generalizes, not just whether it reads well • Experience with diffusion or flow-based generative models is a strong plus — especially if you've thought about how autoregressive and diffusion paradigms can compose How We Work Together We’re a distributed team with real offices that people actually use. Depending on your role, you’ll either join us in Freiburg or SF at least 2 days a week (or one full week every other week), or work remotely with a monthly in-person week to stay connected. We’ll cover reasonable travel costs to make this possible. We think in-person time matters, and we’ve structured things to make it accessible to all. We’ll discuss what this will look like for the role during our interview process. Everything we do is grounded in four values: • Obsessed. We are a frontier research lab. The science has to be right, the understanding deep, the product beautiful. • Low Ego.</st

No apliques sin prepararte

Investigamos quién te entrevista, adaptamos tu CV y te ensayamos en vivo — gratis la primera.

InterviewHack.ai

Preparate para la entrevista exacta: quién te entrevista, tu CV a medida y coach real.

Producto

VacantesRevisar CV (ATS) gratis¿Te pagan bien?Reporte de sueldos LATAMCursos gratisBlogCV a medidaCoach realPrecios

Empleos remotos

ReactPythonFull-StackLATAMMéxicoVer todas →

Preparate

Practicá con coachFrontendBackendAI EngineerPor empresaVendete con tu CV

Empresa

Buscás talentoAcerca deContactoPrivacidadTérminos

© 2026 InterviewHack.ai · Tu CV es tuyo. Nunca se usa para entrenar nada.

Vacantes similares activas

Member of Technical Staff - Image / Video Generation

blackforestlabs

→

Member of Technical Staff - Post Training

blackforestlabs

→

Member of Technical Staff - Pretraining

blackforestlabs

→

Account Executive

blackforestlabs

→