InterviewHack.ai
Empezar gratis
Vacantes / Anthropic

Research Engineer / Research Scientist, RL Frontiers

Anthropic·San Francisco, CA | New York City, NY | Seattle, WAsenior

En corto

  • →Investigar y construir algoritmos de aprendizaje por refuerzo escalables para modelos de lenguaje grandes.
  • →Desarrollar arquitecturas y sistemas que permitan ejecutar experimentos de frontera con alta eficiencia y reproducibilidad.
  • →Trabajar en las corridas de RL más grandes y rápidas de Anthropic, resolviendo problemas que solo aparecen a escala.

Proficiency in English required for research and collaboration.

Postularme en la empresa ↗Compartir por WhatsApp

En ~1 minuto te damos: quién te entrevista, las preguntas probables con respuestas desde tu CV, y tu CV adaptado a esta vacante. Gratis, sin tarjeta.

Las preguntas que te van a hacer

1. ¿Cómo evaluarías la estabilidad numérica de un algoritmo de RL en entrenamiento distribuido a gran escala?

2. Describe un caso donde un experimento pequeño tuvo un comportamiento diferente al escalarlo: ¿qué causas investigarías?

3. ¿Cómo diseñarías un experimento para comparar dos arquitecturas de transformador en un entorno de RL con 1024 GPUs?

🔒 +7 preguntas más

Sin tarjeta. Subís tu CV y en ~1 minuto tenés el dossier completo.

🎧¿Llegás a la entrevista? Llevá el copiloto. Nuestra extensión escucha la entrevista en vivo y te muestra anclas de 3-4 palabras desde tu CV y tu preparación — mirás, conectás, hablás. Gratis. Ver la extensión →

💵 USD · Remote · No visa

¿No encontrás lo que buscás? Probá Micro1

Micro1 te ubica directo en empresas de EE. UU. que pagan en USD. Un solo proceso de vetting, múltiples ofertas — sin aplicar en frío.

Que Micro1 te matchee →
📬Vacantes elegidas para TU CV, cada mañana por WhatsApp. Gratis: escribí “vacantes” y el bot te manda tus matches del día. Suscribirme →

¿Qué piden?

  • ✓Familiaridad profunda con modelos de transformadores y sus dinámicas de entrenamiento.
  • ✓Experiencia práctica entrenando modelos grandes en entornos distribuidos.
  • ✓Historial de trabajo técnico original en ML o sistemas (investigación, open-source o producción).
  • ✓Capacidad para diseñar experimentos rigurosos a gran escala con análisis estadístico sólido.
  • ✓Habilidades sólidas en Python y JAX o PyTorch, con comodidad en el código de toda la pila.
  • ✓Razonamiento cuantitativo sobre costos de cómputo, memoria y comunicación.

¿No cumplís todo? Es lo normal — tu dossier gratis te dice qué gaps tenés y cómo cubrirlos en la entrevista.

PythonJAXPyTorchtransformersdistributed trainingtensor parallelismpipeline parallelismscaling lawsreinforcement learninglarge-scale optimization

¿A quién escribirle en Anthropic?

Tu dossier gratis identifica a las personas que te entrevistarían — con su background, qué valoran y cómo escribirles para destacar antes de aplicar.

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role Reinforcement learning is how Claude learns to reason, write code, and act autonomously over long horizons. The RL Scaling team works on how RL scales: what happens to throughput, stability, and learning efficiency as models get larger, episodes get longer, and compute grows by orders of magnitude, and what has to change in our algorithms and systems to keep getting returns from that scale. This role sits squarely across research and engineering. You'll develop next-generation architectures and RL algorithms, take them from a small-scale result to a frontier-scale run, and understand every place they behave differently along the way. You'll build the systems that set how fast the team can iterate: how many experiments, at what scale, and how quickly we can trust the results. And you'll work on Anthropic's largest and fastest RL runs, where the gap between a good idea and a working one is often a problem no one has solved yet. Key responsibilities • Study how RL training and sampling scale with model size, context length, and compute, and find the algorithmic and systems changes that keep scaling efficient • Develop next-generation model architectures and RL algorithms, and make them run efficiently at frontier scale • Take promising small-scale results to frontier-scale runs, and diagnose why they behave differently when they get there, whether the cause is numerical, algorithmic, or systemic • Build the experimental infrastructure that sets research velocity: fast, reproducible comparisons of architecture and algorithm variants at meaningful scale • Own end-to-end performance of our largest RL runs, from research code down to the hardware • Build performance and cost models for proposed architecture and algorithm changes, and use them to decide which ideas get scaled • Investigate training dynamics at scale, including instabilities, divergence, and throughput regressions, and trace them to root cause Minimum qualifications • Deep familiarity with modern transformer language models, including their architecture, training dynamics, and the behavior of large-scale optimization • Hands-on experience training large models in a distributed setting, including the tradeoffs between data, tensor, and pipeline parallelism • A track record of original technical work in ML training or systems, such as new methods, architectures, or optimizations, demonstrated through research, open-source, or production impact • Ability to design rigorous experiments at scale, including baselines, ablations, and enough statistical care to trust a result that costs real compute • Ability to reason quantitatively about the compute, memory, and communication costs of a model or algorithm • Strong programming skills in Python and JAX or PyTorch, and comfort reading and changing code at every layer of the stack Preferred qualifications • Research experience in reinforcement learning, optimization, or large-scale training, published or otherwise • Experience developing RL algorithms for language models • Experience with scaling laws or other quantitative models of training efficiency • Experience designing or modifying transformer architectures beyond standard configurations • Experience scaling training to large fleets of accelerators and debugging the problems that only appear at scale • Deep understanding of numerics in large-scale training, incl

¿Buscando algo parecido?

Dejá tu email y te avisamos cuando salgan vacantes que coincidan con tu perfil.

No apliques sin prepararte

Investigamos quién te entrevista, adaptamos tu CV y te ensayamos en vivo — gratis la primera.

InterviewHack.ai

Preparate para la entrevista exacta: quién te entrevista, tu CV a medida y coach real.

Producto

VacantesEmpresas contratandoRevisar CV (ATS) gratis¿Cómo suena tu inglés?¿Te pagan bien?Respuesta STAR gratisVeredicto de CV (Jev)Reporte de sueldos LATAMCursos gratisBlogCV a medidaPráctica habladaEs gratis

Empleos remotos

ReactPythonFull-StackLATAMArgentinaMéxicoVer todas →

Preparate

Práctica habladaFrontendBackendAI EngineerPor empresaVendete con tu CV

Empresa

Buscás talentoAcerca deContactoPrivacidadTérminos

© 2026 InterviewHack.ai · Tu CV es tuyo. Nunca se usa para entrenar nada. · Un producto de IA-PTY

Vacantes similares activas

Strategy & Operations, FDE

Anthropic · San Francisco, CA | New York City, NY

→

Applied AI Architects, Partner

Anthropic · Tokyo, Japan

→

Applied AI Architect, International Policy

Anthropic · San Francisco, CA | New York City, NY | Washington, DC

→

Network Deployment and Maintenance Lead - Data Center Operations

Anthropic · Remote-Friendly, United States

→