InterviewHack.ai
Start free
Jobs / Anthropic

Research Engineer / Research Scientist, RL Frontiers

Anthropic·San Francisco, CA | New York City, NY | Seattle, WAsenior

In short

  • →Investigar y construir algoritmos de aprendizaje por refuerzo escalables para modelos de lenguaje grandes.
  • →Desarrollar arquitecturas y sistemas que permitan ejecutar experimentos de frontera con alta eficiencia y reproducibilidad.
  • →Trabajar en las corridas de RL más grandes y rápidas de Anthropic, resolviendo problemas que solo aparecen a escala.

Proficiency in English required for research and collaboration.

Apply on company site ↗Share on WhatsApp

In ~1 minute you get: who interviews you, the likely questions answered from your CV, and your CV tailored to this job. Free, no card.

The questions they'll ask you

1. ¿Cómo evaluarías la estabilidad numérica de un algoritmo de RL en entrenamiento distribuido a gran escala?

2. Describe un caso donde un experimento pequeño tuvo un comportamiento diferente al escalarlo: ¿qué causas investigarías?

3. ¿Cómo diseñarías un experimento para comparar dos arquitecturas de transformador en un entorno de RL con 1024 GPUs?

🔒 +7 more questions

No card. Upload your resume and the full dossier is ready in ~1 minute.

🎧Land the interview? Bring the copilot. Our free extension listens to the live interview and flashes 3-4-word anchors from your resume and prep — glance, connect, talk. Get the extension →

💵 USD · Remote · No visa

Not finding what you want? Try Micro1

Micro1 places engineers directly at US companies paying in USD. One vetting, multiple offers — no cold applying.

Get matched by Micro1 →
📬Jobs picked for YOUR resume, every morning on WhatsApp. Free: text “vacantes” and the bot sends your daily matches. Subscribe →

What they ask for

  • ✓Familiaridad profunda con modelos de transformadores y sus dinámicas de entrenamiento.
  • ✓Experiencia práctica entrenando modelos grandes en entornos distribuidos.
  • ✓Historial de trabajo técnico original en ML o sistemas (investigación, open-source o producción).
  • ✓Capacidad para diseñar experimentos rigurosos a gran escala con análisis estadístico sólido.
  • ✓Habilidades sólidas en Python y JAX o PyTorch, con comodidad en el código de toda la pila.
  • ✓Razonamiento cuantitativo sobre costos de cómputo, memoria y comunicación.

Don't tick every box? That's normal — your free dossier shows your gaps and how to cover them in the interview.

PythonJAXPyTorchtransformersdistributed trainingtensor parallelismpipeline parallelismscaling lawsreinforcement learninglarge-scale optimization

Who should you write to at Anthropic?

Your free dossier identifies the people who'd interview you — their background, what they value, and how to reach out so you stand out before applying.

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role Reinforcement learning is how Claude learns to reason, write code, and act autonomously over long horizons. The RL Scaling team works on how RL scales: what happens to throughput, stability, and learning efficiency as models get larger, episodes get longer, and compute grows by orders of magnitude, and what has to change in our algorithms and systems to keep getting returns from that scale. This role sits squarely across research and engineering. You'll develop next-generation architectures and RL algorithms, take them from a small-scale result to a frontier-scale run, and understand every place they behave differently along the way. You'll build the systems that set how fast the team can iterate: how many experiments, at what scale, and how quickly we can trust the results. And you'll work on Anthropic's largest and fastest RL runs, where the gap between a good idea and a working one is often a problem no one has solved yet. Key responsibilities • Study how RL training and sampling scale with model size, context length, and compute, and find the algorithmic and systems changes that keep scaling efficient • Develop next-generation model architectures and RL algorithms, and make them run efficiently at frontier scale • Take promising small-scale results to frontier-scale runs, and diagnose why they behave differently when they get there, whether the cause is numerical, algorithmic, or systemic • Build the experimental infrastructure that sets research velocity: fast, reproducible comparisons of architecture and algorithm variants at meaningful scale • Own end-to-end performance of our largest RL runs, from research code down to the hardware • Build performance and cost models for proposed architecture and algorithm changes, and use them to decide which ideas get scaled • Investigate training dynamics at scale, including instabilities, divergence, and throughput regressions, and trace them to root cause Minimum qualifications • Deep familiarity with modern transformer language models, including their architecture, training dynamics, and the behavior of large-scale optimization • Hands-on experience training large models in a distributed setting, including the tradeoffs between data, tensor, and pipeline parallelism • A track record of original technical work in ML training or systems, such as new methods, architectures, or optimizations, demonstrated through research, open-source, or production impact • Ability to design rigorous experiments at scale, including baselines, ablations, and enough statistical care to trust a result that costs real compute • Ability to reason quantitatively about the compute, memory, and communication costs of a model or algorithm • Strong programming skills in Python and JAX or PyTorch, and comfort reading and changing code at every layer of the stack Preferred qualifications • Research experience in reinforcement learning, optimization, or large-scale training, published or otherwise • Experience developing RL algorithms for language models • Experience with scaling laws or other quantitative models of training efficiency • Experience designing or modifying transformer architectures beyond standard configurations • Experience scaling training to large fleets of accelerators and debugging the problems that only appear at scale • Deep understanding of numerics in large-scale training, incl

Looking for something similar?

Leave your email and we'll alert you when matching jobs appear.

Don't apply unprepared

We research who's interviewing you, tailor your CV and rehearse you live — first one free.

InterviewHack.ai

Prepare for the exact interview: who's interviewing you, a tailored CV, and a real coach.

Product

JobsCompanies hiringFree ATS checkerInterview-English checkSalary checkFree STAR answerResume verdict (Jev)LATAM salary reportFree coursesBlogTailored CVSpoken practiceIt's free

Remote jobs

ReactPythonFull-StackLATAMArgentinaMexicoSee all →

Prepare

Spoken practiceFrontendBackendAI EngineerBy companySell with your CV

Company

For employersAboutContactPrivacyTerms

© 2026 InterviewHack.ai · Your CV is yours. Never used to train anything. · A product of IA-PTY

Similar open roles

Strategy & Operations, FDE

Anthropic · San Francisco, CA | New York City, NY

→

Applied AI Architects, Partner

Anthropic · Tokyo, Japan

→

Applied AI Architect, International Policy

Anthropic · San Francisco, CA | New York City, NY | Washington, DC

→

Network Deployment and Maintenance Lead - Data Center Operations

Anthropic · Remote-Friendly, United States

→