InterviewHack.ai
Empezar gratis
Vacantes / Scaleai

Staff Software Engineer, RL Environments

Scaleai · San Francisco, CA; New York, NY

En corto

  • ▸Ingeniero de software de staff que construye entornos de aprendizaje por refuerzo a gran escala para evaluar modelos de IA.
  • ▸Diseña plataformas de ejecución aislada, orquestación de lanzamientos y sistemas de verificación de recompensas, además de crear entornos reales para tareas com
  • ▸Destaca por ser un rol técnico hands-on con impacto directo en cómo se entrena y evalúa la inteligencia artificial de vanguardia.

Excellent written and verbal communication; ability to align engineers, researchers, and n

Postularme en la empresa ↗Compartir por WhatsApp

En ~1 minuto te damos: quién te entrevista, las preguntas probables con respuestas desde tu CV, y tu CV adaptado a esta vacante. Gratis, sin tarjeta.

🎧¿Llegás a la entrevista? Llevá el copiloto. Nuestra extensión escucha la entrevista en vivo y te muestra anclas de 3-4 palabras desde tu CV y tu preparación — mirás, conectás, hablás. Gratis. Ver la extensión →

¿Qué piden?

  • ✓8+ años de experiencia en ingeniería de software con fundamentos sólidos en sistemas distribuidos, diseño de sistemas, estructuras de datos
  • ✓Habilidades avanzadas en Python y experiencia en producción; comodidad con al menos otro lenguaje (TypeScript/React, Go, Rust, etc.).
  • ✓Experiencia profunda en contenerización y ejecución aislada (Docker, VMs, gVisor/Firecracker, Kubernetes, etc.).
  • ✓Experiencia construyendo o operando sistemas backend de alto rendimiento: orquestación, programación de tareas, colas y pipelines de datos a
  • ✓Experiencia práctica con LLMs, incluyendo bucles de agentes, llamadas a herramientas, MCP o sistemas de evaluación.
  • ✓Capacidad demostrada para liderar problemas ambiguos desde cero hasta un sistema entregado.

¿No cumplís todo? Es lo normal — tu dossier gratis te dice qué gaps tenés y cómo cubrirlos en la entrevista.

PythonTypeScriptReactGoRustDockerVMsgVisorFirecrackerKubernetes

¿A quién escribirle en Scaleai?

Tu dossier gratis identifica a las personas que te entrevistarían — con su background, qué valoran y cómo escribirles para destacar antes de aplicar.

About Scale AI At Scale, our mission is to develop reliable AI systems for the world's most important decisions. Our products provide the high-quality data and full-stack technologies that power the world's leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact. Scale Frontier Data is the organization behind the training and evaluation data that frontier labs depend on. We build the systems, tooling, and expert workflows that turn hard human expertise into signals that models can learn from, across reasoning, coding, agentic tool use, and domain expertise. Reinforcement learning environments are now the center of gravity for that work: the difference between a model that demos well and a model that reliably completes long-horizon work is almost always the quality of the environments and reward signals it was trained against. Responsibilities As a Staff Software Engineer, RL Environments, you'll own the technical foundation for how Scale builds, runs, verifies, and delivers RL environments at scale. An RL environment is a real piece of software: a containerized world with real dependencies, real state, real tools, and a grader that has to be correct even when the agent is creative about breaking it. Building one is a full-stack engineering problem. Building thousands of them reproducibly, cheaply, with trustworthy reward signals and throughput measured in millions of rollouts is a systems problem that very few people have solved. You'll work on both. You'll design the platform: sandboxed execution, environment packaging and versioning, rollout orchestration, trajectory capture, verifier frameworks, and the authoring surfaces that let engineers and domain experts produce environments without reinventing infrastructure each time. And you'll go deep on the environments themselves by instrumenting real applications, designing task suites that expose specific capability gaps, and building graders that hold up under adversarial optimization. This is a hands-on engineering role. You'll set technical direction across multiple teams, and you'll still be the person who writes the hard part. Required Qualifications • 8+ years of software engineering experience with strong fundamentals in distributed systems, system design, data structures, and algorithms. • Strong Python skills and a track record of shipping production software; comfort in at least one other part of the stack (TypeScript/React, Go, Rust, or similar). • Deep experience with containerization and sandboxed execution, including Docker, VMs, gVisor/Firecracker, Kubernetes, or equivalent. • Experience building or operating high-throughput backend systems: orchestration, job scheduling, queuing, and large-scale data pipelines. • Hands-on experience building with LLMs including agent loops, tool calling, MCP, or eval harnesses, and enough intuition about model behavior to reason about what a training signal actually teaches. • Demonstrated ability to own ambiguous, undefined problems end to end and drive them to a shipped system. • Excellent written and verbal communication; ability to align engineers, researchers, and non-engineering partners on a technical direction. Preferred Qualifications RL & Post-Training • Direct experience building RL environments, agentic benchmarks, or eval harnesses (SWE-bench-style task suites, terminal or browser environments, tool-use benchmarks, or in-house equivalents). • Familiarity with post-training methods: RLHF, RLAIF, RLVR, GRPO/PPO-family algorithms, rejection sampling, reward modeling, and the practical failure modes of each. • Experience designing verifiable reward signals, and firsthand experience with reward hacking and how to defend against it. • Experience with RL t

No apliques sin prepararte

Investigamos quién te entrevista, adaptamos tu CV y te ensayamos en vivo — gratis la primera.

InterviewHack.ai

Preparate para la entrevista exacta: quién te entrevista, tu CV a medida y coach real.

Producto

VacantesRevisar CV (ATS) gratis¿Cómo suena tu inglés?¿Te pagan bien?Reporte de sueldos LATAMCursos gratisBlogCV a medidaPráctica habladaEs gratis

Empleos remotos

ReactPythonFull-StackLATAMArgentinaMéxicoVer todas →

Preparate

Práctica habladaFrontendBackendAI EngineerPor empresaVendete con tu CV

Empresa

Buscás talentoAcerca deContactoPrivacidadTérminos

© 2026 InterviewHack.ai · Tu CV es tuyo. Nunca se usa para entrenar nada. · Un producto de IA-PTY

Vacantes similares activas

Global Payroll Analyst

Scaleai · Mexico City, MX

→

GenAI Compliance Operations & Programs Associate

Scaleai · Mexico City, MX

→

Senior Product Design Manager, Enterprise

Scaleai · London, UK

→

Senior Software Engineer, Platform

Scaleai · San Francisco, CA; New York, NY

→