InterviewHack.ai
Empezar gratis
Vacantes / Nuro

Applied AI Researcher, Agent Systems & Evaluation

Nuro · Mountain View, California (HQ)senior

En corto

  • ▸Construir un sistema de evaluación riguroso para agentes de IA que operen autonomamente dentro de Nuro, enfocado en pruebas reales con datos del entorno real, n
  • ▸Trabajar con modelos de vanguardia (transformers) para maximizar su rendimiento en tareas concretas del código y flujos de trabajo de ingenieros, sin innovar ar
  • ▸El punto diferencial: el equipo se enfoca en la confiabilidad y medición científica del trabajo autónomo de IA, con acceso directo a liderazgo y sistemas crític

Proficiency in English is required for collaboration and documentation.

Postularme en la empresa ↗Compartir por WhatsApp

En ~1 minuto te damos: quién te entrevista, las preguntas probables con respuestas desde tu CV, y tu CV adaptado a esta vacante. Gratis, sin tarjeta.

¿Qué piden?

  • ✓Experiencia en evaluación de modelos de IA en entornos reales, no solo en benchmarks.
  • ✓Conocimiento profundo de modelos de transformadores y su uso en tareas de agente.
  • ✓Capacidad para diseñar pipelines de evaluación cerrados (closed-loop) para flujos de trabajo automatizados.
  • ✓Experiencia con sistemas de agentes de IA en producción o entornos de ingeniería reales.
  • ✓Formulación de hipótesis basadas en evidencia y pruebas rigurosas, no en intuición.
  • ✓Capacidad de trabajo autónomo en un equipo pequeño sin playbook establecido.

¿No cumplís todo? Es lo normal — tu dossier gratis te dice qué gaps tenés y cómo cubrirlos en la entrevista.

Transformer modelsFrontier modelsAI agent systemsClosed-loop evaluationProduction agent fleetSkill libraryPlugin integrationReal-world dataEngineering workflowsEvidence-based decision making

¿A quién escribirle en Nuro?

Tu dossier gratis identifica a las personas que te entrevistarían — con su background, qué valoran y cómo escribirles para destacar antes de aplicar.

Who We Are Nuro is a self-driving technology company on a mission to make autonomy accessible to all. Founded in 2016, Nuro is building the world’s most scalable driver, combining cutting-edge AI with automotive-grade hardware. Nuro licenses its core technology, the Nuro Driver™, to support a wide range of applications, from robotaxis and commercial fleets to personally owned vehicles. With technology proven over years of self-driving deployments, Nuro gives automakers and mobility platforms a clear path to AVs at commercial scale, empowering a safer, richer, and more connected future. About the Team Frontier models are fungible. Any team can rent the same intelligence we can, and the model we build on today will be replaced within a month. What is not fungible is the infrastructure that decides whether an autonomous system's output can be trusted — evaluation, verification, and the discipline to gate on evidence instead of impressions. Nuro has spent a decade building exactly that discipline for a robot that drives on public roads, and this team turns it inward: we build the platform that lets AI agents operate autonomously inside Nuro's own engineering organization, under the same standard of proof we apply to the vehicle. Our mandate is to amplify the output of every engineer and researcher at Nuro by 100x. Not a better IDE, not a faster build — a change in what a single person can attempt. That number is a target, not a claim, and reaching it depends on one thing above all: autonomous work has to be trustworthy enough to run unattended. So our central ambition is to build the most rigorous closed-loop evaluation system for AI work anywhere. Leverage follows from trust, and trust follows from measurement. We operate as a startup inside a company that has already shipped a hard thing. Small team, no established playbook, direct access to compute and to the systems we are automating. You will work directly with engineering leadership and the CEO, and the decisions you make will be yours to make rather than yours to implement. About the Role Most teams building agents make design decisions by intuition and anecdote. Someone tries a new memory scheme, it feels better, it ships. We think that is the central failure of the field right now, and we are building this team to work the other way: every decision about how our agent systems are constructed should be settled by evidence. You would not be starting from zero. We already operate a substantial agent system in production, a fleet of agents with an extensive library of skills and plugins, integrated into the tools our engineers use daily, serving real users with real work. So every hypothesis you form can be tested against genuine production traffic from your first month. And the system is now complex enough that intuition has stopped being sufficient to improve it, which is precisely why this role exists. Everything this team builds is centered on frontier-lab transformer models. We are not inventing architectures. We are extracting the maximum from the best models that exist, and adapting them ourselves in the narrow places where our data gives us an advantage nobody else has. Your charter has two halves. Make the system perform against real-world data, not public benchmarks or tasks we invented to look good, but our codebase, our infrastructure, and our engineers' actual requests, with all the ambiguity that implies. The gap between benchmark performance and real-task performance is where most agent systems quietly fail. And turn any task into a closed loop : for any workflow an agent takes on, you should be able to say what success looks like, where the evaluation data comes from, how signal is collected, and how results feed the next iteration. Own the evaluation pipelin

No apliques sin prepararte

Investigamos quién te entrevista, adaptamos tu CV y te ensayamos en vivo — gratis la primera.

InterviewHack.ai

Preparate para la entrevista exacta: quién te entrevista, tu CV a medida y coach real.

Producto

VacantesRevisar CV (ATS) gratis¿Te pagan bien?Reporte de sueldos LATAMCursos gratisBlogCV a medidaCoach realPrecios

Empleos remotos

ReactPythonFull-StackLATAMMéxicoVer todas →

Preparate

Practicá con coachFrontendBackendAI EngineerPor empresaVendete con tu CV

Empresa

Buscás talentoAcerca deContactoPrivacidadTérminos

© 2026 InterviewHack.ai · Tu CV es tuyo. Nunca se usa para entrenar nada.

Vacantes similares activas

Software Engineer, ML Infrastructure Platform

Nuro · Mountain View, California (HQ)

→

Senior/Staff Software Engineer, ML Inference Platform

Nuro · Mountain View, California (HQ)

→

Senior Software Engineer, ML Infrastructure Platform

Nuro · Mountain View, California (HQ)

→

Software Engineer, Applied AI Infrastructure

Nuro · Mountain View, California (HQ)

→