Senior Site Reliability Engineer — Token Factory (Inference Platform)
Nebius · Berlinsenior
In short
- ▸Ingeniero de confiabilidad de sitios que garantiza la fiabilidad y rendimiento de una plataforma de inferencia de IA a gran escala con miles de GPUs.
- ▸Día a día: diseño de pipelines de telemetría, optimización de Kubernetes, gestión de infraestructura como código y respuesta a incidentes con automatización.
- ▸Destacado: Trabajas en una de las mayores nubes de GPUs del mundo, impactando directamente en el despliegue de modelos de IA multimodal.
Fluent English required for collaboration across global teams.
In ~1 minute you get: who interviews you, the likely questions answered from your CV, and your CV tailored to this job. Free, no card.
🎧Land the interview? Bring the copilot. Our free extension listens to the live interview and flashes 3-4-word anchors from your resume and prep — glance, connect, talk. Get the extension →What they ask for
- ✓Experiencia sólida con Kubernetes y su operación en entornos de producción.
- ✓Conocimiento profundo de Prometheus, Grafana y la creación de dashboards y alertas efectivas.
- ✓Habilidades avanzadas en Terraform para infraestructura como código (IaC).
- ✓Experiencia con stacks de inferencia de IA como vLLM, Triton o Ray en entornos GPU.
- ✓Capacidad para diseñar y mantener SLOs y políticas de alerta para APIs de alto rendimiento.
- ✓Experiencia en MLOps o plataformas de hospedaje de modelos.
Don't tick every box? That's normal — your free dossier shows your gaps and how to cover them in the interview.
Who should you write to at Nebius?
Your free dossier identifies the people who'd interview you — their background, what they value, and how to reach out so you stand out before applying.
More jobs like this
Don't apply unprepared
We research who's interviewing you, tailor your CV and rehearse you live — first one free.