Senior Machine Learning Engineer, LLM Inference Optimization
In short
- →Optimizas inferencia de LLMs/VLMs para reducir latencia y costo por token.
- →Trabajas en equipo con ingenieros de plataforma y kernel en sistemas de inferencia de producción.
- →Destaca por aplicar técnicas avanzadas como speculative decoding y optimización de caché en entornos reales.
Professional proficiency in English required.
In ~1 minute you get: who interviews you, the likely questions answered from your CV, and your CV tailored to this job. Free, no card.
🎧Land the interview? Bring the copilot. Our free extension listens to the live interview and flashes 3-4-word anchors from your resume and prep — glance, connect, talk. Get the extension →What they ask for
- ✓Experiencia comprobada en optimización de inferencia de LLMs/VLMs.
- ✓Dominio de frameworks de inferencia como vLLM, SGLang, TensorRT-LLM, Triton.
- ✓Habilidades en model compression: cuantización, distillación, entrenamiento consciente de cuantización.
- ✓Capacidad para diagnosticar y resolver regresiones en producción.
- ✓Experiencia con técnicas de aceleración como speculative decoding, kv-cache, continuous batching.
- ✓Conocimiento práctico de frameworks de benchmarking y evaluación reproducible.
Don't tick every box? That's normal — your free dossier shows your gaps and how to cover them in the interview.
Who should you write to at Nebius?
Your free dossier identifies the people who'd interview you — their background, what they value, and how to reach out so you stand out before applying.
More jobs like this
Looking for something similar?
Leave your email and we'll alert you when matching jobs appear.
Don't apply unprepared
We research who's interviewing you, tailor your CV and rehearse you live — first one free.