InterviewHack.ai
Start free
Jobs / Nebius

Senior Applied Scientist, Efficient LLM Inference & Model Optimization

Nebius·Zurichsenior

In short

  • →Científico aplicado senior que optimiza inferencia de LLM/VLM con impacto real en producción.
  • →Día a día: investiga, pruebas, codifica prototipos y los entrega a ingenieros para implementar.
  • →Lo destacable: el trabajo se publica, se comparte y se despliega, no solo se queda en un paper.

Proficiency in English required for collaboration and documentation.

Apply on company site ↗Share on WhatsApp

In ~1 minute you get: who interviews you, the likely questions answered from your CV, and your CV tailored to this job. Free, no card.

The questions they'll ask you

1. ¿Cómo diseñarías un experimento para evaluar el impacto de la compresión del KV-cache en latencia y calidad?

2. ¿Qué desafíos enfrentas al implementar QAT en un modelo MoE a escala?

3. ¿Cómo asegurarías la estabilidad numérica al aplicar cuantización a un LLM con 100B parámetros?

🔒 +7 more questions

No card. Upload your resume and the full dossier is ready in ~1 minute.

🎧Land the interview? Bring the copilot. Our free extension listens to the live interview and flashes 3-4-word anchors from your resume and prep — glance, connect, talk. Get the extension →

💵 USD · Remote · No visa

Not finding what you want? Try Micro1

Micro1 places engineers directly at US companies paying in USD. One vetting, multiple offers — no cold applying.

Get matched by Micro1 →
📬Jobs picked for YOUR resume, every morning on WhatsApp. Free: text “vacantes” and the bot sends your daily matches. Subscribe →

What they ask for

  • ✓PhD en ciencia de datos, inteligencia artificial o área relacionada.
  • ✓Experiencia comprobada en optimización de inferencia de modelos grandes.
  • ✓Habilidades sólidas en PyTorch, CUDA y herramientas de bajo nivel.
  • ✓Capacidad para diseñar experimentos rigurosos con métricas medibles.
  • ✓Experiencia colaborando con MLEs y equipos de ingeniería en producción.
  • ✓Publicaciones en conferences relevantes (NeurIPS, ICML, ACL, etc.)

Don't tick every box? That's normal — your free dossier shows your gaps and how to cover them in the interview.

PyTorchTritonCUDAC++PythonTensorRTONNXHugging FaceOpenAIBittensor

Who should you write to at Nebius?

Your free dossier identifies the people who'd interview you — their background, what they value, and how to reach out so you stand out before applying.

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role Nebius Token Factory needs scientists who can turn frontier inference bottlenecks into research problems, publish credible work, and then help ship the results into production. This is not a papers-only research role. The Applied Scientist is expected to design rigorous experiments, write strong code, collaborate with engineers, and convert research into deployed inference capabilities. A Senior Applied Scientist owns well-scoped research and production optimization projects. They can publish or prepare high-quality technical work while also producing code, experiments, and prototypes that engineers can use. Your responsibilities : • Own focused research projects from hypothesis through experiment, ablation, prototype, and production handoff. • Prepare internal reports, technical blogs, or papers when the work is externally credible. • Partner directly with MLEs to ensure research prototypes become usable production components. • Define and execute research programs in efficient LLM and VLM inference with measurable production impact. • Invent, evaluate, and productionize methods for quantization, QAT , distillation, speculative decoding, KV -cache reuse, KV -cache compression, long-context inference, MoE routing, and model/runtime co-optimization. • Build high-quality prototypes in PyTorch, Triton, CUDA -adjacent tooling, or inference-serving frameworks, then work with MLEs and platform engineers to productionize them. • Design rigorous evaluation methodology covering quality, latency, throughput, numerical stability, memory footprint, tail latency, and cost per token. • Publish papers, technical reports, blog posts, and open-source artifacts th

More jobs like this

Remote AI / ML Engineer jobs

Looking for something similar?

Leave your email and we'll alert you when matching jobs appear.

Don't apply unprepared

We research who's interviewing you, tailor your CV and rehearse you live — first one free.

InterviewHack.ai

Prepare for the exact interview: who's interviewing you, a tailored CV, and a real coach.

Product

JobsCompanies hiringFree ATS checkerInterview-English checkSalary checkFree STAR answerResume verdict (Jev)LATAM salary reportFree coursesBlogTailored CVSpoken practiceIt's free

Remote jobs

ReactPythonFull-StackLATAMArgentinaMexicoSee all →

Prepare

Spoken practiceFrontendBackendAI EngineerBy companySell with your CV

Company

For employersAboutContactPrivacyTerms

© 2026 InterviewHack.ai · Your CV is yours. Never used to train anything. · A product of IA-PTY

Similar open roles

Senior Applied ML Engineer (Agentic Search)

Nebius · Zurich

→

Staff / Senior Software Engineer (Agentic Search) - Crawler

Nebius · Zurich

→

Senior Machine Learning Engineer, LLM Inference Optimization

Nebius · Zurich

→

Staff / Senior Software Engineer (Agentic Search) - Runtime

Nebius · Zurich

→