InterviewHack.ai
Start free
Jobs / Nebius

Senior ML Engineer (Token Factory)

Nebius · Berlinsenior

In short

  • ▸Ingeniero de ML sénior enfocado en optimizar inferencia y fine-tuning de modelos grandes a escala en GPU.
  • ▸Días a día: perfilar trabajo en GPU, diseñar pipelines de baja precisión y mejorar motores de inferencia como vLLM o TensorRT-LLM.
  • ▸Destacado: trabajar en una de las mayores nubes de GPU del mundo, con miles de GPUs y un impacto directo en el rendimiento de modelos como GPT-OSS o GLM-5.

Excellent command of the English language, alongside superior writing, articulation, and c

Apply on company site ↗Share on WhatsApp

In ~1 minute you get: who interviews you, the likely questions answered from your CV, and your CV tailored to this job. Free, no card.

🎧Land the interview? Bring the copilot. Our free extension listens to the live interview and flashes 3-4-word anchors from your resume and prep — glance, connect, talk. Get the extension →

What they ask for

  • ✓Conocimiento profundo de arquitecturas de transformadores y ML teórico.
  • ✓Experiencia con herramientas de perfilado de GPU (Nsight, PyTorch profiler).
  • ✓Entendimiento de jerarquía de memoria y tradeoffs de compute/memory en GPU.
  • ✓Familiaridad con conceptos clave de LLM (MHA, RoPE, KV-cache, Flash Attention, cuantización).
  • ✓Experiencia con frameworks modernos de deep learning y alta habilidad en ingeniería de software (Python, CI/CD, testing).
  • ✓Habilidades comunicativas y de liderazgo sólidas.

Don't tick every box? That's normal — your free dossier shows your gaps and how to cover them in the interview.

PythonPyTorchNsightvLLMSGLangTensorRT-LLMTritonCuteCUTLASSCUDA

Who should you write to at Nebius?

Your free dossier identifies the people who'd interview you — their background, what they value, and how to reach out so you stand out before applying.

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role Token Factory is a part of Nebius Cloud, one of the world's largest GPU clouds, running tens of thousands of GPUs. We are building a high-performance inference and fine-tuning platform designed to push foundation models to their hardware limits. Our mission is to maximize throughput, minimise latency, and optimise cost-per-token across tens of thousands of GPUs. Some directions we are currently working on, and which you can be a part of: • Inference Optimization: Identifying LLM inference bottlenecks to drive production speedups. Squeezing the maximum performance for a wide range of LLM architectures at scale (e.g., GPT-OSS, Kimi K2.5, DeepSeek V3.1/V3.2, GLM-5). • Inference engines support: Implement novel speculative decoding architectures, optimise components of various LLM designs (dense/MoE, autoregressive/parallel), and contribute to open-source inference engines. • Low Precision Training & Inference: Design and productionise low-precision (FP8, NVFP4/MXFP4) training and inference pipelines with measurable gains in throughput and cost-efficiency. We expect you to have: • A profound understanding of theoretical foundations of machine learning and transformer architecture. • Experience profiling GPU workloads using Nsight, PyTorch profiler, or similar tools • Understanding of GPU memory hierarchy and compute/memory tradeoffs • Familiarity with important ideas in LLM space, such as MHA, RoPE, KV-cache, Flash Attention, and quantisation • Understanding of performance aspects of large neural network training (sharding strategies, custom kernels, hardware features etc.) • Strong software engineering skills (we mostly use Python) • Deep experience with modern deep learning frameworks • Proficiency in contemporary software engineering approaches, including CI/CD, version control and unit testing • Strong communication and leadership abilities Nice to have: • Experience working with open-source inference engines (vLLM, SGLang, TensorRT-LLM), including contributions • Experience with kernel languages or DSLs such as Triton, Cute, CUTLASS, CUDA • A track record of building and delivering products (not necessarily ML-related) in a dynamic startup-like environment. • Strong engineering skills, including experience in developing large distributed systems or high-load web services. • Open-source projects that showcase your engineering prowess • Excellent command of the English language, alongside superior writing, articulation, and communication skills. Benefits & Perks: • Competitive compensation • Career growth and learning opportunities • Flexibility and ownership • Collaborative and innovative culture • Opportunity to work on impactful AI projects •

Don't apply unprepared

We research who's interviewing you, tailor your CV and rehearse you live — first one free.

InterviewHack.ai

Prepare for the exact interview: who's interviewing you, a tailored CV, and a real coach.

Product

JobsFree ATS checkerInterview-English checkSalary checkLATAM salary reportFree coursesBlogTailored CVSpoken practiceIt's free

Remote jobs

ReactPythonFull-StackLATAMArgentinaMexicoSee all →

Prepare

Spoken practiceFrontendBackendAI EngineerBy companySell with your CV

Company

For employersAboutContactPrivacyTerms

© 2026 InterviewHack.ai · Your CV is yours. Never used to train anything. · A product of IA-PTY

Similar open roles

Senior Software Engineer (Token Factory)

Nebius

→

Senior Software Engineer (Storage Virtualization Team)

Nebius · Berlin

→

Senior Site Reliability Engineer — Token Factory (Inference Platform)

Nebius · Berlin

→

Senior ML Engineer (Token Factory)

Nebius

→