InterviewHack.ai
Empezar gratis
Vacantes / Arago

ML Systems Engineer — Inference Acceleration

Arago · Paris Officessenior

En corto

  • ▸Ingeniero de sistemas de ML que optimiza inferencia en un acelerador óptico-CMOS único.
  • ▸Trabaja en kernels, ejecución distribuida, servidor de inferencia y optimización de memoria en un stack de software en desarrollo.
  • ▸Destaca por ser el único equipo que trabaja con un procesador híbrido óptico-CMOS con rendimiento de orden de magnitud superior.

Proficient level of English

Postularme en la empresa ↗Compartir por WhatsApp

En ~1 minuto te damos: quién te entrevista, las preguntas probables con respuestas desde tu CV, y tu CV adaptado a esta vacante. Gratis, sin tarjeta.

🎧¿Llegás a la entrevista? Llevá el copiloto. Nuestra extensión escucha la entrevista en vivo y te muestra anclas de 3-4 palabras desde tu CV y tu preparación — mirás, conectás, hablás. Gratis. Ver la extensión →

¿Qué piden?

  • ✓Experiencia sólida en inferencia de alto rendimiento de ML.
  • ✓Conocimiento profundo de arquitectura de computadoras y modelos de ejecución en aceleradores.
  • ✓Experiencia con CUDA, Triton, ROCm/HIP o entornos equivalentes para programación de kernels.
  • ✓Habilidades sólidas en C++ y Python para desarrollo en stack personalizado.
  • ✓Experiencia con sistemas de inferencia modernos como vLLM, SGLang o TensorRT-LLM.
  • ✓Nivel de inglés proficient, con trabajo en entornos técnicos en inglés.

¿No cumplís todo? Es lo normal — tu dossier gratis te dice qué gaps tenés y cómo cubrirlos en la entrevista.

CUDATritonROCmHIPMLIRvLLMSGLangTensorRT-LLMC++Python

¿A quién escribirle en Arago?

Tu dossier gratis identifica a las personas que te entrevistarían — con su background, qué valoran y cómo escribirles para destacar antes de aplicar.

Meet Arago and the Aragonians Arago’s mission is to re-engineer the foundations of computing from first principles. The explosive growth of AI is pushing the industry to rethink how processors are built. Arago is meeting that challenge with a proprietary technology that fuses optical and CMOS technologies to deliver an order-of-magnitude increase in performance. Arago is the fastest, and currently the only, company to have built such a processor. It's backed by leading deep-tech investors and some of the most respected figures in semiconductors and computing, including the CEO of Arm, the founder of macOS who worked directly with Steve Jobs at Apple, an Nvidia Fellow, the Head of Optics at Google, and many other industry leaders. Our work is guided by three clear values: do great things, move with high velocity, and operate as one unit. We work in a demanding environment where constant learning, ownership, and execution are expected, and where exceptional people have the opportunity to do their life’s work. What you’ll do Optimize the execution and serving of modern AI models on Arago's custom accelerator. Work across kernels, model execution, multi-device distribution, runtime, and inference serving, while helping shape the software stack around the capabilities of Arago's hardware. Required Skills and Experience Strong experience in high-performance ML inference, GPU/accelerator programming, or ML systems engineering. Deep understanding of computer architecture, accelerator/GPU execution models, memory hierarchies, parallelism, and performance bottlenecks. Experience developing and optimizing custom kernels using CUDA, Triton, ROCm/HIP, or equivalent low-level programming environments. Experience with operator fusion, tiling, scheduling, data movement optimization, graph execution, and profiling of compute- and memory-bound workloads. Strong understanding of distributed model execution, including tensor, pipeline, sequence, and/or expert parallelism and communication/computation overlap. Hands-on experience with modern inference-serving systems such as vLLM, SGLang, TensorRT-LLM, or equivalent, including KV-cache management, continuous batching, paged attention, and prefill/decode scheduling. Strong C++ and Python skills, and comfort working on a custom accelerator stack where compiler, runtime, kernels, and abstractions are actively being developed. Exposure to or experience with MLIR and MLIR dialects is a strong plus. Language: English at a proficient level. Responsibilities Analyze modern AI workloads and identify kernel-, runtime-, memory-, and system-level bottlenecks on Arago's accelerator. Develop and optimize custom kernels, fused operators, and execution strategies to maximize device utilization. Design efficient mappings of models and operators across multiple Arago devices, including communication and synchronization strategies. Develop inference-serving techniques such as continuous batching, paged KV caches, prefix/context caching, chunked prefill, and prefill/decode interleaving or disaggregation. Build profiling, benchmarking, and performance-analysis infrastructure spanning kernels, full models, and serving workloads. Work closely with Arago's hardware, compiler, and runtime teams to co-design software abstractions and influence future hardware features based on real model workloads. Pay and benefits Competitive cash compensation, with final package based on location, experience, and the pay of team members in similar positions. Meaningful stock option plan offered (included in the majority of full time offers). Healthcare coverage (including family-friendly options), pension contributions, professional development support, and 25 days of PTO, in addition to public holidays. Ownership of a key technical domain, with significant vertical and/or horizontal growth opportunities, based on performance and individual drive. We look forward to hearing how you can help shape the future of AI at Arago.

No apliques sin prepararte

Investigamos quién te entrevista, adaptamos tu CV y te ensayamos en vivo — gratis la primera.

InterviewHack.ai

Preparate para la entrevista exacta: quién te entrevista, tu CV a medida y coach real.

Producto

VacantesRevisar CV (ATS) gratis¿Cómo suena tu inglés?¿Te pagan bien?Reporte de sueldos LATAMCursos gratisBlogCV a medidaPráctica habladaEs gratis

Empleos remotos

ReactPythonFull-StackLATAMArgentinaMéxicoVer todas →

Preparate

Práctica habladaFrontendBackendAI EngineerPor empresaVendete con tu CV

Empresa

Buscás talentoAcerca deContactoPrivacidadTérminos

© 2026 InterviewHack.ai · Tu CV es tuyo. Nunca se usa para entrenar nada. · Un producto de IA-PTY

Vacantes similares activas

Senior Hardware Design Engineer

Arago · Paris Offices

→

AI/ML Scientist — Quantization & Numerical Robustness

Arago · Paris Offices

→

Senior Analog Design Engineer

Arago · Paris Offices

→

Mixed-Signal IC Design Engineer

Arago · Paris Offices

→