InterviewHack.ai
Start free
Jobs / Jobgether

AI Research Engineer (Kernel & Inference Optimization)

Jobgether·Francemid

In short

  • →Ingeniero de investigación en IA que optimiza motores de inferencia en dispositivos móviles y bordes.
  • →Desarrolla kernels personalizados en Metal Shading Language (MSL) y técnicas avanzadas como cuantización y Flash Attention.
  • →Destaca por combinar investigación de punta con ingeniería de bajo nivel para mejorar latencia y eficiencia en hardware limitado.

Fluent English required for collaboration with remote teams.

Apply on company site ↗Share on WhatsApp

In ~1 minute you get: who interviews you, the likely questions answered from your CV, and your CV tailored to this job. Free, no card.

The questions they'll ask you

1. ¿Cómo implementarías un kernel MSL para acelerar la atención en un modelo de texto en un dispositivo móvil?

2. ¿Qué técnicas usarías para reducir el consumo de memoria en una inferencia de modelo de difusión en edge?

3. Describe un caso donde optimizaste el rendimiento de una pipeline de inferencia —¿qué métricas mejoraste y cómo?

🔒 +7 more questions

No card. Upload your resume and the full dossier is ready in ~1 minute.

🎧Land the interview? Bring the copilot. Our free extension listens to the live interview and flashes 3-4-word anchors from your resume and prep — glance, connect, talk. Get the extension →

💵 USD · Remote · No visa

Not finding what you want? Try Micro1

Micro1 places engineers directly at US companies paying in USD. One vetting, multiple offers — no cold applying.

Get matched by Micro1 →
📬Jobs picked for YOUR resume, every morning on WhatsApp. Free: text “vacantes” and the bot sends your daily matches. Subscribe →

What they ask for

  • ✓Titulación en Ciencias de la Computación o campo técnico relacionado.
  • ✓Experiencia comprobada con Metal Shading Language (MSL) y desarrollo de kernels personalizados.
  • ✓Demostrado impacto en reducción de latencia, mejora de throughput y eficiencia de memoria.
  • ✓Conocimiento profundo de arquitecturas de inferencia, optimización y motores de modelos.
  • ✓Experiencia práctica en pipelines de inferencia desde optimización hasta producción.
  • ✓Habilidades sólidas en benchmarking, análisis de rendimiento y resolución de cuellos de botella.

Don't tick every box? That's normal — your free dossier shows your gaps and how to cover them in the interview.

Metal Shading Language (MSL)GPU kernelsMobile inferenceQuantizationPruningFlash AttentionKV cachingSpeculative decodingTensor parallelismPipeline parallelism

Who should you write to at Jobgether?

Your free dossier identifies the people who'd interview you — their background, what they value, and how to reach out so you stand out before applying.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Research Engineer (Kernel & Inference Optimization) based in France. You will work at the intersection of AI research, systems engineering, and high-performance model inference. Your focus will be on developing and optimizing model-serving architectures for advanced AI systems across a range of hardware environments. You will tackle challenges involving latency, throughput, memory efficiency, and scalability, including deployment on resource-constrained mobile and edge devices. The role combines hands-on research with low-level engineering, giving you the opportunity to develop novel inference strategies and GPU kernels. You will work with complex architectures spanning text, image, audio, diffusion models, and vision transformers. Your work will involve rigorous benchmarking, production testing, and iterative optimization to translate research into measurable performance improvements. You will collaborate with cross-functional teams in a highly technical, remote environment focused on pushing the boundaries of efficient AI systems. Accountabilities Design and deploy advanced model-serving architectures optimized for high throughput, low latency, and efficient memory utilization. Develop inference pipelines capable of operating effectively across diverse environments, including resource-constrained mobile devices and edge platforms. Establish clear performance targets covering response latency, token generation speed, throughput, memory footprint, and reliability. Build and execute controlled inference benchmarks in simulated and production environments, tracking latency, throughput, memory consumption, and error rates. Create and maintain representative datasets and simulation scenarios for evaluating model performance under real-world and resource-constrained conditions. Identify computational and memory bottlenecks across inference pipelines and implement solutions involving batching, networking, memory management, and other system-level optimizations. Develop custom GPU kernels and compute shaders for mobile hardware, including solutions written in Metal Shading Language (MSL). Apply advanced inference optimization techniques such as pruning, quantization, Flash Attention, KV caching, and speculative decoding. Design and optimize distributed inference systems using approaches such as tensor parallelism, pipeline parallelism, and expert parallelism for large-scale GPU workloads. Work with cross-functional engineering and research teams to integrate optimized inference frameworks into production and edge-device applications. Define evaluation methodologies, document experimental results, compare performance against established benchmarks, and continuously refine optimization strategies. Monitor production performance and use empirical research to identify opportunities for further improvements in scalability, efficiency, and reliability. Requirements: Degree in Computer Science or a related technical field; a PhD in NLP, Machine Learning, or a related discipline is highly relevant, particularly with a strong AI research track record and publications at leading conferences. Proven expertise in Metal Shading Language (MSL), including the ability to write custom compute shaders from scratch. Demonstrated experience with low-level kernel optimization and inference optimization on mobile or other resource-constrained devices. Track record of delivering measurable improvements in inference latency, throughput, and memory footprint for domain-specific applications. Deep understanding of modern model-serving architectures, inference engines, and optimization techniques for high-performance AI deployment. Strong experience writing GPU kernels for mobile devices such as smartphones.

Looking for something similar?

Leave your email and we'll alert you when matching jobs appear.

Don't apply unprepared

We research who's interviewing you, tailor your CV and rehearse you live — first one free.

InterviewHack.ai

Prepare for the exact interview: who's interviewing you, a tailored CV, and a real coach.

Product

JobsCompanies hiringFree ATS checkerInterview-English checkSalary checkFree STAR answerResume verdict (Jev)LATAM salary reportFree coursesBlogTailored CVSpoken practiceIt's free

Remote jobs

ReactPythonFull-StackLATAMArgentinaMexicoSee all →

Prepare

Spoken practiceFrontendBackendAI EngineerBy companySell with your CV

Company

For employersAboutContactPrivacyTerms

© 2026 InterviewHack.ai · Your CV is yours. Never used to train anything. · A product of IA-PTY

Similar open roles

Account Manager (Email Marketing)

Jobgether · Switzerland

→

AI Research Engineer (Kernel & Inference Optimization)

Jobgether · Switzerland

→

Account Manager (Email Marketing)

Jobgether · France

→

(Senior or Staff) Backend Engineer, AI tooling

Jobgether · France

→