InterviewHack.ai
Empezar gratis
Vacantes / Jobgether

AI Research Engineer (Kernel & Inference Optimization)

Jobgether·Switzerlandsenior

En corto

  • →Ingeniero de investigación en IA que optimiza motores de inferencia en hardware móvil y de borde
  • →Diseña kernels personalizados en Metal Shading Language y mejora latencia/eficiencia en modelos de texto, visión y audio
  • →Destaca por aplicar técnicas de optimización como Flash Attention y KV caching en entornos reales de producción

Fluency in English required for technical collaboration

Postularme en la empresa ↗Compartir por WhatsApp

En ~1 minuto te damos: quién te entrevista, las preguntas probables con respuestas desde tu CV, y tu CV adaptado a esta vacante. Gratis, sin tarjeta.

Las preguntas que te van a hacer

1. ¿Cómo optimizarías un modelo de visión transformer para ejecutarse en un móvil con limitaciones de memoria?

2. ¿Qué técnicas usarías para reducir la latencia de generación de tokens en un modelo de texto de gran tamaño?

3. Describe un caso donde implementaste un kernel personalizado en MSL para mejorar el rendimiento de inferencia.

🔒 +7 preguntas más

Sin tarjeta. Subís tu CV y en ~1 minuto tenés el dossier completo.

🎧¿Llegás a la entrevista? Llevá el copiloto. Nuestra extensión escucha la entrevista en vivo y te muestra anclas de 3-4 palabras desde tu CV y tu preparación — mirás, conectás, hablás. Gratis. Ver la extensión →

💵 USD · Remote · No visa

¿No encontrás lo que buscás? Probá Micro1

Micro1 te ubica directo en empresas de EE. UU. que pagan en USD. Un solo proceso de vetting, múltiples ofertas — sin aplicar en frío.

Que Micro1 te matchee →
📬Vacantes elegidas para TU CV, cada mañana por WhatsApp. Gratis: escribí “vacantes” y el bot te manda tus matches del día. Suscribirme →

¿Qué piden?

  • ✓Título en Ciencias de la Computación o campo técnico relacionado
  • ✓Doctorado en NLP, Machine Learning o disciplina relacionada con publicaciones de alto impacto
  • ✓Experiencia probada en Metal Shading Language (MSL) y desarrollo de shaders de computación
  • ✓Habilidades demostradas en optimización de inferencia en dispositivos móviles o de borde
  • ✓Conocimiento profundo de arquitecturas de servido de modelos e inferencia de alto rendimiento
  • ✓Experiencia práctica en pipelines de inferencia completos, desde optimización hasta despliegue

¿No cumplís todo? Es lo normal — tu dossier gratis te dice qué gaps tenés y cómo cubrirlos en la entrevista.

Metal Shading Language (MSL)GPU kernelsMobile inferenceEdge devicesTensor parallelismPipeline parallelismExpert parallelismPruningQuantizationFlash Attention

¿A quién escribirle en Jobgether?

Tu dossier gratis identifica a las personas que te entrevistarían — con su background, qué valoran y cómo escribirles para destacar antes de aplicar.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Research Engineer (Kernel & Inference Optimization) based in Switzerland. You will work at the intersection of AI research, systems engineering, and high-performance model inference. Your focus will be on developing and optimizing model-serving architectures for advanced AI systems across a range of hardware environments. You will tackle challenges involving latency, throughput, memory efficiency, and scalability, including deployment on resource-constrained mobile and edge devices. The role combines hands-on research with low-level engineering, giving you the opportunity to develop novel inference strategies and GPU kernels. You will work with complex architectures spanning text, image, audio, diffusion models, and vision transformers. Your work will involve rigorous benchmarking, production testing, and iterative optimization to translate research into measurable performance improvements. You will collaborate with cross-functional teams in a highly technical, remote environment focused on pushing the boundaries of efficient AI systems. Accountabilities Design and deploy advanced model-serving architectures optimized for high throughput, low latency, and efficient memory utilization. Develop inference pipelines capable of operating effectively across diverse environments, including resource-constrained mobile devices and edge platforms. Establish clear performance targets covering response latency, token generation speed, throughput, memory footprint, and reliability. Build and execute controlled inference benchmarks in simulated and production environments, tracking latency, throughput, memory consumption, and error rates. Create and maintain representative datasets and simulation scenarios for evaluating model performance under real-world and resource-constrained conditions. Identify computational and memory bottlenecks across inference pipelines and implement solutions involving batching, networking, memory management, and other system-level optimizations. Develop custom GPU kernels and compute shaders for mobile hardware, including solutions written in Metal Shading Language (MSL). Apply advanced inference optimization techniques such as pruning, quantization, Flash Attention, KV caching, and speculative decoding. Design and optimize distributed inference systems using approaches such as tensor parallelism, pipeline parallelism, and expert parallelism for large-scale GPU workloads. Work with cross-functional engineering and research teams to integrate optimized inference frameworks into production and edge-device applications. Define evaluation methodologies, document experimental results, compare performance against established benchmarks, and continuously refine optimization strategies. Monitor production performance and use empirical research to identify opportunities for further improvements in scalability, efficiency, and reliability. Requirements: Degree in Computer Science or a related technical field; a PhD in NLP, Machine Learning, or a related discipline is highly relevant, particularly with a strong AI research track record and publications at leading conferences. Proven expertise in Metal Shading Language (MSL), including the ability to write custom compute shaders from scratch. Demonstrated experience with low-level kernel optimization and inference optimization on mobile or other resource-constrained devices. Track record of delivering measurable improvements in inference latency, throughput, and memory footprint for domain-specific applications. Deep understanding of modern model-serving architectures, inference engines, and optimization techniques for high-performance AI deployment. Strong experience writing GPU kernels for mobile devices such as smartphones.

¿Buscando algo parecido?

Dejá tu email y te avisamos cuando salgan vacantes que coincidan con tu perfil.

No apliques sin prepararte

Investigamos quién te entrevista, adaptamos tu CV y te ensayamos en vivo — gratis la primera.

InterviewHack.ai

Preparate para la entrevista exacta: quién te entrevista, tu CV a medida y coach real.

Producto

VacantesEmpresas contratandoRevisar CV (ATS) gratis¿Cómo suena tu inglés?¿Te pagan bien?Respuesta STAR gratisVeredicto de CV (Jev)Reporte de sueldos LATAMCursos gratisBlogCV a medidaPráctica habladaEs gratis

Empleos remotos

ReactPythonFull-StackLATAMArgentinaMéxicoVer todas →

Preparate

Práctica habladaFrontendBackendAI EngineerPor empresaVendete con tu CV

Empresa

Buscás talentoAcerca deContactoPrivacidadTérminos

© 2026 InterviewHack.ai · Tu CV es tuyo. Nunca se usa para entrenar nada. · Un producto de IA-PTY

Vacantes similares activas

AI Research Engineer (Kernel & Inference Optimization)

Jobgether · France

→

Account Manager (Email Marketing)

Jobgether · Switzerland

→

Account Manager (Email Marketing)

Jobgether · France

→

(Senior or Staff) Backend Engineer, AI tooling

Jobgether · France

→