InterviewHack.ai
Empezar gratis
Vacantes / Kog

Agentic Compiler Engineer

Kog·Paris, Francemid

En corto

  • →Construir un compilador autónomo que mejore la ejecución de LLMs en GPUs
  • →Trabajar con IRs, optimizaciones, verificación y bucles de feedback en hardware real
  • →Sistema que explora, compila, prueba y mide mejoras en tiempo real con GPUs estándar

English speaking

Postularme en la empresa ↗Compartir por WhatsApp

En ~1 minuto te damos: quién te entrevista, las preguntas probables con respuestas desde tu CV, y tu CV adaptado a esta vacante. Gratis, sin tarjeta.

Las preguntas que te van a hacer

1. ¿Cómo diseñarías un IR para representar operaciones de atención en LLMs con paralelismo distribuido?

2. ¿Qué estrategias usarías para buscar optimizaciones eficientes en un espacio de transformaciones grandes?

3. ¿Cómo verificarías la corrección de un kernel GPU generado que modifica el orden de operaciones en un grafo de inferencia?

🔒 +7 preguntas más

Sin tarjeta. Subís tu CV y en ~1 minuto tenés el dossier completo.

🎧¿Llegás a la entrevista? Llevá el copiloto. Nuestra extensión escucha la entrevista en vivo y te muestra anclas de 3-4 palabras desde tu CV y tu preparación — mirás, conectás, hablás. Gratis. Ver la extensión →

💵 USD · Remote · No visa

¿No encontrás lo que buscás? Probá Micro1

Micro1 te ubica directo en empresas de EE. UU. que pagan en USD. Un solo proceso de vetting, múltiples ofertas — sin aplicar en frío.

Que Micro1 te matchee →
📬Vacantes elegidas para TU CV, cada mañana por WhatsApp. Gratis: escribí “vacantes” y el bot te manda tus matches del día. Suscribirme →

¿Qué piden?

  • ✓Experiencia sólida en ingeniería de compiladores o sistemas GPU
  • ✓Conocimiento de IRs, optimización, lowering o generación de código
  • ✓Habilidad para profilear y optimizar kernels GPU
  • ✓Experiencia con inferencia de LLMs a nivel de sistema
  • ✓Capacidad para explicar decisiones técnicas y medir resultados
  • ✓Proyectos públicos o escritos técnicos detallados de trabajo relevante

¿No cumplís todo? Es lo normal — tu dossier gratis te dice qué gaps tenés y cómo cubrirlos en la entrevista.

LLVMMLIRCUDAHIPMetalVulkanGPU kernelsattentionMoEspeculative decoding

¿A quién escribirle en Kog?

Tu dossier gratis identifica a las personas que te entrevistarían — con su background, qué valoran y cómo escribirles para destacar antes de aplicar.

ABOUT KOG Kog builds a co-designed inference stack for real-time AI agents on standard datacenter GPUs, spanning model architecture, inference engine, compilers, and low-level GPU kernels. On the model side, we developed Laneformer 2B and Delayed Tensor Parallelism (DTP), a Transformer architecture that overlaps communication with useful computation and weight streaming. On the systems side, the Kog Inference Engine runs this stack on standard AMD and NVIDIA datacenter GPUs. Kog generates 3,500 tokens/s per request on 8 AMD MI300X GPUs and 2,100 tokens/s per request on 8 NVIDIA H200 GPUs, in FP16 at batch size 1, with quantization and speculative decoding disabled. Our next major project is AGCO, our agentic compiler. AGCO is designed to optimize LLMs across different GPUs and optimization targets, including very fast inference. The team has 10 people, including 9 engineers and researchers and 4 PhDs. ai. Read the technical details on the Kog Labs blog. WHAT YOU WILL WORK ON You will work directly on AGCO. The goal is to build a system that can explore ways to optimize LLM execution, generate changes, compile them, check correctness, run them on real hardware, measure the results, and use this feedback to guide the next optimization. You will contribute to areas such as: Compiler and IR design for representing and transforming LLM computations. Optimization passes, lowering, and code generation. Search methods for exploring different implementations and execution strategies. Verification and correctness checks for generated changes. GPU execution, profiling, and performance optimization. LLM inference across operators, memory, parallelism, and communication. Optimization loops that connect generated changes to measurements on real GPUs. One direction we are exploring combines an IR, a verifier, a compiler, and a search optimizer. We plan to start with focused problems, build working prototypes, and extend the system from what we learn. Your main area will depend on your experience, skills, and interests. You may focus more on compilers, GPU systems, or LLM inference while working closely with people across the full stack. WHAT WE LOOK FOR We look for engineers with deep technical expertise and original work in at least one area relevant to AGCO. Relevant experience includes: Compiler engineering, including optimization passes, IRs, lowering, code generation, LLVM, or MLIR. GPU programming with CUDA, HIP, Metal, Vulkan, or similar technologies. GPU performance work involving kernels, memory, synchronization, profiling, or hardware behavior. LLM inference engines and performance optimization. Attention, MoE, parallelism, communication, or other systems-level parts of LLM execution. Formal verification, equivalence checking, SAT/SMT, or related methods. Systems that generate, search, test, benchmark, or optimize code automatically. We care about what you personally built and the technical decisions behind it. Strong candidates can explain the problem, their approach, the alternatives they explored, and how they measured the result. We review technical work during the process. This can be public code, an upstream contribution, a paper, a thesis, a technical project, or a detailed write-up based on work you can share. WHAT WE OFFER You will join a small team building AGCO as a core part of Kog's technology. Work at the intersection of compilers, GPU systems, and LLM inference. Direct access to engineers working across the full inference stack. A fast loop from an optimization idea to compilation, execution, verification, and measurement on real GPUs. The opportunity to go deep in your strongest technical area while expanding into the other parts of the stack. High ownership over technical decisions and systems that will shape how Kog optimizes LLM inference. This role is based in Paris, and we are looking for candidates who can relocate to Paris and work closely with the team.

¿Buscando algo parecido?

Dejá tu email y te avisamos cuando salgan vacantes que coincidan con tu perfil.

No apliques sin prepararte

Investigamos quién te entrevista, adaptamos tu CV y te ensayamos en vivo — gratis la primera.

InterviewHack.ai

Preparate para la entrevista exacta: quién te entrevista, tu CV a medida y coach real.

Producto

VacantesEmpresas contratandoRevisar CV (ATS) gratis¿Cómo suena tu inglés?¿Te pagan bien?Respuesta STAR gratisVeredicto de CV (Jev)Reporte de sueldos LATAMCursos gratisBlogCV a medidaPráctica habladaEs gratis

Empleos remotos

ReactPythonFull-StackLATAMArgentinaMéxicoVer todas →

Preparate

Práctica habladaFrontendBackendAI EngineerPor empresaVendete con tu CV

Empresa

Buscás talentoAcerca deContactoPrivacidadTérminos

© 2026 InterviewHack.ai · Tu CV es tuyo. Nunca se usa para entrenar nada. · Un producto de IA-PTY