InterviewHack.ai
Empezar gratis
Vacantes / Kimchi

Senior ML Engineer | Kimchi (LLM Inference Optimization)

Kimchisenior

En corto

  • →Ingeniero de ML senior enfocado en optimizar inferencias de LLMs para reducir costo y latencia.
  • →Día a día: tunear kernels, probar cuantizaciones y mejorar motores de inferencia como vLLM o TensorRT-LLM en entornos distribuidos.
  • →Lo destacado: tu trabajo impacta directamente en el rendimiento de clientes y en la rentabilidad de la empresa.

Ningún requisito explícito de inglés.

Postularme en la empresa ↗Compartir por WhatsApp
✓ Gratis para empezar✓ Corre en tu navegador✓ Primer dossier sin tarjeta✓ Listo en ~1 minuto

En ~1 minuto te damos: quién te entrevista, las preguntas probables con respuestas desde tu CV, y tu CV adaptado a esta vacante. Tu primer dossier es gratis.

Las preguntas que te van a hacer

1. ¿Cómo medirías el impacto real de una cuantización en un modelo LLM en producción, más allá del benchmark de texto?

2. Describe cómo optimizarías el uso de KV cache en vLLM para cargas con secuencias largas y variadas.

3. ¿Cómo manejarías el desbalanceo de carga en un cluster de inferencia con múltiples GPUs y distribución de lote dinámica?

🔒 +7 preguntas más

Sin tarjeta. Subís tu CV y en ~1 minuto tenés el dossier completo.

🎧¿Llegás a la entrevista? Llevá el copiloto. Nuestra extensión escucha la entrevista en vivo y te muestra anclas de 3-4 palabras desde tu CV y tu preparación — mirás, conectás, hablás. Gratis. Ver la extensión →

💵 USD · Remote · No visa

¿No encontrás lo que buscás? Probá Micro1

Micro1 te ubica directo en empresas de EE. UU. que pagan en USD. Un solo proceso de vetting, múltiples ofertas — sin aplicar en frío.

Que Micro1 te matchee →
📬Vacantes elegidas para TU CV, cada mañana por WhatsApp. Gratis: escribí “vacantes” y el bot te manda tus matches del día. Suscribirme →

¿Qué piden?

  • ✓5+ años construyendo sistemas de ML reales, con experiencia en infraestructura de inferencia o entrenamiento.
  • ✓Experto en Python para servicios de producción, no solo scripts.
  • ✓Experiencia práctica con vLLM, SGLang o TensorRT-LLM y comprensión profunda de su rendimiento en GPU.
  • ✓Conocimiento avanzado de cuantización: medición de pérdida de calidad, no solo ratios de compresión.
  • ✓Familiaridad con sistemas distribuidos: comunicación colectiva, estrategias de sharding y fallos en entornos multi-GPU/nodo.
  • ✓Enfoque basado en medición: instrumentar antes de optimizar, distinguiendo ganancias reales de artefactos de benchmark.

¿No cumplís todo? Es lo normal — tu dossier gratis te dice qué gaps tenés y cómo cubrirlos en la entrevista.

PythonvLLMSGLangTensorRT-LLMPyTorchCUDA-adjacent toolingKubernetesgRPClickHousePostgreSQL

¿A quién escribirle en Kimchi?

Tu dossier gratis identifica a las personas que te entrevistarían — con su background, qué valoran y cómo escribirles para destacar antes de aplicar.

Why Cast AI? Cast AI is an automation platform that operates cloud-native and AI infrastructure at scale. By embedding autonomous decision-making directly into Kubernetes and cloud environments, Cast AI continuously optimizes performance, reliability, and efficiency in production. The old way doesn't work. As Kubernetes and AI environments grow, manual decisions don’t. Cast AI replaces tickets, alerts, and manual tuning with continuous automation that adapts infrastructure as conditions change. Efficiency and cost savings follow naturally from that automation. Over 2,100 companies already rely on Cast AI, including Akamai, BMW, Cisco, FICO, HuggingFace, NielsenIQ, Swisscom, and TGS. Global team, diverse perspectives We're headquartered in Miami, but our impact is international. We take a global and intentional approach to diversity. Today, Cast AI operates across 34 countries spanning Europe, North America, Latin America, and APAC, bringing a wide range of perspectives into how we build and lead. Unicorn momentum In January 2026, we achieved unicorn status with a strategic investment from Pacific Alliance Ventures, the corporate venture arm of Shinsegae Group (a $50+ billion Korean conglomerate). Our valuation now exceeds $1 billion, and we're just getting started. Join us as we build the future of autonomous infrastructure. About the role Throughput. Latency. KV cache utilization. Move those three numbers in the right direction, and two things happen: customers get faster, cheaper inference, and our margins improve. That's the entire thesis of this role. Every kernel you tune, every quantization scheme you ship, every scheduler tweak you land shows up directly in a customer's p99 and on our P&L. This is a high-impact seat. It is also a high-autonomy seat as you'll be given the room to lead the technical direction of inference optimization at Kimchi, not execute someone else's roadmap. The problem: running LLMs in production is a moving target. The "right" model and serving configuration for a workload depend on traffic shape, sequence-length distribution, batch dynamics, GPU SKU, memory bandwidth, quantization tolerance, and a dozen other variables that shift week to week. Most teams pick a model once, over-provision GPUs, and absorb the cost. Kimchi is the system that makes that decision automatically - continuously matching workloads to the most cost-efficient, best-performing LLM and serving configuration on a customer's infrastructure. We're building the optimization layer between the model and the hardware, and we need engineers who understand both sides deeply. Stack Python; vLLM; SGLang; TensorRT-LLM; PyTorch; CUDA-adjacent tooling; Kubernetes; gRP; ClickHouse; PostgreSQL; GCP Pub/Sub; AWS / GCP / Azure; GitLab CI; ArgoCD; Prometheus; Grafana; Loki; Tempo. Requirements: • 5+ years building real ML systems, with a portfolio that shows depth in inference or training infrastructure (not just model training notebooks). • Strong Python - production services, not scripts. • Hands-on experience with at least one of vLLM, SGLang, or TensorRT-LLM, and a working mental model of why an inference engine performs the way it does on a given GPU. • Fluency with quantization tradeoffs - you've measured quality regressions, not just compression ratios. • Comfort with distributed systems: collective communication, sharding strategies, and the practical failure modes of multi-GPU and multi-node setups. • A bias toward measurement. You instrument before you optimize, and you can tell the difference between a real win and a benchmark artifact. • Self-direction. This role comes with a wide mandate; you should be excited by that, not unsettled by it. Responsibilitie

Más empleos como este

Empleos remotos de AI / ML Engineer

¿Buscando algo parecido?

Dejá tu email y te avisamos cuando salgan vacantes que coincidan con tu perfil.

No apliques sin prepararte

Investigamos quién te entrevista, adaptamos tu CV y te ensayamos en vivo — gratis la primera.

InterviewHack.ai

Preparate para la entrevista exacta: quién te entrevista, tu CV a medida y coach real.

Producto

VacantesEmpresas contratandoTodas las herramientasVeredicto de CV (Jev)Cover letter gratisPreguntas de entrevista por rol"Hablame de vos" (respuesta)Revisar CV (ATS) gratis¿Cómo suena tu inglés?¿Te pagan bien?Guion de negociación salarialRespuesta STAR gratisTitular + About de LinkedInReporte de sueldos LATAMCursos gratisBlogCV a medidaPráctica habladaPreciosAfiliados — 30%

Empleos remotos

ReactPythonFull-StackLATAMArgentinaMéxicoVer todas →

Preparate

Práctica habladaFrontendBackendAI EngineerPor empresaVendete con tu CV

Empresa

Buscás talentoAcerca deContactoPrivacidadTérminos

© 2026 InterviewHack.ai · Tu CV es tuyo. Nunca se usa para entrenar nada. · Un producto de IA-PTY

Vacantes similares activas

Staff AI Engineer, Payments Intelligence

Yuno · Europe

→

AI Engineer, Knowledge Infrastructure

Adyen · Amsterdam

→

Machine Learning Engineering Manager, Credit & Underwriting

Adyen · San Francisco

→

Principal AI Engineer - Context - Agents and Context

Elastic · Greece

→