InterviewHack.ai
Empezar gratis
Vacantes / tensordyne

Forward Deployed Inference Engineer

tensordyne·Munichmid

En corto

  • →Ingeniero de inferencia dedicado a optimizar modelos de IA en hardware personalizado de Tensordyne.
  • →Trabajas directamente con clientes para desplegar, benchmarkear y corregir modelos (LLM, VLM, diffusion) en entorno de producción.
  • →Destacado: Eres el puente técnico entre ingeniería interna y clientes en etapas críticas de validación y despliegue en hardware nuevo.

Proficiency in written and spoken English required

Postularme en la empresa ↗Compartir por WhatsApp
✓ Gratis para empezar✓ Corre en tu navegador✓ Primer dossier sin tarjeta✓ Listo en ~1 minuto

En ~1 minuto te damos: quién te entrevista, las preguntas probables con respuestas desde tu CV, y tu CV adaptado a esta vacante. Tu primer dossier es gratis.

Las preguntas que te van a hacer

1. ¿Cómo perfilarías y optimizarías un modelo LLM para inferencia en hardware nuevo, considerando latencia y memoria?

2. ¿Qué diferencias prácticas has observado entre simulación y despliegue real en aceleradores NoC o custom chips?

3. ¿Cómo transformarías un caso de uso de cliente repetitivo en una herramienta o benchmark reutilizable?

🔒 +7 preguntas más

Sin tarjeta. Subís tu CV y en ~1 minuto tenés el dossier completo.

🎧¿Llegás a la entrevista? Llevá el copiloto. Nuestra extensión escucha la entrevista en vivo y te muestra anclas de 3-4 palabras desde tu CV y tu preparación — mirás, conectás, hablás. Gratis. Ver la extensión →

💵 USD · Remote · No visa

¿No encontrás lo que buscás? Probá Micro1

Micro1 te ubica directo en empresas de EE. UU. que pagan en USD. Un solo proceso de vetting, múltiples ofertas — sin aplicar en frío.

Que Micro1 te matchee →
📬Vacantes elegidas para TU CV, cada mañana por WhatsApp. Gratis: escribí “vacantes” y el bot te manda tus matches del día. Suscribirme →

¿Qué piden?

  • ✓Experiencia hands-on con modelos de IA, especialmente LLM densos y MoE (Llama, Qwen, GPT-OSS, etc.)
  • ✓Dominio de Python y PyTorch, capacidad para modificar código de modelos
  • ✓Habilidad para benchmarkear e optimizar inferencia (latencia, throughput, memoria)
  • ✓Experiencia en despliegue o integración con clientes o partners técnicos
  • ✓Capacidad para solucionar problemas complejos entre frameworks, runtime y hardware
  • ✓Uso práctico de herramientas de desarrollo impulsadas por IA (ej. Claude Code, Cursor)

¿No cumplís todo? Es lo normal — tu dossier gratis te dice qué gaps tenés y cómo cubrirlos en la entrevista.

PythonPyTorchLLMsMoE modelsVLMDiffusion modelsvLLMSGLangcompilerruntime

¿A quién escribirle en tensordyne?

Tu dossier gratis identifica a las personas que te entrevistarían — con su background, qué valoran y cómo escribirles para destacar antes de aplicar.

About Tensordyne Tensordyne is building a new class of AI inference system designed for high-performance, power-efficient deployment of the world’s most demanding generative AI workloads. Our platform combines purpose-built silicon, new AI math, optimized scale-up networking, and memory architecture into a tightly integrated system purpose built for large-scale AI inference. We work with hyperscalers, Neoclouds, frontier model developers, enterprises, and infrastructure partners operating at the leading edge of AI. As Tensordyne moves from system development into silicon bring-up, customer validation, beta deployments, and production rollout, we are building the technical customer organization that will sit directly between our engineering teams and the companies deploying the platform. Role summary We are looking for a Forward Deployed Inference Engineer who combines deep AI systems expertise with strong customer instincts. This person will own the path from a customer workload or model request to a technical result and, where needed, to an optimized model running successfully on Tensordyne hardware and software. The role sits at the intersection of model architecture, inference performance, systems optimization, developer tooling, and customer deployment. You will work hands-on with engineering while also acting as a technical bridge to Product, BizDev, Sales, and customers. What you will do • Turn customer workloads into fast, credible performance answers through profiling & benchmarking . Define relevant KPIs, compare against competitive baselines, and keep our evaluation methodology current with external benchmarks. • Model enablement & optimization: convert and bring up customer models on the Tensordyne stack, validate numerical quality, identify performance bottlenecks, and work with compiler, runtime, kernel, and system teams to improve results. • Deployment / forward engineering: work directly with customers and partners on technical PoCs, integration, deployment, and debugging; translate requirements into measurable acceptance criteria for quality, latency, throughput, and other relevant KPIs. • Track profiling-to-hardware accuracy by continuously comparing profiling/simulation results with actual hardware deployments, explain material gaps, and flag missing capabilities in the compiler, SDK, inference server, KV-cache management, or adjacent systems to the owning teams. • Turn repeated customer-specific learnings into reusable tooling, documentation, benchmarks, or product improvements. Core qualifications • Strong hands-on experience with AI models and inference systems , especially dense and MoE LLMs (Llama, DeepSeek, Qwen, GPT-OSS, Kimi, GLM), and VLM, speech and diffusion models. • Strong Python and PyTorch skills and the ability to understand and modify model code. • Experience profiling, benchmarking, or optimizing model inference and reasoning about latency, throughput, memory, and utilization. • Strong problem-solving and communication skills, with the ability to drive ambiguous technical problems across team boundaries. • Proficiency in using AI-powered developer tools (e.g., Claude Code, Cursor). Strong pluses • Experience with LLM serving and deployment stacks such as vLLM, SGLang, or similar systems. • Experience working directly with customers or external technical partners . • Experience bringing models up on new accelerators or non-standard hardware , including performance debugging across framework/runtime/hardware boundaries. • Practical experience with production inference techniques or environments such as <strong

¿Buscando algo parecido?

Dejá tu email y te avisamos cuando salgan vacantes que coincidan con tu perfil.

No apliques sin prepararte

Investigamos quién te entrevista, adaptamos tu CV y te ensayamos en vivo — gratis la primera.

InterviewHack.ai

Preparate para la entrevista exacta: quién te entrevista, tu CV a medida y coach real.

Producto

VacantesEmpresas contratandoTodas las herramientasVeredicto de CV (Jev)Cover letter gratisPreguntas de entrevista por rolSimulador de test técnico"Hablame de vos" (respuesta)Revisar CV (ATS) gratis¿Cómo suena tu inglés?¿Te pagan bien?Guion de negociación salarialRespuesta STAR gratisTitular + About de LinkedInReporte de sueldos LATAMCursos gratisBlogCV a medidaPráctica habladaPreciosAfiliados — 30%

Empleos remotos

ReactPythonFull-StackLATAMArgentinaMéxicoVer todas →

Preparate

Práctica habladaFrontendBackendAI EngineerPor empresaVendete con tu CV

Empresa

Buscás talentoAcerca deContactoPrivacidadTérminos

© 2026 InterviewHack.ai · Tu CV es tuyo. Nunca se usa para entrenar nada. · Un producto de IA-PTY