InterviewHack.ai
Start free
Jobs / Tensordyne

Forward Deployed Inference Engineer

Tensordyne·Munich, Bavaria, Germanymid

In short

  • →Ingeniero de inferencia avanzada que optimiza modelos de IA en hardware especializado para clientes.
  • →Trabaja con modelos reales (LLM, VLM, diffusion) desde benchmarking hasta despliegue en producción.
  • →Destacado: interfaz directa entre ingeniería y clientes en fases críticas de despliegue de hardware de nueva generación.

Experiencia con herramientas de desarrollo impulsadas por IA (ej. Claude Code, Cursor)

Apply on company site ↗Share on WhatsApp
✓ Free to start✓ Runs in your browser✓ First dossier, no card✓ Ready in ~1 minute

In ~1 minute you get: who interviews you, the likely questions answered from your CV, and your CV tailored to this job. Your first dossier is free.

The questions they'll ask you

1. ¿Cómo evaluarías el rendimiento de un LLM MoE en hardware nuevo, y qué métricas priorizarías?

2. Describe un caso donde optimizaste un modelo de inferencia: qué herramientas usaste y qué mejoras lograste.

3. ¿Cómo abordarías un desfase entre resultados de simulación y hardware real en un PoC con cliente?

🔒 +7 more questions

No card. Upload your resume and the full dossier is ready in ~1 minute.

🎧Land the interview? Bring the copilot. Our free extension listens to the live interview and flashes 3-4-word anchors from your resume and prep — glance, connect, talk. Get the extension →

💵 USD · Remote · No visa

Not finding what you want? Try Micro1

Micro1 places engineers directly at US companies paying in USD. One vetting, multiple offers — no cold applying.

Get matched by Micro1 →
📬Jobs picked for YOUR resume, every morning on WhatsApp. Free: text “vacantes” and the bot sends your daily matches. Subscribe →

What they ask for

  • ✓Experiencia práctica con modelos de IA y sistemas de inferencia, especialmente LLM densos y MoE.
  • ✓Habilidades sólidas en Python y PyTorch, incluyendo modificación de código de modelos.
  • ✓Capacidad para benchmarking, análisis de rendimiento y resolución de problemas de latencia/throughput.
  • ✓Experiencia en despliegue de modelos en hardware no estándar o aceleradores nuevos.
  • ✓Habilidades de comunicación y colaboración en equipos multidisciplinarios.
  • ✓Conocimiento de herramientas de desarrollo impulsadas por IA (ej. Claude Code, Cursor).

Don't tick every box? That's normal — your free dossier shows your gaps and how to cover them in the interview.

PythonPyTorchLLMsMoEVLMDiffusion modelsvLLMSGLangRustKubernetes

Who should you write to at Tensordyne?

Your free dossier identifies the people who'd interview you — their background, what they value, and how to reach out so you stand out before applying.

About Tensordyne Tensordyne is building a new class of AI inference system designed for high-performance, power-efficient deployment of the world's most demanding generative AI workloads. Our platform combines purpose-built silicon, new AI math, optimized scale-up networking, and memory architecture into a tightly integrated system purpose built for large-scale AI inference. We work with hyperscalers, Neoclouds, frontier model developers, enterprises, and infrastructure partners operating at the leading edge of AI. As Tensordyne moves from system development into silicon bring-up, customer validation, beta deployments, and production rollout, we are building the technical customer organization that will sit directly between our engineering teams and the companies deploying the platform. Role summary We are looking for a Forward Deployed Inference Engineer who combines deep AI systems expertise with strong customer instincts. This person will own the path from a customer workload or model request to a technical result and, where needed, to an optimized model running successfully on Tensordyne hardware and software. The role sits at the intersection of model architecture, inference performance, systems optimization, developer tooling, and customer deployment. You will work hands-on with engineering while also acting as a technical bridge to Product, BizDev, Sales, and customers. What you will do Turn customer workloads into fast, credible performance answers through profiling & benchmarking . Define relevant KPIs, compare against competitive baselines, and keep our evaluation methodology current with external benchmarks. Model enablement & optimization: convert and bring up customer models on the Tensordyne stack, validate numerical quality, identify performance bottlenecks, and work with compiler, runtime, kernel, and system teams to improve results. Deployment / forward engineering: work directly with customers and partners on technical PoCs, integration, deployment, and debugging; translate requirements into measurable acceptance criteria for quality, latency, throughput, and other relevant KPIs. Track profiling-to-hardware accuracy by continuously comparing profiling/simulation results with actual hardware deployments, explain material gaps, and flag missing capabilities in the compiler, SDK, inference server, KV-cache management, or adjacent systems to the owning teams. Turn repeated customer-specific learnings into reusable tooling, documentation, benchmarks, or product improvements. Core qualifications Strong hands-on experience with AI models and inference systems , especially dense and MoE LLMs (Llama, DeepSeek, Qwen, GPT-OSS, Kimi, GLM), and VLM, speech and diffusion models. Strong Python and PyTorch skills and the ability to understand and modify model code. Experience profiling, benchmarking, or optimizing model inference and reasoning about latency, throughput, memory, and utilization. Strong problem-solving and communication skills, with the ability to drive ambiguous technical problems across team boundaries. , Claude Code, Cursor). Strong pluses Experience with LLM serving and deployment stacks such as vLLM, SGLang, or similar systems. Experience working directly with customers or external technical partners . Experience bringing models up on new accelerators or non-standard hardware , including performance debugging across framework/runtime/hardware boundaries. Practical experience with production inference techniques or environments such as quantization, distributed inference, or Kubernetes . Experience navigating and contributing to Rust codebases. Tensordyne Values Think big. Pursue ambitious technical and business goals. Aim for excellence. Quality matters in everything we build and deliver. Own it and get it done. Take responsibility and drive results. Operate with integrity. Be direct, transparent, and respectful. Win as a team.

Looking for something similar?

Leave your email and we'll alert you when matching jobs appear.

Don't apply unprepared

We research who's interviewing you, tailor your CV and rehearse you live — first one free.

InterviewHack.ai

Prepare for the exact interview: who's interviewing you, a tailored CV, and a real coach.

Product

JobsCompanies hiringAll free toolsResume verdict (Jev)Free cover letterInterview questions by roleTechnical assessment simulator"Tell me about yourself" answerFree ATS checkerInterview-English checkSalary checkSalary negotiation scriptFree STAR answerLinkedIn headline + AboutLATAM salary reportFree coursesBlogTailored CVSpoken practicePricingAffiliates — 30%

Remote jobs

ReactPythonFull-StackLATAMArgentinaMexicoSee all →

Prepare

Spoken practiceFrontendBackendAI EngineerBy companySell with your CV

Company

For employersAboutContactPrivacyTerms

© 2026 InterviewHack.ai · Your CV is yours. Never used to train anything. · A product of IA-PTY

Similar open roles

Jr Software Engineer - ML Runtime - Rust

Tensordyne · Munich, Bavaria, Germany

→