InterviewHack.ai
Empezar gratis
Vacantes / Hcompany

Senior Software Engineer - Paris

Hcompany·Hybrid Parissenior

En corto

  • →Construyes y mantienes el marco de evaluación que ejecuta pruebas de agentes de IA en entornos reales (web, escritorio, CLI).
  • →Trabajas con investigadores y equipos en campo para integrar benchmarks, optimizar rendimiento y garantizar resultados reproducibles.
  • →El sistema maneja 50+ benchmarks ahora y debe escalar a 100-200 con confiabilidad y bajo costo.

No se requiere inglés.

Postularme en la empresa ↗Compartir por WhatsApp
✓ Gratis para empezar✓ Corre en tu navegador✓ Primer dossier sin tarjeta✓ Listo en ~1 minuto

En ~1 minuto te damos: quién te entrevista, las preguntas probables con respuestas desde tu CV, y tu CV adaptado a esta vacante. Tu primer dossier es gratis.

Las preguntas que te van a hacer

1. ¿Cómo garantizarías resultados reproducibles en múltiples ejecuciones de un benchmark en entornos de escritorio?

2. Describe cómo diseñarías un sistema de orquestación para ejecutar benchmarks en diferentes entornos (web, CLI, desktop) con mínima latencia.

3. ¿Qué métricas y herramientas usarías para monitorear el rendimiento de un cluster de evaluación en tiempo real?

🔒 +7 preguntas más

Sin tarjeta. Subís tu CV y en ~1 minuto tenés el dossier completo.

🎧¿Llegás a la entrevista? Llevá el copiloto. Nuestra extensión escucha la entrevista en vivo y te muestra anclas de 3-4 palabras desde tu CV y tu preparación — mirás, conectás, hablás. Gratis. Ver la extensión →

💵 USD · Remote · No visa

¿No encontrás lo que buscás? Probá Micro1

Micro1 te ubica directo en empresas de EE. UU. que pagan en USD. Un solo proceso de vetting, múltiples ofertas — sin aplicar en frío.

Que Micro1 te matchee →
📬Vacantes elegidas para TU CV, cada mañana por WhatsApp. Gratis: escribí “vacantes” y el bot te manda tus matches del día. Suscribirme →

¿Qué piden?

  • ✓5+ años en desarrollo backend con Python en producción.
  • ✓Experiencia construyendo herramientas de prueba, QA o evaluación usadas por otros equipos.
  • ✓Operación de sistemas distribuidos en Kubernetes en AWS.
  • ✓Construcción y despliegue de APIs (REST/GraphQL) e integraciones externas.
  • ✓Conocimiento de bases de datos relacionales y no relacionales, y colas de mensajes (SQS, RabbitMQ, Kafka).
  • ✓Instrumentación con métricas, trazas y monitoreo desde el inicio.

¿No cumplís todo? Es lo normal — tu dossier gratis te dice qué gaps tenés y cómo cubrirlos en la entrevista.

PythonKubernetesAWSRESTGraphQLPostgreSQLSQSRabbitMQKafkaFastAPI

¿A quién escribirle en Hcompany?

Tu dossier gratis identifica a las personas que te entrevistarían — con su background, qué valoran y cómo escribirles para destacar antes de aplicar.

ai . For OSWorld that's 369 desktop tasks, each run 3 times, with the steps, tokens and time of every attempt. This role builds and runs the evaluation framework that produces runs like these. H builds computer-use agents and the models behind them. Developers use them through a managed API , and our forward deployed engineers take them into enterprise workflows. What this team owns The evaluation framework: orchestration, runtimes and observability. Researchers and forward deployed engineers bring the benchmarks, across web apps, desktop applications and the command line. Your job is to make the framework that runs them reliable, fast and cheap, and to make adding a new one quick. Research uses the results to choose checkpoints and decide whether a model ships. Product and the forward deployed engineers use them to measure agents on customer workflows. It carries roughly 50 benchmarks now. That number should be between 100 and 200 soon, and the framework has to keep up. What you'd be doing Integration support for researchers and forward deployed engineers bringing in a benchmark, with a shorter path each time. Setting the standard for how a benchmark enters the framework, and building the checks that enforce it. Scheduling and observability, so cluster capacity isn't left idle while evaluation jobs queue. Reproducible results across trials, so a release decision rests on numbers that hold. Whatever stack a benchmark calls for. One week that's cluster tuning; the next it's a browser extension or desktop environments. Time with customers, from single developers to large companies, to find out what they want measured, then automating it so the results flow back into our harnesses and models. The first few months By 3 months you'll have helped researchers or forward deployed engineers integrate 5 benchmarks, and started fixing what slows the framework down. By 6 months one part of it is yours, for example scaling the runs, observability, or a group of related benchmarks, and a release will have gone out on your numbers. By 12 months you'll know the design and trade-offs of the whole evaluation system, and be the person the rest of H asks about evaluations. Who you'd work with Ceiran Chapman, our VP Engineering, is hiring for this role. You'd join the evaluation team. The people relying on your work day to day are H's researchers and forward deployed engineers. What we think it takes Likely a good fit if you Have spent 5+ years in backend development, with production Python at the core, and use coding agents to go faster without letting quality drop. Have built test, QA or evaluation tooling that other teams depended on, and care whether a number is right. Have operated distributed systems on Kubernetes in a public cloud. AWS experience helps most. Have built and shipped systems end to end, including APIs (REST or GraphQL) and integrations with outside services. Know relational and non-relational databases, and message queues such as SQS, RabbitMQ or Kafka. Instrument what you build, with metrics, tracing and monitoring from the start. Stronger still if you have Measured LLM quality before, or built agents yourself. Packaged and run workloads in Docker and on virtual machines. Used Temporal, Dask, FastAPI, PostgreSQL, Grafana or Datadog. Automated web or desktop software with Playwright, Selenium or a browser extension you wrote. Set standards other engineers follow, through code review, design review or mentoring. You do not need a background in machine learning. We'll work that out with you. If you match most of this but not all of it, apply anyway. How we hire A 30 minute call with our Talent team, a 60 minute technical challenge, a 60 minute system design interview, and a 30 minute final conversation with Ceiran. About 3.5 hours in total. Practicalities Paris posting: Hybrid in Paris.

¿Buscando algo parecido?

Dejá tu email y te avisamos cuando salgan vacantes que coincidan con tu perfil.

No apliques sin prepararte

Investigamos quién te entrevista, adaptamos tu CV y te ensayamos en vivo — gratis la primera.

InterviewHack.ai

Preparate para la entrevista exacta: quién te entrevista, tu CV a medida y coach real.

Producto

VacantesEmpresas contratandoTodas las herramientasVeredicto de CV (Jev)Cover letter gratisPreguntas de entrevista por rol"Hablame de vos" (respuesta)Revisar CV (ATS) gratis¿Cómo suena tu inglés?¿Te pagan bien?Guion de negociación salarialRespuesta STAR gratisTitular + About de LinkedInReporte de sueldos LATAMCursos gratisBlogCV a medidaPráctica habladaPreciosAfiliados — 30%

Empleos remotos

ReactPythonFull-StackLATAMArgentinaMéxicoVer todas →

Preparate

Práctica habladaFrontendBackendAI EngineerPor empresaVendete con tu CV

Empresa

Buscás talentoAcerca deContactoPrivacidadTérminos

© 2026 InterviewHack.ai · Tu CV es tuyo. Nunca se usa para entrenar nada. · Un producto de IA-PTY

Vacantes similares activas

Legal Counsel - Commercial and AI

Hcompany · Hybrid Paris

→