InterviewHack.ai
Empezar gratis
Vacantes / Datadog

Senior Software Engineer, Chaos Engineering

Datadog·Parissenior

En corto

  • →Construyes sistemas de ingeniería de caos para probar fallos en producción con seguridad.
  • →Automatizas evacuaciones y recuperaciones zonales en Kubernetes usando herramientas avanzadas.
  • →Destaca por liderar gamedays y conectar hallazgos a soluciones reales con IA y automatización.

Conocimiento técnico en inglés, para documentación y colaboración.

Postularme en la empresa ↗Compartir por WhatsApp
✓ Gratis para empezar✓ Corre en tu navegador✓ Primer dossier sin tarjeta✓ Listo en ~1 minuto

En ~1 minuto te damos: quién te entrevista, las preguntas probables con respuestas desde tu CV, y tu CV adaptado a esta vacante. Tu primer dossier es gratis.

Las preguntas que te van a hacer

1. ¿Cómo diseñarías un sistema de evacuación zonal segura en Kubernetes con rollback automático?

2. ¿Qué mecanismos usarías para limitar el radio de explosión en una prueba de caos en producción?

3. ¿Cómo integrarías IA para proponer escenarios de fallo basados en métricas de observabilidad?

🔒 +7 preguntas más

Sin tarjeta. Subís tu CV y en ~1 minuto tenés el dossier completo.

🎧¿Llegás a la entrevista? Llevá el copiloto. Nuestra extensión escucha la entrevista en vivo y te muestra anclas de 3-4 palabras desde tu CV y tu preparación — mirás, conectás, hablás. Gratis. Ver la extensión →

💵 USD · Remote · No visa

¿No encontrás lo que buscás? Probá Micro1

Micro1 te ubica directo en empresas de EE. UU. que pagan en USD. Un solo proceso de vetting, múltiples ofertas — sin aplicar en frío.

Que Micro1 te matchee →
📬Vacantes elegidas para TU CV, cada mañana por WhatsApp. Gratis: escribí “vacantes” y el bot te manda tus matches del día. Suscribirme →

¿Qué piden?

  • ✓Experto en sistemas distribuidos y modos de fallo.
  • ✓Experiencia con Kubernetes y ciclos de vida de workloads.
  • ✓Capacidad para diseñar sistemas seguros con controles de radio de explosión.
  • ✓Habilidades en comunicación técnica clara (documentos, postmortems).
  • ✓Colaboración interequipos para mejorar resiliencia.
  • ✓Experiencia en ingeniería de confiabilidad o pruebas de fallo en producción.

¿No cumplís todo? Es lo normal — tu dossier gratis te dice qué gaps tenés y cómo cubrirlos en la entrevista.

KubernetesgRPCPythonGoBashDockerAWSDatadogPostgreSQLRedis

¿A quién escribirle en Datadog?

Tu dossier gratis identifica a las personas que te entrevistarían — con su background, qué valoran y cómo escribirles para destacar antes de aplicar.

Datadog’s Chaos Engineering team builds systems that surface reliability weaknesses before they become outages. As a Senior Software Engineer, you will initially focus on zonal resilience, building automation that helps services safely evacuate and recover from zonal failures, while also contributing to fault injection, incident replay, gameday orchestration, and reliability tooling. You will work across engineering teams to design systems that safely exercise production failure modes and turn findings into verified remediation. You will also help advance the use of AI and automation to identify, test, and close resilience gaps as Datadog’s software and infrastructure evolve. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: • Build zonal-resilience automation that coordinates safe workload evacuations, switchovers, and recovery in partnership with the teams that own affected services. • Design and build fault-injection systems for production environments, including infrastructure- and application-level testing, incident replay, and controlled resilience experiments. • Develop safeguards such as blast-radius controls, kill switches, validation mechanisms, and rollback paths that keep production experiments contained and reversible. • Build agents and automation that help propose failure scenarios, triage experiment results, and connect reliability findings to tracked remediation and verification. • Lead gamedays from hypothesis and scenario design through execution, documented findings, remediation tracking, and validation of completed fixes. • Design and implement reliable distributed systems, including gRPC services, Kubernetes controllers, and shared platform components, while contributing to technical design and mentoring other engineers. Who You Are: • You have strong distributed systems fundamentals and can reason about consistency, failure modes, backpressure, idempotency, quorum, retries, and failure recovery. • You understand Kubernetes workload lifecycles, including how pods, controllers, scheduling, draining, and eviction interact with resilient system design. • You have experience designing, building, or operating production systems where safety, availability, and controlled failure handling are important. • You communicate complex technical decisions clearly through design documents, runbooks, postmortems, and cross-functional technical discussions. • You are comfortable collaborating across engineering teams to understand unfamiliar systems, identify failure modes, and drive resilience improvements. • Experience with reliability engineering, chaos engineering, zonal failover, AI-assisted operational workflows, traffic interception, or large-scale observability systems is beneficial but not required. Datadog values people from all walks of life. We know not everyone will meet all the above qualifications on day one. That’s okay. If you’re passionate about technology and want to grow your experience, we encourage you to apply. Benefits and Growth: • Develop deep expertise in distributed systems, production resilience, Kubernetes, and large-scale infrastructure. • Work on reliability systems that operate across Datadog’s production environment and influence how engineering teams design for failure. • Grow your experience designing safe, automated approaches to fault injection, zonal resilience, and incident reproduction. • Explore practical applications of AI and automation to reliability engineering and operational workflows. • Collaborate with engineers across infrastructure

¿Buscando algo parecido?

Dejá tu email y te avisamos cuando salgan vacantes que coincidan con tu perfil.

No apliques sin prepararte

Investigamos quién te entrevista, adaptamos tu CV y te ensayamos en vivo — gratis la primera.

InterviewHack.ai

Preparate para la entrevista exacta: quién te entrevista, tu CV a medida y coach real.

Producto

VacantesEmpresas contratandoTodas las herramientasVeredicto de CV (Jev)Cover letter gratisPreguntas de entrevista por rol"Hablame de vos" (respuesta)Revisar CV (ATS) gratis¿Cómo suena tu inglés?¿Te pagan bien?Guion de negociación salarialRespuesta STAR gratisTitular + About de LinkedInReporte de sueldos LATAMCursos gratisBlogCV a medidaPráctica habladaPreciosAfiliados — 30%

Empleos remotos

ReactPythonFull-StackLATAMArgentinaMéxicoVer todas →

Preparate

Práctica habladaFrontendBackendAI EngineerPor empresaVendete con tu CV

Empresa

Buscás talentoAcerca deContactoPrivacidadTérminos

© 2026 InterviewHack.ai · Tu CV es tuyo. Nunca se usa para entrenar nada. · Un producto de IA-PTY

Vacantes similares activas

Staff Software Engineer - Security Agent

Datadog · Remote

→

Senior Software Engineer - REDAPL Graph Engine

Datadog · Remote

→

Major Account Manager (EMEA)

Datadog · Remote

→

Enterprise Customer Success Manager (German Speaking)

Datadog · Paris

→