InterviewHack.ai
Empezar gratis
Vacantes / Cloudfactory

Senior Site Reliability Engineer

Cloudfactory · Berlin, Germanysenior

En corto

  • ▸Ingeniero de confiabilidad de sitios que asegura sistemas de IA escalables y seguros en múltiples nubes.
  • ▸Gestiona infraestructura GPU, monitoreo de modelos, autoscalado y respuesta a incidentes con SLOs.
  • ▸Destaca por integrar IA en operaciones diarias, usando herramientas agenticas con juicio humano.

Fluent in English (written and spoken)

Postularme en la empresa ↗Compartir por WhatsApp

En ~1 minuto te damos: quién te entrevista, las preguntas probables con respuestas desde tu CV, y tu CV adaptado a esta vacante. Gratis, sin tarjeta.

🎧¿Llegás a la entrevista? Llevá el copiloto. Nuestra extensión escucha la entrevista en vivo y te muestra anclas de 3-4 palabras desde tu CV y tu preparación — mirás, conectás, hablás. Gratis. Ver la extensión →

¿Qué piden?

  • ✓5+ años en ingeniería de infraestructura, DevOps o SRE en sistemas de alta disponibilidad.
  • ✓Experiencia real operando clústeres de Kubernetes bajo carga productiva.
  • ✓Fluidez con Helm, Terraform o CloudFormation en al menos una nube principal (AWS preferido).
  • ✓Habilidad en Python, Go o scripting para automatización y herramientas.
  • ✓Uso diario de herramientas de IA agenticas (como Claude Code, Codex) en el flujo de trabajo.
  • ✓Razonamiento desde primeros principios: identificación de fallas y limitaciones de sistemas.

¿No cumplís todo? Es lo normal — tu dossier gratis te dice qué gaps tenés y cómo cubrirlos en la entrevista.

KubernetesHelmTerraformAWSPythonGoGrafanaIstioGitOpsModel Registries

¿A quién escribirle en Cloudfactory?

Tu dossier gratis identifica a las personas que te entrevistarían — con su background, qué valoran y cómo escribirles para destacar antes de aplicar.

At CloudFactory, we are a mission-driven team passionate about unlocking the potential of AI to transform the world. By combining advanced technology with a global network of talented people, we make unusable data usable, driving real-world impact at scale. More than just a workplace, we’re a global community founded on strong relationships and the belief that meaningful work transforms lives. Our commitment to earning, learning, and serving fuels everything we do as we strive to connect one million people to meaningful work and build leaders worth following. Our Culture At CloudFactory, we believe in building a workplace where everyone feels empowered, valued, and inspired to bring their authentic selves to work. We are: Mission-Driven: We focus on creating economic and social impact. People-Centric: We care deeply about our team’s growth, well-being, and sense of belonging. Innovative: We embrace change and find better ways to do things together. Globally Connected: We foster collaboration between diverse cultures and perspectives. If you’re passionate about innovation, collaboration, and making a real impact, we’d love to have you on board! Role Summary As a Site Reliability Engineer, you will play a key role in keeping all production systems running smoothly. You will work closely with other engineers and operators to fuse engineering principles, operational knowledge, security, and automation to work towards platform/service production excellence from an angle of infrastructure, reliability, and security. The SRE team owns the foundation of AI Platform’s Core platform - the services and infrastructure that let us deploy to a multitude of public cloud providers and that powers many ML and LLM powered features. We give every other engineering team a reliable base to build on, and we own the software delivery lifecycle end to end: the tooling, patterns, and automation that reduce friction for the whole org. This is an exciting opportunity to grow professionally while contributing to a mission-driven organization. Responsibilities: What you’ll own Reliability of platform(includes ML and LLM workloads) - model serving and inference infrastructure (GPU-backed endpoints, autoscaling, latency and cost tradeoffs), with SLOs, on-call, and incident response that cover models, not just services Observability(includes ML models) - drift and performance monitoring for ML, plus LLM-specific tracing, evals, and guardrails, wired into the same metrics and logging stacks we run everywhere else Company-wide technical direction: shaping the roadmap and building golden paths that raise the baseline for every team Developer tooling and automation that compounds - reusable GitHub Actions, GitOps workflows, Terraform modules - so every engineer ships faster Reusable components packaging common open-source tools (Grafana, Istio, CloudNative stack, and ML tooling such as model registries and feature stores) for teams to deploy in any environment Secure-by-default infrastructure - baking security, compliance audits, cost governance, and audit trails into the platform in close partnership with our lead/backend/staff engineers. Requirements Who you are (must-haves) 5+ years in infrastructure engineering, DevOps, or SRE, operating large-scale, high-availability production systems using Kubernetes Production Operational experience - a live cluster under real load, not a lab. Fluent with Helm, and Terraform or Cloudformation, on at least one major cloud (AWS preferred). Good proficiency in Python or Go or general scripting for automation and tooling(automation with higher language preferred) AI is already in your daily loop - Agentic tooling (Claude Code, Codex, Droid, internal skills) is part of how you ship and not what you are experimenting with. We believe AI tools can be great with human judgement and we want the SRE team to bring the next wave day to day operations.

Más empleos como este

Empleos remotos de DevOps Engineer

No apliques sin prepararte

Investigamos quién te entrevista, adaptamos tu CV y te ensayamos en vivo — gratis la primera.

InterviewHack.ai

Preparate para la entrevista exacta: quién te entrevista, tu CV a medida y coach real.

Producto

VacantesRevisar CV (ATS) gratis¿Cómo suena tu inglés?¿Te pagan bien?Reporte de sueldos LATAMCursos gratisBlogCV a medidaPráctica habladaEs gratis

Empleos remotos

ReactPythonFull-StackLATAMArgentinaMéxicoVer todas →

Preparate

Práctica habladaFrontendBackendAI EngineerPor empresaVendete con tu CV

Empresa

Buscás talentoAcerca deContactoPrivacidadTérminos

© 2026 InterviewHack.ai · Tu CV es tuyo. Nunca se usa para entrenar nada. · Un producto de IA-PTY

Vacantes similares activas

Forward Deployed Engineer

Cloudfactory · Berlin, Germany

→

Senior Product Designer

Cloudfactory · Berlin, Germany

→