Darwoft

1009 - Site Reliability Engineer (AWS/Kubernetes/AI Observability) · Senior · Remoto · LATAM

Darwoft  •  Buenos Aires, AR (Remote)  •  3 hours ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

1009 - Site Reliability Engineer (AWS/Kubernetes/AI Observability) · Senior · Remoto · LATAM

General Information

  • Location: Anywhere in LATAM
  • Work Model: 100% Remote
  • Contract Type: Contractor
  • Project: AI Platform Infrastructure & Observability
  • Industry: Healthtech
  • Time Zone: Availability to collaborate with US and LATAM teams
  • English Level: B2 / C1
  • Seniority: Senior

Get to Know Darwoft

At Darwoft, we build digital products that create real impact. We are a Latin American technology company working with international organizations to develop scalable, innovative, and human-centered solutions.

Our remote-first culture is based on trust, collaboration, continuous learning, and technical excellence.

About the Role

We are looking for a Senior Site Reliability Engineer with strong experience in observability, cloud infrastructure, Kubernetes, and production reliability.

You will be responsible for the operational health and visibility of AI-powered systems, including LLM applications, AI API gateways, token-consumption pipelines, and model-serving infrastructure.

The role combines traditional SRE practices with emerging AI platform operations. You do not need experience with every AI tool available, but you should understand the current AI ecosystem and the challenges of operating LLM-enabled services in production.

What You'll Be Doing

  • Design and operate gateway infrastructure for LLM and AI API traffic.
  • Manage routing, authentication, throttling, rate limiting, retries, and traffic shaping.
  • Build observability for token consumption, model latency, API usage, costs, errors, retries, quotas, and provider availability.
  • Create dashboards and alerts using Prometheus, Grafana, Loki, CloudWatch, or similar tools.
  • Instrument LLM applications to capture prompt, completion, and request telemetry securely.
  • Define and maintain SLIs, SLOs, alerting strategies, and error budgets.
  • Build resilient architectures using fallback models, provider redundancy, caching, circuit breakers, and graceful degradation.
  • Lead incident response for provider outages, gateway saturation, latency issues, quota exhaustion, and unexpected cost increases.
  • Automate infrastructure and operational workflows using Terraform, CI/CD, and scripting.
  • Maintain AWS and Kubernetes environments supporting AI-powered applications.
  • Collaborate with AI, engineering, platform, security, and FinOps teams.
  • Mentor engineers and promote observability and reliability best practices.

What You Bring

  • 7+ years of experience in SRE, DevOps, Platform Engineering, Cloud Infrastructure, or similar roles.
  • Strong experience operating distributed and highly available production systems.
  • Advanced hands-on experience with AWS, ideally including Bedrock, SageMaker, Lambda, API Gateway, EKS, CloudWatch, IAM, and networking services.
  • Strong production experience with Kubernetes and containerized workloads.
  • Experience building observability stacks with Prometheus, Grafana, Loki, CloudWatch, or equivalent technologies.
  • Experience defining application-level and API-level metrics.
  • Solid understanding of API gateway patterns, routing, authentication, throttling, retries, and traffic monitoring.
  • Experience defining SLIs, SLOs, alerts, and error budgets.
  • Understanding of LLM applications, AI APIs, token-based consumption, latency, quotas, and cost monitoring.
  • Proficiency in Python, Go, Bash, or a similar language.
  • Experience with Terraform or another Infrastructure as Code tool.
  • Experience building and maintaining CI/CD pipelines.
  • Strong troubleshooting skills across AWS, Kubernetes, networking, APIs, and distributed systems.
  • Experience participating in or leading production incident response.
  • English proficiency at B2 level or higher.

Nice to Have

  • Experience with Kong AI Gateway, Portkey, LiteLLM, AWS Bedrock, or similar platforms.
  • Familiarity with OpenTelemetry and distributed tracing.
  • Experience monitoring prompts, completions, token usage, model latency, and provider errors.
  • Knowledge of AI FinOps, cost allocation, token budgets, usage forecasting, or anomaly detection.
  • Experience with multi-provider routing, fallback models, semantic caching, or service mesh technologies.
  • Experience in healthtech or another regulated industry.

What Darwoft Offers

  • Contractor agreement with payment in USD
  • 100% remote work
  • Argentina's public holidays
  • English classes
  • Referral program
  • Access to learning platforms

Explore this and other opportunities at:

www.darwoft.com/careers

Darwoft

About Darwoft

Darwoft is the leading nearshore software team accelerating innovation for healthcare companies worldwide. From disruptive startups to global enterprises, we deliver secure, powerful, and user-focused solutions that scale.

Our expertise includes MVP development, application modernization, UX/UI design, staff augmentation, and advanced Data + AI services that empower real-time decision-making and intelligent automation. We work across a variety of industries including healthcare, fintech, telecommunications, retail strategy, and direct sales.

With agile methodologies and cross-functional teams, we deliver value from day one, helping our clients move fast, stay competitive, and build products their users love.

At Darwoft, we foster a collaborative and growth-oriented work culture where technology meets creativity. We believe in strong partnerships, continuous learning, and building meaningful solutions that make an impact.

Let’s connect. Let’s build what’s next, together.

Contact us at www.darwoft.com

Industry
IT & Software
Company Size
51-200 employees
Headquarters
Hillsboro, Oregon
Year Founded
2010
Social Media