Practice by Numbers

DevOps / Site Reliability Engineer – Engineering

Practice by Numbers  •  Kolkata, IN (Onsite)  •  3 hours ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

DevOps / Site Reliability Engineer – Engineering

Experience: 6+ Years

Location: Kolkata, India / Gurgaon, India (On-site)

Shift: Normal or Night — depending on candidate preference and business need Reports to: Lead DevOps / SRE

Employment Type: Full-time

About Practice by Numbers

Practice by Numbers (PbN) is a fast-growing SaaS platform helping healthcare organizations leverage data, automation, patient engagement, and operational intelligence to improve business performance and patient outcomes. We build highly scalable cloud-native products serving thousands of customers across North America.

We are looking for a DevOps/SRE Engineer who writes Terraform and CI/CD pipelines. This is a builder’s role, not a monitoring seat.

Role Overview

You will own Terraform modules across our AWS infrastructure — compute, networking, databases, IAM, DNS — and build the GitHub Actions pipelines that ship our Django, Python, Node.js, and React services.

This role can be staffed on a normal shift or a night shift, depending on your preference and our coverage needs. On a night shift you get an uninterrupted window for the changes that are hardest to make in daylight — database migrations, pipeline rewrites, provider upgrades, and staged cutovers — and you will often be the only engineer on shift. That means real autonomy, real ownership, and the expectation that you can reason through an unfamiliar system on your own and write up clearly what you did.

This role suits someone who wants deep infrastructure work with production ownership and prefers building over ticket-shuffling.

Key Responsibilities

Infrastructure as Code

  • Own and extend Terraform modules across AWS — ECS/Fargate or EKS, RDS, VPC networking, IAM, ALB/NLB, S3, Route 53, CloudWatch.
  • Manage Terraform state safely; write and review plans that reviewers can trust.
  • Eliminate manually created resources by bringing them under code (via terraform import, refactors, and module extraction).
  • Keep environments (dev, QA, production) consistent and reproducible.

CI/CD Engineering

  • Build and maintain GitHub Actions pipelines for build, test, containerization, and deployment.
  • Migrate legacy pipelines (Jenkins, CircleCI, GitLab CI) onto a single, maintainable platform.
  • Design safe deployment paths — staged rollouts, health-gated releases, fast and reliable rollback. Keep pipelines fast: caching, parallelism, test splitting, and honest gating.

Reliability & Operations

  • Own the production change window on your shift: deploys, migrations, cutovers, and maintenance.
  • Run and improve observability — dashboards, SLOs, alerting that fires on real user impact rather than noise.
  • Debug containerized applications in production: logs, metrics, traces, resource limits, networking.
  • Participate in on-call rotation and incident response; drive blameless post-incident reviews. Improve cost efficiency and resource utilization across the AWS footprint.

Databases & Data Services

  • Operate PostgreSQL in production — read query plans, identify slow queries, understand connections, locks, and replication.
  • Plan and execute schema migrations against live systems with minimal disruption.
  • Operate supporting data services: Redis/ElastiCache, message brokers, object storage.

Automation & Tooling

  • Write Python and shell automation to remove repetitive operational work.
  • Build tooling that makes the wider team faster — self-service scripts, runbooks, guardrails.
  • Harden secrets handling, access control, and infrastructure security posture.

Handover & Communication

  • Write clear, complete handover notes at the end of every shift. Where shifts do not overlap, your writing is how the rest of the team learns what happened.
  • Maintain runbooks and infrastructure documentation as systems change.
  • Coordinate with the wider engineering team on planned work and follow-ups.

AI-Enabled Engineering

  • Use modern AI tools to accelerate infrastructure work, scripting, and troubleshooting.
  • Apply AI-assisted practices while maintaining strong engineering, security, and review standards.

Required Qualifications

  • 6+ years in DevOps, SRE, Platform, or Infrastructure engineering with genuine production ownership.
  • Strong AWS experience — container orchestration (ECS/Fargate or EKS), RDS, VPC networking, IAM, load balancing, S3, CloudWatch.
  • Hands-on Terraform: writing modules, managing state, reviewing plans.
  • Practical CI/CD experience, ideally GitHub Actions. GitLab CI, CircleCI, or Jenkins backgrounds are fine if you can migrate.
  • Docker, and comfort debugging containerized applications in production.
  • Solid Linux fundamentals and shell scripting; Python for automation.
  • Working PostgreSQL knowledge — query plans, slow queries, connections, locks, replication. Experience deploying and operating Django/Python and Node.js/React applications. Hands-on use of an observability platform in anger (New Relic, Datadog, Grafana, Prometheus, or similar).
  • Clear written English. Your handover notes are how the rest of the team learns what happened. Openness to either a normal or a night shift, and willingness to participate in an on-call rotation.

Technical Expertise

Cloud & Infrastructure

  • AWS (ECS / Fargate / EKS, EC2, Lambda)
  • VPC, Subnets, Security Groups, NAT, Peering
  • ALB / NLB, Route 53, CloudFront
  • IAM, Roles, Policies, Least-Privilege Design
  • RDS, ElastiCache, S3
  • Terraform / Infrastructure as Code

CI/CD & Automation

  • GitHub Actions
  • Jenkins / CircleCI / GitLab CI
  • Docker, Container Registries
  • Blue-Green & Staged Deployments, Rollback Strategies
  • Python, Bash / Shell Scripting
  • Git and Trunk-Based Workflows

Observability & Reliability

  • New Relic / Datadog / Grafana / Prometheus
  • CloudWatch Metrics, Logs, Alarms
  • OpenTelemetry
  • SLOs, Error Budgets, Alert Design
  • Incident Management & Post-Incident Review

Data & Messaging

  • PostgreSQL Administration & Tuning
  • Schema Migrations on Live Systems
  • Redis / ElastiCache
  • Kafka / Redpanda / Amazon MSK, RabbitMQ, NATS

Security & Compliance

  • Secrets Management (AWS Secrets Manager, SOPS, Vault)
  • Network and Access Hardening
  • Vulnerability and Patch Management
  • Audit Logging

Preferred Qualifications

  • Celery, RabbitMQ, NATS, or Kafka in production.
  • Redis / ElastiCache operations.
  • Secrets management (AWS Secrets Manager, SOPS, Vault).
  • Experience with sharded or multi-tenant database architectures.
  • Healthcare or other compliance-sensitive environments (HIPAA, SOC 2).
  • Telephony / VoIP infrastructure exposure.
  • Kubernetes.
  • Prior experience as the sole engineer on shift, or on a night/off-hours rotation.

What We Look For

  • A builder’s instinct — you would rather codify a fix than repeat it.
  • Comfort with autonomy and sound judgment on when to escalate.
  • Careful, methodical change management on production systems.
  • Strong written communication; your handover notes are a first-class deliverable.
  • Curiosity about how systems actually behave, not just how they are supposed to. Ownership, follow-through, and continuous learning.

Why This Role

  • Real ownership of infrastructure, not a ticket queue.
  • Flexibility on shift, and a protected change window for high-impact work.
  • Modern stack: AWS, Terraform, GitHub Actions, Docker, PostgreSQL, Kafka-compatible messaging.
  • Direct impact on the reliability of a platform used by thousands of healthcare practices.
Practice by Numbers

About Practice by Numbers

Practice by Numbers is an all-in-one software for dental practices, consolidating analytics, patient communication, reputation management, online scheduling and payments, insurance verification, digital forms, and dashboards to streamline daily operations and improve efficiency and profitability.

Industry
IT & Software
Company Size
51-200 employees
Headquarters
Unknown
Year Founded
2015
Social Media