Job Description
Era4 develops, owns and operates AI infrastructure across the UK, powered by renewable energy. Converting legacy industrial and energy sites into modern data-centre facilities, Era4 is combining brownfield regeneration opportunities with cleaner, efficient, scalable compute capacity for healthcare, research, finance, enterprise, and public-sector organisations
We are looking for Platform Engineer (HPC & AI) who canassistin shaping our new Platform team, this role will be customer facing, involve technical troubleshooting, and collaboration with vendor engineering teams to ensure seamless AI platform operations.
Responsibilities:
- Designing, deploying, and managing large‑scale HPC and GPU‑accelerated clusters, including NVIDIAbasedcomputeenvironments.
- Implementing and administering HPC scheduling and resource‑management systems (e.g.,Slurm), including GPU partitioning, workload scheduling, and capacity planning.
- Architecting and optimising InfiniBand and Ethernet network topologies.
- Ensuring high availability and resilience through failover strategies, planned maintenance coordination, and proactive risk mitigation.
- Automating provisioning, configuration, monitoring, and operational workflows across multi‑vendor HPC hardware and software stacks.
- Monitoring real‑time performance and leading troubleshooting efforts across compute, storage, interconnect, drivers, and node failures, engaging vendor support for critical issues.
- Incident response: node failure management, network issues, driver issues, troubleshooting common issues and then working with vendor support to resolve any critical issues.
- Security and access control: Manage user permissions, RBAC, security hardening, data protection.
Required Skills & Experience:
- Experience supporting HPE PCAI or other AI/HPC infrastructure and platforms.
- System administration experience with OS's like RHEL/CentOS, Ubuntu, tuning Linux kernel.
- Proficiencywith Ansible, Nvidia and CUDA toolkits,Kubernetesand container orchestration.
- Understanding of automation,monitoringand security with GPU as a service.
- Extensive experience in system engineering, platformoperationsor SRE.
- Experience with GPU resource allocation (across instances, GPUs count and time).
- Advanced networking skills with High performance networking,troubleshootingand fine tuning.
- Familiarity with cloud-based platforms, APIs, and distributed systems.
- Understanding of AI/ML concepts and tooling (model training, inference, data pipelines basics).
- Experience with monitoring/logging tools (e.g., Grafana, Kibana, Splunk).
- Excellent communication skills to interface with both customers and internal / vendor teams.
- Good understanding of tools requirements for ML engineers and data scientists, and how to optimise the experience.
Why JoinEra4:
You’llbe joining a mission-driven start-up building critical national infrastructure, where operational excellence directly enables growth. This role offers high visibility with leadership, real autonomy, and the chance to shape how a next-generation companyoperatesat scale.
Diversity & Inclusion
Era4 is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
Note:
We appreciate this is arelatively newskill set and we are open to candidates who may not tick all the boxes but are willing to learn and develop their skillset.