Job Description
NextSilicon is revolutionizing high-performance computing. Our innovative coprocessor technology dramatically accelerates supercomputers, propelling them into a new era. Our software-defined hardware architecture empowers HPC/AI to deliver groundbreaking discoveries across all areas of advanced research. We're seeking a dynamic and results-oriented HPC/AI Systems Administrator to join our team.
At NextSilicon, everything we do is guided by three core values:
- Professionalism We strive for exceptional results through professionalism and unwavering dedication to quality and performance.
- Unity Collaboration is key to success. That's why we foster a work environment where every employee can feel valued and heard.
- Impact We're passionate about developing technologies that make a meaningful impact on industries, communities, and individuals worldwide.
Join our Field Deployment & Systems team as an HPC/AI Systems Administrator
As an HPC/AI Systems Administrator at NextSilicon, you will be central to sustaining the successful operation of HPC/AI systems. You will stand-up and maintain HPC/AI hardware and software resources. You will tune and configure systems for high-quality benchmarking efforts. You will ensure that the health and accessibility of the HPC/AI systems is top-notch via cluster management tools and capacity planning efforts.
This is a highly technical, execution-focused individual contributor role with no people management or leadership responsibilities at this time.
Location Hybrid in either our Austin, TX or Minneapolis, MN offices preferred but Remote considered for exceptional candidates.
Requirements
- Bachelor’s degree in engineering, mathematics, computer science, related field, or equivalent experience. Advanced degree is a plus.
- 5-10+ years of experience with HPC/AI system administration.
- Deep understanding of HPC & AI technologies and software ecosystems
- Experience in a fast-paced, entrepreneurial environment is a plus
- Ability to travel within the USA approx. 4 times per year
- US citizenship with eligibility to visit US government research facilities
Responsibilities
- Administer, install, monitor, and maintain HPC/AI systems, including compute nodes, storage, networking, and software stacks.
- Develop and maintain automation tools for system provisioning, configuration management, and monitoring.
- Install, configure, and optimize job scheduling and resource management tools (e.g., Slurm).
- Assist in system security, patch management, and troubleshooting operational issues.
- Contribute to performance benchmarking, system tuning, and capacity planning.
- Deploy and maintain commonly used HPC/AI applications, software stacks, and technologies (e.g., MPI, containers, spack, modules)
- Document system administration procedures and contribute to knowledge-sharing initiatives.
- Support researchers by providing technical expertise and resolving escalated support tickets.
- Participate in vendor coordination, system procurement, and hardware/software lifecycle management.