Senior HPC Systems Administrator

University of Oxford, Oxford

Senior HPC Systems Administrator

£49119-£58265

University of Oxford, Oxford

  • Full time
  • Permanent
  • Onsite working

Posted 1 week ago, 9 Sep | Get your application in now before you're too late!

Closing date: Closing date not specified

Job ref: 542b5ea149a5432a9038898790a44efd

Location ref: Oxford

Full Job Description

We are seeking a full-time Senior HPC Systems Administrator to join the Department of Engineering Science at the University of Oxford, working within the internationally recognised Visual Geometry Group (VGG). This is an exciting opportunity for an experienced systems professional to lead the development, administration, and strategic evolution of the groups high-performance computing (HPC) infrastructure supporting cutting-edge AI, computer vision, and multimodal learning research. The successful candidate will take a leading role in the design, deployment, maintenance, and optimisation of large-scale HPC systems, including CPU and GPU clusters, high-performance storage, and advanced networking technologies. The postholder will work closely with academic researchers, software engineers, and departmental IT teams to ensure robust, scalable, and secure computational infrastructure capable of supporting world-leading research in machine learning and visual computing. You will possess extensive experience in Linux systems administration and HPC environments, together with strong technical expertise in cluster management, storage systems, networking, scripting, and infrastructure automation. Experience with GPU computing environments, containerisation technologies, cloud platforms, and scientific software environments will be highly desirable. The role also requires excellent communication skills and the ability to collaborate effectively with both technical and non-technical stakeholders. The role includes responsibility for:

  • Designing, building, and maintaining HPC clusters and associated infrastructure
  • Managing Linux compute and storage environments
  • Supporting GPU-enabled research computing systems
  • Maintaining high-performance storage and backup solutions
  • Monitoring system performance, security, and availability
  • Supporting and mentoring researchers, software engineers, and postgraduate students
  • Developing documentation, training materials, and operational best practices
  • Collaborating with departmental and university-wide research IT teams
  • Evaluating and deploying new technologies to support evolving research requirements
  • The successful candidate will work under the direction of the departmental HPC Infrastructure Architect, with approximately 20% of their time allocated to broader departmental infrastructure activities.

  • Degree in Computer Science, Engineering, or a related technical discipline
  • Extensive experience managing HPC infrastructure in research or technical environments
  • Strong Linux systems administration expertise
  • Experience with GPU servers and high-performance networking technologies
  • Experience with scripting and automation (e.g. Bash, Python)
  • Knowledge of storage systems and backup/archive procedures
  • Strong troubleshooting and systems integration skills
  • Excellent written and verbal communication abilities
  • Ability to work independently and collaboratively in a research environment
  • Desirable Skills
  • Experience with SLURM or other job scheduling systems
  • Experience with containerisation and virtualisation technologies
  • Knowledge of cloud computing platforms
  • Familiarity with scientific computing tools and AI/ML research workflows
  • Experience supporting research software or academic computing environments
  • Knowledge of high-performance file systems such as BeeGFS or Lustre

Direct job link

https://www.jobs24.co.uk/job/senior-hpc-systems-administrator-127365566