Compute Orchestration & Scheduling

Microsoft, City of Westminster

Compute Orchestration & Scheduling

Salary not available. View on company website.

Microsoft, City of Westminster

  • Full time
  • Permanent
  • Onsite working

Posted today, 10 Oct | Get your application in now to be one of the first to apply.

Closing date: Closing date not specified

Job ref: 9494ae5dae16451d8537930d88ca537e

Location ref: City of Westminster

Full Job Description

Microsoft AI is looking for engineers to build the compute infrastructure powering frontier-model development. The Compute Orchestration & Scheduling team owns cluster orchestration, workload scheduling, resource allocation, quota management, and the systems that enable fast initialization and reliable fault recovery on next-generation GPU supercomputers. Our team prepares infrastructure for the next generation of distributed training and inference across multiple locations and a rapidly growing fleet of accelerators, including NVIDIA Grace Blackwell, Vera Rubin, and AMD GPU platforms, alongside the CPU, network, and storage systems they depend on. Key challenges include topology-aware placement, large-scale distributed training, fast and reliable initialization of training runtimes, rapid recovery for long-running AI jobs, observability into cluster behavior, and improved utilization across increasingly large and diverse accelerator fleet. Better scheduling efficiency and platform reliability translate directly into more effective compute for AI research and products. You will work closely with researchers, model engineers, hardware architects, and infrastructure teams to turn frontier-model requirements into scalable platform capabilities. We value engineers who navigate ambiguity, remove roadblocks, and deliver improvements to users quickly and iteratively. Responsibilities

  • Develop and tune the compute infrastructure stack allocating NVIDIA Grace Blackwell (GB) and Vera Rubin (VR) resources to AI workloads.
  • Scale GPU clusters across hardware generations to thousands of accelerators and beyond.
  • Use operational data and workload insights to inform the compute (GPU and CPU) roadmap for large-scale AI research.
  • Partner with model-development teams to improve the infrastructure used to train and serve AI models.
  • Find practical ways around roadblocks and deliver improvements rapidly, iterating with users in a fast-paced, design-driven environment.
  • Embody Microsoft's culture and values.

    A bachelor's degree in computer science or a related technical field and at least six years of engineering experience writing code in languages such as C, C++, Python, Go, or JavaScript; or equivalent practical experience., Significant additional engineering experience, with a master's degree or equivalent practical experience, building production software and distributed systems.
  • Experience with Ray, Kubernetes, Kueue, Volcano, or another AI-focused system for scheduling, scaling, or fault tolerance is especially relevant.
  • Software Engineering IC5 - The typical base pay range for this role across United Kingdom is £ 93,500.00 - £ 161,800.00 per year. Certain roles may be eligible for benefits and other compensation.

Direct job link

https://www.jobs24.co.uk/job/compute-orchestration-scheduling-127560283