Senior Researcher in Interpretability and AI Safety

University of Oxford, Oxford

Senior Researcher in Interpretability and AI Safety

£49119-£58265

University of Oxford, Oxford

  • Full time
  • Temporary
  • Onsite working

Posted 2 weeks ago, 30 Aug | Get your application in now before you miss out!

Closing date: Closing date not specified

Job ref: 6f9ce31d10b040c8b250a18757fad4dd

Location ref: Oxford

Full Job Description

We are seeking a full time Senior Researcher to join the Technical AI Governance programme at the Department of Engineering Science in central Oxford. The post is funded by the Oxford Martin AI Governance Initiative and is fixed-term to 1 year, with the possibility of an extension for an additional year.
The Senior Researcher will work on Interpretability, evaluations and AI safety for continuously learning systems. They will evaluate systems whose behaviour keeps moving, and track, from the inside, whether the structures that carry their capabilities and safety properties hold up under the change.
You will be responsible for the follow : (full details of duties available from the Job Description)

Collaboration and engagement
You will have completed a PhD in machine learning, computer science, or a closely related field and with several years of experience and have research experience in at least one of: mechanistic or representational interpretability; model evaluation and benchmarking; continual or lifelong learning; or AI safety (alignment, scalable oversight, adversarial robustness). A strong publication record at relevant venues (NeurIPS, ICML, ICLR, ACL, EMNLP, or established AI safety venues and workshops) together with the ability to design, run, and make sense of large-scale experiments on foundation models.

Direct job link

https://www.jobs24.co.uk/job/senior-researcher-in-interpretability-ai-safety-127299800