Senior Researcher in Interpretability and AI Safety
University of Oxford, Oxford
Senior Researcher in Interpretability and AI Safety
£49119-£58265
University of Oxford, Oxford
- Full time
- Temporary
- Onsite working
Posted 2 weeks ago, 30 Aug | Get your application in now before you miss out!
Closing date: Closing date not specified
Job ref: 6f9ce31d10b040c8b250a18757fad4dd
Location ref: Oxford
Full Job Description
We are seeking a full time Senior Researcher to join the Technical AI Governance programme at the Department of Engineering Science in central Oxford. The post is funded by the Oxford Martin AI Governance Initiative and is fixed-term to 1 year, with the possibility of an extension for an additional year.
The Senior Researcher will work on Interpretability, evaluations and AI safety for continuously learning systems. They will evaluate systems whose behaviour keeps moving, and track, from the inside, whether the structures that carry their capabilities and safety properties hold up under the change.
You will be responsible for the follow : (full details of duties available from the Job Description)
Collaboration and engagement
You will have completed a PhD in machine learning, computer science, or a closely related field and with several years of experience and have research experience in at least one of: mechanistic or representational interpretability; model evaluation and benchmarking; continual or lifelong learning; or AI safety (alignment, scalable oversight, adversarial robustness). A strong publication record at relevant venues (NeurIPS, ICML, ICLR, ACL, EMNLP, or established AI safety venues and workshops) together with the ability to design, run, and make sense of large-scale experiments on foundation models.