Postdoctoral Researcher in Large Language Model Development

University of Oxford, Norham Manor, Oxford

Postdoctoral Researcher in Large Language Model Development

£39424-£41636

University of Oxford, Norham Manor, Oxford

  • Full time
  • Temporary
  • Onsite working

Posted 2 weeks ago, 23 Sep | Get your application in now before you miss out!

Closing date: Closing date not specified

Job ref: c27f039ecae840918c7c104ed930a456

Location ref: Norham Manor, Oxford

Full Job Description

Humanities Divisional Office, Stephen A. Schwarzman Centre for the Humanities, Radcliffe Observatory Quarter, Woodstock Road, Oxford, OX2 6GG
Contract
Fixed term for 12 months
Hours
Full-time (37.5 per week)
About the role
The postholder is expected primarily to deliver the core objectives of the BodleianLLM project: curating, cleaning and documenting a multilingual humanities training corpus drawn from Oxford's GLAM collections; performing continual pre-training and fine-tuning of an open-weight foundation model on the University's own computing infrastructure; and, working with subject specialists and curators, designing and validating the first dedicated benchmark suite for evaluating language models on humanities research tasks.
You will also be expected to publish the model weights, code, corpus documentation and benchmarks as fully open-source resources; to contribute to the technical roadmap for a major follow-on funding bid; to collaborate on research publications and present at conferences; and to help deliver the project's one-day workshop for the GLAM, digital humanities and AI communities. You will report to the Principal Investigator, Professor Glenn Roe, with day-to-day technical supervision from a Senior Research Software Engineer at Digital Scholarship at Oxford (DiSc).

You will hold a relevant PhD/DPhil or have substantial experience of machine learning in a research or industry environment, and have the ability and willingness to combine machine learning research with sustained engagement with historical and multilingual source material. You will have a strong understanding of deep learning and natural language processing, with practical experience of training, adapting and deploying large language models.
You will also be able to process and transform historical, multilingual or otherwise non-standard textual data and to design and validate evaluation benchmarks; have strong Python and software engineering skills; have a strong publication record commensurate with experience; and have excellent communication and organisational skills. Experience of cultural heritage collections, reading knowledge of historical languages, and familiarity with HPC and MLOps practice would be desirable.

Direct job link

https://www.jobs24.co.uk/job/postdoctoral-researcher-in-large-language-model-development-127443315