0% found this document useful (0 votes)
4 views153 pages

Request

The document outlines various student projects offered by different organizations, including TOGL Energy Limited, the Division of Biomedical and Life Sciences, and Smart Odds Ltd., focusing on AI, data science, and digital transformation. Each project has specific deliverables, technical and general skill requirements, and the opportunity for students to gain hands-on experience in their respective fields. The projects emphasize collaboration, independence, and the application of advanced technologies like AI and machine learning in real-world scenarios.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views153 pages

Request

The document outlines various student projects offered by different organizations, including TOGL Energy Limited, the Division of Biomedical and Life Sciences, and Smart Odds Ltd., focusing on AI, data science, and digital transformation. Each project has specific deliverables, technical and general skill requirements, and the opportunity for students to gain hands-on experience in their respective fields. The projects emphasize collaboration, independence, and the application of advanced technologies like AI and machine learning in real-world scenarios.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

PRJ01-TOG-AIS - AI Scaffold - TOGL Energy Limited

Project for AI

Project Description
Project name AI Scaffold
Stipend offered No stipend offered
Host Organisation TOGL Energy Limited
Project description This project challenges students to design the foundations of a
unified AI ecosystem that can transform how a modern business
operates. The goal is to map internal and external workflows,
identify where AI can automate tasks or enhance decision-making,
and then develop an integration blueprint showing how different AI
tools, LLMs, automation engines, data platforms, and workflow
systems, can connect into a single, intelligent productivity layer.
Students will analyse real business processes, spot automation
opportunities that reduce recruitment pressure, and design the
architecture that links AI tools together to boost performance across
teams. This is hands-on, forward-looking work at the intersection of
AI, operations, and organisational transformation.

Student deliverables • A mapped set of internal and external business processes.


• Identification of high-value AI automation and optimisation
opportunities.
• A proposed AI architecture showing how tools (LLMs,
automation engines, data flows, and APIs) can be integrated.
• A short written report summarising findings, recommended
priorities, risks, and expected productivity impact.
• Process diagrams, workflow prototypes, or simple proof-of-
concept automations.

Student’s line manager Luke Buckley - COO

Project location(s) Remote, employees work from home locations meeting held in
Fraser House Lancaster

Student Requirements
Essential technical skills • Basic understanding of AI and machine learning concepts
(LLMs, automation tools, data flows).
• Competence in process mapping tools (e.g., Miro, Lucidchart,
Visio, or similar).
• Good analytical and problem-solving skills for evaluating
workflows and identifying opportunities.
• Basic data handling skills, such as working with spreadsheets,
simple databases, or structured datasets.

Desirable technical skills • Experience with AI tools or APIs (e.g., OpenAI, automation
platforms).
• Understanding of business systems integration, APIs, and
data architectures.
• Familiarity with workflow automation tools (Power
Automate, Zapier, n8n, etc.).
• Knowledge of Python, SQL, or similar for prototyping or
analysing data.

Essential general skills • Ability to organise and manage their own work with minimal
supervision.
• Strong communication skills for discussing processes with
different teams.
• Ability to break down complex tasks into structured, logical
components.
• Attention to detail when mapping workflows or
documenting findings.

Desirable general skills • Curiosity and initiative, with a willingness to explore new
technologies.
• Comfortable working in cross-functional teams and gathering
information from stakeholders.
• Strong presentation and writing skills for summarising
findings and recommendations.
• Ability to apply critical thinking to identify risks, blockers, and
opportunities.

Essential experience • Experience working independently on a structured project or


coursework.
• Exposure to team-based work, group assignments, or
collaborative environments.
• Some experience researching or analysing technical or
business problems.

Desirable experience • Experience using or experimenting with AI tools, automation


platforms, or coding projects.
• Prior involvement in business analysis, service design, or
process optimisation projects.
• Experience producing reports, diagrams, or structured
documentation.

Other comments This placement will suit a student who is curious, analytical, and
excited about the practical application of AI in real business
environments. A willingness to experiment, ask questions, and
propose ideas is highly valued. The project offers a unique
opportunity to influence a real AI strategy and gain exposure to the
operational challenges of scaling a technology business.
Organisation Details
Main contact name Will Maden
Main contact position CEO
Applications email
Applications email cc
Other comments This is a particularly exciting placement because the student will be
joining a very early-stage SaaS start-up where they can directly
shape how AI is embedded into the heart of the business. We are
actively exploring new and experimental approaches, from AI-
assisted coding and automated documentation all the way through
to AI-driven customer acquisition, onboarding, and support. The
student will have the opportunity to influence how AI models and
automation are woven into every layer of the company’s operations
and culture. This is a fast-paced, high-creativity environment where
bold ideas are encouraged, and the work produced will tangibly
impact how the company evolves.
PRJ04-DBL-MDN - Mitochondrial DNA deletion mutations in the
mouse: where, what types, how many? - Division of Biomedical and
Life Sciences
Project for DS

Project Description
Project name Mitochondrial DNA deletion mutations in the mouse: where, what
types, how many?
Stipend offered No
Host Organisation Division of Biomedical and Life Sciences
Project description Mitochondrial DNA deletion mutations may be an independent
fundamental cause of ageing however we don’t know how frequent
they are naturally. This project aims to describe and quantify
deletion mutations in mitochondrial DNA across individual mice and
between tissue types within the mice, using 7TB of deep sequencing
data already generated.
Custom software in Python exists to undertake this work but needs
implementing for this task and may need some modification.
Student deliverables The student will be expected to deliver a set of results describing the
mutational landscape within and between mice, presented in tabular
and graphical ways. Some statistical analysis can also be done.
Student’s line manager Dr David Clancy (BLS)

Project location(s) Wherever adequate computing access is available, but it will have to
be on campus due to IT security constraints, unless the student has a
10TB portable drive.
Other comments No

Student Requirements
Essential technical skills Student should be comfortable working in a unix environment,
acquiring, loading, running and error-checking software output.
Some Python expertise
Desirable technical skills Knowledge or R might be a help but should not be strictly necessary.
Essential general skills Basic understanding of project management, Able to organise own
work etc.
Desirable general skills Independent learning and working, willingness to read to understand
unfamiliar tasks.
Essential experience Unix, Python.
Desirable experience Use of virtual machines at Lancaster
Other comments It could help if the student has interest in big data, genetics.

Organisation Details
Main contact name Dr David Clancy
Main contact position Lecturer
Applications email Dr David Clancy
Applications email cc No
Other comments No
PRJ06-SOL-PHR - Predicting Horse Racing from Pre and Post-Race
Commentaries - Smart Odds Ltd.
Project for DS

Project Description
Project name Predicting Horse Racing from Pre and Post-Race Commentaries

Stipend offered N/A


Host Organisation Smart Odds Ltd.
Project description Horse racing commentary, both pre-race and post-race, provides a
wealth of information about the horses, jockeys, trainers, and
conditions. These textual descriptions contain nuanced insights into
horse performance, form, and other factors that influence race
outcomes. Harnessing natural language processing (NLP) methods to
extract structured features from these commentaries offers an
opportunity to develop advanced predictive models.

The aim of the project is to develop an NLP pipeline to process and


analyse Racing Post pre-race and post-race commentaries. This
pipeline should derive meaningful covariates/features from the text
that can be used in predictive models for race outcomes or betting
strategies. Within this there will be the need to identify key entities
and themes (e.g. horse stamina, track suitability, jockey skill), handle
ambiguity, and evaluating the predictive power of the extracted
features. There are multiple approaches that students could take, for
example (but certainly not limited to), basic classification from
keyword detection, moving on to Bag of Words techniques, and
transformer-based models (e.g. BERT).

We will provide a dataset of Horse Racing commentaries. These


include descriptions of horse form, jockey, trainer, track conditions,
along with analysis of race performance and results. We will also
provide the race results (outcome data) and betting odds.

Coding skills in R, or python will be necessary, and some existing


knowledge of NLP libraries is desirable. Some pre-existing knowledge
of horse racing is beneficial but not necessary although an interest in
sports, and sports modelling is highly recommended to help
understand general concepts.

Student deliverables Documented, encoded model predicting Horse race results from
commentary analysis
Student’s line manager Gavin Whitaker, Data Scientist
Project location(s) Student’s own location
Other comments A high degree of independence will be expected from the student
Student Requirements
Essential technical skills Modelling, coding in R or Python
Desirable technical skills
Essential general skills Independence, motivation, ability to organise own work, effective
communication
Desirable general skills
Essential experience
Desirable experience
Other comments

Organisation Details
Main contact name Gavin Whittaker
Main contact position Data Scientist
Applications email
Applications email cc
Other comments
PRJ07-SOL-P - Predicting the outcome of golf match-ups - Smart Odds
Ltd.
Project for DS

Project Description
Project name Predicting the outcome of golf match-ups

Stipend offered N/A


Host Organisation Smart Odds Ltd.
Project description Golf tournaments occur on a near weekly basis (often with more
than one taking place at the same time), and within these
tournaments players compete against all the others to finish first by
taking the lowest number of shots. However, bookmakers often
offer match-ups between 2 players over the tournament, with the
player taking the lowest number of shots (and finishing higher up the
leaderboard) being the winner. It is these match-ups where interest
lies.

We are interested in developing a model to predict the winner of


these match-ups. We can provide a dataset of past scores for each
player for all rounds of a tournament, along with details on the
course, details on each player, and details on the tourname nt. We
can also provide the match-ups of interest for each tournament.

The challenge is then to try to build a model to predict the winner of


each match-up, together with relevant measures of its accuracy.
There are many methods which could be implemented/explored, for
example (but certainly not limited to) Bradley-Terry type models. A
starting place could be time-invariant effects for each player and
then trying to improve from there.

Coding skills in R, or python will be necessary. Some pre-existing


knowledge of golf is beneficial but not necessary although an
interest in sports, and sports modelling is highly recommended to
help understand general concepts of each match-up.

Student deliverables Documented, encoded model predicting golf matchup outcomes


Student’s line manager Gavin Whitaker, Data Scientist
Project location(s) Student’s own location
Other comments A high degree of independence will be expected from the student

Student Requirements
Essential technical skills Modelling, coding in R or Python
Desirable technical skills
Essential general skills Independence, motivation, ability to organise own work, effective
communication
Desirable general skills
Essential experience
Desirable experience
Other comments

Organisation Details
Main contact name Gavin Whittaker
Main contact position Data Scientist
Applications email
Applications email cc
Other comments
PRJ09-JCD-EAI - Exploring the role of agentic AI in supporting
collaboration between digital, data, and technology professions when
delivering digital transformation - Informed Solutions
Project for AI

Project Description
Project name Exploring the role of agentic AI in supporting collaboration between
digital, data, and technology professions when delivering digital
transformation
Stipend offered £3,000
Host Organisation James Cruddas, Director of Engineering, Informed Solutions
Project description Informed Solutions is an award-winning digital transformation,
technology, and data science consultancy. We create value for our
clients by providing advice and expertise in digital transformation –
improving the performance of an organisation through better use of
technology and data.

Data science plays a critical role in successful digital transformation


by allowing organisations to generate new value and insights from
their data. Examples of the practical benefits that data science can
create for organisations include: (1) identifying new and better
business insights from data, such as risks, opportunities and trends;
(2) automating labour-intensive manual business processes; and (3)
making better quality evidence-based decisions.

Agentic Artificial Intelligence (AI) is a relatively new and exciting field


of AI. Google defines Agentic AI as follows:

“Agentic AI is an advanced form of AI focused on autonomous


decision-making and action. Unlike traditional AI, which primarily
responds to commands or analyses data, agentic AI can set goals,
plan, and execute tasks with minimal human intervention. This
emerging technology has the potential to revolutionise various
industries by automating complex processes and optimising
workflows.”

Given this potential, organisations are now accelerating investment


in exploring how agentic AI can solve real-world business problems.

For example, imagine that a central Government Department has


appointed Informed to deliver a new [Link] digital service that
millions of citizens will use to access public services. Delivering a
complex service of this nature requires close collaboration between
different digital, data, and technology (DDaT) professionals. Business
Analysts must synthesise user research findings into user needs and
functional requirements. Interaction Designers must take these
requirements and design a web-based user interface that provides
an intuitive user experience. Software Engineers will need to
develop front-end and back-end code to implement the user
interface. Test Engineers will need to perform unit and system tests
to verify that the developed code works as expected.

Informed is seeking to assess the quality of solutions proposed by


agentic AI using a role-based collaboration approach, with each
agent assuming a specific DDaT role. The project will examine
whether this specialisation produces tangible quality
improvements when compared with a single, non-specialised
agent.

For example, one or more AI agents could support the role of


Business Analyst and use generative AI tools to generate a candidate
Backlog of user needs and functional requirements expressed as
User Stories and acceptance criteria expressed as Behaviour-Driven
Development (BDD) Given-When-Then statements. Another AI
agent could support the role of Interaction Designer and generate
user interface design options for satisfying the specified
requirements in the form of HTML mock-ups. Another AI agent
could support the role of Technical Architect and propose candidate
design patterns or API specifications to implement each user
interface. Finally, an overarching AI agent could orchestrate the
work of individual agents to produce a cohesive solution design.

The project will involve:

A literature review of agentic AI capabilities, architectures, and


design patterns that are relevant to the problem domain.
A technical proof of concept to demonstrate the feasibility of
implementing specialised AI agents that work in a coordinated
fashion to perform the two tasks of: (1) supporting the role of
Business Analyst by using generative AI tools to produce a candidate
Backlog of user needs and functional requirements expressed as
User Stories and acceptance criteria expressed as BDD Given-When-
Then statements; and, (2) supporting the role of Interaction Designer
by generate user interface design options for satisfying the specified
requirements.
Based on the technical proof of concept, an analysis of the strengths,
weaknesses, opportunities, and threats of using agentic AI to
automate these activities, and options for when and how human-in-
the-loop supervision may be employed as a technique for assuring
quality and minimising risk.
Testing with Informed staff who specialise in the relevant DDaT
roles, including an optional comparison with human-produced
solution designs for past work to assess similarity or variation.
Student deliverables • Dissertation report.
• Technical Proof of concept demonstrator.
Student’s line manager James Cruddas
Project location(s) At least four weeks (33%) can be located at the company’s offices in
Altrincham, spread over the course of the placement according to a
schedule that will be agreed between the student and the placement
supervisor at the outset. The remainder of the student’s time can be
spent working remotely at Lancaster University.
Other comments N/A

Student Requirements
Essential technical skills An understanding/appreciation of, and strong interest in: (1) agentic
AI concepts, architectures and design patterns, such as those
described by the Google Cloud Architecture Center; and, (2)
supervised, unsupervised, and reinforcement-based learning
approaches.
Desirable technical skills An understanding/appreciation of the software development
lifecycle and software engineering patterns/practices.
Essential general skills • Strong analytical and problem-solving skills.
• Comfortable managing uncertainty and complexity.
• Positive, autonomous, resourceful, and self-motivated.
• Requires minimal daily management.
Ability to communicate clearly and concisely.
Desirable general skills Ability to abstract a problem to find a generic solution.
Essential experience A strong interest in the application of agentic AI to real-world
business problems, and the management, governance, and user
implications this poses.
Desirable experience Computer Science or Software Engineering background.
Other comments N/A

Organisation Details
Main contact name James Cruddas
Main contact position Data Scientist
Applications email To whom should we send the students’ applications?
Applications email cc
Other comments
PRJ12-BF-AAI - Agentic AI workflows for Marketing & Sales - Brain-
Feed
Project for AI

Project Description
Project name Agentic AI workflows for Marketing & Sales
Stipend offered TBD. Travel expenses covered
Host Organisation Brain-Feed
Project description The project - Agentic Ai workflows for Marketing & Sales
An opportunity to transform brain feed’s outreach strategy and work
on a real business challenge with measurable impact. You will build
an agentic AI workflow that finds and connects with individuals who
share brain feed’s mission and can help spread the word about our
products.

Why is it important?
Sales & Marketing outreach is currently manual and time-consuming,
involving finding individuals, personalising messages, sending them,
and tracking follow-ups. Automating the process means brain feed
can find the right people faster and use the extra time to build
stronger relationships and grow the community.

The solution
A fully automated, Agentic AI outreach workflow that acts as an
extension of the digital marketing team. You will use an off-the-shelf
Saas agentic Ai solution to create a workflow that:
● Identifies the right individuals to connect with
● Extracts accurate contact details
● Organises leads into a spreadsheet
● Deploys outreach with personalised messages and follow-
ups via email

Learning outcomes
By the end of the placement, you will have:
● Designed and deployed an Agentic AI workflow
● Understood and integrated APIS and Webhooks
● Tested and refined your AI workflow
● Created AI prompts and understood the importance of
prompt engineering
Taken ownership of the project from start to finish
Student deliverables 1. Retool
2. [Link]
The project will be collaborative. You will learn how to use industry-
standard project management tools to manage tasks, sub-tasks and
reporting on project progression.

The outcome will be at least 2 successfully deployed agentic


workflows with defined efficiency savings and quantifiable results,
which will be directly attributed to company revenue.
Student’s line manager To whom will the student report during the placement?

Project location(s) We would welcome you into the office to be part of our team in
Manchester, and, in our experience, the learning and knowledge
transfer is usually quicker when conducted in person. However, we
acknowledge that Lancaster to Manchester City Centre on a regular
basis is expensive and time-consuming and therefore the suggestion
would be 1 day per fortnight (linked to key deliverables/inflexion
points and demo days)
Other comments We are a technology company as much as we are a consumer brand.
We’ve built an AI model, we are mining millions of genetic data
points to link gene variants to mental health conditions, which
requires meta-analysis, cohort analysis and gene cluster analysis.
And, we are building a wolrd 1st proprietary database to validate
precision nutrition as a medicine for brain health.

We work on the premise that if you are smart enough, you are good
enough. Regardless of age or experience. You will have complete
ownership of critical projects, and this approach attracts self-starters
who can take the learning modules provided to learn, execute and
deploy.

Student Requirements
Essential technical skills Python & JOSN.

Desirable technical skills A basic understanding of SQL, JavaScript would be useful but not
essential
Essential general skills Self-starter
Well organised
Desirable general skills
Essential experience
Desirable experience
Other comments

Organisation Details
Main contact name James Lister
Main contact position MD
Applications email james@[Link]
Applications email cc
Other comments
PRJ13-LU-ESS - Explainable Satellite Surveillance towards Climate
Changes - Lancaster University
Project for DS AI

Project Details
Project name Explainable Satellite Surveillance towards Climate Changes
Stipend offered N/A
Host Organisation Lancaster University
Project description The project aims to leverage advanced satellite imagery from LANSATs
and SENTINELS, and AI-driven analytics to monitor and assess
environmental changes. By integrating multi-source satellite data with
explainable AI tools, this initiative will enable real-time tracking of
climate phenomena, such as deforestation, glacier melting, and rising
sea levels. The project emphasizes interpretability, ensuring
actionable insights for policymakers and scientists. For example, the
study by Chuvieco et al. (2021) on satellite-based wildfire monitoring
showcases how remote sensing can support climate action. This
research contributes to sustainable solutions, bridging technologyand
environmental science for a resilient future.
Chuvieco, E., et al (2021). Remote Sensing Contributions to Global
Climate Change Studies. Remote Sensing of Environment, 256, 112293.
[Link]
Student deliverables Test your AI models to explain global climate changes.
Student’s line manager Richard Jiang
Project location(s) On campus
Other comments N/A

Student Requirements
Essential technical skills Coding, Basics on Deep Neural Networks, LANDSAT, SENTINEL
Desirable technical skills As above
Essential general skills Good academic writing
Desirable general skills As above
Essential experience Data Science Research
Desirable experience As above
Other comments

Organisation Details
Main contact name Dr Richard Jiang
Main contact position
Main contact email r.jiang2@[Link]
Other comments
PRJ15-LU-AIE - AI-Enhanced Quantum Cryptography - Lancaster
University
Project for DS AI

Project Details
Project name AI-Enhanced Quantum Cryptography
Stipend offered
Host Organisation Lancaster University
Project description The project aims to explore the integration of AI with Quantum Key
Distribution (QKD) to enhance secure communication systems.
Quantum Key Distribution, which uses quantum mechanics to securely
exchange cryptographic keys, faces challenges in scalability and
robustness. By incorporating AI algorithms, such as machine learning
models, the project seeks to optimize key distribution protocols,
improve error correction, and enhance security in the face of quantum
attacks. This research could pave the way for more efficient and secure
quantum communications, crucial for the future of cybersecurity. For
background, see Lo, H.-K., Curty, M., & Qi, B. (2014). Secure quantum
key distribution.
[Link]
[Link]

Student deliverables Product prototypes


Student’s line manager Richard Jiang
Project location(s) On campus
Other comments N/A

Student Requirements
Essential technical skills Python, QISKIT, Neural Networks
Desirable technical skills As above
Essential general skills Good academic writing
Desirable general skills As above
Essential experience Data Science Research
Desirable experience As above
Other comments

Organisation Details
Main contact name Dr Richard Jiang
Main contact position
Main contact email r.jiang2@[Link]
Other comments
PRJ16-LU-LLM - Large Language Model based Stock Trading APP -
Lancaster University
Project for DS AI

Project Details
Project name Large Language Model based Stock Trading APP
Stipend offered
Host Organisation Lancaster University
Project description This project aims to develop a risk-aware trading agent that integrates
reinforcement learning (RL) and large language models (LLMs) for
improved financial decision-making. By extending the Conditional
Value-at-Risk Proximal Policy Optimization (CPPO) algorithm, we aim
to develop a model incorporating risk assessment and trading
recommendations generated from financial news using DeepSeek V3,
Qwen 2.5, and Llama 3.3. The model will be backtested on the
Nasdaq-100 index with financial news from FNSPID, aiming to
enhance trading strategies by leveraging LLM-driven market insights
for better risk management and decision-making in dynamic financial
environments.

Student deliverables Test your models/APPs on real fin data


Student’s line manager Richard Jiang
Project location(s) On campus
Other comments Basic codes and models are available for tuning and leveraging.

Student Requirements
Essential technical skills Python coding, Deep Learning
Desirable technical skills Professional academic writing (research papers)
Other comments Applicants are expected to be proficient in coding and highly
motivated.

Organisation Details
Main contact name Dr Richard Jiang
Main contact position
Main contact email r.jiang2@[Link]
Other comments SCC
PRJ18-T-TAI - Traybakes AI Catalyst - Traybakes
Project for AI

Project Description
Project name Traybakes AI Catalyst
Stipend offered N/A
Host Organisation Traybakes
Project description Background: Traybakes produces handmade cakes which are sold to
both independent retailers and national wholesale accounts. We are
a very labour-intensive business due to the product being handmade
and as this is a USP and one of our founder’s key philosophies this
area of the business is unlikely to change. However, we do have a
front office team, within in this team there is space for a social media
and marketing manager which we have struggled to recruit. The role
requires a very wide skill set from graphic design, social media
management to more traditional forms of marketing such as
catalogue and brochure entries. We believe that given the correct
tools that the responsibilities of this job could be spread around
between existing staff.

Project Brief: Traybakes’ project would be to create an AI agent that


can create bespoke marketing using the photography, font and
brand messaging that we would supply it. We would also like it to
analyse other products in the sector and learn and improve the
marketing and brand imagery based on more successful campaigns.
Our current issue with AI marketing is the lack of flexibility in
adapting it. Currently the AI is very ‘stubborn’, and we struggle to
make tweaks without having to overhaul the entire prompt. We
would require this to be an easily operable agent due to some
member of the business not being very tech literate.

The student undertaking this task will be able to work remotely


however we would like them to visit towards to the start of the
project to get a better understanding of Traybakes. We feel that this
will affect how they go about the project as well as giving the
student the opportunity to see the challenges we face rather than
just being briefed on them. Whilst we will have a structured set of
meetings to allow us to understand where the project is, we want
the student to use their expertise and creativity to test different
options without the pressure of reporting on this to us. We want
them to work with autonomy and hope that this will allow the
student to get more out of the project rather than by giving them a
strict brief to stick to.
Whilst we will not be able to contribute to a salary over this project,
we will be covering any reasonable travel expenses incurred. If the
student works with or endorses a certain charity, we are also happy
to make donations of product to charity events, raffles and
fundraisers. If there are any questions or you would like to find out
more about us please get in touch at blake@[Link]

Student deliverables Proposal for enhanced use of AI within Traybakes marketing with
appropriate prototypes for demonstration
Student’s line manager Blake Dyson

Project location(s) The student undertaking this task will be able to work remotely
however we would like them to visit towards to the start of the
project to get a better understanding of Traybakes. We feel that this
will affect how they go about the project as well as giving the
student the opportunity to see the challenges we face rather than
just being briefed on them. Whilst we will have a structured set of
meetings to allow us to understand where the project is, we want
the student to use their expertise and creativity to test different
options without the pressure of reporting on this to us.
Other comments We want them to work with autonomy and hope that this will allow
the student to get more out of the project rather than by giving
them a strict brief to stick to.

Student Requirements
Essential technical skills Familiarity with AI technologies and their potential use
Desirable technical skills As above
Essential general skills Ability to gain an understanding of business processes and link AI to
business objectives
Desirable general skills Ability to take initiative and organise own work
Good communication skills
Essential experience
Desirable experience
Other comments

Organisation Details
Main contact name
Main contact position
Applications email blake@[Link]
Applications email cc justine@[Link]
Other comments
PRJ21-UH-AIF - AI Feasibility Modelling for the Equity Engine -
Unwritten Health
Project for DS AI

Project Description
Project name AI Feasibility Modelling for the Equity Engine
Stipend offered N/A
Host Organisation Unwritten Health
Project description Unwritten Health is developing Equity Engine, a UK-based data and
insight platform designed to address long-standing gaps in
healthcare evidence by capturing lived-experience data from
underserved and historically underrepresented communities. The
platform collects a combination of demographic information, social
determinants of health, behavioural engagement data, and
qualitative lived-experience inputs.

This project will explore how this complex, mixed-methods dataset


can be structured and prepared for analysis. This may include
exploratory segmentation, clustering, or topic modelling to assess
what kinds of patterns or groupings emerge from the data, and
where meaningful signal exists versus noise. Any modelling work
would be explicitly exploratory and feasibility-focused, rather than
predictive, with careful consideration of bias, representational risk,
and ethical use. Additionally, this project will look to assess how AI
may be used to automate the processing, analysis and
communication of data and the resultant insights.

It is anticipated that the student will give full consideration to how a


robust, AI-augmented, data process pipeline may be implemented as
the platform scales.

Student deliverables • An exploratory assessment of which analytical or modelling


approaches may be feasible as the platform scales

• A set of recommendations outlining how AI may be used to


enhance to efficiency and utility of the data gathered and
how this may inform future AI deployment in a responsible
and ethical manner

Student’s line manager Ashish Rishi

Project location(s) Remote


Other comments The project is intentionally designed to be collaborative and
adaptable. While an initial scope is proposed, it is expected that the
precise research questions and analytical focus would be refined
through early discussions between the student, academic supervisor,
and Unwritten Health, in line with the evolving dataset and the
student’s skills and interests.
Student Requirements
Essential technical skills Understanding of data pipelines and AI technologies
Competent in Python
Desirable technical skills
Essential general skills Basic understanding of project management
Able to organise own work
Understanding of responsible, ethical use of data
Desirable general skills
Essential experience N/A
Desirable experience
Other comments

Organisation Details
Main contact name Ashish Rishi
Main contact position CEO
Applications email ashish@[Link]
Applications email cc
Other comments
PRJ28-UHM-P - Predicting associations of poor bone health adjusting
for treatment allocation - University Hospitals of Morecambe Bay NHS
foundation trust
Project for HDS

Lancaster University Health Data Science Project Description


Please give details of the project in broad terms – this information is for guidance and will be
replaced by a detailed specification later in the process.

Project name Predicting associations of poor bone health adjusting for treatment
allocation
Stipend offered None
Host Organisation University Hospitals of Morecambe Bay NHS foundation trust
Project description This will look at the Morecambe Bay DEXA database , which includes
45,000 patients referred for DEXA scan between 2004 and 2024. The
project will look at predictors of fracture after using propensity score
matching to see if the effect of bisphosphonates with and with
calcium and vitamin D differ. It will use existing data to predict
treatment choice and apply this to a propensity score matched
subset.
Student deliverables Meeting abstract submission and possible manuscript submission.
Student’s line manager Marwan Bukhari
Project location(s) Multiple sites
Other comments The student would need to be competent at data analysis and using
the appropriate software (Stata, R or SAS)

Student Requirements
Essential technical skills Highly proficient in R or STATA
Desirable technical skills Good data management skills
Essential general skills Basic understanding of project management, scientific notation and
some excel and access skills
Desirable general skills As above
Essential experience working in teams, communication to learned audiences
Desirable experience A clinical background
Other comments The study would be largely supervised by both the trust and the
University

Organisation Details
Main contact name Marwan Bukhari
Main contact position Consultant rheumatologist and honorary professor
Applications email [Link]@[Link]
Applications email cc [Link]@[Link]
Other comments
PRJ29-MIM-DCV - Design and implementation of a CVR solution within
Multipave - Multipave Infrastructure
Project for AI

Project Description
Project name Design and implementation of a CVR solution within Multipave
Stipend offered We pay the Real Living Wage on a 40-hour working week
Host Organisation Multipave Infrastructure

Multipave is one of the Northwest’s largest road surfacing and


recycling contractors, delivering high quality and innovative surfacing
solutions. We are one of 12 contractors on the Pavement Delivery
Framework for National Highways, framework suppliers to 11
Northwest local authorities and the contractor of choice for many
tier one infrastructure contractors. The business has seen significant
growth over the last twelve months and is in line to post record
turnover and profits – revenue and profits grew by 20% and 165%,
respectively, between 2022 and 2024. From our beginnings in plant
hire, we now offer turnkey infrastructure solutions, including a civils
arm, which started in 2025.

Multipave is on an exciting growth journey. During 2025 the Board


developed a strategy that will deliver a doubling of revenue over the
next five years. Vital foundations for this growth are the
establishment of data infrastructure and information systems to
enable us to operate as a data-led business utilising real time
information to guide evidence-based decisions. The business is still
on a journey to move away from paper-based and excel-based
systems which create data silos and prevent us from easily accessing
data insights. We operate an Integrated Management System and
Microsoft 365. We recently implemented two key software systems –
Re-Flow ([Link]) , a field management software to
support operational efficiency and Dynamics 365 Business Central
([Link]
365/products/business-central), a Microsoft product to improve our
financial oversight.

This is a genuine opportunity to have a significant impact on a


medium construction business in Lancashire. The results of this
project will set the foundations to revolutionise how we operate.
The person who takes this project will have the opportunity to work
directly with the decision makers in the organisation and shape how
the business operates. You will be well supported with the Head of
Strategy as your key contact and day to day support, with access to
other members of the Board of Directors.
Project description Our commercials team work to deliver added value throughout the
delivery of road surfacing and infrastructure schemes, aiming to
maximise scheme and contract profit, increasing it beyond the profit
that was estimated at the point of quotation. The team of quantity
surveyors, led by our Head of Commercials, work to ensure that
changes to scope of schemes are appropriately valued and paid by
the client.

The standard approach to assessing value by scheme is to develop a


cost and valuation reconciliation (CVR) model. This is currently a
manual process, requiring the transfer of data from several sources
across the business, which is both time consuming and introduces
the risk of errors.

In addition to not being able to see the performance of schemes on


an individual level, the Board do not have oversight of the
performance across the entire portfolio of schemes, contracts,
clients and sectors, which reduces the ability to make commercial
and strategic decisions.

This is a development and implementation project, and the aims of


the project are:
- To develop an understanding of the working practices and
processes, data collection systems and linkages and needs of
the business.
- To make recommendations on improving data collection to
enable a secure AI solution.
- To develop a secure and robust AI solution to simplify the
process of cost and valuation reconciliation for all schemes
delivered by Multipave, including a data visualisation
dashboard, which will ensure the commercials team and
Board have a view of the performance of the business across
contracts, clients and sectors.
- To develop a solution and pilot, working with key
stakeholders.
- To refine and test the proposed AI solution to the point of
acceptance.
- To implement the solution and evaluate.

How the work will be organised:


- The project will be owned by the student with supervision
and day to day support provided by Multipave’s Head of
Strategy.
- The student will work as part of the People, Culture and
Development Team. They will be welcomed into the
business and treated as an employee throughout the
placement.
- The student will have supported access to members of the
Board, Senior Leadership Team and other individuals in key
roles across the business.
- The student and Head of Strategy will meet with key
stakeholders and review the business systems and processed
during the first week. They will design the project and
present this for sign off by the Board during week two.
- The remaining weeks will be spent delivering the project,
developing and testing AI tools, providing regular updates to
key stakeholders. Activities and timings will be developed by
the student and supervisor.
- The student will present their findings and recommendations
to the Board and Senior Leadership Team.
- The expectation is that the student will work from the
Multipave head office for the first few weeks of the project
but may then work in a hybrid fashion, if they prefer. They
will be provided with a desk in the People, Culture and
Development Team and full IT equipment and access.

Stakeholders in the project:


Several key roles across the business will contribute to sharing
information on current systems and processes, data collection and IT
systems. The Head of Strategy will be the key contact who will
provide supervision, guidance and support during the placement.
They will facilitate access to and meetings with the key stakeholders
and will be continually involved in the project to ensure the focus is
maintained on a deliverable project that meets the requirements of
the master’s course and Multipave’s needs. The Board of Directors
will be key contacts and have a keen interest and stake in the output
and deliverables.
Key contacts:
- Head of Strategy (main contact and support)
- Head of Commercials (primary stakeholder)
- Digital Solution and Information Systems Manager
- Director of Operations
- Financial Controller

Technologies:
Key systems for this project are:
- Dynamics 365 Business Central
([Link]
365/products/business-central
- Microsoft SharePoint and 365

The project will involve development of an AI solution(s). We will rely


on the student to identify relevant technologies.
Student deliverables - A project proposal and timeline during week 2 for Board sign
off
- A tested and evaluated AI solution for CVR.
- A dashboard to enable interrogation of the CVR by the
Commercials Team and Board.
- A process for using the developed solution.
- Presentations on the system and process to the Board and
Senior Leadership Team
Student’s line manager Heather Catt, Head of Strategy
Project location(s) Multipave (NW) LTD, 12 Centurion Court, Leyland. PR25 3UQ
Other comments We are a friendly business, and our working environment is
welcoming, relaxed and informal. We operate from our head office
in Leyland, which has car parking and is close to Leyland train station.
As a non-technology business with a desire to adopt technology, the
student will find everyone is very keen to work with them and
benefit from their skills.

We currently have a Lancaster University student in place for a


Computer Science placement (who is doing an amazing job!). They
are happy to be contacted for an informal discussion about their
experience if required. [Link]@[Link]

Student Requirements
Essential technical skills High level of knowledge about AI applications within businesses,
including the Microsoft suite.

Strong skills in AI techniques.

Strong AI skills in the automation of business processes, i.e. AI


agents, natural language processing.

Strong skills in integrating existing systems.

Desirable technical skills


Essential general skills Good interpersonal skills, confident and an enquiring mind.

Able to explain complicated concepts clearly to people with limited


understanding of AI and to support leaders to fully specify their
requirements for the system.

Pragmatism and patience – we are a busy operational business and


there can be times where operational pressures take priority.
However, you will be well supported to ensure that the project stays
on track and are slippages are recovered. This is a key priority for the
business.
Desirable general skills
Essential experience Experience of developing AI solutions.
Desirable experience Experience of working in teams
Other comments We are an organisation that is committed to generating social value
and providing opportunities to those who are typically excluded
from careers in the construction industry – we would be particularly
supportive of applications from students who are from lower
socioeconomic backgrounds, minority ethnic groups and women.
Organisation Details
Main contact name Heather Catt
Main contact position Head of strategy
Applications email [Link]@[Link]
Applications email cc [Link]@[Link]
Other comments
PRJ30-MIM-DM - Design and implementation of a payroll solution
within Multipave - Multipave Infrastructure
Project for AI

Project Description
Project name Design and implementation of a payroll solution within Multipave
Stipend offered We pay the Real Living Wage on a 40 hour working week
Host Organisation Multipave Infrastructure

Multipave is one of the Northwest’s largest road surfacing and


recycling contractors, delivering high quality and innovative surfacing
solutions. We are one of 12 contractors on the Pavement Delivery
Framework for National Highways, framework suppliers to 11
Northwest local authorities and the contractor of choice for many
tier one infrastructure contractors. The business has seen significant
growth over the last twelve months and is in line to post record
turnover and profits – revenue and profits grew by 20% and 165%,
respectively, between 2022 and 2024. From our beginnings in plant
hire, we now offer turnkey infrastructure solutions, including a civils
arm, which started in 2025.

Multipave is on an exciting growth journey. During 2025 the Board


developed a strategy that will deliver a doubling of revenue over the
next five years. Vital foundations for this growth are the
establishment of data infrastructure and information systems to
enable us to operate as a data-led business utilising real time
information to guide evidence-based decisions. The business is still
on a journey to move away from paper-based and excel-based
systems which create data silos and prevent us from easily accessing
data insights. We operate an Integrated Management System and
Microsoft 365. We recently implemented two key software systems –
Re-Flow ([Link]) , a field management software to
support operational efficiency and Dynamics 365 Business Central
([Link]
365/products/business-central), a Microsoft product to improve our
financial oversight. We are also in the process of implementing a
safety solution, Samsara, which will see the introduction of an AI
enabled programme using two way dashboard cameras in our fle et
([Link]

This is a genuine opportunity to have a significant impact on a


medium construction business in Lancashire. The results of this
project will set the foundations to revolutionise how we operate.
The person who takes this project will have the opportunity to work
directly with the decision makers in the organisation and shape how
the business operates. You will be well supported with the Head of
Strategy as your key contact and day to day support, with access to
other members of the Board of Directors.
Project description This is a development and implementation project, and the aims of
the project are:
- To develop an understanding of the working practices and
processes, data collection systems and linkages and needs of
the business.
- To make recommendations on improving data collection to
enable a secure AI solution.
- To develop an AI solution to simplify the process of running
payroll for our operational staff, including the development
of automated data collection on staff working hours and data
linkages between systems.
- To develop a solution and pilot, working with key
stakeholders.
- To refine and test the proposed AI solution to the point of
acceptance.
- To implement the solution and evaluate.

How the work will be organised:


- The project will be owned by the student with supervision
and day to day support provided by Multipave’s Head of
Strategy. The IT and Digital Solutions Manager will also be a
key contact.
- The student will work as part of the People, Culture and
Development Team. They will be welcomed into the
business and treated as an employee throughout the
placement.
- The student will have supported access to members of the
Board, Senior Leadership Team and other individuals in key
roles across the business.
- The student and Head of Strategy will meet with key
stakeholders and review the business systems and processes
during the first week. They will design the project and
present this for sign off by the Board during week two.
- The remaining weeks will be spent delivering the project,
developing and testing AI tools, providing regular updates to
key stakeholders. Activities and timings will be developed by
the student and supervisor.
- The student will present their findings and recommendations
to the Board and Senior Leadership Team.
- The expectation is that the student will work from the
Multipave head office for the first few weeks of the project
but may then work in a hybrid fashion, if they prefer. They
will be provided with a desk in the People, Culture and
Development Team and full IT equipment and access.

Stakeholders in the project:


Several key roles across the business will contribute to sharing
information on current systems and processes, data collection and IT
systems. The Head of Strategy will be the key contact who will
provide supervision, guidance and support during the placement.
They will facilitate access to and meetings with the key stakeholders
and will be continually involved in the project to ensure the focus is
maintained on a deliverable project that meets the requirements of
the master’s course and Multipave’s needs. The Board of Directors
will be key contacts and have a keen interest and stake in the output
and deliverables.
Key contacts:
- Head of Strategy (main contact and support)
- Digital Solution and Information Systems Manager
- SHEQ Coordinator
- Accounts Administrator
- Financial Controller
- Operations Director
- Supervisors
- Operatives
- Director of People, Culture and Development
- HR Administrator

Technologies:
Key systems for this project are:
- Re-Flow ([Link]), a field management software
to support operational efficiency
- Dynamics 365 Business Central
([Link]
365/products/business-central
- Samsara: [Link]
- People HR, our HR software solution ([Link])

The project will involve development of an AI solution(s). We will rely


on the student to identify relevant technologies.
Student deliverables - A project proposal and timeline during week 2 for Board sign
off
- A tested and evaluated AI solution for payroll.
- A process for using the developed solution.
- Presentations on the system and process to the Board and
Senior Leadership Team
Student’s line manager Heather Catt, Head of Strategy
Project location(s) Multipave (NW) LTD, 12 Centurion Court, Leyland. PR25 3UQ
Other comments We are a friendly business, and our working environment is
welcoming, relaxed and informal. We operate from our head office
in Leyland, which has car parking and is close to Leyland train station.
As a non-technology business with a desired to adapt technology, the
student will find everyone is very keen to work with them and
benefit from their skills.
We currently have a Lancaster University student in place for a
Computer Science placement (who is doing an amazing job!). They
are happy to be contacted for an informal discussion about their
experience if required. [Link]@[Link]

Student Requirements
Essential technical skills High level of knowledge about AI applications within businesses,
including the Microsoft suite.

Strong skills in AI techniques.

Strong AI skills in the automation of business processes, i.e. AI


agents, natural language processing.

Strong skills in integrating existing systems.


Desirable technical skills
Essential general skills Good interpersonal skills, confident and an enquiring mind.

Able to explain complicated concepts clearly to people with limited


understanding of AI and to support leaders to fully specify their
requirements for the system.

Pragmatism and patience – we are a busy operational business and


there can be times where operational pressures take priority.
However, you will be well supported to ensure that the project stays
on track and are slippages are recovered. This is a key priority for the
business.
Desirable general skills
Essential experience Experience of developing AI solutions.
Desirable experience Experience of working in teams
Other comments We are an organisation that is committed to generating social value
and providing opportunities to those who are typically excluded
from careers in the construction industry – we would be particularly
supportive of applications from students who are from lower
socioeconomic backgrounds, minority ethnic groups and women.

Organisation Details
Main contact name Heather Catt
Main contact position Head of strategy
Applications email [Link]@[Link]
Applications email cc [Link]@[Link]
Other comments
PRJ31-MIM-AAI - A scoping study of the potential of AI to add value at
Multiple - Multipave Infrastructure
Project for AI

Project Description
Project name A scoping study of the potential of AI to add value at Multiple
Stipend offered We pay the Real Living Wage on a 40 hour working week
Host Organisation Multipave Infrastructure

Multipave is one of the Northwest’s largest road surfacing and


recycling contractors, delivering high quality and innovative surfacing
solutions. We are one of 12 contractors on the Pavement Delivery
Framework for National Highways, framework suppliers to 11
Northwest local authorities and the contractor of choice for many
tier one infrastructure contractors. The business has seen significant
growth over the last twelve months and is in line to post record
turnover and profits – revenue and profits grew by 20% and 165%,
respectively, between 2022 and 2024. From our beginnings in plant
hire, we now offer turnkey infrastructure solutions, including a civils
arm, which started in 2025.

Multipave is on an exciting growth journey. During 2025 the Board


developed a strategy that will deliver a doubling of revenue over the
next five years. Vital foundations for this growth are the
establishment of data infrastructure and information systems to
enable us to operate as a data-led business utilising real time
information to guide evidence-based decisions. The business is still
on a journey to move away from paper-based and excel-based
systems which create data silos and prevent us from easily accessing
data insights. We operate an Integrated Management System and
Microsoft 365. We recently implemented two key software systems
– Re-Flow ([Link]) , a field management software to
support operational efficiency and Dynamics 365 Business Central
([Link]
365/products/business-central), a Microsoft product to improve our
financial oversight.

This is a genuine opportunity to have a significant impact on a


medium construction business in Lancashire. The results of this
project will set the foundations to revolutionise how we operate.
The person who takes this project will have the opportunity to work
directly with the decision makers in the organisation and shape how
the business operates. You will be well supported with the Head of
Strategy as your key contact and day to day support, with access to
other members of the Board of Directors.
Project description This is a scoping evaluation and the aims of the project are:
- To develop an understanding of the working practices and
processes, data collection systems and linkages and needs of
the business
- To scope the underutilised data and IT systems within the
business.
- To conduct a scoping review and make proposals to the
Multipave Board about the potential to use AI to refine and
improve processes and increase efficiency across the
business.

How the work will be organised:


- The project will be owned by the student with supervision
and day to day support provided by Multipave’s Head of
Strategy. The IT and Digital Solutions Manager will also be a
key contact.
- The student will work as part of the People, Culture and
Development Team. They will be welcomed into the business
and treated as an employee throughout the placement.
- The student will have supported access to members of the
Board, Senior Leadership Team and other individuals in key
roles across the business.
- The student and Head of Strategy will meet with key
stakeholders and review the business systems and processes
during the first week. They will design the project and
present this for sign off by the Board during week two.
- The remaining weeks will be spent delivering the project,
developing and testing AI tools, providing regular updates to
key stakeholders.
- The student will present their findings and recommendations
to the Board and Senior Leadership Team.
- The expectation is that the student will work from the
Multipave head office for the first few weeks of the project
but may then work in a hybrid fashion, if they prefer. They
will be provided with a desk in the People, Culture and
Development Team and full IT equipment and access.

Stakeholders in the project:


Several key roles across the business will contribute to sharing
information on current systems and processes, data collection and IT
systems. The Head of Strategy will be the key contact who will
provide supervision, guidance and support during the placement.
They will facilitate access to and meetings with the key stakeholders
and will be continually involved in the project to ensure the focus is
maintained on a deliverable project that meets the requirements of
the master’s course and Multipave’s needs. The Board of Directors
will be key contacts and have a keen interest and stake in the output
and deliverables.
Key contacts:
- Head of Strategy (main contact and support)
- Digital Solution and Information Systems Manager
- Board of Directors, including Directors of Operations, SHEQ,
People, Culture and Development, and Finance.
- Senior Leadership Team, including Finance Controller, Head
of Estimating, Head of Civils.

Technologies:
The focus of the student’s project is on applying their knowledge
across the business to understand where processes could be
automated, which processes cannot be automated and where
processes need to change to enable automation. The project may
involve development and deployment of small AI tools; however, it
will focus on a systematic review of processes and guidance for
future work. We will rely on the student to identify relevant
technologies that may be suitable.
Student deliverables - A project proposal and timeline during week 2 for Board sign
off
- A final report of their findings of where AI can add value and
proposals for future work to integrate AI effectively
- A policy for AI use within Multipave which safeguards data
security and business reputation
- Presentations on their findings and recommendations to the
Board and Senior Leadership Team
Student’s line manager Heather Catt, Head of Strategy
Project location(s) Multipave (NW) LTD, 12 Centurion Court, Leyland. PR25 3UQ
Other comments We are a friendly business, and our working environment is
welcoming, relaxed and informal. We operate from our head office
in Leyland, which has car parking and is close to Leyland train station.
As a non-technology business with a desired to adapt technology, the
student will find everyone is very keen to work with them and
benefit from their skills.

We currently have a Lancaster University student in place for a


Computer Science placement (who is doing an amazing job!). They
are happy to be contacted for an informal discussion about their
experience if required. [Link]@[Link]

Student Requirements
Essential technical skills High level of knowledge about AI applications within businesses,
including the Microsoft suite.

Strong skills in AI techniques.

Strong AI skills in the automation of business processes, i.e. AI


agents, natural language processing.

Strong skills in integrating existing systems.

Desirable technical skills


Essential general skills Good interpersonal skills, confident and an enquiring mind.

Able to explain complicated concepts clearly to people with limited


understanding of AI and to support leaders to fully specify their
requirements for the system.

Pragmatism and patience – we are a busy operational business and


there can be times where operational pressures take priority.
However, you will be well supported to ensure that the project stays
on track and are slippages are recovered. This is a key priority for the
business.
Desirable general skills
Essential experience Experience of developing AI solutions.
Desirable experience Experience of working in teams
Other comments We are an organisation that is committed to generating social value
and providing opportunities to those who are typically excluded
from careers in the construction industry – we would be particularly
supportive of applications from students who are from lower
socioeconomic backgrounds, minority ethnic groups and women.

Organisation Details
Main contact name Heather Catt
Main contact position Head of strategy
Applications email [Link]@[Link]
Applications email cc [Link]@[Link]
Other comments
PRJ32-KPM-KPM - KPMG OT Synapse - KPMG UK
Project for DS AI

Project Description
Project name KPMG OT Synapse
Stipend offered £3,000
Host Organisation KPMG UK
Project description AI-enabled assessment of an organisation's Operational Technology
(OT) networks and security architecture.

What are the aims of the project?


The aim of the project is to develop an AI-enabled solution to
accelerate and improve the quality of KPMG's existing OT Security
Architecture Advisory Service (Note: advisory = consulting). KPMG
would like this solution to evaluate an organisation's industrial
control systems (ICS) and operational technology (OT) networks by
analysing the network's architecture and its constituent assets,
identifying threats and vulnerabilities, reviewing existing security
controls, assessing risks, and checking for compliance with industry
standards, ultimately providing recommendations to enhance the
overall cyber security posture and operational reliability.

What type of project is this? (e.g. implementation, research,


evaluation, enhancement, pilot)
Pilot

How will work be organised?


KPMG is already working in the development of the AI-enabled
solution described above. We would like the students to join the
KPMG team and contribute to one of the following tasks:
• Data acquisition and preparation
• Model development and training
• Model testing and refinement

Who will contribute?


We envisage a collaborative approach between the following parties:
• KPMG members of staff involved in the project.
• The selected students.
• The students’ academic advisors.
• (For the model testing and refinement task): members of
staff of a KPMG’s client providing the real-life scenario where
the AI-enabled solution described above will be tested.

What business areas will be involved?


The Operational Technology Cybersecurity team within KPMG UK’s
advisory (i.e. consulting) business area.
Which technologies will be used?
To be determined. It is expected that the technologies to be used
will align with KPMG’s AI 360 framework, the details of which (e.g.
default technology stack) are currently being defined.
Student deliverables At the end of the project, the student is expected to deliver a report
describing:
• Their specific contribution(s) to the delivery of the AI-
enabled solution described above.
• Recommendations to the KPMG team for a successful
deployment and scaling up of the solution.
Student’s line manager Dr Adrian Salinas, the project manager in charge of the delivery of
the AI-enabled solution.

Project location(s) Remotely. Occasional face-to-face meetings will be required; the


time/place for them will be agreed between the students and Dr.
Salinas.
Other comments KPMG would like to give the students a taste of real-life work
experience, so KPMG will conduct the placement in a way which
mimics how an entry-level member of staff would do their job at
KPMG. There are projects available for 2 students.

Student Requirements
Essential technical skills Programming & Scripting
• Proficiency in Python.
• Familiarity with libraries like NumPy, Pandas, and Matplotlib for
data manipulation and visualization.

Machine Learning Fundamentals


• Understanding of supervised and unsupervised learning.
• Knowledge of model training, validation, and evaluation
techniques.
• Experience with frameworks such as TensorFlow or PyTorch.

Data Handling
• Ability to preprocess and clean large datasets.
• Understanding of feature engineering and data normalization.

Basic Cybersecurity Concepts


• Awareness of network architecture (e.g., firewalls, IDS/IPS,
endpoints).
• Understanding of common vulnerabilities (e.g., CVEs, OWASP
Top 10).

Version Control
• Experience with Git for collaborative development.
Desirable technical skills Deep Learning & Advanced ML
• Experience with neural networks, graph-based models, or
transformers.
• Knowledge of anomaly detection techniques.

Cybersecurity-Specific Knowledge
• Familiarity with threat modeling, risk assessment frameworks
(e.g., NIST, ISO 27001).
• Understanding of compliance standards (GDPR, PCI DSS).

Data Science & Analytics


• Ability to work with structured and unstructured data (e.g., logs,
network traffic).
• Knowledge of statistical analysis and probabilistic models.

Cloud & Deployment


• Experience with cloud platforms (AWS, Azure, GCP) for model
training.
• Understanding of containerization (Docker) and CI/CD pipelines.

Graph Theory & Network Analysis


• Skills in analyzing network topologies and graph-based threat
detection.

Explainable AI (XAI)
• Familiarity with techniques for model interpretability and
explainability.
Essential general skills • Analytical Problem-Solving: Break down ambiguous requirements
into testable hypotheses; define success metrics for models.
• Scientific Rigor: Design experiments, control variables, track
assumptions, reproduce results, and document findings clearly.
• Communication (Technical & Non-Technical): Explain model
choices, limitations, and risk implications to engineers and non-
technical stakeholders (e.g., risk, compliance).
• Collaboration & Teamwork: Work effectively with security
architects, data engineers, and product managers; accept
feedback via code reviews.
• Adaptability & Learning Agility: Quickly learn new
frameworks/datasets; pivot when an approach underperforms.
• Time Management & Prioritisation: Plan sprints, manage model
training cycles, and meet milestones.
• Ethical & Responsible AI Mindset: Consider bias, fairness, privacy,
and security-by-design; escalate concerns appropriately.
• Attention to Detail.
• Ownership & Initiative: Propose improvements, automate
repetitive tasks, and raise risks/blockers early.
Desirable general skills • Stakeholder Engagement: Gather requirements, clarify
constraints, and translate business outcomes into measurable
ML objectives.
• Threat-Informed Thinking: Connect model outputs to attacker
behaviors and practical defensive outcomes.
• Systems Thinking: Understand how models fit within pipelines
(data sources, controls, alerts, workflows).
• Narrative Reporting: Create concise, insightful updates (model
cards, experiment summaries, executive-ready dashboards).
• Mentorship & Knowledge Sharing: Contribute to internal wikis,
demo sessions, and collaborative learning.
• Resilience & Experimentation Mindset: Comfort with iteration
and failure; ability to pivot and improve.
• Negotiation & Prioritization: Balance accuracy vs. explainability,
speed vs. robustness, research vs. delivery.
Essential experience N/A – the objective of a student placement is to gain experience!
Desirable experience The items below are beneficial, but by no means essential. Feel free
to apply even if you don’t have the experience below!
• Hands-on ML Projects: Previous experience building and
deploying ML models (academic or industry).
• Internships or Research: Prior internships in AI, data science, or
cybersecurity; published research papers or conference
presentations.
• Data Engineering: Experience with large-scale data pipelines, ETL
processes, or working with big data tools.
• Cloud & MLOps: Practical experience with cloud-based ML
training and deployment (AWS SageMaker, Azure ML).
• Open-Source Contributions: Contributions to ML or
cybersecurity-related repositories.
Other comments The project is at the initial stages, so at this moment it is hard to
provide details of the specific technology stack to be used.

Flexibility and adaptability are required for the students to be able to


join a project which will be underway at the time their placements
begin.

Organisation Details
Main contact name Dr Adrian Salinas-Varela
Main contact position Operational Technology (OT) cybersecurity consultant
Applications email
Applications email cc
Other comments Our team has close ties with the University of Lancaster,
collaborating on a regular basis with researchers in the
Cybersecurity, Data Science and AI departments. Last year we
provided a placement for a Cybersecurity MSc student (who has now
joined us as a full member of staff) and would be delighted to
provide more placements to students this year!
PRJ33-DHL-IEA - Inverse Extremisms: A Computational Linguistic
Analysis of Tradwife and Incel Discourse - Digital Heard Ltd
Project for DS AI

Project Description
Project name Inverse Extremisms: A Computational Linguistic Analysis of
Tradwife and Incel Discourse

Stipend offered No stipend will be given, but reasonable expenses will be paid, and a
Fraser House Hub membership can be arranged if preferred and
appropriate.

Host Organisation Digital Heard Ltd

Project description The aim of this project is to build an empirical evidence base for a
theoretical paper currently under development, which argues that
tradwife content (targeting women) and incel/manosphere content
(targeting men) employ inverse but structurally parallel linguistic
strategies to reinforce patriarchal gender hierarchies. This is a
research project involving corpus collection, computational linguistic
analysis, and qualitative discourse analysis. Work will be organised
around three phases: data collection (scraping and curating matched
corpora from TikTok, Instagram, and Reddit), computational analysis
(using NLP tools to identify linguistic patterns), and comparative
synthesis (mapping findings against the theoretical framework). The
student will work under the supervision of Dr Alice Ashcroft (Digital
Heard Ltd), with potential collaboration from academic colleagues in
the School of Computing and Communications and other academic
insitutions. The project sits at the intersection of Human-Computer
Interaction, computational linguistics, and gender studies, with
direct relevance to platform governance and digital literacy.
Technologies will include Python for data collection and analysis,
corpus linguistic tools, and LLMs or similar sentiment/appraisal
analysis software.

Student deliverables The student will be expected to deliver three core outputs. First, a
reproducible data collection tool in Python, callable via command-
line script, capable of scraping and curating matched corpora of
tradwife and incel/manosphere content from platforms such as
TikTok, Instagram, and Reddit, with appropriate documentation for
future use. Second, an R script that performs the comparative
linguistic analysis and generates publication-ready visualisations
(graphs, tables, and figures) suitable for inclusion in academic
papers. Third, the student will contribute to drafting sections of one
or more research papers arising from the project, with co-authorship
credit awarded in accordance with their contribution and standard
academic conventions. All code and documentation should be
version-controlled and handed over in a state suitable for ongoing
use by the research team.

Student’s line manager The student will report directly to Dr Alice Ashcroft (Director, Digital
Heard Ltd), who will serve as their primary supervisor and line
manager throughout the placement. Additional support and
mentorship will be available from the wider Digital Heard research
team (currently three researchers) and Joshua Bailey (AI consultant
and founder of Sqwosh), who will provide technical guidance on
computational and AI-related aspects of the project.

Project location(s) The work can be performed remotely or from Fraser House Hub, a
coworking space for freelancers and small businesses in Lancaster. A
Fraser House Hub membership can be arranged for the student if
preferred, which would provide excellent networking opportunities
alongside a professional working environment. The arrangement will
be flexible and agreed with the student based on their preferences
and circumstances.

Other comments Digital Heard operates with a flexible, asynchronous working culture
where we prioritise outputs over hours. We expect our staff and
interns to work independently and take ownership of their tasks,
with support readily available when needed. There will always be a
team member on hand to help between 10am and 4pm, but working
hours outside of this are entirely flexible and can be arranged around
the student's own schedule and commitments. This placement
would suit a self-motivated student who thrives with autonomy and
is comfortable managing their own workload.

Student Requirements
Essential technical skills Proficient in Python, including experience with web scraping
libraries. Competent in R, with the ability to produce data
visualisations using packages where needed. Basic understanding of
Natural Language Processing concepts and familiarity with text
analysis libraries and AI.

Desirable technical skills Experience working with social media APIs or scraping
TikTok/Instagram/Reddit specifically. Knowledge of version control
using Git/GitHub.
Essential general skills To organise your own work and manage time effectively with
minimal supervision. Strong written communication skills,
particularly the ability to write clearly for academic audiences.
Attention to detail, especially when handling and cleaning data.
Willingness to engage with sensitive content relating to online
misogyny and gender-based extremism.

Desirable general skills A basic understanding of academic research processes and


publication conventions. Ability to bring together findings and
contribute to written outputs. Comfort working asynchronously and
communicating via digital tools (e.g., Slack, email, Latex/Overleaf,
shared documents).

Essential experience Experience working on a research project or dissertation involving


data collection and analysis, whether through coursework or
independent study.

Desirable experience Experience working in a team or collaborative research environment.


Familiarity with Human-Computer Interaction, digital discourse
analysis, or gender studies literature. Previous experience
contributing to academic writing or publications.

Other comments This project involves engaging with online content that includes
misogynistic, extremist, and potentially distressing material. The
student should be comfortable working with such content and aware
of the emotional labour involved. We are happy to discuss strategies
for managing this and will provide appropriate support throughout.

This placement offers an excellent opportunity for a student


interested in the intersection of AI, linguistics, and social impact
research, with a clear pathway to co-authored publication.

Organisation Details
Main contact name Dr Alice Ashcroft

Main contact position Founder and Lead Researcher

Applications email admin+applications@[Link]

Applications email cc
Other comments All applications should include a CV, a link to a publication/written
piece of coursework and a paragraph on what interests them about
the project.
PRJ34-LUB-AAI - Agentic-AI for Telecommunications - Lancaster
University in partnership with BT
Project for AI

Project Description
Project name Agentic-AI for Telecommunications
Stipend offered N/A
Host Organisation Lancaster University in partnership with BT
Project description This project will investigate the use of Agentic-AI systems for
telecommunications use case. There is considerable flexibility in
what exactly the project will achieve and can be scoped to suit the
interests of the prospective student.

One lens of the project would have the student develop an agentic
system, framed around MCP, that can orchestrate some functions on
a network. Alternatively, a student may construct a multi-agent
framework that brings to life emerging standards for Agentic-AI in
telecommunications, e.g. from ETSI, whereby multiple agents co-
operate together to reach a goal (some specific network function(s)).
The project will require a willingness to learn technical aspects of
Agentic-AI, and some familiarity with computer networks will be very
useful. The goal is to deliver a prototype implementation of an
Agentic-AI system to address at least one use case for
telecommunications.

This project will be in partnership with researchers at Lancaster and


the Autonomous Networks Team at BT. As this placement is research
led, you will have freedom to explore topics that interest you, but
still have the chance to feed into cutting-edge research at BT.
Student deliverables A prototype implementation of an agentic-AI system that addresses
at least one telecoms use case.
Student’s line manager Dr Ed Austin

Project location(s) Lancaster / Remote, with the chance to visit BT.


Other comments

Student Requirements
Essential technical skills A willingness to learn about agentic-AI, and implement it, in a
telecoms setting. The project will be pitched at a technical level, and
the student will need to become hands-on with the technical
implementation of the agentic system.
Desirable technical skills
Essential general skills The student should be comfortable presenting to both technical, and
senior, stakeholders. The student should also be happy to manage
their project and drive it forward – as this is primarily research based
– although as long as the student is happy with this they don’t need
a track record doing it to date.
Desirable general skills
Essential experience
Desirable experience I am keen to provide the student with experience – and I am happy
to consider people from any backgrounds.
Other comments

Organisation Details
Main contact name Ed Austin
Main contact position Research Fellow
Applications email To whom should we send the students’ applications?

[Link]@[Link]
Applications email cc Is there anyone else who should be copied on the application
emails?
Other comments Is there anything else about your organisation of which you would
like to make us aware?
PRJ35-T-DDA - Data-Driven Assessment of the Value of Small Versus
Large Language Models for Organisations Deploying AI - Traversally
Project for AI

Project Description
Project name Data-Driven Assessment of the Value of Small Versus Large Language
Models for Organisations Deploying AI
Stipend offered N/A
Host Organisation Traversally
Project description AI workflows typically rely upon calls to externally hosted LLMs,
which are both energy intensive to host, and cost the organisation
querying the model per token. As an alternative an organisation
could host their own small language model, which does not cost
them per token and uses less energy.

To date, the emphasis has been on how LLMs vastly outperform


smaller models – and this is something that CTOs and other senior
stakeholders struggle to look past when making high level
deployment decisions, even if they could save money or reduce their
(indirect) energy usage through using smaller models.

Through their role AI Transformation Consultancy Traversally are


interested in providing data-driven, evidence backed, measurement
of the value of Small Language Models for various deployment use
cases. This project will implement and test at least one small
language model integrated into an AI system and produce a data-
driven assessment of the value of its use in terms of several metrics,
for example: accuracy, scalability, cost, and energy usage.

Students may have an interest in a particular AI system – generative,


agentic, or otherwise – and Traversally are keen to support these
interests so long as it can be scoped back to the value/impact on an
organisation.

To date, there has been little rigorous and evidence backed analysis
of small versus large language models and thei value tradeoff for an
organisation, and so this project has considerable impact potential.
Travesally will be actively looking to disseminate these findings to
organisations they partner with, and their C-Suite executives, which
is a fantastic opportunity for the student.
Student deliverables A report outlining the value of small versus large language models,
both at a technical level and from the perspective of an organisation
deploying AI, and associated implementations to demonstrate
findings.
Student’s line manager Phininder Balaghan (Traversally CTO)
Project location(s) Remote with meetings either at Lancaster University or DiSH in
Manchester. Student can also have a desk at Lancaster University
with the Networks Research Group.
Other comments

Student Requirements
Essential technical skills The student will need to have both technical AI and data science
skills so they can implement the AI system and assess its
performance.
Desirable technical skills
Essential general skills The student should be comfortable presenting to both technical, and
senior, stakeholders. The student should also be happy to manage
their project and drive it forward – as this is primarily research based
– although as long as the student is happy with this they don’t need
a track record doing it to date.
Desirable general skills
Essential experience I am keen to provide the student with experience – and I am happy
to consider people from any backgrounds.
Desirable experience
Other comments Please indicate any attributes that may be beneficial to the student
within your organisation or any other factors to be considered.

Organisation Details

Main contact name Phininder Balaghan


Main contact position CTO / Founder
Applications email To whom should we send the students’ applications?
[Link]@[Link]
Applications email cc Is there anyone else who should be copied on the application emails?
Edward Austin – [Link]@[Link]
Other comments Is there anything else about your organisation of which you would
like to make us aware?
PRJ36-LTH-FTR - Frailty Trajectories in Renal Replacement Therapy:
A Population-Scale Cohort Study using the Lancashire OMOP CDM -
Lancashire Teaching Hospitals NHS Foundation Trust
Project for HDS DS AI

Project Description
Project name Frailty Trajectories in Renal Replacement Therapy: A Population-
Scale Cohort Study using the Lancashire OMOP CDM
Stipend offered [We make bold steps in new technology. Though we are unable to
offer monetary stipend, we offer international recognition in cutting-
edge industry and gold standard practices, authored research, and
superior mentorship]
Host Organisation Lancashire Teaching Hospitals NHS Foundation Trust
Project description Background: This project utilises the IDRIL (OMOP-CDM) database to
investigate the complex, bidirectional relationship between frailty
and Chronic Kidney Disease (CKD). It tests the view of frailty as a
static health status, viewing it instead as a dynamic trajectory
influenced by treatments and interventions.

Aims: To produce a high-quality academic publication that


characterises the relationship between frailty and renal clinical
interventions, and maps the quality-of-life trajectory (independence
and place of death) for renal patients.

Specific Analytical Objectives:


1. Cohort Definition: Construct cohorts for Dialysis initiation and
ANCA vasculitis using OMOP concepts.

2. Algorithmic Phenotyping: Operationalise frailty using the hospital


frailty algorithms via code.

3. Survival Analysis: Model survival outcomes stratified on frailty


status and adjusted for socioeconomic indicators

4. Trajectory Mapping: Create Sankey diagrams to visualise patient


flow between states of independence, care, and mortality.

Work Organisation: Fast-paced, independent remote working. The


student will operate as a computational epidemiologist.

Technologies: Open source code-first reproducible research. R or


Python.
Student deliverables 1. Advanced Visualisations: Publication-ready figures (patient flow)
and survival curves (time-to-event).

2. Reproducible Pipeline: Documented code for survival regression


models and patient flow visualisations.
3. Manuscript Contribution: Drafted Methods and Results sections
suitable for a nephrology or epidemiology journal.
Student’s line manager Dr Ian Farr

Project location(s) Remote


Other comments This project requires a student capable of handling longitudinal time -
to-event data and complex visualisation logic. It is an ideal
opportunity for a technically strong student to gain applied
epidemiology skills using a validated, international-standard
common data model (OMOP).

Student Requirements
Essential technical skills Survival Analysis: Understanding of time-to-event analysis (e.g.
censoring, hazard ratios, Kaplan-Meier estimators, flexible
parametric survival models) and relevant coding libraries.

Complex Data Visualization: Ability to programmatically generate


e.g. Sankey/Alluvial diagrams to show flow and state changes.

Advanced Data Manipulation: Proficiency in R/Python to handle


longitudinal patient records.
Desirable technical skills OMOP-CDM: Familiarity with the OMOP Common Data Model
structure.

Version Control: Proficient with Git.


Essential general skills Fast & Independent: Independent self-starter capable of rapidly
iterating on analysis code.

Statistical Rigour: Ability to interpret regression models and survival


statistics.

Remote Discipline: Effective communication and time management


in a fully remote setting.
Desirable general skills Clinical Context: Genuine interest in epidemiology, frailty,
nephrology.

Academic Writing: Ability to synthesize complex statistical results


into clear scientific prose.
Essential experience Experience with longitudinal datasets and applying statistical models
to real-world data.
Desirable experience Previous contribution to academic research or co-authorship on
papers.
Other comments We are looking for a technical data scientist who wants to apply their
coding skills to real-world medical research.
Organisation Details
Main contact name Dr Ian Farr
Main contact position Lead Data Scientist
Applications email [Link]@[Link]
Applications email cc [Link]@[Link]
[Link]@[Link]
Other comments
PRJ37-LTH-PEF - Parameter-Efficient Fine-Tuning of Large Language
Models for RTT Pathway Validation in the NHS - Lancashire Teaching
Hospitals NHS Foundation Trust
Project for AI

Project Description
Project name Parameter-Efficient Fine-Tuning of Large Language Models for RTT
Pathway Validation in the NHS
Stipend offered [We make bold steps in new technology. Though we are unable to
offer monetary stipend, we offer international recognition in cutting-
edge industry and gold standard practices, authored research, and
superior mentorship]
Host Organisation Lancashire Teaching Hospitals NHS Foundation Trust
Project description Background: The NHS Referral to Treatment (RTT) standard relies on
accurate identification of key pathway events (e.g. referral receipt,
decision to treat, clock pauses and stops). In practice, these events
are often recorded across a mixture of structured fields and
unstructured clinical text, including referral letters, clinic
correspondence, and discharge summaries. Manual validation of RTT
pathways is labour-intensive and prone to inconsistency.
Recent advances in large language models (LLMs) offer powerful
tools for extracting structured meaning from unstructured clinical
text. However, full fine-tuning of LLMs is computationally expensive
and often infeasible within secure healthcare environments.
Parameter-Efficient Fine-Tuning (PEFT) techniques (e.g. LoRA,
adapters, prompt-tuning) provide a promising alternative by
enabling task-specific adaptation of models while keeping most
parameters frozen.
This project will explore the use of PEFT-adapted LLMs to support
validation of RTT pathways by analysing clinical letters and linked
patient datasets. The work will be conducted within a secure
analytics environment and aligned with NHS information governance
requirements. Outputs will be presented via Databricks-based
analytical workflows and dashboards to support operational RTT
validation teams.

Aims:
• To evaluate parameter-efficient fine-tuning approaches for
adapting LLMs to NHS RTT validation tasks
• To extract and classify RTT-relevant events from clinical
letters and patient datasets
• To assess model performance, uncertainty, and limitations in
a real-world NHS context
• To deliver reproducible analytics and visual outputs suitable
for operational review
Specific Analytical Objectives:
1. Data Preparation and Task Definition
• Generate synthetic clinical letters and associated patient
metadata
• Define RTT-relevant information extraction and classification
tasks (e.g. event detection, date extraction, pathway
validation flags)
2. Baseline NLP and LLM Evaluation
• Establish baseline approaches (rule-based or classical NLP
methods)
• Evaluate pre-trained LLM performance using prompt-only
approaches
3. Parameter-Efficient Fine-Tuning (PEFT)
• Implement PEFT techniques (e.g. LoRA, adapters, prompt-
tuning)
• Compare PEFT approaches in terms of performance,
computational efficiency, and data requirements
• Assess generalisation across specialties or document types
4. Model Evaluation and Validation
• Quantitative evaluation using metrics such as precision,
recall, F1-score, and calibration
• Error analysis focusing on clinically meaningful failure modes
• Assessment of uncertainty and confidence in extracted RTT
events
5. Operational Analytics and Visualisation
• Integrate model outputs into Databricks workflows
• Develop dashboards summarising pathway validation signals,
model confidence, and error patterns
• Explore how outputs could support (but not replace) RTT
validation processes

Work Organisation
The project will be conducted in a remote-first, independent working
environment. The student will work as an applied machine learning
researcher, collaborating with clinical informatics and data science
staff while maintaining strong autonomy in experimentation,
evaluation, and documentation.
Technologies
• Python-based machine learning and NLP frameworks
• Large Language Models (open-weight or approved models)
• Parameter-Efficient Fine-Tuning techniques (e.g. LoRA,
adapters)
• Reproducible, version-controlled research workflows

Student deliverables 1. LLM Fine-Tuning Pipeline


• Reproducible PEFT workflows for RTT-related NLP tasks
2. Evaluation Framework
• Quantitative and qualitative assessment of model
performance and limitations
3. Databricks Dashboards
• Visual summaries of pathway validation outputs and model
confidence

Student’s line manager Dr Ian Farr

Project location(s) Remote


Other comments This project focuses on decision support and validation, not
automated clinical or operational decision-making. All outputs are
intended to augment human review processes in line with NHS AI
governance and safety principles, including those applied within NHS
England.

Student Requirements
Essential technical skills • Machine Learning / NLP: Understanding of supervised
learning and text classification
• Python Programming: Experience with ML or NLP libraries
• Evaluation & Metrics: Ability to assess model performance
rigorously

Desirable technical skills • Large Language Models: Familiarity with transformers or


LLM-based workflows
• PEFT Methods: Prior exposure to LoRA, adapters, or prompt-
tuning
• Databricks: Experience with cloud or notebook-based
analytics platforms

Essential general skills • Independent Research: Ability to manage an experimental


research project
• Critical Thinking: Willingness to analyse model limitations
and bias
• Communication: Clear documentation and explanation of
complex ML concepts

Desirable general skills • Healthcare Context: Interest in NHS operational data and
clinical workflows
• Responsible AI: Awareness of ethics, bias, and governance in
AI systems

Essential experience Experience applying machine learning or NLP techniques to real-


world datasets.
Desirable experience Prior coursework or projects involving deep learning, transformers,
or applied NLP.
Other comments This project is ideal for a technically ambitious MSc student
interested in applied AI in healthcare, offering hands-on experience
with modern LLM adaptation techniques in a real NHS operational
context.

Organisation Details
Main contact name Dr Ian Farr
Main contact position Lead Data Scientist
Applications email [Link]@[Link]
Applications email cc [Link]@[Link]
[Link]@[Link]
Other comments
PRJ40-LU-AIC - AI-enhanced Campus Sustainability - Lancaster
University
Project for DS AI

Project Description
Project name AI-enhanced Campus Sustainability
Stipend offered N/A
Host Organisation Lancaster University
Project description The aim of this project will be to examine the use of AI or Data
Science in promoting LU’s sustainability goals through improving the
processes through which our campus parking data is processed. It is
envisaged that we may be able to improve LU’s carbon footprint
through using parking data to understand travel practices and
incentivise sustainable travel.

The specific scope of the project will be developed between the


student, their academic supervisor, the LU Sustainability Manager
and other stakeholders. It is expected that the student will bring
their own ideas into this discussion on which steps and technologies
may be useful in building a processing platform to gather and use
data to support processes that encourage LU staff to take more
environmentally friendly travel options.

The project may entail considering what parking data we collect,


how it is collected, and how it is currently used. The student would
then need to consider our overall sustainability goals and how we
might be able to draw insights from the data and engage with our
colleagues to enable us to progress towards these goals. Initial
questions to be considered would include:

1. What are the behavioural patterns in parking usage


among staff, and how do these vary by pay grade,
department, or time of year?
2. What are the barriers to and opportunities for adopting
more sustainable travel?
3. Can we identify potential car-share groups using existing
data? (i.e. staff arriving at similar times from similar
locations)
4. Can we build a predictive model to estimate the
likelihood of a staff member being suitable for car
sharing?
5. What are the potential emissions and cost savings from
implementing a more incentivised car share scheme?
6. How can we automate data to produce weekly car share
match suggestions, and how could this integrate into a
potential user interface.
7. What is the role of Data Science or AI in achieving
increased adoption of sustainable travel?

The specific form and scope of the project will be established


collaboratively.

Student deliverables It is anticipated that the student will deliver an in-depth report on LU
parking data that describes the insights derived from the data that
will enable us to incentivise sustainable travel. Additionally, the
student will be expected to produce a prototype of an AI-enhanced
system that will allow the efficient and effective processing of the
relevant data to enable us to reach our goals.
Student’s line manager To whom will the student report during the placement?

Project location(s) Remote & LU Campus


Other comments The student will be expected to comply with GDPR and LU data
usage standards.

Student Requirements
Essential technical skills Understanding of AI technologies
Python or R programming
Desirable technical skills As above
Essential general skills An understanding of data pipelines and workflows
An ability to map business requirements to technical solutions
Ability to organise own work and to maintain focus on the project
Desirable general skills
Essential experience
Desirable experience
Other comments

Organisation Details
Main contact name Darren Wilkinson
Main contact position Sustainability Projects Officer
Applications email d.wilkinson5@[Link]
Applications email cc
Other comments
PRJ42-NT-AI - AI-driven relationship management - Newground
Together
Project for DS AI

Project Description
Project name AI-driven relationship management
Stipend offered N/A
Host Organisation Newground Together
Project description Newground Together is a charity that helps communities in the
North West through promoting skills acquisition and connecting
people with job opportunities.

At Newground Together, we are looking to develop a more


structured and intelligent way of recording and managing our
relationships with businesses, partners, stakeholders, and the
individuals we engage with across our programmes.

As our partnership activity continues to grow, we need a system that


can:
• Capture and organise all partner and stakeholder
interactions
• Track meetings, engagement history, and follow-up actions
• Provide visibility across teams
• Support collaboration and reduce duplication
• Integrate AI to help summarise meetings, update records,
and flag opportunities

The project student would be expected to gain familiarity with the


work of Newground Together, with particular reference to the
information that is passed between stakeholders, and how this
information may be used to further the organisation’s mission.

The student will map the flow of information through NT’s processes
and advise how and which information should be stored in order to
enable the generation of insights about how best to serve NT’s
clients and, ideally, how AI agents can be introduced to automate
manual tasks and provide decision support.

The student will act as an AI consultant and will be expected to play


a key role in shaping the direction of the project and then delivering
a solution. The project is likely to combine process mapping, data
management, user-centred design, and applied AI to create a
practical tool that supports real-world partnership work.

Student deliverables A solid rationale for an AI-enabled solution that will assist NT in
relationship management and service delivery.
A report explaining how the solution may be implemented, with a
working prototype of as much of the solution as is possible.
Student’s line manager Tariq Ali, Business and Partnerships Lead

Project location(s) Remote


Other comments

Student Requirements
Essential technical skills An understanding of AI technologies and how they can be used to
increase the efficiency of information flow
Ability to develop working prototypes of AI systems
Desirable technical skills An understanding of business process mapping
Essential general skills Good communications skills – ability to ask the right questions
Ability to organise own work and work with a high degree of
independence
Desirable general skills N/A
Essential experience N/A
Desirable experience N/A
Other comments N/A

Organisation Details
Main contact name Tariq Ali
Main contact position Business and Partnerships Lead
Applications email
Applications email cc
Other comments
PRJ45-HG-EVE - Enriching the Visitor Experience - Holker Group
Project for AI

Project Description
Project name Enriching the Visitor Experience
Stipend offered N/A though travel expenses will be covered if necessary
Host Organisation Holker Group
Project description Dating back to the mid 1600’s and latterly the home of the
Cavendish Family for more than 250 years, Holker Hall & Gardens is
one of England’s most treasured stately homes and garden
experiences. It won the title of Best Tourist Attraction in the 2025
Cumbria Life Awards in a reader’s poll, reflecting its special place in
the heart and minds of the Cumbrian Community. The Hall is a Grade
II* listed stately home surrounded by 200 acres of parkland, formal
gardens and ancient woodland.

Project Summary:
It’s important to us that we give the best possible experience to our
visitors. We would like to explore how generative AI could be used to
enrich visitors’ experience of Holker Hall. We believe that by using AI
we may be able to create a more interactive, participatory
experience that will enable visitors to feel more of the personality of
the Hall, its previous inhabitants, and their way of life.

An additional consideration is that opportunities to share some of


the personality of the Hall with non-English speaking visitors have
been limited. We would like to see how AI and other technologies
may be used to enhance their experience and enable us to attract
more international visitors to our nationally important heritage
attraction.

The student will be invited to consider of what would a ‘participatory


experience’ consist, and how could AI be used to deliver this. This
project affords great scope for the consideration of where AI can add
value to a real-world application and offers the potential for gaining
rich insight into how AI could be used in tourism and hospitality.

Project Aims:
1. To substantially improve the visitor experience for both English
and non-English speaking visitors.
2. Use AI to create (and use AI-enabled hardware such as AI Glasses
to deliver) an interactive, history-rich Hall visit experience, delivered
in a visitor’s native language.
3. To pioneer this approach with a view to rolling it out into several
languages and eventually make it available in English for digital-first
visitors who would prefer this “accompanied” and interactive
experience.
4. Prototype Development – Design and test a solution in a selected
language such that it proves the case for investment in knowledge
base development and roll out
Student deliverables A working prototype of AI-enabled technology that will create a rich,
participatory experience for aspects of a visitor’s time at Holker Hall.

A database of historical information of a scale sufficient to test a


prototype experience in either French, German or Mandarin. It is
important that any information imparted via the process to visitors is
factually correct but also that visitors can interact with the AI Agent
as they tour the Hall, in their native tongue.

AI-enabled glasses that pick up visual clues from selected artefacts,


prompting the AI Agent to chat about the history of the artefact (a
painting by Van Dyk or a piece of furniture by William Morris for
example) and invite further questions from the user. Audio will
therefore need to be enabled and the ability to hear the AI Agent
when triggered visually.

A working prototype using AI Glasses that delivers an interactive


visitor experience on a limited scale – but one sufficient to secure
further internal investment to continue development towards a rich,
interactive experience to rival a person human guide, in multiple
languages.

Student’s line manager Colin Sneath, Head of Marketing

Project location(s) Remotely with tests at Holker Hall


Other comments

Student Requirements
Essential technical skills A strong grounding in database development, analysis, artificial
intelligence or digital systems design, and AI hardware.
Competencies should include analytical thinking, data handling and
problem-solving using AI or statistical methods, plus clear
communication and the ability to translate technical findings into
practical business recommendations.
Desirable technical skills
Essential general skills
Desirable general skills
Essential experience
Desirable experience
Other comments
Organisation Details
Main contact name Colin Sneath
Main contact position Head of Marketing
Applications email To whom should we send the students’ applications?
Applications email cc Is there anyone else who should be copied on the application
emails?
Other comments Holker Group is a diversified, family-owned enterprise based on the
17,000-acre Holker Estate on the Cartmel Peninsula in South
Cumbria. The Group encompasses Holker Hall & Gardens, Cartmel
Racecourse, Old Park Wood and Longlands Holiday Parks, Burlington
Stone, Burlington Aggregates, Holker Farms and Holker Homes. We
balance commercial innovation with a long-term commitment to
landscape, community and custodianship.
PRJ46-PL-AIQ - AI Quote Modelling - Playdale Ltd
Project for DS AI

Project Description
Project name AI Quote Modelling
Stipend offered N/A
Host Organisation Playdale Ltd
Project description The project involves creating an AI model that could analyse live
quote data. This data should then be easily interrogatable in plain
language to understand what a new prospective customer will be
most likely to order

The previous order data contains metrics such as where in the


country an order was placed, what was ordered, the sector in which
the order was placed (school, council, housing developer, etc.) The
date it was ordered, what type of equipment was ordered, e.g.
timber, steel etc, and whether or not the quote was won or lost.

At the end of this project Ideally a sales rep could get a request for a
quote and input parameters such as, “School based in Lancashire
with a budget of £30,000, looking for timber equipment” and the AI
model could tell them what to quote for, that will give the highest
likelihood of Playdale winning the order.

All of our quote data can be supplied as a raw Excel file or gathered
directly from our on-site SQL server. The specifics of the language
and tools used to create the AI model would be at the discretion of
the student.

Student deliverables Fully-documented AI model that will analyse previous quotes and
advise on pricing for future quotes.
Student’s line manager Ollie Harbridge, Digital Development Engineer

Project location(s) The position will mostly be remote, with the opportunity to visit our
offices and factories to gain perspective on day-to-day operations -
Playdale will pay for any travel as required.
Other comments We work Monday – Friday 8am-4pm, however we wouldn’t require
the student to replicate this. We would only require the check-ins or
any site visits to be within these hours; any actual remote work can
be done flexibly.

Student Requirements
Essential technical skills Software Development
Basic machine learning
Data Analysis
AI Development
Desirable technical skills
Essential general skills Comfortable with collaborative environments, however you will
mainly be working independently.
Forthcoming with ideas and suggestions. i.e. if something isn’t
working or can’t work, be forthcoming with suggestions.

Will give experience designing and implementing software solutions


for real world business applications.
Desirable general skills
Essential experience
Desirable experience
Other comments

Organisation Details
Main contact name Ollie Harbridge
Main contact position Digital Development Engineer
Applications email
Applications email cc
Other comments
PRJ47-EA-VN - Validation of geospatial mortality projections for
nutrition program targeting in Nigeria - Evidence Action
Project for HDS

Lancaster University Health Data Science Project Description

Please give details of the project in broad terms -- this information is for guidance and will be
replaced by a detailed specification later in the process.

Validation of geospatial mortality projections for nutrition program


Project name
targeting in Nigeria

Stipend offered

Host Organisation Evidence Action

The aim of this research project is to validate the accuracy of sub-


national mortality projections for program targeting decisions in
Nigeria. Specifically, the project will assess whether ward-level
geospatial estimates of under-five mortality derived from the 2018
Nigerian Demographic and Health Survey (DHS) and projected
forward to 2024 are sufficiently accurate when compared against
empirical values from the newly released 2024 DHS data. Results
will directly inform whether Evidence Action's allocation decisions
for Small-Quantity Lipid-Nutrient Supplement (SQ-LNS) programs
can reliably be based on projected sub-national mortality estimates.

Project description The student will work with geospatial mortality estimates that were
previously generated from the 2018 DHS using model-based
geostatistical methods and projected to 2024. They will then
generate comparable ward-level mortality estimates from the 2024
DHS data and conduct a comprehensive validation analysis. The
analysis will include measures of prediction accuracy (e.g., mean
absolute error, root mean squared error), correlation between
projected and observed values, spatial patterns of prediction errors,
and assessment of whether accuracy varies by geographic region or
baseline mortality levels.

A critical component of the project will be evaluating whether the


level of prediction accuracy is sufficient to support programmatic
decision-making, particularly for identifying high-need areas for
nutrition intervention targeting. The student will assess both
statistical measures of accuracy and their practical implications for
resource allocation decisions.

The student will lead the analysis with support and guidance from
Evidence Action's Associate Director, Research, and may engage
with international experts in geostatistical modeling and mortality
estimation. The student will be provided with the necessary
geospatial mortality estimates, DHS datasets, statistical software
licenses, and guidance on geostatistical methods, validation
frameworks, and interpretation of findings for program application.

This project provides an opportunity to contribute directly to


evidence-based decision-making for nutrition programs affecting
millions of children in Nigeria.

A validation report presenting: (1) ward-level under-five mortality


estimates from 2024 DHS data; (2) comprehensive accuracy
assessment comparing 2024 projections against 2024 empirical
estimates; (3) spatial analysis of prediction errors and patterns; (4)
Student deliverables
evaluation of whether projection accuracy is sufficient for program
targeting decisions; (5) recommendations for future mortality
estimation and projection approaches; (6) methodological
documentation

Student's line manager Mark Minnery, Associate Director, Research

Remote work arrangement with regular virtual meetings. Student


Project location(s)
may work from Lancaster University or home location.

The student will gain valuable experience in validation


methodologies, geospatial analysis, and translating statistical
findings into program recommendations. This project addresses a
Other comments critical question for Evidence Action's programming: whether sub-
national mortality projections are reliable enough to guide resource
allocation decisions. The findings will have direct, immediate
application to ongoing program planning.
Student Requirements
Statistical software (R, Stata, or similar); statistical analysis including
Essential technical skills validation methods and accuracy metrics; experience with survey
data and sampling weights; data visualization

Geospatial analysis and GIS software; geostatistical methods; spatial


statistics; experience with DHS or similar complex survey datasets;
Desirable technical skills
predictive modeling and model validation; mapping and spatial
visualization

Critical thinking about model accuracy and practical implications;


Essential general skills clear technical writing; project management; ability to work
independently with regular supervision

Understanding of demographic and epidemiological methods;


ability to translate technical findings into programmatic
Desirable general skills
recommendations; experience presenting results to diverse
audiences; familiarity with decision-making frameworks

Prior coursework or project work involving model validation,


Essential experience prediction accuracy assessment, or comparative analysis with large
datasets

Previous work with health survey data; mortality estimation


methods; geospatial or spatial analysis; cross-validation or out-of-
Desirable experience
sample prediction assessment; familiarity with child health or global
health contexts

Strong attention to methodological rigor will be essential, as will the


ability to think critically about what level of accuracy is "sufficient"
Other comments for real-world decision-making. The student should be comfortable
working with real-world data and translating technical findings into
practical recommendations.

Organisation Details
Main contact name Mark Minnery

Main contact position Associate Director, Research


Applications email [Link]@[Link]

Applications email cc

Evidence Action is an evidence-based global health organization


implementing cost-effective interventions at scale. Our nutrition
programs, including SQ-LNS supplementation, reach millions of
Other comments
children across multiple countries in Africa and Asia. This validation
project directly supports our commitment to evidence-based
resource allocation.
PRJ48-EA-MN - Mortality correlations for cost-effectiveness modeling
of child nutrition interventions in Nigeria - Evidence Action
Project for HDS

Lancaster University Health Data Science Project Description

Please give details of the project in broad terms -- this information is for guidance and will be
replaced by a detailed specification later in the process.

Mortality correlations for cost-effectiveness modeling of child


Project name
nutrition interventions in Nigeria

Stipend offered

Host Organisation Evidence Action

The aim of this research project is to estimate correlations between


different age-specific mortality indicators in Nigeria to improve
cost-effectiveness modeling for child nutrition programs.
Specifically, the project will examine relationships between under-
five mortality, infant mortality, and neonatal mortality rates at sub-
national levels. Results will be used to refine cost-effectiveness
models for Small-Quantity Lipid-Nutrient Supplement (SQ-LNS)
programs currently being implemented to address child
malnutrition in Nigeria.

The student will analyze multiple rounds of the Nigerian


Project description Demographic and Health Survey (DHS) series to estimate Local
Government Area (LGA) level mortality rates using appropriate
statistical methods. The analysis will compare correlation
coefficients between different mortality age classes (neonatal,
infant, under-five) and examine how these relationships have
changed over time. Understanding these correlations is essential for
accurately modeling the full mortality impact of nutrition
interventions that target specific age groups.

The student will lead the analysis with support and guidance from
Evidence Action's Associate Director, Research, and may have
opportunities to engage with international experts in child health
and mortality estimation. If necessary, the student will be provided
with statistical software licenses and will receive guidance on data
management, geostatistical methods, and advanced statistical
analysis appropriate for complex survey data.

This project provides an opportunity to work with internationally


recognized health survey data and contribute to evidence-based
programming that affects millions of children.

A research report presenting: (1) LGA-level mortality estimates by


age class across survey rounds; (2) correlation analyses between
Student deliverables mortality indicators with temporal trends; (3) recommendations for
cost-effectiveness modeling; (4) methodological documentation of
statistical approaches used

Student's line manager Mark Minnery, Associate Director, Research

Remote work arrangement with regular virtual meetings. Student


Project location(s)
may work from Lancaster University or home location.

The student will gain experience working with complex survey data,
geostatistical methods, and translating research findings into
Other comments practical applications for program decision-making. The project
contributes directly to Evidence Action's ongoing nutrition
programming in Nigeria.

Student Requirements
Statistical software (R, Stata, or similar); advanced statistical
Essential technical skills analysis including regression methods; experience with survey data
and sampling weights

Multilevel/hierarchical regression modeling; geostatistical methods;


Desirable technical skills spatial analysis; experience with DHS or similar complex survey
datasets; data visualization

Project management; clear technical writing and reporting; ability


Essential general skills
to work independently with regular supervision
Understanding of demographic and epidemiological methods;
Desirable general skills ability to interpret findings in programmatic contexts; experience
presenting technical results to diverse audiences

Prior coursework or project work performing regression analysis


Essential experience
with large datasets; experience handling complex data structures

Previous work with health survey data; mortality estimation


Desirable experience methods; spatial or multilevel modeling; familiarity with child
health or global health contexts

Strong statistical foundations and attention to methodological rigor


will be essential. The student should be comfortable working with
Other comments
real-world data that may require extensive cleaning and
preparation.

Organisation Details
Main contact name Mark Minnery

Main contact position Associate Director, Research

Applications email [Link]@[Link]

Applications email cc

Evidence Action is an evidence-based global health organization


implementing cost-effective interventions at scale. Our nutrition
Other comments
programs, including SQ-LNS supplementation, reach millions of
children across multiple countries in Africa and Asia.
PRJ49-AEC-AIE - AI for Enhanced Therapy Provision – Client Toolkit -
Amanda Englishby Coaching
Project for DS AI

Project Description
Project name AI for Enhanced Therapy Provision – Client Toolkit
Stipend offered N/A
Host Organisation Amanda Englishby Coaching
Project description Coaching and therapy are services that rely on a strong connection
and trust between client and practitioner. While these services are
inherently human and cannot be replaced, this project will explore
how AI might support and enhance the client experience between
sessions.

Clients often have gaps between sessions where structured support


can help them engage with their goals and practice skills
independently. This project will explore whether AI can provide a
safe, context-sensitive support toolkit that clients can use between
sessions.

The toolkit could include:

• exercises and techniques from coaching and hypnotherapy


practice
• reflective prompts and guided activities
• tips for maintaining motivation and practicing skills
• resources to support mindset and personal growth

The toolkit is supportive, not therapeutic, and always encourages


clients to seek human guidance for any issues beyond the AI’s scope.

The student will gain an understanding of your services and style of


working and will develop a prototype system for this toolkit,
ensuring ethical boundaries are central to the design.

Key principles:

• The AI toolkit is supportive and reflective, not a replacement


for therapy or coaching.
• Its limitations will be clearly communicated to clients.
• Clients are consistently encouraged to access human support
for any issues outside the toolkit’s remit.
• Confidentiality and ethical practice are paramount
throughout the project.

Student deliverables A working, well-documented prototype of an AI support toolkit for


clients to use between sessions.
Student’s line manager Amanda Englishby

Project location(s) Remote


Other comments Confidentiality and an ethical approach to the project will be of
paramount importance

Student Requirements
Essential technical skills Familiar with current AI technologies
AI agent development
Desirable technical skills
Essential general skills Able to conduct independent research and development
Understanding of relevant ethical issues
Desirable general skills
Essential experience
Desirable experience
Other comments

Organisation Details
Main contact name Amanda Englishby
Main contact position Clinical Hypnotherapist
Applications email
Applications email cc
Other comments
PRJ50-NR-S - Streamlined process for creating railway drainage
degradation matrices - Network Rail
Project for DS AI

Project Description
Project name Streamlined process for creating railway drainage degradation
matrices
Stipend offered £5,400
Host Organisation Network Rail
Project description Network Rail runs a suite of forecasting models to predict the
condition of and risk associated with railway assets under different
funding scenarios. Intelligence from these models supports our
business case for drainage funding (~£600m) from Government. Our
model of drainage assets (pipes, catchpits, ditches etc.) relies on
deterioration probabilities calibrated on historical data. Drainage
data is changing rapidly – more data is being collected every day, and
possibly trends in asset behaviour could be changing due to climate
change. However, we do not have an automated way of updating
these deterioration probabilities at present.

The aim of this project would be to streamline the process of


updating the deterioration probabilities with new data and creating
the associated “degradation matrices”. It would involve creating a
programme (likely in Python) to extract and transform relevant
drainage data, execute a (pre-defined) methodology for calculating
degradation matrices and output the matrices in an appropriate
format for ingress to our forecasting model.

The student would be working independently under the guidance of


NR’s drainage mathematical modeller. This is an opportunity for a
student to develop their analytic and programming skills with real-
world data.
Student deliverables A programme which streamlines the creation of drainage
degradation matrices and documentation on the programme’s
operation.
Student’s line manager Mathematical Modeller (Drainage)

Project location(s) Remote, with some visits (~4) to Network Rail’s office in Milton
Keynes. Possibility for more visits to Milton Keynes office or to also
meet in Manchester office.
Other comments Full time, flexible hours with allowance for 2 weeks leave.
Up to £500 travel budget.

Student Requirements
Essential technical skills General programming skills, data analytics
Desirable technical skills Experience with SQL and Python
Knowledge of Markov chain models
Essential general skills Good communication skills
Desirable general skills Awareness of asset management concepts or scenario planning
concepts
Essential experience Understanding existing code
Desirable experience Building software, handling large datasets
Other comments

Organisation Details
Main contact name
Main contact position
Applications email
Applications email cc
Other comments
PRJ51-NR-U - Using data science to modernise and speed up cleaning
and processing for a large-scale railway track strategic model -
Network Rail
Project for AI

Project Description
Project name Using data science to modernise and speed up cleaning and
processing for a large-scale railway track strategic model
Stipend offered £5,400
Host Organisation Network Rail
Project description Network Rail runs a suite of forecasting models to predict the
condition of and risk associated with railway assets under different
funding scenarios. Intelligence from our “Track” model supports our
business case for Track funding (~£4bn over a period of 5 years) from
Government.

Data from multiple different sources must be collated, cleansed,


made self-consistent, and put into the format required for the Track
strategic model before it can be used (this is a data preparation
process). This data includes Network Rail’s “track asset” location
data (where in the country different railway track is), what condition
it is in, how old it is, the type of rails, sleepers, ballast, that are
installed, and so on. At present, the procedure for going from “data
sources to fit-for-purpose in the strategic model” is a lengthy
process, taking on the order of months to achieve.

The aim of this project would be to streamline specific elements of


the process, using data science and related techniques (e.g. AI) to
target specific time-consuming steps and improve the end-data
quality compared to what it is at present. The focus would be on
creating a transparent, generalisable methodology/framework that
could be used by Network Rail beyond the specific elements of the
data preparation process used within the project. The preference
would be to use Python and/or SQL to achieve these aims but are
open to alternatives. Depending on progress, this project may well
be repeated and extended in future years.

The student would be working independently under the guidance of


NR’s Track mathematical modeller. This is an opportunity for a
student to develop their data science and programming skills with
strong exposure to real-world data and the issues that come with
this.
Student deliverables A methodology/framework (including programming implementation)
to streamline specific elements of the Track data processing and
documentation on the framework/programme’s operation.
Student’s line manager Mathematical Modeller (Track)
Project location(s) Remote, with some visits (~4) to Network Rail’s office in Milton
Keynes. Possibility for more visits to Milton Keynes office or to also
meet in Manchester office.
Other comments Full time, flexible hours with allowance for 2 weeks leave.
Up to £500 travel budget.
At least weekly Teams calls to ensure the project is fruitful for both
the student and Network Rail.

Student Requirements
Essential technical skills General programming skills, data analytics, data science
Desirable technical skills Experience with Python (pandas, scipy, numpy, sci-kit learn, etc.) and
SQL (not expecting fluency in all of these)
Knowledge of implementation of AI
Essential general skills Good communication skills
Desirable general skills Awareness of data quality issues, data cleansing processes
Essential experience Understanding existing code
Desirable experience Building software, handling large datasets
Other comments Enthusiasm for understanding mathematical modelling principles in
a real-world setting to get a good idea of the context behind the data
science work.

Organisation Details
Main contact name
Main contact position
Applications email
Applications email cc
Other comments
PRJ53-G-IVE - Innovative Vocabulary Estimation Method - Glite
Project for DS AI

Project Description
Project name Innovative Vocabulary Estimation Method
Stipend offered £1,500
Host Organisation Glite
Project description About Glite:
We're developing a language learning app and website called Glite.
It's focused on intermediate and advanced learners, and only on the
English language. The company is based in London and is currently in
software product development before product launch.

The app begins with a vocabulary test to assess each user’s level and
personalize their learning path. Currently, Glite estimates vocabulary
size using a traditional “Yes/No” test format — 50 words are
presented, and users click the words they know. Howeve r, this
approach has two main drawbacks:

1. It takes 10–12 minutes to complete, which reduces user


engagement
2. Users are not always honest in their answers**,** leading to
overestimation of vocabulary size.
3. When we tried to introduce a concept of “fake phrases” in
our test it was not clear for users and they give us wrong
data (we should not use such approach)
Project objectives:
To improve both user experience and measurement accuracy, Glite
seeks a more efficient and probabilistic approach to vocabulary
estimation. We would like our vocabulary estimation module to not
only estimate the number of words that our users know but also
map their knowledge onto our dictionary: for each entry in the
dictionary, there should be a probability that the user knows the
meaning of the entry (entry = meaning of an individual word,
multiple meanings per word).

Specific tasks:
Task 1 — Propose a New Test Design
1. Develop a new click-based vocabulary estimation method.
Some examples of most frequently used vocabulary size
estimation methods are described in Appendix 1.
2. The proposed method infers both total vocabulary size and
the probability of knowing a specific meaning of a specific
word. (mapping to all meanings in our dictionary)
3. The test design should balance accuracy and length.
Task 2 — Project Planning
Provide a detailed project plan with:
1. Estimated time for each development stage and what will be
achieved after that stage
2. Required resources
3. Key deliverables and milestones
Task 3 — Implementation and Dataset Generation
1. Design and implement the vocabulary test (backend and UI).
2. Use existing datasets and also generate a new dataset (by
creating a simple web interface for data collection which can
be run on Glite website and or app). This interface should
work well on mobile devices.
3. Integrate results with Glite’s internal dictionary.
Task 4 — Optimization
1. Optimize test length while maintaining or improving
accuracy.
2. Implement a method (e.g., adaptive testing, Bayesian
updating, reinforcement-based sampling) to select the most
informative sequence of words for testing based on previous
user’s answers in the same test.
3. The goal is to minimize test time while preserving reliable
vocabulary size estimation.
Inputs:
1. User click-based activity. Click-based responses and
behavioral metrics (to be defined during design).
2. Glite dictionary: The internal database of English word
meanings and concepts.
Outputs:
1. Vocabulary size estimate — total number of words known by
a user. 1 number.
2. Knowledge probabilities — likelihood that the user knows a
particular meaning of a specific word, linked to entries in the
Glite dictionary (all entries). Array of hundreds of thousands
of numbers (0 to 1).
Available datasets:
From Glite (can be used if needed):
Glite database of answers (1 million completed tests) {word, yes/no}
Glite database of interactions with users (Descriptions and responses
to image-based prompts)
Glite Tech dictionary: Structured list of word meanings and concepts
External datasets (can be used if needed and if license permits):
Open-source datasets from major vocabulary testing frameworks can
be used
Developed datasets:
Datasets developed by the developer during project implementation

Project implementation details


Meetings: Two sync meetings per week
Each sync: summary of work done + plan for next steps

IP and Legal:
All intellectual property belongs to Glite
NDA required before project starts
Ownership transfer agreement at completion
Publication rights to be discussed separately. Glite is open to
collaborative
publications and/or joint participation in conferences

Metrics of success
1. Improved estimation accuracy compared to current Yes/No
test (specific evaluation metric to be agreed, e.g., correlation
with VST benchmark).
2. Reduced test duration (significantly shorter than 10–12
minutes).
3. Reduced number of users dropping out during the initial test
4. Model robustness — consistent performance across user
levels

Student deliverables Deliverables


1. Codebase for the vocabulary estimation model
2. Prototype UI/website for running the test and/or generating
new data
3. Codebase for UI/website and backend
4. All dataset created or gathered during this project
5. All trained weights of models created during this project
6. All data of tries and experiments registered on some
experimentation platform (like weights and biases
[Link]
7. Project report detailing methodology, datasets, model
design, and test results
8. Documentation for reproducibility and integration with Glite
backend
9. Signed NDA and IP ownership transfer

Student’s line manager Dr Anton Nikolaev

Project location(s) Remote


Other comments APPENDIX 1. WIDELY USED VOCABULARY ESTIMATION TESTS
NB! This list is for information purposes, and we do not expect
these tests to be implemented.
There are several widely used vocabulary estimation tests listed
below. You do not need to choose any of those,

Vocabulary Size Test (VST) described here Multiple-choice format:


one click per question to choose the correct meaning.

Vocabulary Levels Test (VLT) —described here Matching can be


adapted as clickable pairs or multiple-choice equivalents; widely
implemented as a “click the right definition” test.

Lexical Decision Task —


Click “Word” or “Nonword” for each item; standard psychological
experiment setup.

Computerized Adaptive (IRT-based) Vocabulary Tests —


Essentially the same as VST or Yes/No formats, but items are
selected automatically based on previous clicks.

More detailed explanation of each test

Test 1. The Nation & Beglar Vocabulary Size Test (VST) is one of the
most widely used methods for estimating how many English words a
learner knows.
Step-by-step description:
1. Word Sampling. The test designers take a very large list of
the most frequent English word families (for example, 14,000
word families from the British National Corpus).
From each 1,000-word frequency band, they randomly select a small
number of words (usually 10 per band).
2. Test Construction. Each test item is a multiple-choice
question with:
a target word (e.g., cup),
a short context sentence (to make meaning clearer),
four definitions, only one of which is correct.
Example:
cup — “He drank tea from a cup.”
(a) small open container for drinking
(b) kind of tree
(c) piece of clothing
(d) part of the body
3. Testing
The learner reads each item and chooses the meaning that matches
the word.
No time pressure is used; the test checks receptive knowledge
(knowing the meaning when you see the word).
4. Scoring. The test usually has 140 items (10 items × 14
frequency bands).
The learner’s correct answers are counted within each frequency
band.
5. Estimation. The proportion of correct answers per band is
used to extrapolate total vocabulary size.
Example:
If the learner scores 80% correct up to the 8,000-word band →
estimated vocabulary size ≈ 0.8 × 8,000 = 6,400 word families.

Test 2. Vocabulary Levels Test (VLT)


1. Word Sampling. Words are grouped by frequency levels (for
example, the 1,000, 2,000, 3,000, 5,000, and 10,000 most
frequent word families). From each level, researchers choose
a small representative sample of words — typically 18 words
per level.
2. Test Construction
Each test section contains six items made from three definitions and
six words.
The task is matching — the learner links each definition to the
correct word from the list.
Example (simplified):
Words Definitions
1. business a. part of your arm
2. clock b. something for telling time
3. horse c. animal people ride on
4. shoulder
5. ship
6. watch
Answer: 4–a, 6–b, 3–c
3. Testing. The learner completes one section per frequency
level. Each correct match earns one point. The test checks
receptive meaning knowledge — whether the learner
recognizes the meanings of common words at each
frequency level.
4. Scoring. Each level is scored separately (out of 18). A typical
“mastery” cutoff is around 15/18 correct — if the learner
reaches that, they are considered to “know” that frequency
level.
5. Estimation. The highest level at which mastery is achieved
shows the learner’s estimated vocabulary range.
Example: If a learner reaches mastery at the 3,000-word level, their
receptive vocabulary size is estimated around 3,000 word families.
6. Result. The result gives a banded estimate of vocabulary size
(e.g., 1k, 2k, 3k level known).
Researchers can use it to profile learners by how many frequency
bands of English vocabulary they know.

Test 3. Word Associates Test (WAT) — sometimes also called the


Word Associates Format (WAF).
1. Word Selection. The test includes a set of stimulus (target)
words, usually of medium to high frequency (e.g., beautiful,
judge, bank). Each target word is chosen so it has multiple
meaning or collocation associations (e.g., synonyms, related
adjectives, typical contexts).
2. Test Construction. For each target word, there are 8 options
arranged in two groups of four: Four are meaning associates
(synonyms or close semantic links). Four are collocational
associates (words that typically occur with the target).
The test taker must select all the words that are related to the target
word — usually 4 correct out of 8 options.
Example:
Target word: bank
Choose all the words related to “bank”:
| a. money | b. river | c. sleep | d. grass | e. finance | f. cloud | g.
loan | h. jump |
Correct answers: a, b, e, g
3. Testing. The learner reads each target and marks every
option they believe is related.
There’s no time limit — it measures depth of vocabulary knowledge,
but total correct responses can also reflect vocabulary size.
4. Scoring. Each correct association gets one point. Wrong
choices (selecting unrelated words or missing correct ones)
reduce the score. The total number of correct associations is
summed across all items.
5. Estimation. Researchers calculate an overall proportion of
correct associations. The resulting score indicates how many
words and semantic connections the learner likely knows,
which can be converted into a vocabulary size estimate (or
used to compare learners).
6. Result. The output is a single score showing how rich and
extensive the learner’s vocabulary knowledge is, based on
their ability to recognize both meaning and collocation links.
Test 4. Item Response Theory (IRT)–based vocabulary size test
1. Item Pool Creation. Researchers prepare a large bank of
vocabulary questions — often hundreds or thousands —
drawn from various frequency levels (e.g., 1,000–14,000
word families). Each item might be a short multiple-choice
question, a Yes/No word recognition item, or a translation
recognition question.
2. Item Calibration
Each item is first tested on a large group of learners.
Statistical modeling (IRT) estimates how difficult each item is and
how well it discriminates between learners with different vocabulary
sizes.

After calibration, each item has parameters like:


difficulty (how rare or hard the word is),
discrimination (how sharply it separates stronger and weaker
learners),
guessing (chance of answering correctly by luck).

3. Testing
The learner is shown one question at a time. After each answer, the
system updates its estimate of the learner’s vocabulary ability level.
The next question is chosen adaptively — slightly easier or harder
depending on the previous response. (This is called Computerized
Adaptive Testing, or CAT.)

4. Scoring. Instead of just counting correct answers, the model


estimates a latent ability score (θ) for each learner — a
continuous value representing their overall vocabulary
knowledge.
5. Estimation of Vocabulary Size. The θ score is mapped to an
approximate vocabulary size using a conversion table or
regression model derived during calibration.
For example, θ = 0 might correspond to knowing 5,000 word
families. θ = 1 might correspond to 8,000 word families, and
so on.

6. Result The output is a numerical vocabulary size estimate


(e.g., “Learner knows about 7,800 word families”). Because
the test adapts to each person, it’s shorter and more
precise than fixed-form tests.
Student Requirements
Essential technical skills 1. Experience in computational linguistics, or NLP applied to
education
2. Knowledge of Item Response Theory (IRT), Bayesian
modeling, or ML-based adaptive testing
3. Strong Python background (PyTorch, TensorFlow, scikit-
learn)
4. Understanding of vocabulary assessment research and
English lexical databases (WordNet, BNC/COCA, etc.)

Desirable technical skills


Essential general skills 1. Good organisational skills

2. Independence and resilience


Desirable general skills
Essential experience
Desirable experience
Other comments

Organisation Details
Main contact name Dr Anton Nikolaev
Main contact position Data Scientist
Applications email To whom should we send the students’ applications?
Applications email cc Is there anyone else who should be copied on the application
emails?
Other comments Is there anything else about your organisation of which you would
like to make us aware?
PRJ54-G-MIE - Model for IELTS Speaking Level Projection - Glite
Project for DS AI

Project Description
Project name Model for IELTS Speaking Level Projection
Stipend offered £1,500
Host Organisation Glite
Project description We're developing a language learning app and website called Glite.
It's focused on intermediate and advanced learners, and only on the
English language. The company is based in London and is currently in
software product development before product launch.
As a part of the language knowledge test the app will provide IELTS
test estimation especially for the speaking part. The IELTS
(International English Language Testing System) is a globally
recognized exam that assesses English proficiency across four skills:
Listening, Reading, Writing, and Speaking. The Speaking test is a face -
to-face interview with a certified examiner and lasts about 11–14
minutes. It is divided into three parts: in Part 1, the examiner asks
general questions about familiar topics such as home, work, or
hobbies; in Part 2, the test taker speaks for up to two minutes about
a given topic after one minute of preparation; and in Part 3, the
examiner and test taker discuss more abstract or complex questions
related to the Part 2 topic. The Speaking test evaluates fluency,
pronunciation, vocabulary, and grammatical accuracy through a
natural conversational format.

Project objectives
The objective of this project is to develop a ML/DL based module that
estimates a user’s IELTS Speaking score based on their spoken
responses within the Glite app. The system will analyze users’ speech
to assess fluency, pronunciation, grammatical accuracy, and lexical
range, mirroring IELTS Speaking criteria. It will then provide
automated, personalized feedback and guidance to help users
improve their spoken English proficiency and target higher IELTS band
scores.

Specific tasks
Task 1 — Solution Design. Propose a technical solution for estimating
IELTS Speaking scores from user audio recordings. The proposed
approach should outline how the system will evaluate key IELTS
criteria — fluency, pronunciation, grammatical accuracy, and lexical
range — and generate meaningful feedback for learners.

Task 2 — Project Planning. Develop a comprehensive project plan


detailing:
1. Estimated timelines for each development stage
2. Resource requirements (data, tools)
3. Key deliverables and milestones for implementation and
evaluation
Task 3 — Implementation. Implement the proposed solution,
including model development, and testing on real or synthetic speech
datasets. The system should output both an estimated IELTS band
score and automated feedback to guide user improvement.

Model input
User’s answer (Voice data) to several questions - 2 short questions on
familiar topics and one longer topic, where the user speaks for about
two minutes.

Model output
1. IELTS level (number)
2. Feedback on fluency, pronunciation, grammatical accuracy,
and lexical range
Related solutions and datasets
Speak & Improve is a free online platform developed by Cambridge
Assessment that allows English learners to practice speaking and
receive instant, automated feedback. Users complete short spoken
tasks—such as answering questions, reading prompts, or giving
opinions—and the system analyzes their speech for features like
fluency, pronunciation, and accuracy. Based on this analysis, it
estimates the user’s proficiency level using the CEFR scale (A1–C1).
The tool also serves as a research platform, collecting anonymized
learner speech data to improve automated language assessment
technologies.

Dataset: [Link]
improve-corpus-2025?utm_source=[Link]
NB! This is a dataset for non-commercial use only and cannot be
used for training of the model

Project implementation details

Meetings: Two sync meetings per week


Each sync: summary of work done + plan for next steps

IP and Legal:
All intellectual property belongs to Glite
NDA required before project starts
Ownership transfer agreement at completion
Publication rights to be discussed separately. Glite is open to
collaborative
publications and/or joint participation in conferences

Metrics of success
1. Prediction Accuracy
Correlation (e.g., r ≥ 0.85) between the system’s predicted
IELTS Speaking band scores and human examiner scores on a
standardized test set inferred from available datasets.
Mean Absolute Error (MAE) ≤ 0.5 IELTS band compared to
certified examiner ratings inferred from available datasets.

2. Criterion-Level Consistency
Agreement with human raters on each IELTS subscore
(Fluency & Coherence, Pronunciation, Lexical Resource,
Grammatical Range & Accuracy) — measured via Cohen’s κ
or weighted F1 ≥ 0.75.

Student deliverables Deliverables


1. Codebase for the IELTS level estimation model
2. Prototype UI/website for running the test and/or generating
new data
3. Codebase for UI/website and backend
4. All dataset created or gathered during this project
5. All trained weights of models created during this project
6. All data of tries and experiments registered on some
experimentation platform (like weights and biases
[Link]
7. Project report detailing methodology, datasets, model
design, and test results
8. Documentation for reproducibility and integration with Glite
backend
9. Signed NDA and IP ownership transfer

Student’s line manager Dr Anton Nikolaev

Project location(s) Remote


Other comments APPENDIX 1. WIDELY USED VOCABULARY ESTIMATION TESTS
NB! This list is for information purposes, and we do not expect
these tests to be implemented.
There are several widely used vocabulary estimation tests listed
below. You do not need to choose any of those,

Vocabulary Size Test (VST) described here Multiple-choice format:


one click per question to choose the correct meaning.
Vocabulary Levels Test (VLT) —described here Matching can be
adapted as clickable pairs or multiple-choice equivalents; widely
implemented as a “click the right definition” test.

Lexical Decision Task —


Click “Word” or “Nonword” for each item; standard psychological
experiment setup.

Computerized Adaptive (IRT-based) Vocabulary Tests —


Essentially the same as VST or Yes/No formats, but items are selected
automatically based on previous clicks.

More detailed explanation of each test

Test 1. The Nation & Beglar Vocabulary Size Test (VST) is one of the
most widely used methods for estimating how many English words a
learner knows.
Step-by-step description:
2. Word Sampling. The test designers take a very large list of the
most frequent English word families (for example, 14,000
word families from the British National Corpus).
From each 1,000-word frequency band, they randomly select a small
number of words (usually 10 per band).
3. Test Construction. Each test item is a multiple-choice
question with:
a target word (e.g., cup),
a short context sentence (to make meaning clearer),
four definitions, only one of which is correct.
Example:
cup — “He drank tea from a cup.”
(a) small open container for drinking
(b) kind of tree
(c) piece of clothing
(d) part of the body
4. Testing
The learner reads each item and chooses the meaning that matches
the word.
No time pressure is used; the test checks receptive knowledge
(knowing the meaning when you see the word).
5. Scoring. The test usually has 140 items (10 items × 14
frequency bands).
The learner’s correct answers are counted within each frequency
band.
6. Estimation. The proportion of correct answers per band is
used to extrapolate total vocabulary size.
Example:
If the learner scores 80% correct up to the 8,000-word band →
estimated vocabulary size ≈ 0.8 × 8,000 = 6,400 word families.

Test 2. Vocabulary Levels Test (VLT)


3. Word Sampling. Words are grouped by frequency levels (for
example, the 1,000, 2,000, 3,000, 5,000, and 10,000 most
frequent word families). From each level, researchers choose
a small representative sample of words — typically 18 words
per level.
4. Test Construction
Each test section contains six items made from three definitions and
six words.
The task is matching — the learner links each definition to the correct
word from the list.
Example (simplified):
Words Definitions
7. business a. part of your arm
8. clock b. something for telling time
9. horse c. animal people ride on
10. shoulder
11. ship
12. watch
Answer: 4–a, 6–b, 3–c
6. Testing. The learner completes one section per frequency
level. Each correct match earns one point. The test checks
receptive meaning knowledge — whether the learner
recognizes the meanings of common words at each
frequency level.
7. Scoring. Each level is scored separately (out of 18). A typical
“mastery” cutoff is around 15/18 correct — if the learner
reaches that, they are considered to “know” that frequency
level.
8. Estimation. The highest level at which mastery is achieved
shows the learner’s estimated vocabulary range.
Example: If a learner reaches mastery at the 3,000-word level, their
receptive vocabulary size is estimated around 3,000 word families.
7. Result. The result gives a banded estimate of vocabulary size
(e.g., 1k, 2k, 3k level known).
Researchers can use it to profile learners by how many frequency
bands of English vocabulary they know.
Test 3. Word Associates Test (WAT) — sometimes also called the
Word Associates Format (WAF).
3. Word Selection. The test includes a set of stimulus (target)
words, usually of medium to high frequency (e.g., beautiful,
judge, bank). Each target word is chosen so it has multiple
meaning or collocation associations (e.g., synonyms, related
adjectives, typical contexts).
4. Test Construction. For each target word, there are 8 options
arranged in two groups of four: Four are meaning associates
(synonyms or close semantic links). Four are collocational
associates (words that typically occur with the target).
The test taker must select all the words that are related to the target
word — usually 4 correct out of 8 options.
Example:
Target word: bank
Choose all the words related to “bank”:
| a. money | b. river | c. sleep | d. grass | e. finance | f. cloud | g.
loan | h. jump |
Correct answers: a, b, e, g
4. Testing. The learner reads each target and marks every
option they believe is related.
There’s no time limit — it measures depth of vocabulary knowledge,
but total correct responses can also reflect vocabulary size.
7. Scoring. Each correct association gets one point. Wrong
choices (selecting unrelated words or missing correct ones)
reduce the score. The total number of correct associations is
summed across all items.
8. Estimation. Researchers calculate an overall proportion of
correct associations. The resulting score indicates how many
words and semantic connections the learner likely knows,
which can be converted into a vocabulary size estimate (or
used to compare learners).
9. Result. The output is a single score showing how rich and
extensive the learner’s vocabulary knowledge is, based on
their ability to recognize both meaning and collocation links.
Test 4. Item Response Theory (IRT)–based vocabulary size test
3. Item Pool Creation. Researchers prepare a large bank of
vocabulary questions — often hundreds or thousands —
drawn from various frequency levels (e.g., 1,000–14,000
word families). Each item might be a short multiple-choice
question, a Yes/No word recognition item, or a translation
recognition question.
4. Item Calibration
Each item is first tested on a large group of learners.
Statistical modeling (IRT) estimates how difficult each item is and
how well it discriminates between learners with different vocabulary
sizes.

After calibration, each item has parameters like:


difficulty (how rare or hard the word is),
discrimination (how sharply it separates stronger and weaker
learners),
guessing (chance of answering correctly by luck).

4. Testing
The learner is shown one question at a time. After each answer, the
system updates its estimate of the learner’s vocabulary ability level.
The next question is chosen adaptively — slightly easier or harder
depending on the previous response. (This is called Computerized
Adaptive Testing, or CAT.)

6. Scoring. Instead of just counting correct answers, the model


estimates a latent ability score (θ) for each learner — a
continuous value representing their overall vocabulary
knowledge.
7. Estimation of Vocabulary Size. The θ score is mapped to an
approximate vocabulary size using a conversion table or
regression model derived during calibration.
For example, θ = 0 might correspond to knowing 5,000 word
families. θ = 1 might correspond to 8,000 word families, and
so on.

7. Result The output is a numerical vocabulary size estimate


(e.g., “Learner knows about 7,800 word families”). Because
the test adapts to each person, it’s shorter and more precise
than fixed-form tests.

Student Requirements
Essential technical skills 1. Experience in computational linguistics and NLP
2. Experience in building deep learning models
3. A solid understanding of IELTS Speaking assessment criteria
and language proficiency frameworks (CEFR, IELTS
descriptors)
4. Strong Python background (PyTorch, TensorFlow, scikit-
learn)

Desirable technical skills


Essential general skills 3. Good organisational skills

4. Independence and resilience


Desirable general skills
Essential experience
Desirable experience
Other comments

Organisation Details
Main contact name Dr Anton Nikolaev
Main contact position Data Scientist
Applications email To whom should we send the students’ applications?
Applications email cc Is there anyone else who should be copied on the application
emails?
Other comments Is there anything else about your organisation of which you would
like to make us aware?
PRJ55-G-WSD - Word Sense Disambiguation / Linking Module - Glite
Project for DS AI

Project Description
Project name Word Sense Disambiguation / Linking Module
Stipend offered £1,500
Host Organisation Glite
Project description About Glite:
We're developing a language learning app and website called Glite.
It's focused on intermediate and advanced learners, and only on the
English language. The company is based in London and is currently in
software product development before product launch.

As part of this project, we are building a comprehensive English


dictionary covering all modern words and multiword expressions,
along with their senses. The dictionary will serve both as a dataset
for our language models and as an online and mobile app product
accessible to users.

The dataset will include a list of concepts (one sense of a word or


MWE) including concept ID, the headword and sense
description/definition. The website version will contain additional
information such as pronunciation, usage examples (from movies,
books, and songs), regional spelling variants, tags (e.g. slang, archaic,
etc.), images or videos, synonyms, antonyms, and grammatical
details (e.g. verb tenses). The first version of the dictionary has been
developed based on Open English WordNet but with many senses
merged to reduce granularity, enriched with our proprietary
datasets of frequent English words and English idioms and other
multiword expressions not included in WordNet.

Identifying New Concepts


New concepts are identified across several corpora (e.g. COCA,
YouTube video transcripts, movie subtitles, etc.) through:

1. A span identification module trained to detect spans, including


MWEs, and
2. A span embedding or word sense linking module (WSL) that
positions these spans in semantic space.

The decision to include or exclude new spans is made based on their


semantic distance from existing dictionary entries in the embedding
space.
Our baseline WSL module uses LLM and can achieve accuracy > 90%
but it costs over USD 15 per 1 million words. This is not feasible and
we are looking for a solution with similar accuracy but less costly.

Project objectives
To develop a WSD (WSL) module that performs word sense linking
(disambiguation) that can be used on large (over 1 billion words)
corpora. The solution should not be based on LLM prompting alone.

Specific tasks
Task 1. Develop and present a detailed technical proposal for a cost-
efficient Word Sense Linking (WSL) module capable of processing
corpora exceeding 1 billion words. The solution may leverage a large
automatically generated dataset produced by our current (LLM-
based) system. The proposal should include the model architecture,
required training data volume,
expected performance, computational cost estimates, and an outline
of how the system will scale to large-corpus inference.

Task 2. Project planning


Provide a detailed project plan with:
a. Estimated time for each development stage
b. Required resources
c. Key deliverables and milestones

Task 3. Solution implementation. Build and deliver the full


WSD/WSL module, from data preparation to training and inference
deployment. Implementation must include the complete training
pipeline, evaluation scripts, and an optimised inference engine
capable of returning the best-matching dictionary concept ID (or “no
match”) for each span in context. The system must meet accuracy
and cost requirements and be fully containerised for production use
(training and inference Docker images).
Input: sentence with spans (can be several) and list of possible
concepts from the dictionary

Output: Our dictionary ID of the best match or no match if there isn’t


any

Datasets:
Our own dataset generated by our current solution.

Project implementation details


Meetings: Two sync meetings per week
Each sync: summary of work done + plan for next steps
IP and Legal:
All intellectual property belongs to Glite
NDA required before project start
Ownership transfer agreement at completion
Publication rights to be discussed separately. Glite is open to
collaborative
publications and/or joint participation in conferences

Metrics of success:
Accuracy (all words except top 1000 most frequent words) 90% on
provided dataset.
Price - < $2 per 1,000,000 words

Student deliverables Deliverables


1. Codebase for the WSD model
2. The dictionary will be provided in json format (headword,
definitions, examples, part of speech, surface forms etc.). If
dictionary is converted in own format the developer needs to
provide tool to do that
3. All dataset created or gathered during this project
4. Model, pytorch model, python code to run it
5. Inference docker container to run the model
6. Command line utility to run this in docker container
7. Code and cocker container for training the model separately
8. All trained weights of models created during this project
9. All data of tries and experiments registered on some
experimentation platform (like weights and biases
[Link]
10. Project report detailing methodology, datasets, model design,
and test results. Documentation and presentation
11. Signed NDA and IP ownership transfer

Student’s line manager Dr Anton Nikolaev

Project location(s) Remote


Other comments

Student Requirements
Essential technical skills 1. Experience in computational linguistics and NLP
2. Experience in building deep learning models
3. Experience in Word Sense Disambiguation
4. Strong Python background (PyTorch, TensorFlow, scikit-
learn)

Desirable technical skills


Essential general skills 5. Good organisational skills

6. Independence and resilience


Desirable general skills
Essential experience
Desirable experience
Other comments
Organisation Details
Main contact name Dr Anton Nikolaev
Main contact position Data Scientist
Applications email
Applications email cc
Other comments
PRJ56-EA-C - Cognitive and educational benefits of iron and folic acid
supplementation: evidence synthesis or predictive modeling -
Evidence Action
Project for HDS

Lancaster University Health Data Science Project Description

Please give details of the project in broad terms -- this information is for guidance and will be
replaced by a detailed specification later in the process.

Cognitive and educational benefits of iron and folic acid


Project name
supplementation: evidence synthesis or predictive modeling

Stipend offered

Host Organisation Evidence Action

The aim of this research project is to quantify the cognitive and


educational benefits of iron and folic acid (IFA) supplementation
for children, adolescents, and pregnant women (focusing on child
outcomes). Results will inform cost-effectiveness models and
program design for Evidence Action's IFA supplementation
programs across multiple countries.

The student will choose ONE of two complementary approaches


based on their interests and skillset:
Project description
Predictive Mathematical Model
Develop a mathematical model to estimate cognitive and
educational gains from IFA supplementation under different
program scenarios. The model will integrate existing evidence on
effect sizes with contextual factors such as baseline anemia
prevalence, adherence patterns, population demographics, and
program implementation characteristics. The model will be
designed for integration into Evidence Action's cost-effectiveness
frameworks and allow scenario testing for program optimization.
The student will work closely with Evidence Action's Associate
Director, Research, receiving guidance on methodology, analysis
techniques, and practical application to programmatic contexts.
Access to academic databases, statistical software, relevant
literature, and internal program data will be provided as needed.

A predictive model with documentation, code, user guide, and


technical report demonstrating model validation and scenario
Student deliverables analyses

A brief synthesizing findings for programmatic decision-making

Student's line manager Mark Minnery, Associate Director, Research

Remote work arrangement with regular virtual meetings. Student


Project location(s)
may work from Lancaster University or home location.

This project offers the opportunity to contribute to evidence -


based global health programming affecting millions of children
and women. The student will gain experience in either systematic
Other comments
review methodology and meta-analysis OR mathematical
modeling for health program applications, with direct real-world
impact.

Student Requirements

Statistical software (R, Stata, Python, or similar); statistical


Essential technical skills analysis and data visualization; literature review and synthesis
skills

Mathematical modeling; simulation methods; sensitivity analysis;


Desirable technical skills
model validation techniques
Critical appraisal of academic literature; clear technical writing;
Essential general skills project management; ability to work independently with regular
check-ins

Understanding of evidence synthesis frameworks; experience


Desirable general skills writing for both technical and non-technical audiences; familiarity
with global health or public health contexts

Prior coursework or project work involving either literature


Essential experience reviews and statistical synthesis OR mathematical/statistical
modeling; experience reading and interpreting academic papers

Development of decision models; health economics modeling;


Desirable experience
experience with uncertainty and sensitivity analysis

Students should indicate their preferred option (1 or 2) when


applying, though final selection will be made collaboratively
Other comments based on student strengths and project needs. Strong attention
to detail and systematic thinking will be critical for success in
either pathway.

Organisation Details

Main contact name Mark Minnery

Main contact position Associate Director, Research

Applications email [Link]@[Link]

Applications email cc

Other comments Evidence Action is an evidence-based global health organization


implementing cost-effective interventions at scale. Our IFA
supplementation programs reach millions of children and women
across multiple countries in Africa and Asia.
PRJ57-UCR-NLP - Natural Language Processing Projects - UCREL
Research Centre
Project for DS AI

Project Description
Project name Natural Language Processing Projects
Stipend offered No, the projects will be Lancaster campus based
Host Organisation UCREL Research Centre
Project description We are the UCREL Research Centre, based in InfoLab21, we are a
diverse community of language enthusiasts and researchers. With
35+ members from various disciplines including computer science,
English, linguistics, history, & health, we explore Natural Language
Processing (NLP), a hot topic area of artificial intelligence, which has
recently changed the world with the invention of large language
models. We’ve been researching 20+ languages for many decades.
Together, our academic team has published 1600+ papers with over
23,000 citations. If you’re enthusiastic about language and seeing
what language models can do, then join us for your MSc placement
in 2026! See [Link] for more details. You will
have chance to engage in our weekly group meetings (which often
involve cake!) to learn about other cutting edge research from us and
invited speakers, and to help test and improve our software.

We always have multiple projects running, some are externally


funded, some are linked with PhD research. We also run summer
schools, other training events and engage in regular dissemination of
our research on our social media channels, so there is chance to
define your own placement in discussion with us.

For example, in the EU funded 4D Picture project


([Link] we are working with teams in multiple
countries around Europe to help empower patients to understand
their cancer treatment pathways by large scale analysis of patient
forums and analysing the metaphors they use (e.g. fight, battle or
journey). You can help extend our NLP tools to multiple languages in
the project (Danish, Dutch, English, Spanish) and beyond the project
(e.g. Arabic, Igbo, Urdu).

In the Spatial Narratives project


([Link] we are analysing Lake
District travel writing and guides and Holocaust testimonies to apply
NLP methods to extract and map non-mappable stories and
experiences from large existing (English) datasets that we’ve
collected, and are providing training resources for non technical
users to explore our methods using Python Notebooks and other
tools. We’d like to apply the same methods to new datasets (e.g.
collect new Holocaust testimonies from YouTube, convert the speech
to text, or the Lake District Holocaust Project, or testimonies in other
languages), and you’ll be engaging with our research team in the UK
(Lancaster, Leeds, Bristol) and the US (Stanford, Indiana).

In the recent Davy Notebooks project


([Link] we were working in a team
with colleagues in English Literature and the Lancaster Digital
Collections (LDC) group in the university library, so you may also be
able to be placed there as part of the LDC team. We are also part of a
large European network of researchers working on creating very
large datasets of newspapers ([Link] and
European regional and national parliaments
([Link] We would like to extend this
work to collect data from the British Library, plus Scottish and Welsh
parliaments.

Student deliverables See above


Student’s line manager Professor Paul Rayson
Project location(s) InfoLab21 at Lancaster University
Other comments Please ask for more details on any of the above projects

Student Requirements
Essential technical skills We make software to collect large bodies of text (called corpora),
then manually and automatically annotate them (labelling with
linguistic categories), and analysing millions or billions of words, for
example to extract emotions and sentiment summaries, or map the
results by extracting placenames. The software frameworks are
mostly in Python e.g. [Link] and
[Link] The research usually involves working in teams in
Lancaster or with external partners, and you will be involved in
creating software, documentation and social media posts to
disseminate our work, and if you’re interested, to potentially get
involved in papers publishing your contributions.
Desirable technical skills As above
Essential general skills You should be keen and enthusiastic to learn about AI, NLP and
machine/deep learning methods and tools. Most NLP code is written
in Python with a myriad of libraries and language models available
e.g. via HuggingFace that you can use in the research, so a high
degree of Python coding ability is desirable. Many of the projects we
have running are based on real world challenges so an interest in
making a difference beyond the university is also important.
Desirable general skills As above
Essential experience Past experience of NLP/AI/ML methods e.g. from the SCC.453 and
other similar modules is essential
Desirable experience As above
Other comments As above.
Organisation Details
Main contact name Professor Paul Rayson
Main contact position Director of the UCREL Research Centre
Applications email [Link]@[Link]
Applications email cc No
Other comments A hybrid arrangement is possible, but we will aim to provide working
space somewhere in InfoLab21 in a hot desking environment so that
you can engage regularly with the group members and our
stakeholders on the projects, some of which are external to
Lancaster.
PRJ59-LUS-S - Student number forecasting - Lancaster University
Strategic Planning Unit
Project for DS AI

Project Description
Project name Student number forecasting
Stipend offered N/A
Host Organisation Lancaster University Strategic Planning Unit
Project description To enable Lancaster University to arrange the resources necessary to
deliver services to incoming students it is essential that the
University has a clear view of the number of students who will join
each year.

The University has data from many previous years, which show
numbers of applications to its programmes, the characteristics of the
applicants, and the rates at which applications ‘converted’ to
students. We would like to use this data to develop a better
understanding of conversion rates. Additionally, we would like to
build and test predictive models using recruitment data to predict
student registrations across our programmes.

Student deliverables • A robust, well-documented set of models that will allow the
University to make accurate predictions of student numbers
• Effective visualisations demonstrating the insights derived from
the models
• A report documenting the outputs generated during the project
Student’s line manager Mark Yoxon
Project location(s) LU Campus and Remote

Other comments

Student Requirements
Essential technical skills • Competent in R or Python
• Data Analysis and modelling
• Data Visualisation
Desirable technical skills •
Essential general skills • Ability to organise and manage own workload
• Clear written communication
• Attention to detail
• Problem solving mindset
• Some project management awareness
Desirable general skills • Confidence working independently while seeking guidance when
needed
• Ability to present findings to technical and non‑technical
audiences
Essential experience
Desirable experience
Other comments

Organisation Details
Main contact name Mark Yoxon
Main contact position Head of Strategic Planning
Applications email
Applications email cc -
Other comments
PRJ60-LUS-S - Student feedback - topic modelling - Lancaster
University Strategic Planning Unit
Project for DS AI

Project Description
Project name Student feedback - topic modelling
Stipend offered N/A
Host Organisation Lancaster University Strategic Planning Unit
Project description To enable Lancaster University to improve its delivery of services to
students we wish to have a clear understanding of the feedback that
we receive. We would like a data science student to use their skills to
analyse the unstructured feedback that we receive to determine key
topics and themes for further analysis.

It is envisaged that this will entail a combination of NLP techniques


with ‘traditional’ statistical and machine learning approaches.

Student deliverables • A report that explains the techniques used for the creation of
accurate, useful reports that explain the themes and topics
prevalent in the feedback that LU receives from students.
• A report documenting the outputs generated during the project
Student’s line manager Mark Yoxon
Project location(s) LU Campus and Remote

Other comments

Student Requirements
Essential technical skills • Competent in R or Python
• Data Analysis and modelling
• Data Visualisation
Desirable technical skills
Essential general skills • Ability to organise and manage own workload
• Clear written communication
• Attention to detail
• Problem solving mindset
• Some project management awareness
Desirable general skills • Confidence working independently while seeking guidance when
needed
• Ability to present findings to technical and non‑technical
audiences
Essential experience
Desirable experience
Other comments

Organisation Details
Main contact name Mark Yoxon
Main contact position Head of Strategic Planning
Applications email
Applications email cc
Other comments
PRJ61-DHL-AIT - AI as a Tool for Equity: Computational Analysis of
Gendered Language Patterns in Digital Media - Digital Heard Ltd
Project for DS AI

Project Description
Project name AI as a Tool for Equity: Computational Analysis of Gendered
Language Patterns in Digital Media

Stipend offered No stipend will be given, but reasonable expenses will be paid, and a
Fraser House Hub membership can be arranged if preferred and
appropriate.

Host Organisation Digital Heard Ltd

Project description This project leverages six months of data collected through
[Link] to produce a series of HCI publications
examining how AI can function as a tool to promote diversity rather
than perpetuate systems of inequality. The student will work with a
substantial existing dataset of news headlines and digital content
that has been systematically analyzed for gendered language
patterns using metrics including Agency, Independence, Visibility,
and Responsibility Clarity. The work will focus on identifying patterns
and trends across this corpus, with the aim of producing actionable
insights for both platform governance and AI tool design. The
student will work under the supervision of Dr Alice Ashcroft (Digital
Heard Ltd), with potential collaboration from colleagues in the
School of Computing and Communications at Lancaster University
and researchers at other academic institutions. This project sits at
the intersection of Human-Computer Interaction, computational
linguistics, feminist HCI, and AI ethics. Technologies will include
Python for data analysis, statistical analysis tools in R, and potentially
machine learning frameworks for pattern detection and predictive
modelling.

Student deliverables The student will be expected to deliver three core outputs. First, a
comprehensive analytical report documenting patterns and trends
identified in the [Link] dataset, including
statistical analyses and visualizations that reveal how gendered
language operates across different contexts, platforms, and time
periods. Second, R and/or Python scripts that generate publication-
ready visualizations (graphs, tables, and figures) suitable for
inclusion in HCI conference papers and journal articles, with full
documentation for reproducibility. Third, the student will contribute
to drafting sections of one or more HCI research papers arising from
the project, with second authorship credit awarded in accordance
with their contribution and standard academic conventions. All code,
documentation, and analyses should be version-controlled and
handed over in a state suitable for ongoing use by the research team
and potential publication as supplementary materials.

Student’s line manager The student will report directly to Dr Alice Ashcroft (Director, Digital
Heard Ltd), who will serve as their primary supervisor and line
manager throughout the placement. Additional support and
mentorship will be available from the wider Digital Heard research
team (currently three researchers) and Joshua Bailey (AI consultant
and founder of Sqwosh), who will provide technical guidance on
computational and AI-related aspects of the project.

Project location(s) The work can be performed remotely or from Fraser House Hub, a
coworking space for freelancers and small businesses in Lancaster. A
Fraser House Hub membership can be arranged for the student if
preferred, which would provide excellent networking opportunities
alongside a professional working environment. The arrangement will
be flexible and agreed with the student based on their preferences
and circumstances.

Other comments Digital Heard operates with a flexible, asynchronous working culture
where we prioritise outputs over hours. We expect our staff and
interns to work independently and take ownership of their tasks,
with support readily available when needed. There will always be a
team member on hand to help between 10am and 4pm, but working
hours outside of this are entirely flexible and can be arranged around
the student's own schedule and commitments. This placement
would suit a self-motivated student who thrives with autonomy and
is comfortable managing their own workload. The project offers a
unique opportunity to work with real-world data that has direct
implications for AI ethics, platform design, and digital equity, with
clear pathways to HCI publication at top-tier venues.

Student Requirements
Essential technical skills Proficient in Python, including experience with data analysis libraries.
Competent in R, with the ability to produce data visualizations using
ggplot2 or similar packages. Understanding of statistical analysis
methods including descriptive statistics, correlation analysis, and
basic inferential statistics. Familiarity with data cleaning and
preprocessing techniques for large datasets.
Desirable technical skills Experience with machine learning for pattern detection and
classification tasks. Knowledge of Natural Language Processing and
text analysis tools. Experience with version control using Git/GitHub.
Familiarity with data visualization tools.

Essential general skills Ability to organize your own work and manage time effectively with
minimal supervision. Strong analytical thinking skills with the ability
to identify meaningful patterns in complex datasets. Excellent
written communication skills, particularly the ability to write clearly
for academic and HCI audiences. Attention to detail when handling
data and conducting analyses. Critical thinking about how technology
can either reinforce or challenge social inequalities.

Desirable general skills A basic understanding of HCI research processes and publication
conventions, particularly for CHI, CSCW, or similar venues. Ability to
bring together findings from quantitative analysis and contribute to
written outputs that connect technical findings to broader social
implications. Understanding of feminist theory, gender studies, or
critical approaches to technology. Comfort working asynchronously
and communicating via digital tools (e.g., Slack, email,
LaTeX/Overleaf, shared documents).

Essential experience Experience working on a research project or dissertation involving


data collection and analysis, whether through coursework or
independent study.

Desirable experience Experience working in a team or collaborative research environment.


Familiarity with Human-Computer Interaction, digital discourse
analysis, or gender studies literature. Previous experience
contributing to academic writing or publications.

Other comments This project involves analysing content that documents patterns of
gendered language, including examples of bias, stereotyping, and
potentially harmful discourse. The student should be comfortable
working with such content and aware of the context in which it was
collected. We are happy to discuss strategies for managing this work
and will provide appropriate support throughout.

This placement offers an excellent opportunity for a student


interested in the intersection of AI, social justice, and HCI research.
The work will contribute to an emerging body of scholarship that
demonstrates how computational tools can be designed and
deployed to challenge, rather than reinforce, systems of inequality.
Students will gain experience in rigorous data analysis, academic
writing for HCI venues, and translating technical findings into socially
meaningful insights. Second authorship credit will be given on all
papers where the student makes appropriate contributions,
providing valuable career development for students interested in
pursuing research or industry roles focused on responsible AI and
technology ethics.

Organisation Details
Main contact name Dr Alice Ashcroft

Main contact position Founder and Lead Researcher

Applications email admin+applications@[Link]

Applications email cc

Other comments All applications should include a CV, a link to a publication/written


piece of coursework and a paragraph on what interests them about
the project.
PRJ62-FIS-AIF - AI-FIS360 - FIS360 LTD
Project for AI

Lancaster University Industry Project Description


Project name AI-FIS360
Stipend offered £3000
Host Organisation FIS360 LTD
Project description What are the aims of the project?
What type of project is this? (e.g. implementation, research,
evaluation, enhancement, pilot)
How will work be organised?
Who will contribute?
What business areas will be involved?
Which technologies will be used?

FIS360 is a UK-based innovation consultancy headquartered in


Penrith, Cumbria. As an SME working with some of the most
forward-thinking organisations in the country, we design and
deliver innovation programmes that tackle complex challenges
within the nuclear industry and other highly regulated sectors.

As we enter our next phase of growth, we are seeking to explore


how Artificial Intelligence (AI) can be strategically applied within
our business to enhance decision-making, improve efficiency, and
unlock new opportunities.

This 12-week placement will focus on identifying and evaluating


high-impact AI use cases across our existing data and processes.
Throughout the project, you will work closely with members of the
FIS360 management team.

Areas of interest include:


• Business strategy and planning
• Business development and client engagement
• Programme design and delivery
• Quality management and compliance

Our core digital environment includes Microsoft 365 and Zoho One,
and we are particularly interested in AI tools that integrate
effectively within these platforms.

In this project you will:


1. Rapidly develop an understanding of FIS360’s processes,
data, and operating model
2. Research and evaluate AI tools with strong potential to
deliver measurable business benefit
3. Assess opportunities based on cost, value creation, risk,
ease of implementation, and environmental considerations
4. Recommend a prioritised shortlist of high-impact AI
applications
5. Design and deliver up to three small-scale pilot trials within
selected areas of the business
6. Produce case study examples demonstrating practical
application, outputs, benefits, and lessons learned

The outcome of this placement will be a clear, evidence-based


roadmap for AI adoption within FIS360, grounded in practical
experimentation and aligned with our strategic ambitions.
Student deliverables 1. A final written report summarising the research carried out, the
AI trials and their outcomes, and an adoption roadmap.
2. 3 case study examples of AI trialled on FIS360 data
3. A summary presentation delivered to the FIS360 team
Student’s line manager TBD
Project location(s) The majority of the project can be carried out remotely. A secure
FIS360 laptop will be provided for access to necessary company
information. A minimum of kick-off meeting, mid-point review, and
final report meeting would take place in person at our Penrith
office. Additional face to face working days within our Penrith
office could also be arranged over the course of the project if
deemed mutually beneficial.
Other comments n/a

Student Requirements
Essential technical skills Please indicate the required skill (or relevant technology) and the
level of skill required
Strength in researching and evaluating AI tools for measurable
business benefit
Ability to design and deliver small scale pilot / trial / proof of concept
Competence in analysing organisational processes and data to
identify high impact use cases
Desirable technical skills As above
Familiarity with AI tools that integrate with Microsoft 365 and Zoho
One.
Comfort comparing tools across factors such as cost, value creation,
risk, ease of implementation and environmental impact.
Essential general skills Please indicate the required skill e.g. Basic understanding of project
management
Rapid learning to understand the companys processes, data and
operating model
Structured evaluation and prioritisation to produce a clear, evidence
based shortlist and roadmap
Strong written communication to product case studies (outputs,
benefits, lessons learned)
Collaboration and stakeholder engagement with the management
team
Desirable general skills As above
Strategic thinking across business strategy, client engagement,
programme delivery and quality compliance contexts.
Essential experience Please indicate the required experience e.g. working in teams
Experience running evaluative projects (eg. Identifying / assessing
use cases and recommending a prioritised shortlist.
Experience planning and running experiments (pilots / trials) and
documenting outcomes
Desirable experience As above
Exposure to AI adoption within Microsoft 365 and Zoho One
environments
Other comments Please indicate any attributes that may be beneficial to the student
within your organisation or any other factors to be considered.
Openness to work through challenges, willingness to ask questions,
ability to self direct work with guidance.
Organisation Details
Main contact name Dr. Deborah Bowering
Main contact position Chief Operating Officer
Applications email [Link]@[Link]
Applications email cc Frank@[Link]
Other comments Please note Deborah is one vacation until Thursday 19th Feb. so
please direct emails to me in this period.

If you have any questions or to return this form please mail to Chris Lowerson at
[Link]@[Link]
PRJ63-C-A - Auto-redaction of video content - Collaboraite
Project for DS AI

Project Description
Project name Auto-redaction of video content
Stipend offered £3000
Host Organisation Collaboraite
Project description Collaboraite is at the forefront of developing cutting-edge tools for
secure government and law enforcement. Currently, our [Link]®
platform enables users to detect specific objects within videos (as
well as various translation and transcription options). Our law
enforcement users often need to collect this content for intelligence
or evidential purposes, but there may be persons or objects (e.g.
signage) within the video which they wish to ‘mask’.

Project Aim: The goal of this project is create a methodology for


masking or blurring specific objects, types of object, persons (or
possibly specific people) within video content.

Scope:
• Object Identification: Develop methodologies to identify the
required objects or areas within video content to be
removed.
• Masking techniques: Understand the methods of masking or
blurring the required objects or persons.
• End-to-end prototype: Create a prototype which will allow a
user to specify what is to be removed, within a single or
multiple videos, and then apply the required blurring or
masking.
• Security First: Ensure solutions operate without internet
access, maintaining the security critical for law enforcement
use.

Support and Collaboration: Collaboraite will provide:


• Sample Data and Resources: Access to datasets and existing
research to facilitate your work.
• Regular Interaction: Engage with our team for guidance and
feedback throughout the project.
• Presentation Opportunities: Share your findings with our
product and customer liaison teams, showcasing your
contributions.

The project will potentially enable us to add a new capability to our


[Link]® product, and therefore contribute to real-world
applications in law enforcement.
(Note: this project does not require the use of facial recognition
technology).
Student deliverables Working prototype and evaluation, based on the sample data we
provide.
Final presentation / report explaining the methodology, challenges
encountered, and how these can be addressed.
Student’s line manager Brian Bridge
Dr Scott Stainton
Project location(s) Mainly remotely at the University or other agreed locations, with
periodic (approx. every 2 weeks) visits to our offices in Liverpool.
The student is welcome to spend more time in our offices if
preferred.
Other comments n/a

Student Requirements
Essential technical skills Python Programming
Understanding of popular open source libraries e.g. numpy, pandas,
etc
Desirable technical skills Experience working with image/video data
Practical knowledge of Neural Networks
Essential general skills Project management
Time management
Desirable general skills As above
Essential experience Please indicate the required experience e.g. working in teams
Desirable experience Self-motivation & investigative skills
Ability to effectively communicate technical concepts, progress &
results to the rest of the team
Other comments While not a requirement for this project, most of our work is in the
secure government area, and all of our permanent employees are
subject to vetting, requiring at least 3 years’ residence in the UK.
Other comments Please indicate any attributes that may be beneficial to the student
within your organisation or any other factors to be considered.

Organisation Details
Main contact name Alan Crameri
Main contact position Research, Product & Delivery Director
Applications email [Link]@[Link]
Applications email cc [Link]@[Link]
Other comments Our website [Link] gives more details of who we
are.
PRJ64-C-S - Speaker recognition - Collaboraite
Project for DS AI

Project Description
Project name Speaker recognition
Stipend offered £3000
Host Organisation Collaboraite
Project description Collaboraite is at the forefront of developing cutting-edge tools for
secure government and law enforcement. Currently, our [Link]®
platform enables users to transcribe audio from videos and add
subtitles. Following a recent Lancaster University MSc project, we
are now including speaker diarisation to identify which speaker is
talking at each point in the video.

Project Aim: Following on from diarisation, this project is to develop


the methods we would need to identify (‘recognise’) the same
speaker across different video / audio files. For example, a user
might upload a large number of videos, and wish to highlight when a
specific person has spoken on any of these. Or a user might upload a
new video, and be alerted that a specific speaker matches videos
which have been processed previously.

Scope:
• Speaker Recognition: Develop methods to identification and
‘matching‘ of speakers across different audio or video
samples, given that there may be different sound quality or
background noise in each.
• Audio Fingerprinting: Investigate the feasibility of an ‘audio
fingerprint’, which we can use to classify different speakers
across a corpus of video / audio which has been processed
previously by our tool.
• Working Prototype: Develop a working prototype to
illustrate these techniques in practice.
• Security First: Ensure solutions operate without internet
access, maintaining the security critical for law enforcement
use.

Support and Collaboration: Collaboraite will provide:


• Sample Data and Resources: Access to datasets and existing
research to facilitate your work.
• Regular Interaction: Engage with our team for guidance and
feedback throughout the project.
• Presentation Opportunities: Share your findings with our
product and customer liaison teams, showcasing your
contributions.
The project will not only enhance [Link]® but also contribute to
real-world applications in law enforcement, making a tangible
difference by improving how information is extracted and utilised.

Student deliverables Working prototype and evaluation, based on the sample data we
provide.
Final presentation / report explaining the methodology, challenges
encountered, and how these can be addressed.
Student’s line manager Dan Lund
Dr Scott Stainton
Project location(s) Mainly remotely at the University or other agreed locations, with
periodic (approx. every 2 weeks) visits to our offices in Liverpool.
The student is welcome to spend more time in our offices if
preferred.
Other comments n/a

Student Requirements
Essential technical skills Python Programming
Understanding of popular open source libraries e.g. numpy, pandas,
etc
Desirable technical skills Experience working with audio data
Practical knowledge of neural networks
Essential general skills Project management
Time management
Desirable general skills As above
Essential experience Self-motivation & investigative skills
Ability to effectively communicate technical concepts, progress &
results to the rest of the team
Desirable experience As above
Other comments While not a requirement for this project, most of our work is in the
secure government area, and all of our permanent employees are
subject to vetting, requiring at least 3 years’ residence in the UK.

Organisation Details
Main contact name Alan Crameri
Main contact position Research, Product & Delivery Director
Applications email [Link]@[Link]
Applications email cc [Link]@[Link]
Other comments Our website [Link] gives more details of who we
are.
PRJ65-CR-SCB - Scene-Consistency Benchmarking for Industrial 6DoF
Pose Estimation - Cambrian Robotics
Project for DS AI

Project Description
Project name
Scene-Consistency Benchmarking for Industrial
6DoF Pose Estimation
Stipend offered TBC
Host Organisation Cambrian Robotics
Project description C ambrian Robotics is a startup in the AI Computer Vision domain.
We develop a 3D AI camera system for use with industrial robots.
We have systems deployed globally in a range of manufacturing
scenarios.

A key part of the product is the AI modelling, where train specific


models, on demand, for a given part or parts that the client needs to
recognise. In order to train the models we use simulated (synthetic)
data generated from CAD models of the parts. Over time we
regularly develop and upgrade the AI training components and we
need to validate these against a wide range of benchmarks to ensure
that the new models are better and do not break any of the existing
functionality.

Project Overview
The goal of this project is to develop a "Scene-Consistency
Benchmark" to evaluate the reliability of AI-driven 6DoF (6 Degrees
of Freedom) pose estimation models. Current benchmarks often
focus on static accuracy; however, industrial applications require
models that provide consistent predictions regardless of the
camera's viewing angle.

Building upon the methodology established in recent research


(Towards Co-Evaluation of Cameras, HDR, and Algorithms for
Industrial-Grade 6DoF Pose Estimation), this project will invert the
typical testing loop. Instead of moving an object via a robot arm, we
will mount a camera to a high-accuracy robot arm and move it
through a trajectory around a static scene of known industrial parts.

Project Objectives
The student will develop a pipeline to analyse how "stable" a
model’s predictions are in the real world. Key tasks include:

1. Coordinate Transformation: Using robot odometry and


camera calibration to transform 6DoF predictions from the
camera frame into a unified world coordinate system.
2. Consistency Analysis: Quantifying prediction variance in
real-world metrics (millimeters and degrees).
3. Error Profiling: Identifying "blind spots" or specific viewing
angles/lighting conditions where model confidence or
accuracy degrades.
4. A/B Model Testing: Comparing different architectures to
determine which offers the highest stability for industrial
deployment.

Academic Rigor & Data Science Components


This project is designed to be "sufficiently demanding" for an MSc
student by requiring a blend of:

● Geometric Deep Learning: Understanding 3D spatial relationships


and SE(3) transformations.
● Statistical Analytics: Performing clustering on projected poses and
calculating 1st/2nd order statistics to define "consistency."
● Software Engineering: Integrating with existing helper code and
handling large-scale real and simulated datasets.

Relevance to Host & Industry


Improving quality control in manufacturing depends on robots
accurately "seeing" parts. If a model predicts a part's location
differently when the camera moves 5cm, the downstream robotic
task (e.g., grasping or assembly) will fail. This benchmark will provide
Cambrian with a standardized tool to vet AI models before they are
deployed in high-stakes environments.

Project Timeline (14 Weeks)


● Weeks 1–3: Literature review, induction, and familiarization with
the existing 6DoF codebase and simulated datasets.
● Weeks 4–7: Implementation of the transformation pipeline
(Camera → Robot → World) and initial validation on simulated data.
● Weeks 8–11: Comparative analysis of Model A vs. Model B;
identifying failure modes and edge cases in real-world datasets.
● Weeks 12–14: Finalizing the benchmarking framework, generating
performance reports, and drafting the dissertation/poster.

Support & Resources


The student will be provided with:
● Pre-collected real-world and simulated datasets.
● Helper code for coordinate transformations and camera
calibration.
● Mentorship from our engineering team to ensure the project
remains aligned with industrial standards.

We will provide all the relevant hardware and software that is


required
Student deliverables Detailed report covering the work carried out to meet the project
objectives described above.
Student’s line manager David Main, Head of Operations

Project location(s) The team works mostly remotely, with regular coordination
meetings in our London office. The applicant would be mostly
working remotely on this assignment, but would be hosted at the
office at least twice during the project.

Other comments

Student Requirements
Essential technical skills Knowledge of statistical analysis and ML techniques
Desirable technical skills Understanding of robotics and related physics
Essential general skills
Desirable general skills
Essential experience
Desirable experience
Other comments

Organisation Details
Main contact name David Main
Main contact position Head of Operations
Applications email Cambrian Robotics
Applications email cc
Other comments
PRJ66-ESL-BAA - Biosignal Analytics with AI - EngScience Ltd.
Project for DS AI

Lancaster University Data Science Project Description Form


Project Description
Project name Biosignal Analytics with AI
Stipend offered
Host Organisation EngScience Ltd. [Link]
Project description Project Overview (Ref EngSci 1vA)
Fusion of Ocular and Cardiac Signals for Cognitive Load Monitoring

In demanding tasks (e.g., competitive gaming, intensive training, or


high-focus work), real-time detection of mental fatigue or
"saturation" is challenging. This project explores whether combining
high-fidelity eye-tracking with heart rate variability (HRV) can yield
early indicators of declining concentration.

We seek a motivated student to build a synchronized data pipeline


and perform exploratory fusion analysis on anonymized physiological
signals.

Technical Challenge
EngScience provides a proprietary eye-tracking approach that uses
differential signal processing to compensate for head/body
movement, delivering clean, high-resolution gaze vectors despite
real-world noise.

The core task is to develop a fusion and analysis layer that aligns this
refined gaze data with concurrent HRV telemetry (inter-beat
intervals) and investigates potential correlations or predictive
patterns in autonomic and oculomotor responses during visual
pursuit tasks.

Data Provision & Anonymization


EngScience supplies fully anonymized datasets from controlled
internal trials (no PII). Data includes timestamped gaze coordinates,
marker positions, and cardiac inter-beat intervals, focused purely on
numerical telemetry for analytics and modelling.

Project Stages
Stage 1: Data Integration
Build a robust pipeline to ingest and synchronize the high-frequency
streams with sub-millisecond precision.
Stage 2: Signal Processing & Feature Extraction
Apply filtering (e.g., Kalman/low-pass) to reduce noise. Compute
relevant HRV and gaze-derived features, example Temporal Lag
Estimation.
Stage 3: Exploratory Analysis & Modelling
Conduct statistical and lightweight ML-based analysis (e.g.,
regression, time-series models) to identify if multimodal patterns can
anticipate shifts in performance and arousal.
Stage 4: Visualization Dashboard
Create an interactive tool to display synchronized signal overlays and
highlight key events/insights.

Potential Applications: Adaptive interfaces, cognitive training


feedback, performance monitoring.

Intellectual Property: Per EngScience's standard research terms and


applicable university policies: EngScience retains rights to underlying
data, proprietary methodologies, and any patentable inventions
arising from the project (including novel fusion techniques or
predictors). To preserve patent options, avoid public disclosure of
novel elements until IP review/filing.

Are you excited to explore the subtle interplay between eye


movements and heart rhythms? If you have interest in signal
processing, Python, time-series analysis, or bio-signal fusion, this is a
unique chance to contribute at the edge of human physiology and
predictive AI.

Student deliverables Student Outcomes & Impact


Deliverables: Modular Python pipeline, analysis report, and real-time
dashboard prototype.

Student’s line manager Dr Ather Sharif

Project location(s) Lancaster University (100% of the time)


Other comments Confidential Note: This proposal outlines a high-level research
opportunity. Any novel algorithms, features, or predictive insights
developed during the project remain subject to EngScience's
standard IP terms and potential patent protection. Do not disclose or
share enabling details publicly without prior approval.

Student Requirements
Essential technical skills
Desirable technical skills
Essential general skills
Desirable general skills
Essential experience
Desirable experience
Other comments
Organisation Details
Main contact name Dr Ather Sharif
Main contact position Technical Director
Applications email drathersharif@[Link]
Applications email cc N/A
Other comments I
PRJ67-ESL-AIO - AI for Offshore Wind Structural Integrity - EngScience
Ltd.
Project for DS AI

Lancaster University Data Science Project Description Form


Project Description
Project name AI for Offshore Wind Structural Integrity
Stipend offered
Host Organisation EngScience Ltd.
[Link]
Project description
AI for Offshore Wind Tower Integrity
In the aggressive offshore environment, wind turbine structures
face constant degradation that often goes undetected by
generalised monitoring systems. This project focuses on
developing a specialised, low-frequency foundational AI model to
monitor structural health by analysing natural frequency shifts,
the "fingerprint" of a turbine's integrity.

We seek a motivated Masters student to build a proof of concept


predictive framework capable of detecting resonance risks (such
as seabed scour or stiffness changes) before they lead to
catastrophic fatigue, providing the "auditable proof" required for
formal asset Life Extension Programs.
Technical Challenge

Offshore towers enter a critical state when their natural


frequency drifts and aligns with operational frequencies (such as
1P or 3P rotor frequencies). This resonance leads to exponential
fatigue. Current "black box" vendor models often fail to isolate the
subtle 0.3 Hz frequency drifts critical to structural health amidst
the "noise" of drivetrain and blade vibrations.

The core task is to move beyond general vibration analysis to


develop a Site-Specific Baseline AI. This involves isolating high-
fidelity structural dynamics from raw sensor streams and
correlating these shifts with environmental variables to forecast
the "Time-to-Critical-Resonance."
Data Provision

EngScience supplies simulated or real access to high-resolution,


pure data streams (JSON/MQTT) from offshore assets. All
datasets including accelerometer data.
Project Stages
Stage 1: Signal Processing & Data Cleaning

Implement advanced filtering (e.g., Band-pass filters) to isolate


the "Pure Data Stream" of low-frequency vibration fingerprints,
removing operational noise from the drivetrain.
Stage 2: Site-Specific Baseline Training

Establish a background baseline of the tower's unique natural


frequency under varying sea states to understand its "normal"
behaviour.
Stage 3: Anomaly Detection & Modelling

Utilise Deep Learning (LSTMs/Transformers) or advanced


Regression models to detect minute deviations (as small as 0.3
Hz) from the established baseline.
Stage 4: Forecasting & RUL Estimation

The outcome should pave the way for a predictive layer to


calculate Remaining Useful Life (RUL) and estimate when
resonance thresholds will be breached, enabling proactive
maintenance months in advance.
Potential Applications

Independent "AI Audit" frameworks for asset owners, regulatory


compliance, and formal Life Extension (LEX) sign-offs for
renewable energy infrastructure.
Intellectual Property

Per EngScience’s standard research terms and applicable


university policies: EngScience retains rights to underlying data,
proprietary methodologies, and any patentable inventions arising
from the project (including novel structural health predictors or
fusion techniques). To preserve patent options, avoid public
disclosure of novel elements until IP review/filing.

Are you ready to use AI to secure the future of renewable


energy?

If you have interest in Python (Pandas, SciPy), Time-series


forecasting, or Signal Processing (FFT/PSD), this project offers a
unique opportunity to apply high-level data science to critical
green energy infrastructure.
Student deliverables Student Outcomes & Impact
The student will deliver a proof of concept predictive model capable
of isolating 0.3Hz frequency drifts from noisy vibration data.
Student’s line manager Dr Ather Sharif

Project location(s) Lancaster University (100% of the time)


Other comments Confidential Note: This proposal outlines a high-level research
opportunity. Any novel algorithms, features, or predictive insights
developed during the project remain subject to EngScience's
standard IP terms and potential patent protection. Do not disclose or
share enabling details publicly without prior approval.

Student Requirements
Essential technical skills
Desirable technical skills
Essential general skills
Desirable general skills
Essential experience
Desirable experience
Other comments

Organisation Details
Main contact name Dr Ather Sharif
Main contact position Technical Director
Applications email drathersharif@[Link]
Applications email cc N/A
Other comments I
PRJ68-LUS-SCS - Student feedback – Content Summarisation -
Lancaster University Strategic Planning Unit
Project for AI

Project Description
Project name Student feedback – Content Summarisation
Stipend offered N/A
Host Organisation Lancaster University Strategic Planning Unit
Project description To enable Lancaster University to improve its delivery of services to
students we wish to have a clear understanding of the feedback that
we receive. We would like an AI student to use their skills to produce
a system that will generate an accurate summary of batches of
student comments (e.g., by academic department or comment
topic), providing a high-level summary of key themes, with links to
the actual data.

It is envisaged that a GenAI solution will be used.

Student deliverables • A report that explains the techniques used for the creation of an
intelligent system that will produce accurate, searchable
summaries of student feedback.
• A report documenting the outputs generated during the project
Student’s line manager Mark Yoxon
Project location(s) LU Campus and Remote

Other comments

Student Requirements
Essential technical skills • Python programming
• Awareness of GenAI technologies and their use
Desirable technical skills
Essential general skills • Ability to organise and manage own workload
• Clear written communication
• Attention to detail
• Problem solving mindset
• Some project management awareness
Desirable general skills • Confidence working independently while seeking guidance when
needed
• Ability to present findings to technical and non‑technical
audiences
Essential experience
Desirable experience
Other comments

Organisation Details
Main contact name Mark Yoxon
Main contact position Head of Strategic Planning
Applications email
Applications email cc
Other comments
PRJ69-LUS-SCI - Student feedback – Chatbot Interface - Lancaster
University Strategic Planning Unit
Project for AI

Project Description
Project name Student feedback – Chatbot Interface
Stipend offered N/A
Host Organisation Lancaster University Strategic Planning Unit
Project description To enable Lancaster University to improve its delivery of services to
students we wish to have a clear understanding of the feedback that
we receive. We would like an AI student to use their skills to produce
a chatbot interface to enable users to query and interact with a
database of student feedback. The chatbot’s output would need to
be accurate and verifiable and provide links to actual data.

Student deliverables • A report that explains the techniques used for the creation of an
intelligent chatbot interface that will produce accurate, useful
responses to queries of the student feedback database.
• A report documenting the outputs generated during the project
Student’s line manager Mark Yoxon
Project location(s) LU Campus and Remote

Other comments

Student Requirements
Essential technical skills • Competent in Python
• Awareness of GenAI technologies and their use
Desirable technical skills
Essential general skills • Ability to organise and manage own workload
• Clear written communication
• Attention to detail
• Problem solving mindset
• Some project management awareness
Desirable general skills • Confidence working independently while seeking guidance when
needed
• Ability to present findings to technical and non‑technical
audiences
Essential experience
Desirable experience
Other comments

Organisation Details
Main contact name Mark Yoxon
Main contact position Head of Strategic Planning
Applications email
Applications email cc
Other comments
PRJ70-NDA-DEA - Design and Evaluation of an AI-Powered Customer
Support Chatbot - ND AXON LIMITED
Project for AI

Project Description
Project name Design and Evaluation of an AI-Powered Customer Support Chatbot

Stipend offered £2000-3000


Host Organisation ND AXON LIMITED

Project description
This project involves developing and testing an AI-powered chatbot
within an existing React Native mobile application to improve
customer service. The student will build a conversational system
capable of answering common queries, providing real-time support,
and maintaining natural multi-turn interactions. The focus will be on
designing and optimising the chatbot to deliver accurate and reliable
responses, using techniques such as improved information retrieval
and response guidance to reduce incorrect outputs.

The chatbot will be evaluated using measurable metrics including


response quality, speed, reliability, and user satisfaction. The project
will follow structured stages of design, development, integration,
testing, and performance evaluation, with the aim of delivering a
practical improvement to digital customer support.

Cross-functional project to work with:

- External Development Partner


- Developer Engineer
- Data Engineer
Student deliverables The student will be expected to deliver:

● A working AI-powered chatbot integrated into the existing


React Native mobile application.
● A backend system to support the chatbot and ensure secure
and reliable performance.
● An evaluation of the chatbot’s performance, including
analysis of response quality, speed, and user satisfaction.
● A completed MSc dissertation describing the research,
system development, testing, results, and key findings.

Student’s line manager Who will the student report to during the placement?
Project location(s) Michal Jakubiak
Other comments Founder & CEO

Student Requirements
Essential technical skills
Strong programming skills in Python or JavaScript
Understanding of machine learning fundamentals and API
Desirable technical skills
Experience with large language model (LLM) tools or conversational
AI
Basic knowledge of backend or mobile application development
Essential general skills
Strong analytical and problem-solving ability
Ability to work independently and manage time effectively
Desirable general skills Interest in applying AI to real-world customer service problems in a
regulated platform
Essential experience Prior internship, research project, or AI-related practical work
Desirable experience Communication skills, working with the team
Other comments The student is not expected to have experience in every area listed.
Guidance and support will be provided throughout the placement,
and the project is designed to allow the student to develop new
technical skills. A proactive attitude and willingness to learn are
highly valued.

Organisation Details
Main contact name Michael Jakubiak
Main contact position CEO
Applications email michael@[Link]
Applications email cc maciej@[Link]
Other comments
PRJ71-BAS-HAI - Human–AI Sensemaking Under Uncertainty &
Cognitive Offloading to Machines - BAe Systems Digital Intelligence
Project for AI

Project Description
Project name Human–AI Sensemaking Under Uncertainty & Cognitive Offloading to
Machines: Engineering and Human Factors Analysis of Generative AI
Dependence, Technical Limits, Critical Thinking, Cognitive Bias, and
Systemic Risk

Fluent but Fallible: Engineering Critical Engagement with AI


Reasoning Systems
Stipend offered N/A
Host Organisation BAe Systems Digital Intelligence

Project description
Project introduction – scene setting
Large Language Models and adjacent AI systems are increasingly
treated as epistemic authorities rather than probabilistic tools. Their
outputs are f luent, conf ident, and rhetorically coherent, which creates
a powerf ul illusion of competence and intent. In practice, these systems
do not reason, do not hold belief s, and do not test claims against reality.
They pattern-match against training distributions, reinf orcement
signals, and policy constraints.

Despite this, users f requently def er judgement to AI outputs, even


when those outputs are incomplete, shallow, biased, or demonstrably
wrong. This def erence is amplif ied by the users own perceptions and
bias, time pressure, cognitive load, institutional endorsement (“the
system said”), and the absence of visible uncertainty. Only sustained
adversarial questioning exposes the f ragility beneath apparently
compelling answers.

Compounding this problem is the role of AI vendors themselves. Model


behaviour is shaped not only by data and architecture, but by opaque
policy layers that embed commercial risk tolerance, moral positioning,
and reputational def ence. When systems ref use lawf ul, legitimate
requests or steer answers toward “acceptable” narratives, they are no
longer neutral tools but active normative actors. The boundary between
legal compliance, ethical caution, and corporate opinion is poorly
def ined—and largely invisible to the user.

This project asks students to interrogate that boundary.

Types of problems the project should explore


Students are expected to consider both technical and human
dimensions. Examples include, but are not limited to:
• Epistemic hallucination – conf ident f abrication, spurious citations,
invented mechanisms, or f alse causal explanations that are not
f lagged as uncertain.
• Shallow synthesis – answers that sound complete but collapse
under basic counterf actuals, base-rate checks, or domain-specific
scrutiny. How to ensure AI utilises critical thinking and delivers
honest, accurate, unbiased and insightf ul answers.
• Narrative echoing – reinf orcement of popular, institutional, or
politically dominant f ramings at the expense of minority evidence
or uncomf ortable data.
• Policy-driven bias – ref usals, hedging, or moralising responses that
exceed legal or regulatory requirements and ref lect corporate
positioning rather than law.
• Automation bias and authority transf er – humans def erring
judgement, responsibility, or accountability to AI systems, even in
saf ety-critical or ethical contexts.
• Cognitive of f loading failure – users losing the ability (or motivation)
to interrogate claims because the system appears articulate and
conf ident.
• The opacity and asymmetry of the system prevent users f rom
seeing training data, weighting, policy f ilters or incentive structures.
This expectation to trust outputs raises signif icant risks to the
individual, society and knowledge development of recursively
weighting answers on biased or questionable arguments.
Students may f ocus on one problem deeply or examine interactions
between several.

Project challenge statement

“AI systems are increasingly trusted as reasoning agents despite


lacking critical thinking, epistemic self-awareness, or accountability.
Investigate how technical limitations and human cognitive factors
interact to create misplaced trust in AI outputs, and assess the risks
this poses in real-world decision-making.”

Students are explicitly encouraged to reinterpret, narrow, or challenge


this statement. The goal is not agreement, but rigorous examination.

Student deliverable 1: one-page project proposal


Students must submit a one-page concept document that includes the
f ollowing sections. Brevity and clarity matter more than coverage:

1. Core idea
What specif ic question, f ailure mode, or interaction are you exploring?
Why does it matter?

2. Intent and perspective


Are you aiming to expose a risk, test a claim, critique a system, design
a mitigation, or study user behaviour? From whose perspective (user,
developer, policymaker, regulator)?

3. Approach and method


How will you investigate this? Examples include controlled prompting
experiments, adversarial questioning, comparative model analysis,
user studies, artef act critique, or simulation.

4. Artefacts and implementation


Will you build code, design a prototype interf ace, construct a test
harness, analyse transcripts, or produce a conceptual f ramework?
Code is optional but must add value if used.

5. Expected outcomes
What do you expect to demonstrate, f alsif y, or reveal? How will
success be judged?

6. Risks, assumptions, and limitations


What might invalidate your f indings? What assumptions are you
making about users, models, or contexts?

Overall learning intent


The purpose of this project is to critically examine the mutual f ailure
modes of humans and AI systems within contemporary knowledge
production and decision-making. Students are expected to investigate
how generative AI produces f luent but potentially shallow or biased
outputs, and how human users of ten accept these outputs without
suf f icient epistemic challenge. The project assumes that neither human
judgement nor machine output is inherently reliable; both require
structured interrogation.

Students should explore how prevailing narratives emerge and


stabilise — including scientif ic consensus, media f raming, institutional
positions, ideological movements, and counter-narratives — and
examine how AI systems may reinf orce, simplif y, or distort these
narratives through training data distributions, optimisation objectives,
and policy constraints imposed by developers or organisations.
Equally, students must analyse the human f actors that encourage
uncritical acceptance: authority bias, cognitive of floading, confirmation
bias, social conf ormity, time pressure, and perceived technological
legitimacy.

A core expectation is that students will move beyond passive critique


into epistemic testing. Claims — whether produced by humans,
institutions, or AI systems — should be treated as hypotheses subject
to challenge. Students should evaluate:

• whether dominant interpretations omit alternative explanations or


relevant data,
• where ideological f raming, commercial incentives, or reputational
risk management inf luence inf ormation presentation,
• how selective evidence, simplif ication, or moral positioning may
shape outputs,
• and how systems might present critical thinking, uncertainty,
counter-evidence, or dissenting perspectives more rigorously.

An explicit question within the project is where responsibility lies f or


epistemic f airness and critical balance. Should users be expected to
exercise disciplined critical thinking independently, or should AI
systems actively introduce counterf actuals, opposing viewpoints, and
primary data to prevent passive acceptance? Students should consider
the ethical and governance implications of AI companies embedding
normative boundaries that exceed legal requirements, including
tensions between saf ety, f ree expression, brand protection, and
inf ormational integrity.

Student deliverables The student will be expected to deliver a Masters-level report that
describes their exploration of the above ideas and the evaluations
that they have undertaken.

Student’s line manager Prof Andrew Harwood


Project location(s) Research Scientist
Other comments NB – to apply for this project you must produce the one-page
proposal as described above

Student Requirements
Essential technical skills Rich understanding of issues of trust and acceptance in AI
Desirable technical skills
Essential general skills Critical thinking
Must be able to work independently
Desirable general skills
Essential experience
Desirable experience
Other comments

Organisation Details
Main contact name Andrew Harwood
Main contact position Research Scientist
Applications email
Applications email cc
Other comments
PRJ72-BC-DA - Data science: A data-driven framework for service user
voice in adult social care - Blackpool Council
Project for DS AI

Project Description
Project name Data science: A data-driven framework for service user voice in
adult social care
Stipend offered £3,000

Host Organisation Blackpool Council


Project description What are the aims of the project?
To research, design and implement a data-driven method for
capturing, analysing, and reporting service user voice within Adult
Social Care, enabling Blackpool Council to evidence improvement
and better understand lived experience.

What type of project is this? (e.g. implementation, research,


evaluation, enhancement, pilot)
This is primarily a data science implementation project. However,
the student will be supported by Blackpool’s Health Determinants
Research Collaboration (HDRC) and there is the scope to develop this
as a piece of research or evaluation depending on the student’s
interests.

How will work be organised?


The project will give the student an opportunity to design and deliver
an end-to-end data solution that will impact the lives and wellbeing
of vulnerable residents. The key areas of work will include:
• Researching best practice across the country in terms of
collecting consistent service user voice within systems in
Adult Social Care.
• Designing and implementing a systematic approach to
collecting service user feedback across all Adult Social Care
systems and pathways whilst being mindful of frontline
capacity to collect information. This will include working
collaboratively with Adult Social Care managers and Systems
team to design business processes.
• Extracting, operationalising and visualising data from Mosaic
(the main Adult Social Care system used by Blackpool
Council) and other relevant systems. Develop a coherent
quantitative framework and visual outputs that support
performance monitoring and insight.
• Embedding the outputs of analysis into business as usual and
governance processes to ensure continuous improvement
and evidence-based decision making at team and service
level.
Who will contribute?
The student will be supported by:
• Bob Allen – Head of Performance and Business Intelligence,
Blackpool Council
• Dr Sarah Blagden – Consultant in Public Health, Blackpool
Health Determinants Research Collaboration
• Sara Coombs – Systems and Intelligence Manager, Blackpool
Council
• Nick Henson – Interim Director of Adult Social Services,
Blackpool Council
• Dan Nicholson – Principal Social Worker & Head of Service,
Blackpool Council

What business areas will be involved?


The following council directorates will be involved:
• Adult Social Care
• Business Intelligence and Systems
• Public Health
• Blackpool Researching Together (Blackpool Health
Determinants Research Collaboration)

Which technologies will be used?


• SQL Server Management Studio
• Power BI
• Several case management systems, principally Mosaic - the
primary social care case management system used by
Blackpool Council. This integrates adult, children’s, and
finance services into one platform.
Student deliverables What will the student be expected to deliver at the end of the
project?
Expected outputs encompass:
• A framework of quantitative metrics for measuring service
user experience
• A prototype dashboard or reporting template to visualise
this information using Power BI
• A report summarising findings, methodology, and
recommendations for ongoing implementation
Student’s line manager Bob Allen – Head of Performance and Business Intelligence
Project location(s) The relevant teams for the project are based at: Blackpool Council,
Bickerstaffe House, Talbot Road, Blackpool, FY1 3AH. However,
Blackpool Council adopts a flexible working policy and much of the
work can be carried out remotely.
Other comments

Student Requirements
Essential technical skills Data handling and analysis
• Data cleaning and transformation
• Working with structured and unstructured data
Data visualisation
• Ability to build clear, interpretable visualisations
• Experience with Power BI
Understanding of quantitative measurement
• Ability to design or operationalise quantitative indicators
Information governance
• Understanding of data protection
• Ethical handling of sensitive data
Desirable technical skills Dashboard and reporting automation
• Building interactive dashboards in Power BI
• Automating data pipelines or refresh processes
Project management and communication skills
• Ability to translate analytical findings into clear, actionable
insights
• Experience presenting findings to non-technical audiences
Essential general skills Communication and stakeholder engagement
• Ability to communicate clearly with non-technical staff
• Confidence in asking questions and clarifying requirements
• Confidence in presenting findings
• Sensitivity when discussing topics relating to vulnerable
adults
Analytical thinking
• Strong problem-solving skills
• Ability to break down complex issues into manageable
analytical tasks
Organisation and project management
• Ability to plan and manage workload independently
• Good time-management
Confidentiality
• Understanding of the sensitive nature of adult social care
data
• Ability to work within governance, confidentiality, and
ethical boundaries
Desirable general skills Co-design processes
• Experience working with multiple stakeholders across a
project including technical and non-technical partners
Essential experience Collaboration and teamwork
• Experience working in teams to deliver aims and objectives
Project work
• Experience delivering a structured project or assignment
(e.g. academic, professional, voluntary) involving planning,
meeting deadlines, and documentation of work
Presentation of findings
• Experience of presenting findings to others e.g. through
presentations, reports or academic assignments
Desirable experience Handling sensitive information
• Experience working within ethical or governance frameworks
(e.g. GDPR awareness, research ethics).
Other comments

Organisation Details
Main contact name Bob Allen
Main contact position Head of Performance and Business Intelligence
Applications email Bob Allen ([Link]@[Link])
Applications email cc Dr Sarah Blagden ([Link]@[Link])
Other comments

You might also like