0% found this document useful (0 votes)
3 views19 pages

Pharma Analytics Python AI Programme

The document outlines a 12-hour intensive training program on Pharma Analytics with Python and AI scheduled for June 2–6, 2025, aimed at pharmaceutical professionals. Participants will learn to apply AI and Python tools to real-world datasets, generate statistical reports, and create visualizations for regulatory submissions through hands-on projects. The course covers various topics including AI fundamentals, data cleaning, and visualization techniques, with no prior Python experience required.

Uploaded by

Neetu Kamra
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views19 pages

Pharma Analytics Python AI Programme

The document outlines a 12-hour intensive training program on Pharma Analytics with Python and AI scheduled for June 2–6, 2025, aimed at pharmaceutical professionals. Participants will learn to apply AI and Python tools to real-world datasets, generate statistical reports, and create visualizations for regulatory submissions through hands-on projects. The course covers various topics including AI fundamentals, data cleaning, and visualization techniques, with no prior Python experience required.

Uploaded by

Neetu Kamra
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

PHARMA ANALYTICS WITH PYTHON & AI | 12-Hour Intensive Programme | June 2–6, 2025

PHARMA ANALYTICS
WITH PYTHON & ARTIFICIAL INTELLIGENCE
12-Hour Intensive Training Programme
June 2–6, 2025 | 2.5 Hours Per Day

12
Total Hours
5
Days
5
Projects
6+
AI Tools

Python • Claude AI • Gemini • Pandas • Plotly • Smartwatch APIs

Confidential Course Material | Page


PHARMA ANALYTICS WITH PYTHON & AI | 12-Hour Intensive Programme | June 2–6, 2025

Table of Contents
TOC \h \o "1-3" \t "Heading1,1,Heading2,2,Heading3,3"

Confidential Course Material | Page


PHARMA ANALYTICS WITH PYTHON & AI | 12-Hour Intensive Programme | June 2–6, 2025

1. Programme Overview
This 12-hour intensive training programme equips pharmaceutical professionals, clinical data scientists,
and biostatisticians with practical skills at the intersection of Python programming, artificial intelligence,
statistical analysis, and digital health technology. Delivered over five days (June 2–6, 2025) in sessions
of 2.5 hours each, the programme progresses systematically from AI fundamentals to hands-on data
pipelines and real-world capstone projects.

1.1 Programme Objectives


• Understand the landscape of AI and machine learning in pharmaceutical R&D
• Apply Python and AI tools to real pharma datasets with confidence
• Generate professional statistical reports using Claude AI and Gemini
• Build end-to-end data pipelines for smartwatch and wearable health data
• Create publication-quality visualisations for regulatory submissions
• Demonstrate 5 capstone projects integrating Python, AI, and digital health

1.2 Target Audience


• Pharmaceutical data scientists and bioinformaticians
• Clinical research associates and biostatisticians
• Regulatory affairs professionals adopting AI tools
• Pharma IT professionals managing health data pipelines
• Medical affairs teams seeking data literacy and AI reporting skills

1.3 Prerequisites
• Basic familiarity with data concepts (spreadsheets, statistics)
• No prior Python experience required (Day 3 starts from foundations)
• Laptop with internet connection and ability to install Python/Anaconda

2. Course Schedule at a Glance


The following table provides a structured overview of all five sessions, timings, topics, and deliverables
for the complete 12-hour programme:

Day /
Time Topic Key Deliverables Hours
Date

Mon Jun 09:00 – Introduction to Artificial Jupyter ready, AI 2.5 hrs

Confidential Course Material | Page


PHARMA ANALYTICS WITH PYTHON & AI | 12-Hour Intensive Programme | June 2–6, 2025

Day /
Time Topic Key Deliverables Hours
Date

2 11:30 Intelligence in Pharma landscape map


AI/ML overview, regulatory
frameworks, Python setup

Tue Jun 09:00 – AI Applied in Pharma AI use case matrix, demo 2.5 hrs
3 11:30 NLP, computer vision, drug notebook
discovery AI, ethics

Wed Jun 09:00 – Hands-On Python for Clean pharma dataset, 2.5 hrs
4 11:30 Pharma Data EDA report notebook
Pandas, EDA, cleaning
clinical datasets, feature
engineering

Thu Jun 09:00 – Data Visualisation + Interactive dashboard, 2.5 hrs


5 11:30 Descriptive Statistics publication charts
Matplotlib, Seaborn, Plotly,
PK parameters, correlation

Fri Jun 6 09:00 – Statistical Reports + AI-generated stat report, 2.5 hrs
11:30 Smartwatch Projects 3 wearable project
Claude/Gemini reporting, demos
wearable data, project
showcase

TOTAL PROGRAMME DURATION 12 hrs

Confidential Course Material | Page


PHARMA ANALYTICS WITH PYTHON & AI | 12-Hour Intensive Programme | June 2–6, 2025

3. Detailed Session Content


3.1 Day 1 — Monday, June 2, 2025
Topic Introduction to Artificial Intelligence in Pharma

Duration 09:00 – 11:30 (2.5 Hours)

Session Overview
Day 1 lays the conceptual foundation for the entire programme. Participants will develop a clear
understanding of the AI/ML landscape in pharmaceutical contexts, regulatory expectations, and the
Python environment they will use throughout the week.

Learning Objectives
1. Define artificial intelligence, machine learning, deep learning, and generative AI in
pharmaceutical context
2. Map the AI adoption landscape across drug discovery, clinical trials, pharmacovigilance, and
regulatory affairs
3. Identify relevant regulatory guidance (FDA, EMA, ICH E9(R1)) on AI/ML use
4. Set up a fully configured Python environment (Anaconda, Jupyter, key libraries)

Detailed Content Breakdown


Module 1 — What is Artificial Intelligence? (45 min)
• Definitions: AI vs Machine Learning vs Deep Learning vs Generative AI
• Key algorithms overview: regression, classification, clustering, neural networks
• Supervised vs unsupervised vs reinforcement learning — pharma examples of each
• Large Language Models (LLMs) explained: how Claude and Gemini work
• Demonstration: querying a pharma dataset with a generative AI model

Module 2 — AI in Pharmaceutical R&D (45 min)


• Drug discovery: AlphaFold, molecular property prediction, virtual screening
• Clinical trials: patient recruitment AI, protocol optimisation, dropout prediction
• Pharmacovigilance: adverse event signal detection, MedDRA coding automation
• Manufacturing: predictive maintenance, quality control vision AI
• Case study: how AstraZeneca uses AI in oncology pipeline (overview)

Module 3 — Regulatory & Ethical Landscape (30 min)


• FDA's Action Plan for AI/ML-Based Software as a Medical Device (SaMD)

Confidential Course Material | Page


PHARMA ANALYTICS WITH PYTHON & AI | 12-Hour Intensive Programme | June 2–6, 2025

• EMA reflection paper on AI use in medicinal product development


• ICH E9(R1): estimands and sensitivity analysis — AI implications
• Bias in AI models: protected characteristics, training data limitations
• GDPR and patient data privacy when using AI on clinical datasets

Module 4 — Python Environment Setup (30 min)


• Installing Anaconda and creating a pharma-analytics conda environment
• Installing key libraries: pandas, numpy, scipy, matplotlib, seaborn, plotly, scikit-learn
• Launching Jupyter Notebook and navigating the interface
• Introduction to Python syntax: variables, lists, dictionaries, functions
• Loading the first pharma dataset (CSV) into a Pandas DataFrame

Hands-On Exercise
Participants will set up their environment, run their first Jupyter notebook, and perform a simple
descriptive summary of a provided clinical trial dataset using [Link]().

Deliverables
• Configured Python/Jupyter environment ready for remaining sessions
• AI in pharma landscape mapping worksheet completed
• First notebook: basic pandas summary of sample pharma dataset

3.2 Day 2 — Tuesday, June 3, 2025


Topic AI Applied in Pharma

Duration 09:00 – 11:30 (2.5 Hours)

Session Overview
Day 2 moves from theory to applied examples, exploring how AI is actively deployed across key
pharmaceutical workflows. Participants will see live demonstrations of NLP, computer vision, and AI
agent tools, and will critically evaluate AI outputs for regulatory and scientific accuracy.

Learning Objectives
5. Explain how NLP detects adverse events and extracts insights from unstructured pharma text
6. Describe how computer vision is applied to histopathology and cell imaging
7. Use Claude and Gemini APIs to generate intelligent summaries of pharma data
8. Critically evaluate AI model outputs for bias, accuracy, and regulatory acceptability

Confidential Course Material | Page


PHARMA ANALYTICS WITH PYTHON & AI | 12-Hour Intensive Programme | June 2–6, 2025

Detailed Content Breakdown


Module 1 — NLP for Pharmacovigilance (40 min)
• Text preprocessing: tokenisation, stop word removal, stemming on patient narratives
• Named entity recognition (NER) for drug names, adverse events, disease terms
• Sentiment analysis on patient-reported outcomes (PRO) data
• MedDRA coding automation using transformer-based NLP models
• Practical: extract adverse event terms from a sample clinical narrative dataset

Module 2 — Computer Vision in Pathology & Imaging (30 min)


• CNN architectures for histopathology slide classification
• Cell counting and tumour detection in microscopy images
• Retinal scan analysis for diabetic retinopathy screening
• FDA-cleared AI imaging tools: overview of PathAI, [Link], [Link]

Module 3 — AI-Driven Drug Discovery & Repurposing (30 min)


• Molecular property prediction: ADMET (absorption, distribution, metabolism, excretion, toxicity)
• Graph Neural Networks for molecular interaction modelling
• Drug repurposing using knowledge graphs and AI
• Generative AI for novel molecule design: overview of approaches

Module 4 — Claude & Gemini for Pharma Intelligence (40 min)


• Accessing Claude API via Python: authentication, prompt engineering
• Accessing Gemini API via Python: setup, key parameters
• Prompt design for pharma contexts: specificity, safety constraints, structured output
• Live demo: ask Claude to summarise a clinical study report section
• Live demo: use Gemini to query a pharmacokinetics dataset in natural language
• Comparison: when to use Claude vs Gemini for pharmaceutical tasks

Hands-On Exercise
Participants will use the Claude API to generate a one-paragraph adverse event summary from a
provided safety dataset, and use Gemini to answer natural language questions about the same data.

Deliverables
• AI use case matrix: list of AI applications mapped to pharma workflow stages
• Demo notebook: Claude and Gemini API calls with pharma prompts
• Critical evaluation checklist for AI outputs in regulated environments

3.3 Day 3 — Wednesday, June 4, 2025


Topic Hands-On Python for Pharma Data

Confidential Course Material | Page


PHARMA ANALYTICS WITH PYTHON & AI | 12-Hour Intensive Programme | June 2–6, 2025

Duration 09:00 – 11:30 (2.5 Hours)

Session Overview
Day 3 is the most technically intensive session, providing structured hands-on training in Python data
manipulation using real pharmaceutical datasets. Participants move from raw data loading through to a
cleaned, analysis-ready dataset.

Learning Objectives
9. Load, inspect, and profile pharmaceutical datasets using Pandas
10. Clean missing values, outliers, and inconsistencies in clinical trial data
11. Merge multiple data sources (demographics, lab results, dosing records)
12. Engineer features relevant to pharmacokinetic and pharmacodynamic analysis
13. Export a clean, annotated dataset ready for statistical analysis

Detailed Content Breakdown


Module 1 — Loading & Profiling Pharma Datasets (40 min)
• Reading clinical trial data formats: CSV, Excel, SAS7BDAT, CDISC SDTM
• [Link](), [Link](), [Link](), [Link] — dataset first look
• Data profiling with pandas-profiling / ydata-profiling: auto EDA report
• Identifying data quality issues: duplicates, impossible values, unit mismatches
• Understanding CDISC SDTM domain structure (DM, AE, LB, EX, VS domains)

Module 2 — Data Cleaning & Missing Value Handling (35 min)


• Strategies for missing data: MCAR, MAR, MNAR classification
• Imputation methods: mean, median, LOCF, BOCF for pharma data
• Detecting and handling outliers: IQR method, Z-score, clinical plausibility checks
• Standardising units: mg/kg vs mg, mmol/L vs mg/dL conversions
• Correcting date/time fields: treatment start, sample collection, visit windows

Module 3 — Merging & Joining Datasets (30 min)


• [Link]() for inner, left, outer joins on USUBJID (subject identifier)
• Combining demographics (DM) with lab results (LB) and dosing (EX)
• Handling one-to-many relationships: multiple lab visits per subject
• Validating merges: checking row counts, duplicate checks post-merge

Module 4 — Feature Engineering for Pharma (35 min)


• Calculating derived PK parameters: AUC trapezoid rule, Cmax, Tmax, T½ in Python
• Creating dose-normalised variables for cross-study comparisons
• Biomarker categorisation: quartile binning, clinically meaningful thresholds

Confidential Course Material | Page


PHARMA ANALYTICS WITH PYTHON & AI | 12-Hour Intensive Programme | June 2–6, 2025

• Encoding categorical variables: treatment arm, gender, race, study site


• Time-from-dose calculations for PK sample timing alignment

Hands-On Exercise
Participants will work through a complete data cleaning pipeline on a provided Phase II clinical trial
dataset: load raw data, profile it, clean it, merge demographics with laboratory results, and export a final
analysis-ready CSV with documentation.

Deliverables
• Cleaned, analysis-ready pharma dataset (CSV format)
• Data cleaning log documenting all transformations applied
• EDA profiling report generated automatically

3.4 Day 4 — Thursday, June 5, 2025


Topic Data Visualisation & Descriptive Statistics for Pharma

Duration 09:00 – 11:30 (2.5 Hours)

Session Overview
Day 4 covers the full spectrum of pharmaceutical data visualisation and descriptive statistics — from
exploratory analysis through to publication-ready charts and regulatory statistical summaries.
Participants will produce visualisations directly applicable to CSRs, regulatory submissions, and clinical
team presentations.

Learning Objectives
14. Create publication-quality statistical visualisations using Matplotlib and Seaborn
15. Build interactive pharma dashboards using Plotly and Streamlit
16. Compute key descriptive statistics and pharmacokinetic parameters using Python and SciPy
17. Interpret dose-response curves, bioequivalence plots, and safety data visualisations
18. Produce chart outputs formatted for regulatory submissions (ICH E3, CTD)

Detailed Content Breakdown


Module 1 — Matplotlib & Seaborn for Pharma (45 min)
• Histograms and KDE plots for biomarker and PK parameter distributions
• Box plots and violin plots: comparing treatment arms, visit timepoints
• Heatmaps: correlation matrices for multi-biomarker panels

Confidential Course Material | Page


PHARMA ANALYTICS WITH PYTHON & AI | 12-Hour Intensive Programme | June 2–6, 2025

• Scatter plots with regression lines: dose-concentration relationships


• Kaplan-Meier survival curves for time-to-event analysis
• Publication formatting: figure size, DPI, font sizes, axis labels, legends, colour palettes
• Exporting to SVG, PNG 300 DPI for regulatory document embedding

Module 2 — Interactive Dashboards with Plotly (30 min)


• Plotly Express for rapid interactive charts from DataFrames
• Multi-panel dashboard: adverse event profile, dose-response, PK curves
• Plotly Dash/Streamlit: deploying a pharma monitoring dashboard
• Adding dropdown filters by treatment arm, visit, patient subgroup

Module 3 — Descriptive Statistics for Pharma Data (40 min)


• Central tendency: mean, median, geometric mean for log-normal PK data
• Variability: SD, SEM, CV%, IQR — when to use each in pharma reporting
• Normality testing: Shapiro-Wilk, Q-Q plots, Anderson-Darling test
• Correlation analysis: Pearson, Spearman for biomarker-efficacy relationships
• Confidence intervals: t-distribution vs bootstrap CIs for small pharma samples
• ANOVA and Kruskal-Wallis for comparing treatment groups

Module 4 — Pharmacokinetic Analysis in Python (35 min)


• Calculating Cmax, Tmax, AUC0-t, AUC0-inf, T½, CL/F from concentration-time data
• Non-compartmental analysis (NCA) using Python [Link]
• Mean concentration-time profiles: semi-log scale plotting
• Bioequivalence: calculating GMR and 90% CI for AUC and Cmax
• Use case: complete NCA walkthrough on a provided PK dataset

Hands-On Exercise
Participants will produce a complete visualisation package for a clinical trial dataset: distribution plots
for key endpoints, a correlation heatmap, a dose-response scatter plot, and summary statistics tables
ready for inclusion in a Clinical Study Report (CSR).

Deliverables
• Visualisation package: 6+ charts formatted for regulatory submission
• Interactive Plotly dashboard (HTML) showing trial safety and efficacy data
• Summary statistics table: n, mean, SD, median, IQR, 95% CI by treatment arm

3.5 Day 5 — Friday, June 6, 2025


Topic Statistical Reports with AI + Smartwatch Digital Data Projects

Confidential Course Material | Page


PHARMA ANALYTICS WITH PYTHON & AI | 12-Hour Intensive Programme | June 2–6, 2025

Duration 09:00 – 11:30 (2.5 Hours)

Session Overview
The final session integrates all skills from the week into two parallel streams: (1) generating automated,
AI-assisted statistical reports using Claude and Gemini, and (2) three end-to-end capstone projects
focused on digital health data from smartwatches and wearable devices. The session concludes with a
project showcase and programme close.

Learning Objectives
19. Automate statistical analysis and report generation using Python with Claude and Gemini APIs
20. Build a wearable cardiac monitoring data pipeline with anomaly detection
21. Develop a medication adherence prediction system using smartwatch sensor data
22. Create a multi-device remote patient monitoring dashboard
23. Present a complete end-to-end pharma analytics project to peers

Detailed Content Breakdown


Module 1 — Automated Statistical Reporting with Claude & Gemini (50 min)
• Architecture of an automated reporting pipeline: Python analysis → AI narration → formatted
output
• Claude API: generating clinical language descriptions of statistical results
• Gemini API: cross-referencing findings with published literature via search
• Combining Python-computed stats with AI-generated interpretation in one document
• Auto-generating DOCX statistical analysis reports with python-docx
• Auto-generating PDF reports with ReportLab
• Building a reusable reporting template for standard pharma endpoints
• Quality assurance: verifying AI-generated text for accuracy and scientific correctness

Module 2 — Smartwatch Project Showcase (80 min — 3 projects in rotation)


Project demonstrations and peer review of the three smartwatch capstone projects:
• Project 1: Cardiac Monitoring Pipeline — live demo of ECG analysis and alert generation
• Project 2: Medication Adherence Tracker — demo of ML classifier and AI patient report
• Project 3: Remote Patient Monitoring Dashboard — demo of multi-metric Streamlit dashboard

Module 3 — Programme Wrap-Up & Next Steps (20 min)


• Review of all 5 capstone projects and key learning from each day
• Recommended resources: books, courses, pharma data science communities
• Building a pharma analytics portfolio with GitHub and Jupyter Book
• Overview of advanced topics: NLP clinical text mining, ML model validation, survival analysis
• Q&A and individual feedback session

Confidential Course Material | Page


PHARMA ANALYTICS WITH PYTHON & AI | 12-Hour Intensive Programme | June 2–6, 2025

Confidential Course Material | Page


PHARMA ANALYTICS WITH PYTHON & AI | 12-Hour Intensive Programme | June 2–6, 2025

4. Capstone Projects
Five capstone projects form the applied core of this programme. Projects 1–3 focus on smartwatch and
wearable digital health data. Projects 4–5 cover pharma statistical analysis and AI-powered reporting.
All projects use Python as the primary language and integrate Claude AI and/or Gemini for intelligent
output generation.

Project 1 — Cardiac Monitoring via Smartwatch


Category Digital Health | Signal Processing | Anomaly Detection

Objective Build a real-time ECG and heart rate data pipeline using smartwatch APIs to
detect cardiac arrhythmia patterns and generate clinician-facing alerts.
Data Source Apple Watch HealthKit (ECG, HR), Garmin Connect API (HRV, RR intervals),
Polar H10 chest strap (reference ECG)

Key Libraries NeuroKit2, [Link], pandas, numpy, plotly, scikit-learn (IsolationForest),


anthropic (Claude API)

AI Integration Claude AI generates a natural language clinician alert report for each flagged
cardiac episode, describing the anomaly, confidence level, and recommended
action.

Detailed Project Tasks


24. Connect to Apple HealthKit and Garmin Connect API to retrieve ECG time-series and HRV data
25. Preprocess ECG signals: bandpass filtering (0.5–40 Hz), baseline wander removal using
NeuroKit2
26. R-peak detection using Pan-Tompkins algorithm; compute RR intervals and HRV metrics
(SDNN, RMSSD, pNN50)
27. Feature extraction: mean HR, HR variability, QRS duration, ST segment deviation, ectopic beat
frequency
28. Apply Isolation Forest anomaly detection model to identify arrhythmia episodes from feature
vectors
29. Generate interactive Plotly visualisation: 24-hour HR trend with flagged anomaly windows
30. Call Claude API with flagged episodes to produce a formatted clinical alert report (DOCX)
31. Unit test the pipeline on a benchmark cardiac dataset (MIT-BIH Arrhythmia Database)

Project 2 — Medication Adherence Tracker


Category Wearable Sensing | Machine Learning | AI Reporting

Objective Use wearable accelerometer and reminder data to predict missed medication
doses and generate personalised AI-generated patient adherence reports.

Data Source Smartwatch accelerometer (wrist motion), medication reminder logs, patient

Confidential Course Material | Page


PHARMA ANALYTICS WITH PYTHON & AI | 12-Hour Intensive Programme | June 2–6, 2025

self-report check-in timestamps, HealthKit step/activity data

Key Libraries pandas, scikit-learn (RandomForestClassifier), SHAP, matplotlib, google-


generativeai (Gemini API), python-docx

AI Integration Gemini generates personalised weekly patient summaries in plain language.


Claude generates a structured weekly adherence PDF report for the clinical
team with statistical adherence metrics.

Detailed Project Tasks


32. Parse raw smartwatch accelerometer logs (CSV) and synchronise with dose schedule
timestamps
33. Extract wrist motion features at ±30 min windows around each scheduled dose time: mean
acceleration, peak-to-peak amplitude, gesture count
34. Label training data: confirmed doses (based on reminder confirmation + motion signature) vs
missed doses
35. Train Random Forest classifier; tune hyperparameters with 5-fold cross-validation; evaluate F1,
AUC-ROC
36. SHAP analysis: identify the top accelerometer features predictive of missed doses
37. Compute weekly adherence rate (%), streak lengths, and time-of-day miss patterns
38. Gemini API call: generate a plain-language personalised summary ('You took 85% of your
doses this week...')
39. Claude API call: generate a structured clinical team report as DOCX with tables, trend charts,
and recommendations

Project 3 — Remote Patient Monitoring Dashboard


Category Multi-Device Integration | Clinical Dashboard | AI Narrative

Objective Aggregate multi-modal wearable data (SpO2, HRV, steps, sleep stages) from
multiple device brands into a unified pharma-grade monitoring dashboard for
clinical trial participants.

Data Source Fitbit Web API, Apple HealthKit, Samsung Health, Garmin Connect — all
providing SpO2, HR, HRV, steps, sleep stage, activity data

Key Libraries pandas, numpy, streamlit, plotly, [Link] (Shewhart control charts),
anthropic (Claude API), google-generativeai (Gemini), fpdf2

AI Integration Claude generates a daily narrative summary for each patient ('Patient vitals
today were within normal range except for one SpO2 dip at 02:14...'). Gemini
flags any values that diverge from reference ranges in the literature.

Detailed Project Tasks


40. Build a data ingestion pipeline supporting Fitbit, Apple, Samsung, and Garmin APIs with OAuth2
authentication

Confidential Course Material | Page


PHARMA ANALYTICS WITH PYTHON & AI | 12-Hour Intensive Programme | June 2–6, 2025

41. Normalise data across device brands: unified schema for SpO2 (%), HR (bpm), HRV (ms),
steps, sleep stages
42. Implement Shewhart individual control charts for each vital sign to detect out-of-control signals
43. Compute daily summary statistics: mean, min, max, SD for each metric per patient
44. Build Streamlit dashboard: patient selector sidebar, multi-metric trend charts, alert flags, data
quality indicators
45. Claude API integration: call daily to produce a one-paragraph patient narrative for each
participant
46. Gemini API integration: flag any value outside published reference ranges with literature
citations
47. Export daily summary PDF report per participant using fpdf2

Project 4 — Descriptive Statistics on Pharma Datasets


Category Clinical Statistics | PK Analysis | Regulatory Reporting

Objective Perform end-to-end descriptive statistical analysis on CDISC-formatted clinical


trial and pharmacokinetics datasets, producing regulatory-submission-ready
statistical tables and figures.

Data Source Simulated Phase I PK dataset (CDISC SDTM format), real-world clinical
endpoints dataset, bioequivalence study data

Key Libraries pandas, numpy, [Link], [Link], pingouin, seaborn, matplotlib,


python-docx, anthropic (Claude API)

AI Integration Claude API generates the statistical methods section and results interpretation
in ICH E3 Clinical Study Report language, ready for copy-paste into a CSR
template.

Detailed Project Tasks


48. Load and parse a CDISC SDTM dataset (DM, LB, EX, VS, AE domains) using Pandas
49. Compute primary endpoint descriptive statistics: n, mean, SD, SE, median, IQR, 95% CI, min,
max by treatment arm and visit
50. Pharmacokinetic NCA: calculate Cmax, AUC0-t, AUC0-inf, T½, Kel, CL/F, Vd/F using
[Link]
51. Normality testing: Shapiro-Wilk and Q-Q plots for all continuous endpoints
52. Outlier detection: Grubbs test and IQR-based flagging with clinical plausibility review
53. Subgroup analysis: stratify by age group, gender, renal function (eGFR category), treatment
arm
54. Generate formatted statistical tables (ICH E3 Table 14.1.x format) using python-docx
55. Claude API call: generate the Results section narrative for the CSR in clinical language

Project 5 — AI-Powered Statistical Report Generator


Category Pipeline Automation | AI Integration | Document Generation

Confidential Course Material | Page


PHARMA ANALYTICS WITH PYTHON & AI | 12-Hour Intensive Programme | June 2–6, 2025

Objective Build a fully automated one-click pipeline that accepts a raw pharma dataset,
runs a pre-specified statistical analysis plan, and produces a formatted
Word/PDF statistical analysis report (SAR) combining Python-computed results
with AI-generated interpretation.

Data Source Any cleaned pharma CSV dataset (compatible with projects 3 and 4 outputs)

Key Libraries pandas, [Link], pingouin, matplotlib, python-docx, reportlab, anthropic


(Claude API), google-generativeai (Gemini API), jinja2

AI Integration Claude: writes statistical results narrative, methods section, conclusion. Gemini:
retrieves and summarises relevant published literature for the discussion
section. Both APIs are called sequentially in the pipeline.

Detailed Project Tasks


56. Design a modular Python pipeline with configurable statistical analysis plan (SAP) as JSON
input
57. Auto-run: descriptive stats, normality tests, t-tests / Mann-Whitney, ANOVA / Kruskal-Wallis,
correlation analysis
58. Auto-generate all required charts: distribution plots, boxplots, forest plots, correlation heatmaps
59. Call Claude API with statistical outputs and generate: Introduction, Methods, Results, and
Conclusion sections
60. Call Gemini API to retrieve 3–5 relevant literature citations and generate a Discussion
paragraph
61. Assemble full report in DOCX using python-docx with table of contents, section headers, auto-
numbered tables and figures
62. Convert DOCX to PDF using LibreOffice headless for final PDF output
63. Validate: run the pipeline on 3 different datasets; review AI output for scientific accuracy

Confidential Course Material | Page


PHARMA ANALYTICS WITH PYTHON & AI | 12-Hour Intensive Programme | June 2–6, 2025

5. Technology Stack
The following table describes all tools, libraries, and platforms used throughout the programme, their
specific role, and which sessions they are used in:

Tool /
Role in Programme Key Libraries / Features Used In
Technology

Python 3.11+ Primary Language Pandas, NumPy, SciPy, All sessions


Matplotlib, Seaborn, Plotly,
Scikit-learn, NeuroKit2,
Streamlit

Claude AI Report Generation & Automated statistical report Day 2, 5


(Anthropic) NLU drafting, clinical narrative
generation, DOCX/PDF
output

Gemini Multimodal AI & Dataset Q&A, literature Day 2, 5


(Google) Search retrieval, cross-validation,
multimodal analysis

Pandas / Data Wrangling Clinical CSV loading, Day 3, 4, 5


NumPy missing value handling,
CDISC SDTM/ADaM
parsing, feature engineering

Matplotlib / Static Visualisation Histograms, boxplots, Day 4


Seaborn heatmaps, dose-response
curves, publication-quality
export

Plotly / Interactive Real-time patient monitoring, Day 4, 5


Streamlit Dashboards adverse event dashboards,
interactive charts
Scikit-learn Machine Learning Random Forest, Isolation Day 2, 5
Forest, classification,
anomaly detection on
pharma data

Smartwatch Digital Health Data Apple HealthKit, Garmin Day 5


SDKs Connect API, Fitbit Web API,
Samsung Health for
wearable data ingestion

Jupyter Interactive Session-by-session All sessions


Notebook Computing notebooks, reproducible
analysis, markdown
documentation

Confidential Course Material | Page


PHARMA ANALYTICS WITH PYTHON & AI | 12-Hour Intensive Programme | June 2–6, 2025

6. Assessment & Certification


6.1 Assessment Criteria
Participants are assessed on the quality and completeness of their capstone project deliverables.
Assessment is formative (feedback-based) rather than graded, with a focus on practical competence.

Assessment Criterion Weighting Evaluated On

Working Python code: runs without errors 30% All projects

Correctness of statistical analysis and 25% Projects 4 & 5


interpretation

Quality and clarity of visualisations 20% Day 4 deliverable


AI integration: appropriate prompt design and 15% Projects 1–5
output quality

Documentation and code comments 10% All notebooks

6.2 Certification
• Certificate of Completion issued to all participants who attend all 5 sessions
• Certificate of Achievement issued to participants who submit all 5 project deliverables
• Both certificates are digitally signed and suitable for inclusion in a professional portfolio or
LinkedIn profile

7. Recommended Resources
7.1 Books
• Python for Data Analysis (Wes McKinney) — pandas and data wrangling reference
• Hands-On Machine Learning with Scikit-Learn, Keras & TensorFlow (Aurelien Geron)
• An Introduction to Statistical Learning (James, Witten, Hastie, Tibshirani) — free PDF
• Deep Learning for the Life Sciences (Bharath Ramsundar et al.) — pharma/bio ML

7.2 Online Courses & Platforms


• Coursera: Machine Learning Specialization (Andrew Ng)
• [Link]: Practical Deep Learning for Coders
• Anthropic Console documentation ([Link])
• Google AI Studio ([Link]) — Gemini API playground
• Kaggle: pharmaceutical and clinical trial public datasets

Confidential Course Material | Page


PHARMA ANALYTICS WITH PYTHON & AI | 12-Hour Intensive Programme | June 2–6, 2025

7.3 Pharma-Specific Communities


• CDISC Standards ([Link]) — pharmaceutical data standards documentation
• PhUSE (Pharmaceutical Users Software Exchange) — pharma data science network
• PSI (Statisticians in the Pharmaceutical Industry) — biostatistics resources
• RStudio/Posit Pharmaverse — R and Python pharma data science ecosystem

Pharma Analytics with Python & AI


12-Hour Intensive Programme | June 2–6, 2025 | Confidential Course Material

Confidential Course Material | Page

You might also like