Problem Statements
Problem Statements
Problem Statements
Digital Lending:
Portfolio Optimization
Industry Context
Digital lending institutions in emerging markets have witnessed rapid growth by leveraging technology-
driven customer acquisition, alternative data sources, and faster credit decisioning. These lenders cater
to a diverse customer base across urban and semi-urban regions, including first-time borrowers and
underbanked segments, through products such as unsecured personal loans, SME working capital loans,
and buy-now-pay-later (BNPL) offerings.
While scale and speed have enabled strong topline growth, evolving borrower behavior, macroeconomic
volatility, and rising acquisition costs have increased pressure on portfolio quality and profitability. Many
lenders now face the challenge of balancing growth ambitions with prudent risk management, while
effectively leveraging the rich data generated across the customer lifecycle.
Business Challenge
Despite having access to extensive customer, loan, repayment, and behavioral data, lenders often
struggle to translate this information into actionable insights for strategic decision-making. Key issues
include limited visibility into risk drivers, delayed identification of potential delinquencies, suboptimal
pricing and product structures, and fragmented performance monitoring for leadership.
The central challenge is to design a data-driven approach that enables lenders to understand customer
risk and value holistically, detect early signs of financial stress, optimize growth strategies without
compromising portfolio quality, and support leadership with clear, forward-looking insights.
Problem Statement
How can a digital lending institution leverage end-to-end customer and transaction data to strengthen
credit risk assessment, enable early detection of delinquencies, optimize pricing and product strategies,
and drive sustainable, risk-adjusted portfolio growth?
Geography, income proxy, employment type, Weaker profiles tend to exhibit higher
Customer Profile credit quality indicator risk
Product type, loan size, tenure, pricing, origination Higher risk may result in tighter terms or
Loan & Product risk grade higher pricing
Behavioral Cash flow consistency, balance volatility, Higher volatility may result in higher
Signals spending shocks delinquency likelihood
Time Dimension Origination and repayment periods Newer cohorts may behave differently
Deliverables
The final submission consists of two components: a Credit Policy Recommendation Report and a
Presentation deck.
1. The report should be 4 to 6 pages in length, addressed to a fictional Chief Risk Officer. It must
include an executive summary, risk segmentation findings, early warning signal logic, and a specific
policy recommendation accompanied by a projected impact assessment. This format mirrors what a
strategy analyst at a fintech or a consulting associate on a financial services engagement would be
expected to produce.
2. The presentation deck should complement the report and communicate the team's solution to a
stakeholder audience. Its structure and design are intentionally left to the participants' discretion; we
are more interested in the quality of thinking than in adherence to a prescribed format.
Across both components, the standard being applied is this: the work should lead with strategic
recommendations that are substantiated by data, not data outputs in search of a conclusion.
Point of Contact
Sanchit Pendharkar - 6363679144
Prakhar Gupta - 9235527628
SQL | consulting
Core Problem
The brand has data but no intelligence built on top of it. Specifically, it cannot answer:
• Who are the customers likely to still be buying two years from now, and what do they look like today?
• Is the discount and promo program actually building a loyal customer base, or just attracting one-time
bargain hunters?
• Which product categories are associated with lower purchase history, and which ones appear most
among high-frequency, high-tenure customers?
• Are there cities or regions where the brand has strong traction that it has not yet deliberately targeted,
and does that traction look different in terms of category preference, promo sensitivity, or average spend
per customer?
• What does the brand's best customer actually look like in terms of age, purchase habits, payment
preferences, and satisfaction?
Without answers to these questions, the brand is making marketing and product decisions based on gut
feel rather than evidence.
Problem Statement
Using only transactional and behavioral data, can the brand identify what its most valuable customers
look like, measure how much of its current revenue depends on promotions, and build a data-backed
retention strategy that reduces discount dependency without hurting sales?
Scope of Analysis
1) Data Preparation & Feature Engineering (Python):
Clean and prepare the raw dataset. Then build the customer-level metrics the brand needs — but here
is the constraint: you are not told which metrics to build. You need to decide what to measure and
justify why those choices make sense for this business.
As a starting point, think about what signals in the data could indicate value, satisfaction, and
promotional reliance. But do not stop at computing numbers, explain the logic behind each metric you
create and what it is actually trying to capture.
One thing to keep in mind: metrics that sound analytical but do not lead to a decision are not useful.
Every engineered feature should answer a question the brand actually cares about.
Dataset
Link
SQL | consulting
Expected Outcome
A comprehensive customer intelligence report and interactive dashboard designed to provide the brand
with a clear and actionable answer to a critical strategic question:
“Is the business successfully building a loyal customer base, or is it reliant on continuous promotional
activity? Furthermore, what strategic actions should be taken under either scenario?”
Deliverables
Cleaned dataset + engineered features (dependency score, value
Python
tier, satisfaction flag)
Point of Contact
Achyuth - 6381774762
Dhairya Nisar - 8928149400
Resources
Link
strategy| data analytics
The Challenge
You are part of the Business Analytics team of a mid-sized airline. You will work with data for around
16,700 Canadian loyalty members from 2012 to 2018.
Your goal is to build a solution that helps the marketing team identify customers likely to disengage,
understand which members are most valuable, and recommend actions to improve retention. The main
challenge is not only building a model, but also turning insights into clear business actions.
Core Objectives
1) Churn Prediction
The first objective is to identify members who are likely to stop engaging with the airline in the coming
[Link] dataset contains formal cancellation records, but cancellation is not the only form of churn
that matters. Some customers may stop flying for a long period, while others may remain enrolled in the
loyalty program but stop earning or redeeming points.
You are expected to define what churn means in your analysis and explain the reasoning behind your
definition. There is no single correct answer, but your approach should be logical, realistic, and easy to
defend.
An important challenge in this task is avoiding the use of future information. Your model should only use
data that would have been available at the time of prediction. How you define churn, create features, and
avoid data leakage will be an important part of the evaluation.
3) Smart Retention
For each segment you identify, propose specific interventions. Vague recommendations ("offer bonus
points to at-risk members") are not appreciated. A strong recommendation is one that a non-technical
manager could hand to an operations team tomorrow; it would specify who receives it, when, why, and in
what form.
How you structure the logic connecting churn risk to retention action is an open design problem. Elegant
solutions will be recognized.
Deliverables
1) Working Prototype
A functional interface (any tool: Streamlit, Power BI, Tableau, Excel, or other) that a non-technical
marketing manager can open and use without guidance. It must close the gap between your model's
output and a concrete business action. What that interface looks like, and what it prioritizes, is your
design call.
Standard: A first-time user should be able to identify who needs attention and what to do about it
without reading a manual.
2) Technical Report (6–8 pages)
Cover: how you framed the problem, what cleaning decisions you made and why, the logic behind your
model and segmentation choices, the most interesting things you found, and three business
recommendations that a CFO and a CMO could both get behind. The report is written for a smart non-
technical reader, not a peer reviewer.
Point of Contact
Nalin goel - 8289031644
Keerthana - 9019647256
machine learning | consulting
The Challenge
The strategic question: Delhivery's OSRM system underestimates actual delivery time on a significant
fraction of routes. Can a graph-based model, one that treats the logistics network as a connected graph
of facilities and corridors, not a collection of independent point-to-point estimates, produce more
accurate ETAs and identify which corridors and hubs are systematically causing delays?
When ETA is wrong, SLAs are missed and customers are unhappy
Downstream capacity planning breaks down across the network
There is currently no systematic way to identify which hubs and corridors are the biggest contributors to
delays
Route-type decisions (FTL vs Carting) are made without accounting for graph position or structural risk of a
facility
Your Mission
As a data science team, your goal is to build a graph-based intelligence system for Delhivery's logistics
network. You will model the entire network as a directed graph, facilities as nodes, corridors as edges, and
use this structure to produce smarter ETA predictions, surface bottleneck hubs, and generate actionable
recommendations for the Head of Network Operations.
Your final output is not just a model, it is a consulting deliverable: a strategy memo that a real operations
leader could act on.
Expected Impact
By the end of this project, your solution should help Delhivery's operations team:
Make smarter ETA predictions that reduce the gap between promised and actual delivery time
Identify and prioritize hub upgrades based on their structural risk and SLA breach contribution
Reduce SLA breaches on chronic delay corridors through targeted interventions
Improve route-type decisions with a data-backed FTL vs Carting framework
Recover revenue at risk by quantifying and acting on the highest-impact bottlenecks
Point of Contact
Saksham Gupta - 7009410295
Yashvi Mehta - 6363679144