0% found this document useful (0 votes)
9 views21 pages

Optimizing RAG Applications for Success

The document outlines a structured approach to improving Retrieval-Augmented Generation (RAG) applications through a series of sessions led by Jason Liu. It emphasizes the importance of systematic problem-solving, experimentation, and user feedback in AI development. The course includes a flipped classroom model, focusing on practical applications and continuous improvement of RAG systems.

Uploaded by

Moi ChezMoi
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
9 views21 pages

Optimizing RAG Applications for Success

The document outlines a structured approach to improving Retrieval-Augmented Generation (RAG) applications through a series of sessions led by Jason Liu. It emphasizes the importance of systematic problem-solving, experimentation, and user feedback in AI development. The course includes a flipped classroom model, focusing on practical applications and continuous improvement of RAG systems.

Uploaded by

Moi ChezMoi
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

jxnl.

co
@jxnlco

Systematically Improving
RAG Applications
Session 0
Introduction
Jason Liu
Over the next few weeks, we’re going to dismantle the usual
guesswork in AI development and replace it with something
structured, measurable, and repeatable.

@jxnlco [Link]/applied-llms/rag-playbook
@jxnlco [Link]/applied-llms/rag-playbook
2
Agenda

Who am I?
Who are you?

Who are you?


What does success look like?

What does success look like?


The RAG playbook

The RAG playbook

@jxnlco [Link]/applied-llms/rag-playbook
@jxnlco [Link]/applied-llms/rag-playbook
3
Who am I?
Jason Liu

Data scientist Staff MLE Applied AI Consultant

• Graph and content analysis to • Creation of robust recommendation framework • Strategic consulting to solve problems related
identify cyber crime and observability tools, handing 350M+ to RAG, query understanding, prompt
• “RAG apps to fight internet evil” recommendations per week. engineering, embedding finetuning, MLOps
• In 2017-2018, implemented vision, multimodal observability
search, and VAE-GANs (boosting revenue by • Work with leading companies on AI products
$50M+) such as Limitless AI, Sandbar, Raycast,
• Oversaw a $400k budget for data curation Tensorlake, Dunbar, Bytebot, Naro, Trunk
Tools, New Computer, [Link], Modal
• Implemented upstream search systems for Labs, Pydantic, and Weights & Biases.
similar items, complementary items, outfits,
curated collections

@jxnlco [Link]/applied-llms/rag-playbook
@jxnlco [Link]/applied-llms/rag-playbook
4
Who am I?

Why don’t you build stuff rather than The focus on my career:
be a consultant?”
• In 2018: “Make vision work”
• Hand injury prevents me from
• In 2022: “Make text work”
coding too much

My highest leverage work now is in I have 6+ years of experience:


advising and education
• Building visual product search with embeddings
• Building recommendation systems
• Improving ML-based product reliability
I want to share what I know...

@jxnlco [Link]/applied-llms/rag-playbook
@jxnlco [Link]/applied-llms/rag-playbook
5
Course overview

Week 1 Week 2 Week 3 Week 4 Week 5 Week 6

Kickstart the Data If you’re not fine- The Art of RAG UX: Split: When to Map: Navigating Apply: Function
Flywheel: Fake it till tuning, you’re Turning Design into double down vs. Multimodal RAG Calling Done Right
you make it Blockbuster, not Data when to fold
Netflix

@jxnlco [Link]/applied-llms/rag-playbook
@jxnlco [Link]/applied-llms/rag-playbook
6
Weekly schedule (All times are TBF)

Tuesday
Office hours (1hr Thursday
each) Office hours (1hr each)
9AM EST / 3PM CET / 7:30PM IST / 10PM SGT 1PM EST / 10AM PST / 7PM CET
1PM EST / 10AM PST / 7PM CET 2PM EST / 11AM PST / 8PM CET

Wednesday Friday
Guest lecture (1hr) Following week’s
1PM EST / 10AM PST / 7PM CET lecture released on
Maven
@jxnlco [Link]/applied-llms/rag-playbook
@jxnlco [Link]/applied-llms/rag-playbook
7
What is a flipped classroom?
A flipped classroom is a pedagogical approach where traditional lecture and homework elements are
reversed.

How will our flipped classroom work?


• Pre-work: On Fridays, lecture content is released through videos that students can engage with
• Office hours: On Tuesdays and Thursdays, students will ask questions about lecture material,
specific problems they are facing related to implementing their changes, and any other questions
• Guest lectures: On Wednesdays, we will host guest lectures relevant to that week's topic

What are the benefits of a flipped classroom?


• Increased active learning time through office hours
• More personalized guidance from instructors
• Improved student engagement and cross-pollination
• Self-paced learning opportunities

@jxnlco [Link]/applied-llms/rag-playbook
@jxnlco [Link]/applied-llms/rag-playbook
8
Agenda

Who am I?

Who am I?

Who are you?


What does success look like?

What does success look like?


The RAG playbook

The RAG playbook

@jxnlco [Link]/applied-llms/rag-playbook
@jxnlco [Link]/applied-llms/rag-playbook
9
@jxnlco [Link]/applied-llms/rag-playbook
@jxnlco [Link]/applied-llms/rag-playbook
10
@jxnlco [Link]/applied-llms/rag-playbook
@jxnlco [Link]/applied-llms/rag-playbook
11
Agenda
Who am I?

Who am I?
Who are you?

Who are you?


What does success look like?

What does success look like?

Build a “system” to improve RAG


The RAG playbook

The RAG playbook

@jxnlco [Link]/applied-llms/rag-playbook
@jxnlco [Link]/applied-llms/rag-playbook
13
What is a system? Why do we need a system?

A structured approach to solving problems that guide how we Systems free up mental energy for innovation and for
think about and tackle challenges tacking unique challenges

In the context of building an AI product: • Companies often waste time guessing instead of testing
new technologies.
1 A framework for evaluating different technologies and
• Quantitative results and qualitative assessments are
tools
both necessary.

2 A decision-making process for prioritizing development • Securing resources for developing new capabilities is
efforts challenging without being data-driven.

3 A methodology for diagnosing and improving application


performance

4 A set of standardized metrics and benchmarks to


measure success
@jxnlco [Link]/applied-llms/rag-playbook
@jxnlco [Link]/applied-llms/rag-playbook
14
What does success look like? Where have I done this?

Personal agents

A At the next stand-up, have lower anxiety


when told “make the AI better” Construction

B Feel less overwhelmed when told “look at the data”,


instead leverage data to: Workflow automations
• Identify high-impact tasks
• Prioritize effectively to make informed tradeoffs
• Choose relevant metrics that correlate with Sales and marketing
business outcomes

Drive better outcomes: Transcript and data mining


C
• Higher user satisfaction
• Higher retention
Private equity / due diligence
• Higher and more frequent usage
@jxnlco [Link]/applied-llms/rag-playbook
@jxnlco [Link]/applied-llms/rag-playbook
15
What does the path to success look like?

A I focus on experimentation rather than a vague “make the AI better”

B When I hear “look at the data”, I ask


• Why am I looking at my data?
• What am I looking for?
• When I see the signals:
• What can I do with this information to improve my application?
• Is the juice worth the squeeze?

C Once the flywheel is in place, success is defined by consistency - doing the most obvious thing over and over
• This will be boring… because a better app is a result, not an effort
• Your only job is to apply consistent effort
• For example, if you want to build muscle, you track calories and workouts (which is boring)

@jxnlco [Link]/applied-llms/rag-playbook
@jxnlco [Link]/applied-llms/rag-playbook
16
Agenda
Who am I?

Who am I?
Who are you?

Who are you?


What does success look like?

What does success look like?


The RAG playbook

The RAG playbook


RAG vs. Recsys

@jxnlco [Link]/applied-llms/rag-playbook
@jxnlco [Link]/applied-llms/rag-playbook
18
RAG is more than just R-A-G
Engineers focus on
generation without even
What you think RAG is: knowing if the right
information is retrieved

Retrieval Prompting Generation


Only through improved
search can you improve
retrieval which will
What RAG actually is: ultimately improve
generation

Retrieval index
L L The LLM chooses from
L Retrieval index Filtering Scoring Ordering L different search backends
M M based on the query
Retrieval index
Retrieval is not only multi-
Good Retrieval in RAG resembles a four-step recommendation model step, but also multi-index

@jxnlco [Link]/applied-llms/rag-playbook
@jxnlco [Link]/applied-llms/rag-
[Link]/applied-llms/rag-playbook
playbook 19
Instead of implementing R-A-G, implement RecSys in an LLM sandwich

Retrieval index
Input L L Output
L Retrieval index Filtering Scoring Ordering L
UI M M UI
Retrieval index

Step 1: Step 2: Step 3: Step 4:


• Chat interface or web UI • Based on original routing • Produce representations • Build a special UI to
intakes original query steps, retrieval, filtering, for the chunks render responses beyond
• Prepare a list of queries scoring, and ordering is • Choose a generation text streams
that are routed to specific done prompt to reply to the • For example: profile cards,
search engines LLM’s pdf viewers, citations,
o Planning, CoT, Parallel feedback buttons
Tool use

@jxnlco [Link]/applied-llms/rag-playbook
@jxnlco [Link]/applied-llms/rag-playbook
21
Example

@jxnlco [Link]/applied-llms/rag-playbook
@jxnlco [Link]/applied-llms/rag-playbook
22
The RAG Flywheel
Thesis: The principles we’ve applied in search are highly relevant to what we want to do with RAG
Step Description
1 Initial implementation Start with a basic RAG system setup

2 Synthetic data generation Create synthetic questions to test the system’s retrieval abilities

Conduct quick, unit test-like evaluations to assess basic retrieval capabilities (e.g., precision,
3 Fast evaluations
recall, mean reciprocal rank), and explain why each matters

4 Real-world data collection Gather real user queries and interactions. Ensure feedback is aligned with business outcomes
or correlated with important qualities that predict customer satisfaction

5 Classification and analysis Categorize and analyze user questions to identify patterns and gaps

6 System improvements Based on analysis, make targeted improvements to the system

7 Production monitoring Implement ongoing monitoring to track system performance

8 User feedback integration Continuously incorporate user feedback into the system
@jxnlco [Link]/applied-llms/rag-playbook
@jxnlco [Link]/applied-llms/rag-playbook
23
The RAG Playbook process

Step 1: Address cold start Step 2: Instrument and observe user data Step 3: Implement new routers, tools, and
a. Generate synthetic data to identify strengths and weaknesses in capabilities
b. Establish general baselines for: clusters a. Evaluate experiments for new
i. Recall a. Implement UI to collect user data tools/indices against both offline
to start fine-tuning and product metrics
ii. Precision
b. Use topic modeling and identify b. Explore and Exploit given
iii. Lexical search
missing topics and capabilities with distributions of capabilities
iv. Semantic search domain experts c. Iterate and monitor continuously
v. Re-rankers c. Target underperforming, high-
c. Use this to test all other impact query segments
hypotheses Rinse and repeat:
d. Convert offline analysis to online
a. Continuously conduct exploratory
classifiers and query routers to
data analysis, partner with users
monitor production data
and domain experts, and extend
capabilities
b. Understand and improve the
process of evaluating candidate
selection
@jxnlco [Link]/applied-llms/rag-playbook
@jxnlco [Link]/applied-llms/rag-playbook
24

You might also like