100% found this document useful (1 vote)
62 views83 pages

Evolution from GenAI to Agentic AI

Generative AI (GenAI) has evolved into Agentic AI systems, which are capable of autonomous action and goal pursuit, moving beyond static content generation. These intelligent agents utilize large language models (LLMs) to reason, plan, and interact with their environments, allowing for dynamic problem-solving and adaptability. The transition signifies a fundamental shift in AI capabilities, unlocking new applications and efficiencies across various sectors.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd
100% found this document useful (1 vote)
62 views83 pages

Evolution from GenAI to Agentic AI

Generative AI (GenAI) has evolved into Agentic AI systems, which are capable of autonomous action and goal pursuit, moving beyond static content generation. These intelligent agents utilize large language models (LLMs) to reason, plan, and interact with their environments, allowing for dynamic problem-solving and adaptability. The transition signifies a fundamental shift in AI capabilities, unlocking new applications and efficiencies across various sectors.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd

Understanding Generative AI: From

Static to Dynamic
Defining GenAI Transforming Applications

Generative AI, or GenAI, encompasses advanced GenAI marks a significant leap from traditional static
systems capable of generating novel content. This input-output applications. It enables the creation of
includes a wide array of media such as text, images, dynamic systems that possess the ability to reason
music, and more. These systems leverage large-scale within a given context. This allows for more
machine learning models like GPT-4 and Claude to sophisticated and adaptive responses, moving beyond
create content that can be contextually relevant and simple predefined outputs to truly generated and
highly creative. original content.
The Inherent Limitations of Traditional
GenAI
While powerful, traditional GenAI models, when used in isolation, exhibit several key limitations that prevent them from
operating autonomously or persistently. These constraints highlight the necessity for more advanced architectures that
integrate memory, iterative action, and environmental interaction.

No Goal Memory GenAI systems do not inherently remember past goals or


adjust their behavior based on a historical sequence of
actions. Each query is treated as an independent

One-Shot Responses instance.


These models are primarily designed to provide a single,
immediate response to a prompt. They are not structured
to act sequentially over extended periods or across
multiple related tasks.
No Environment Interaction Traditional GenAI cannot directly operate external tools,
interact with APIs, or perceive and manipulate real-world
or digital environments. Their functionality is confined to
language generation.
The Emergence of Intelligent AI Agents
The paradigm shift to AI agents addresses the limitations of standalone GenAI. These agents represent a new
frontier in AI, moving beyond simple content generation to embody more complex cognitive functions, enabling
them to navigate and perform actions within dynamic environments.

Cognitive LLM-Powered Dynamic


Functionality
Agents don't just answer; they Evolution
Inspired by traditional Interaction
These agents can maintain
are designed to think, plan, cognitive agents and software context, make autonomous
act, and adapt based on their bots, modern AI agents are decisions, interact with
observations and goals. now supercharged by Large various tools and APIs, and
Language Models (LLMs), adjust their behavior based on
giving them advanced real-time feedback from their
reasoning capabilities. environment.
Key Characteristics of Agentic AI Systems
Agentic AI systems are defined by their ability to operate with a degree of independence and pursue specific objectives.
This autonomy, combined with their capacity for iterative refinement, sets them apart from previous AI generations.

Autonomy Goal-Driven Behavior


The ability to operate independently without constant Agents are designed to pursue and achieve specific
human intervention, making decisions and executing goals, often breaking down complex tasks into smaller,
actions based on internal reasoning and environmental manageable sub-goals. This contrasts with one-shot
feedback. query responses.

Environmental Interaction Learning and Adaptation


Unlike static GenAI, agents can interact with external Agentic systems can learn from their experiences,
tools, APIs, databases, and even physical environments, adapting their strategies and knowledge over time to
extending their capabilities beyond text generation. improve performance and achieve goals more
effectively.
The Agentic Loop: A Continuous Cycle of
Intelligence
The core of an Agentic AI system lies in its continuous operational loop, which allows for dynamic interaction and iterative
refinement. This loop mirrors human cognitive processes, enabling agents to learn and adapt over time.

Planning
Observation
Based on observations and goals, the
The agent perceives its environment,
agent formulates a strategy and
gathers information, and interprets
sequences of actions.
input.

Reflection
Action
The agent evaluates the outcome of its
The agent executes the planned actions,
actions and updates its internal state or
often interacting with tools or APIs.
knowledge for future planning.

This iterative cycle ensures that agents can operate effectively in complex and changing environments, continuously improving
their performance and problem-solving abilities.
Why Autonomy is Critical for Next-Gen AI
The move towards autonomous and goal-driven AI systems is not merely an incremental improvement; it is a fundamental shift that unlocks
unprecedented capabilities and addresses the growing complexity of modern challenges. Autonomy enables AI to tackle problems that require
sustained effort and dynamic decision-making.

Solving Complex
Problemsagents can break down intricate tasks and execute multi-step solutions, far beyond what single-shot GenAI can
Autonomous
achieve.
Operational
Efficiency
By acting independently, agents reduce the need for constant human oversight, leading to greater efficiency and scalability in
various applications.

Enhanced
Adaptability
Goal-driven behavior allows agents to adapt to new information or changing conditions, ensuring robustness in dynamic
environments.
Unlocking New
Applications
Autonomy opens doors to novel AI applications in areas like scientific discovery, personalized education, and advanced robotics,
where continuous action and learning are vital.
Key Takeaways and Future Outlook
The transition from isolated Generative AI to integrated Agentic AI systems represents a significant leap forward in
artificial intelligence. This evolution empowers AI to move beyond mere content generation towards proactive
problem-solving and autonomous operation.

Looking ahead, Agentic AI is poised to revolutionize various sectors by enabling more intelligent, adaptive, and
efficient systems. As these systems become more sophisticated, they will redefine the boundaries of what AI can
achieve, paving the way for truly autonomous and beneficial AI in the future.

From Static to Interaction & Future of AI


Agentic
GenAI transformed content creation; Adaptation
Key to agents are their ability to Autonomy and goal-driven behavior
Agentic AI empowers autonomous interact with tools and adapt to are critical for the next generation of
action and goal pursuit. environments. intelligent systems.
Unlocking Autonomy: The
Evolution of Agentic AI
Systems
This presentation delves into the transformative journey from
traditional Generative AI (GenAI) to the sophisticated realm of
Agentic AI systems. We will explore the inherent limitations of
isolated GenAI, the compelling emergence of intelligent agents,
and the paramount importance of autonomy and goal-driven
behavior in shaping the next generation of artificial intelligence.
The Dawn of Agentic
AI: Autonomous
Systems Reshaping the
Discover how agentic AI systems are revolutionizing automation,
Future
moving beyond reactive responses to deliver goal-oriented
autonomy. This presentation delves into their core mechanics, the
factors driving their emergence, and their profound business
implications.
Anatomy of Agentic
Systems
Agents are autonomous goal-seeking systems powered by LLMs. They can
use tools, retain memory, and reason over time.

1 Perceive 2 Plan Actions 3 Execute via Tools


Environment
Inputs and current state inform the agent. LLMs guide the agent's next steps. Agents interact with external systems.

4 Store/Recall 5 Adapt Through


Memory
Relevant information is retained for future use. Loops
Continuous feedback refines agent behavior.
Why Agentic AI
Now?
LLM Maturity Tool Developer
GPT-4 and Claude Integration
Frameworks like Accessibility
offer advanced LangChain enable Python and low-code

reasoning across vast agents to invoke tools platforms simplify

contexts. and share tasks. agent development.

Business
Impact
Rising adoption in
research, customer
service, and
productivity roles.
Evolution of Automation: From Scripts to
Agents Manual 1
Automation
Early scripts and macros handled repetitive tasks.

2 RPA Platforms
Robotic Process Automation streamlined workflows.

iPaaS 3
Integration Platform as a Service connected diverse systems.

4 LLM-Powered
GenAI
Generative AI began to understand and create content.

Agentic AI 5
Systems
Current systems reason, plan, and use tools autonomously.

This evolution highlights why older automation architectures struggle to meet today’s expectations for intelligence and adaptability.
GenAI vs. Agentic AI: A
Comparison
Generative AI
(GenAI)
• Reactive and content-focused.
• Stateless, answers queries directly.
• Output is typically text or images.

Agentic AI
• Goal-oriented and autonomous.
• Plans and acts in a continuous loop.
• Interacts with tools and users.

While GenAI answers, an agent plans and acts in a loop of reflection,


memory, and interaction with tools and users.
Types of AI
Agents
LLM-Enhanced
Agents
Utilize large language models for core reasoning.

Tool-Enabled
Agents
Integrate with external tools for diverse actions.

Memory-Based
Agents
Retain and recall information for context.

Self-Reflecting
Agents
Evaluate their own performance and refine actions.

Environment-Controlling
Agents
Actively modify and interact with their surroundings.

Each type offers specific capabilities for different use cases, leading to a diverse and powerful agent ecosystem.
Future Vision: Digital
Collaborators
Beyond
Assistants
AI agents will move beyond simple tasks.

Autonomous Problem-
Solving
They will tackle complex challenges independently.

Task Chaining
Agents will chain multiple steps to achieve goals.

Seamless
Integration
Working alongside humans in various roles.

This shift fundamentally changes how we interact with technology,


making AI a true partner in innovation.
Key Takeaways & Next
Steps
Agentic AI's Digital
Autonomy Collaborators
Goal-oriented and task- Agents become partners,
chaining capabilities. not just tools.

Strategic
Adoption
Embrace agentic systems for future growth.

Explore further resources and consider how agentic AI can transform


your operations.
The Evolution of
Automation: From
Scripts to Agents
This presentation examines the fascinating journey of automation,
from its humble beginnings with simple scripts to the cutting-edge
realm of intelligent, self-adapting AI agents. We'll explore the
limitations of traditional fixed automation and delve into how Large
Language Models (LLMs) have revolutionized the field by
introducing dynamic, contextually aware capabilities. Join us as we
uncover the transformative shift that is reshaping industries and
redefining how we interact with technology.
Historical Context: How Automation
Began
Script-Based Robotic Process iPaaS (Integration
Automation Automation (RPA) Platforms as a Service)
Early automation relied on simple
rule-based scripts, such as shell RPA emerged to mimic user actions Platforms like Zapier and Workato

scripts or macros, designed to on screens, such as clicking took automation a step further by

repeat predictable tasks. While fast buttons or reading PDFs. Tools like chaining APIs and cloud tools. This

and low-cost, these scripts offered UiPath and Automation Anywhere allowed for rule-driven automation

no flexibility and would often break enabled automation of semi- workflows across various services

when faced with unexpected structured processes, but they like Gmail, Google Sheets, and

inputs, limiting their utility to remained stateless and fragile, Slack, connecting disparate

highly structured environments. unable to adapt to minor changes systems but still operating on
in interfaces or workflows. predefined rules.
What Changed with Large Language
Models?
When Large Language Models (LLMs) entered the scene, automation experienced a paradigm shift, elevating capabilities to an
unprecedented level. LLMs introduced contextual understanding, the ability to chain multiple tasks, and a natural language
interface, fundamentally altering how automation systems operate.

Reasoning Rule-based Reasoning-based

Workflows Hardcoded Dynamic task planning

Learning No learning Contextual adaptation

Task Execution One task per bot Multistep task chains


Enter the Age of AI
Agents
AI Agents build upon the foundation of LLMs by integrating a suite of advanced capabilities, enabling them to
operate autonomously and intelligently. These agents are no longer just executing commands; they are capable of
complex problem-solving.

Planning Tool Calling


Utilizing advanced logic frameworks like ReAct or Seamlessly interacting with external APIs,
Tree-of-Thoughts to strategize and break down databases, and other software tools to gather
complex problems into manageable steps. information or perform specific actions.

Long-Term Reflection &


Memory
Employing memory stores (e.g., LangChain, Retryingtheir own outputs and processes based
Evaluating
CrewAI) to retain context, learned information, on task feedback, allowing them to self-correct
and past interactions for improved performance and improve their strategies iteratively.
over time.
Applications of AI
Agents
These sophisticated AI agents are rapidly transforming various sectors by replacing traditional workflows
and offering unprecedented levels of assistance. Their ability to understand context, plan, and adapt
makes them invaluable across diverse applications.

Customer Service R&D Assistance Education


Replacing rigid scripted Empowering research and Creating personalized
flows with dynamic, development teams by learning experiences,
empathetic responses automating multi- adapting content, and
that understand customer document search, providing tailored
intent and provide summarization, and data feedback to students
personalized solutions. synthesis, accelerating based on their individual
discovery. progress and needs.

Healthcare
Developing intelligent
assistants for patient
support, data analysis,
and administrative tasks,
enhancing efficiency and
patient care.
Example Transition
Tasks
The shift from traditional script-based automation or RPA to agentic AI is best
illustrated through how common tasks are now handled. Agents bring a level of
intelligence and adaptability that previous automation methods simply couldn't
achieve.
Respond to
Emails
Script/RPA: Regex + template

Agentic AI: Understand query, summarize docs, write contextual reply

Generate Report
Script/RPA: Schedule Excel script

Agentic AI: Pull data, interpret trends, format summary

Book a Meeting
Script/RPA: iCal + reminder

Agentic AI: Negotiate time, reschedule conflicts, follow-up

This evolution represents a move from mere execution to intelligent decision-making


and proactive problem-solving, dramatically improving efficiency and user experience.
Automation Maturity
Workflow
The journey through automation maturity can be visualized as a progression through distinct stages, each building
on the capabilities of the last. This model highlights the increasing levels of intelligence and adaptability in
automation systems, as detailed in "Mastering AI Agents" (Figure 1.2 and 1.3).

ReAct Agent
LLM-Enhanced
Loop of Reasoning ↔ Action until
Fixed Automation Agent
Input → Contextual Reasoning → outcome
Agent
Trigger → Rule-based Action → Output Represents a significant leap, where
Output More flexible due to contextual agents continuously reason, take
The most basic form, where a understanding, but still constrained action, and reflect on the results to
predefined trigger leads to a fixed, by underlying rules or limited achieve a desired outcome.
rule-based action, producing a adaptability.
predictable output.
Summary of the
Shift
The transition in automation is profound, moving from rigid, pre-programmed instructions to systems capable of understanding intent
and adapting dynamically. This fundamental change promises to unlock new levels of efficiency and innovation across all industries.

From If-Then rules to Intent-Driven autonomy.

From macro-level scripting to micro-cognitive task handling.

As Galileo famously noted, "Agents don’t just automate—they think, plan, and adapt." This marks a new era where machines can
collaborate with humans in more intelligent and intuitive ways, redefining the future of work. The capabilities of AI agents are
continuously expanding, promising even greater advancements in the years to come.
Unveiling AI Agents:
Definitions and
Architectures
Welcome to Chapter 3, where we delve into the fundamental
concepts of AI agents. This presentation will define what
constitutes an AI agent, explore their core characteristics, and
examine the sophisticated architectures that enable them to
operate autonomously. Understanding these building blocks is
crucial for anyone looking to grasp the next generation of AI
systems.
Defining the AI Agent: Core
Traits
Perception Reasoning Action Memory & Learning

AI agents actively They process perceived Agents execute actions They utilize memory to
perceive their information, reason within their environment store past experiences
environment, gathering about their observations, to achieve predefined and learn from
relevant information and determine objectives, transforming outcomes, continuously
through various sensors appropriate actions their understanding into improving their
or data inputs. based on their goals. tangible outcomes. performance and

An AI agent is a sophisticated system designed to interact with its environment [Link]-making over
Unlike traditional
time.
generative AI that provides one-shot responses, agents operate in continuous feedback loops, constantly refining
their behavior. This continuous cycle of perception, reasoning, action, memory, and learning is what differentiates
true AI agents, enabling them to adapt and evolve autonomously.
The Core Agent Cycle: A Continuous Loop
Input
The agent receives data from its environment or user commands.

Perception
Raw input is processed to understand the current state of the environment.

Memory Recall & Reasoning


Relevant information is retrieved from memory, and decisions are made based on goals and perceived state.

Action
The agent executes a chosen action, impacting the environment.

Feedback & Memory Update


Outcomes of actions are observed, and memory is updated to learn from the experience, closing the loop.

Output
The final result or modified environment state is presented.

The operational heart of an AI agent lies in its continuous cycle. This iterative process begins with an input, leading to perception, followed by memory recall and sophisticated
reasoning. The agent then takes action, observes feedback, and updates its memory to refine future actions. This cyclical nature, depicted in Figure 3.1 from Mastering AI
Agents, is crucial for their ability to learn and adapt, making them far more dynamic than traditional generative AI models.
Evolution of AI Agent Types
Reflex Agents 1
React solely to the current input without maintaining an internal state
(stateless).
2 Model-Based Reflex
Maintain a simple internal state or model of the world to inform
decisions.
Goal-Based 3
Use explicit goals to guide their behavior, planning actions to achieve
specific objectives.
4 Utility-Based
Optimize decisions by evaluating potential actions based on a utility
function, aiming for the best outcome.
Learning Agents 5
Adapt and improve their performance over time by learning from
feedback and environmental changes.

The evolution of AI agents, as beautifully detailed in Artificial Intelligence: A Modern Approach, Chapter 2, showcases a progression from simple reactive systems
to highly adaptive learning machines. Starting with basic reflex agents that only respond to immediate stimuli, the journey moves through model-based agents
that incorporate internal state, to goal-based agents driven by objectives, and utility-based agents that optimize decisions. The pinnacle of this evolution is the
learning agent, capable of continuous adaptation and self-improvement based on real-world interactions and feedback.
Self-Learning Architectures for Advanced Agents
Reasoning
Perception
Processing information to make
Observing and understanding the
decisions.
environment.
Action
Executing decisions and interacting
with the environment.

Evolution
Adapting and improving based on Feedback
feedback, refining the internal model. Receiving results and evaluating
performance.

The most advanced AI agents incorporate sophisticated self-learning architectures that allow them to continuously improve. This
involves a cyclical process where agents perceive their environment, reason about it, take actions, receive feedback on those
actions, and then use that feedback to evolve their internal models and strategies. This iterative self-improvement, often depicted as
a "Feedback-Reason-Act-Evolve" loop, is central to creating truly autonomous and intelligent systems capable of mastering complex
Specialized Agent Architectures
1 ReAct Agent 2 ReAct + RAG Agent
Combines Reasoning and Acting in a tight, iterative Integrates the ReAct loop with Retrieval-Augmented
loop, allowing for intermediate evaluations and dynamic Generation (RAG), enabling agents to query external
task adaptation. This is crucial for tasks requiring careful knowledge bases for enhanced factual grounding,
thought before execution. particularly vital for complex, knowledge-intensive
tasks.
3 Tool-Enhanced Agent 4 Self-Reflecting Agent
Designed to dynamically select and execute external Utilizes a "Reflexion" technique, allowing the agent to
tools via APIs, significantly expanding an agent's analyze its past actions and revise its plans accordingly.
capabilities and interaction possibilities with various This meta-cognition leads to more robust and reliable
systems and data sources. autonomous behavior.
Beyond the core agent cycle, various specialized architectures empower AI agents to tackle more intricate challenges. The ReAct
agent, for instance, tightly couples reasoning and acting, facilitating dynamic adjustments. Adding RAG capabilities allows for
robust factual grounding by integrating external knowledge. Tool-enhanced agents can dynamically interact with external
systems, while self-reflecting agents demonstrate meta-cognition by learning from their own past actions, leading to advanced
autonomous capabilities.
Key Takeaways: The Essence of AI
Agents
Beyond Chatbots Feedback-Driven Loops
AI agents are sophisticated, structured systems that Their continuous operation relies on perception,
operate with intent, unlike simpler conversational memory, reasoning, and iterative feedback to learn
models. and adapt.

Tool Utilization Modular & Evolving Architectures


Agents integrate and employ various tools and APIs to Their designs are layered and adaptable, allowing for
extend their capabilities and interact with the real continuous improvement and integration of new
world. functionalities.

In summary, AI agents are far more than mere chatbots. They are intelligent, structured systems built on core principles
of perception, memory, and reasoning. Their effectiveness stems from operating within continuous, feedback-driven
loops, allowing them to learn and refine their behavior. Furthermore, their ability to utilize external tools and their
modular, ever-evolving architectures position them as powerful, adaptable systems capable of autonomously navigating
and impacting complex environments. This foundational understanding is vital for appreciating their potential across
Looking Ahead: The
Future of Autonomous
Systems
As we conclude our exploration of AI agent definitions and
architectures, it's clear that these systems represent a significant
leap forward in artificial intelligence. Their ability to perceive,
reason, act, and learn positions them at the forefront of innovation,
promising transformative applications across industries.
The continued development of more sophisticated architectures,
enhanced learning mechanisms, and seamless tool integration will
unlock unprecedented levels of autonomy and intelligence. We
encourage you to delve deeper into the practical applications and
ethical considerations surrounding AI agents, as their impact on
our future will undoubtedly be profound.
Building AI Agents:
Foundational Technologies
and Practical Examples
This presentation delves into the core technologies that power AI agents,
exploring how Large Language Models (LLMs), APIs, external tools, and
memory systems are integrated to create intelligent, autonomous systems.
We will use the "Fred the Finance Agent" example from Mastering AI Agents
to demonstrate a hands-on, working AI agent. This agent accepts a
question, breaks it down into sub-questions, utilizes external tools like web
search, collects results, plans a final answer, and evaluates its own quality.
The AI Agent Workflow: From Question to Evaluation
Accept Question
The agent receives an initial query or task from the user.

Break Down
Complex questions are deconstructed into smaller, manageable sub-questions or steps.

Utilize External
Tools
The agent leverages APIs and web search to gather necessary information.

Collect Results
Information retrieved from tools is compiled and processed.

Plan Final Answer


Based on collected data, the agent formulates a comprehensive response.

Self-Evaluate
The agent assesses the quality and accuracy of its own answer.

This structured workflow ensures that AI agents can tackle complex problems systematically, mimicking human problem-solving approaches. By breaking down tasks and utilizing external
resources, agents can provide more accurate and comprehensive solutions.
Defining the Agent: Introducing
Fred the Finance Agent
Agent Name: Fred
Our illustrative AI agent for this demonstration.

Role: Financial Research


Assistantin providing insights for investment decisions.
Specializes

Core Technology: ReAct


Agent
Built using LangChain, LangGraph, OpenAI, and Tavily tool integration.

Key Tools:
• Web Search (Tavily)
• Memory for past interactions
• GPT-4 for advanced reasoning

Fred is designed to simulate the capabilities of a human financial analyst, but with the speed and
access to vast amounts of data that only an AI can provide. Its integration with specialized tools
allows for dynamic and context-aware information retrieval.
Building the Plan: Structured Multi-Step
Breakdown
Initial Query
"Should I invest in Tesla given the current EV market?" This is the starting point for Fred's analysis.

Planning Prompt Template


Fred uses a specialized prompt to ensure a logical breakdown of the question.

Step 1: Stock
Performance
Search for historical and current Tesla stock performance data.

Step 2: Market
Trends
Analyze the broader Electric Vehicle (EV) market trends and outlook.

Step 3: Competitor
Analysis
Compare Tesla's position against key competitors in the EV sector.

This systematic approach, guided by the PlannerPrompt, ensures that Fred addresses all relevant aspects of the user's query comprehensively,
much like a human expert would.
State Management: Maintaining Context and
Progress

Original Query Current Step


The initial question that triggered the The specific action or sub-task the agent
agent's process. is currently executing.

Final Answer Completed Steps +


The conclusive response generated after Results
A log of all actions performed and their
processing all steps. corresponding outcomes.

Effective state management is crucial for complex AI agents. It allows Fred to keep track of its progress, refer back to previous steps, and
maintain coherence throughout the multi-step reasoning process. This dynamic state graph ensures efficient and accurate execution.
Action Execution: The ReAct Loop in
Action
Input
The agent receives the user's query or a sub-task from the planner.

Plan
Fred uses its reasoning capabilities (GPT-4) to determine the next logical action.

Execute
The agent performs the planned action, e.g., searching Tavily or accessing internal memory.

Summarize Results
Information retrieved is summarized and integrated into the context.

Replan / Next Step


Based on the results, Fred either refines its plan or moves to the next defined step.

The ReAct (Reasoning and Acting) loop is fundamental to how Fred operates. It allows the agent to iteratively reason about its environment, take actions, and
update its internal state, leading to a dynamic and adaptive problem-solving process.
Evaluation: LLM-as-Judge and Galileo AI Integration
Supervisor: GPT-4o Logging & Visualization: Galileo AI

A sophisticated LLM acts as an impartial judge, evaluating Fred's responses. All evaluation steps and scores are meticulously logged and visualized for
performance analysis.
• Context Adherence: How well the response stays on topic.
• Accuracy: Precision and correctness of the information provided.
• Completeness: Whether all aspects of the query are addressed.

This self-evaluation mechanism is a cornerstone of advanced AI agents, allowing for continuous improvement without constant human intervention. The integration with
Galileo AI provides crucial insights into the agent's performance, enabling developers to fine-tune its capabilities and ensure reliable outcomes.
Summary: The Strategic,
Searchable, Self-Aware Agent
Strategic Thinking External Tool Utilization
AI agents can break down complex Agents can effectively use APIs, web
problems and devise multi-step plans, search, and other external tools to gather
mimicking human cognitive processes. and process information.

Self-Evaluation & Improvement


The ability to evaluate their own performance and learn from outcomes drives continuous
enhancement.

"Fred is not just smart—he’s strategic, searchable, and self-aware."

This chapter demonstrates that AI agents are evolving beyond simple task execution to become
truly intelligent entities capable of complex reasoning, resource utilization, and autonomous
improvement. The foundational technologies discussed are pivotal in building the next generation
of AI applications.
Frameworks for AI Agents:
LangGraph, AutoGen,
CrewAI
This chapter provides a comprehensive, side-by-side comparison of
leading AI agent development frameworks and large language
models (LLMs).

We will explore tool compatibility, agent loop flexibility, memory


support, and real-world industry use cases to guide your selection
process.

by SADANAND
RAJPUROHIT
Leading Agent-Capable LLMs in 2024–
25
OpenAI (GPT-4/GPT- Anthropic (Claude 3 DeepSeek (Open Source
4o) Opus/Haiku) LLM + Tools)
Powerful, versatile, and strong in
reasoning. Excellent for complex Known for its long context window An open-source alternative offering

tasks and tool integration. Widely and safety. Ideal for applications cost-efficiency and transparency.

adopted with robust developer requiring extensive conversational Great for experimentation and

support. memory and ethical custom deployments where


considerations. flexibility is key.

Choosing the right base model for your agent depends on specific needs, including tool integration, reasoning
strength, context length, cost, and open-source preferences.
Side-by-Side Comparison of LLM
Features
Context Window 128K 200K 128K

Tool Calling (functions) ✅ Native & flexible ✅ Via API prompts ✅ Experimental, code-based

JSON Structured Output ✅ Best-in-class ✅ Strong ⚠️Needs post-processing

Memory ✅ Available (ChatGPT Pro) 🟡 Roadmap/limited ❌ Manual only

Open Source? ❌ No ❌ No ✅ Yes (MIT Licensed)

This table highlights key differences in context window size, tool calling capabilities, JSON output quality, memory support, and open-
source availability across the three leading LLMs.
Key Capabilities: Tool Use &
Planning
OpenAI GPT-4o
Natively calls tools via Assistants API and LangChain. Supports
ReAct, tool loops, and self-evaluation for robust planning.

Claude 3
Performs well with API wrappers. Excellent for language-heavy
reasoning, but lacks integrated function calling for tools.

DeepSeek
Can call tools using LangChain/OpenDevin-style frameworks.
Lacks a seamless abstraction layer for integrated tool use.

Understanding how each LLM handles tool use and planning loops is crucial
for designing effective and efficient AI agents.
Key Capabilities: Contextual Memory
Systems
Short-Term Memory Long-Term Memory
Retains recent interactions for Stores persistent knowledge over
immediate context. Essential for extended periods. Enables agents
fluid conversational flows. to learn and recall past
information.

Contextual Memory
Entity Memory
Combines various memory types
Tracks specific entities like users
for comprehensive understanding.
or objects. Helps maintain
Supported by LangGraph and
consistent understanding of key
CrewAI.
subjects.
Effective memory management is vital for agents to maintain context across interactions, especially in complex and
long-running tasks.
Cost & Performance Optimization
Strategies
Model Cache
Reduce re-inference costs by caching previously generated responses for repeated
queries.

Small Model API


Utilize smaller, cost-effective models like DeepSeek for simple and routine tasks, minimizing token
usage.

Serverless Strategy
Implement smart routing with Lambda-based architectures to dynamically select the
optimal model based on task complexity.

Medium/Large Model
API
Reserve powerful models like GPT-4 or Claude only for complex
logic and critical reasoning tasks.

Optimizing AI agent deployment involves strategic model selection and efficient architectural patterns to balance performance with cost-
effectiveness.
Summary Evaluation
Table
GPT-4o Claude DeepSeek
Best for advanced Best for large Best for open-source
reasoning and multi- documents and safe projects and
tool integration. reasoning. Consider learning. Consider if
Consider if you need if you are building you prioritize
high accuracy and long Q&A agents or transparency, low
robust plugin calls in require strict safety cost, and hands-on
your agent. protocols. experimentation.
Each LLM offers distinct advantages, making the choice dependent
on the specific requirements and constraints of your AI agent
project.
Choosing Your Agent's
Engine
1 2
OpenAI: The Ferrari Claude: The Lexus
Powerful, high-performance, and excellent Smooth, safe, and reliable. Ideal for
for development. However, it comes with applications requiring stability and
a higher cost. extensive conversational memory.

3
DeepSeek: The DIY Tesla
Affordable, flexible,Kit
and hands-on. Perfect
for experimentation, learning, and cost-
conscious deployments.

Prototype with GPT-4o for its development support, test Claude 3 for memory-intensive
applications, and utilize DeepSeek for cost-effective experimentation.
Mastering AI Agents:
Build and Deploy on a
Budget
Welcome to a comprehensive guide on building powerful AI agents
without breaking the bank. This presentation is designed for
aspiring AI enthusiasts, budget-conscious entrepreneurs, and
professionals looking to transition into the exciting world of agentic
AI. We'll provide a zero-to-one roadmap, focusing on practical skills
and cost-effective strategies to help you become proficient in AI
agent development.
Your Roadmap to AI Agent
Proficiency
Concept & Theory Build & Develop Evaluate & Refine Deploy & Scale
Understand the foundational Implement agents using Test agent performance and Launch agents for real-world
principles of AI agents and frameworks like LangChain optimize for efficiency and applications and manage
their capabilities. and LangGraph. accuracy. their lifecycle.

This roadmap outlines the essential stages you'll navigate on your journey to becoming proficient in AI agent development. Each phase
builds upon the last, ensuring a structured and effective learning experience. From theoretical understanding to practical deployment,
we'll equip you with the knowledge and tools needed to succeed.
Essential Tools & Frameworks for Agent
Development
Free & Low-Cost Community & Utility
Solutions Tools
Google Colab: Run LangChain/LangGraph code CrewAI: Role-based agent development for
directly in the cloud. Free collaborative AI. Free
LangGraph: Agent flow and memory management for OpenDevin: An open-source development agent for
complex interactions. Free coding tasks. Free
LangChain Hub: Access toolkits and templates to Tavily API: Fast and efficient web search capabilities
jumpstart your projects. Free for agents. Free tier
PromptLayer: Track and visualize your prompt
interactions for optimization. Free tier

Building powerful AI agents doesn't require expensive software. We'll explore a suite of free and low-cost tools and
frameworks that enable robust development. These resources provide the foundation for creating sophisticated
agents, managing their workflows, and integrating essential functionalities like web search and prompt tracking.
30-Day Weekly Learning
Plan
Understand Agents (Theory + LangChain +
Demo)
Read "Mastering AI Agents" (Ch. 1–3) and Tools
Learn to create planner and executor agents
experiment with demo agents in Google Colab to using LangChain's powerful tool integration
grasp core concepts. features.

Memory + Multi-Agent Build + Evaluate Own


Flow
Implement state and memory management with Agent your own AI agent, then evaluate its
Develop
LangGraph, and explore collaborative multi-agent performance using LLM-as-a-judge techniques
roles using CrewAI. and quality metrics from tools like Galileo.

This structured 30-day plan is designed to rapidly accelerate your AI agent development skills. Each week focuses
on critical aspects, moving from foundational theory to hands-on building and evaluation. By the end of this month,
you'll have the practical experience to build and assess your own AI agents, like the research agent "Fred."
Practical Templates for Agent
Experiments
Finance Agent
Breaks down complex stock evaluations into manageable, stepwise tasks for in-depth analysis.

Resume Review
Agent user resumes and provides actionable recommendations for improvement.
Analyzes

Research Agent
Gathers information from multiple sources, synthesizes data, and provides concise summaries.

Dev Agent
Analyzes code from GitHub repositories, identifies potential issues, and recommends optimizations.

These practical agent templates offer a starting point for your experiments. Each template
demonstrates how agentic capabilities can improve efficiency in diverse domains, from financial analysis
to software development. The Resume Review Agent, for instance, can transform a typical resume into a
powerful document, as shown in the before-and-after comparison.
Smart Cost-Saving Tips for AI Agent
Development

Optimize LLM Off-Peak Local Agent Limit Chain


Usage
Use gpt-4o-mini instead of Processing
Schedule heavy processing Deployment Length
Control API token usage by
GPT-4 turbo during tasks to run overnight on Utilize open-source models limiting chain length and
development to free Colab instances, like OpenDevin or recursion (e.g., 5 cycles
significantly reduce API leveraging available DeepSeek as local agents max) in your agent's
costs without sacrificing computing resources when for tasks that don't require operation, preventing
quality for initial demand is lower. cloud-based solutions, runaway costs.
prototyping. enhancing privacy and cost
Developing AI agents efficiently also means being smart about costs. These tips will help you manage your expenses
control.
effectively, ensuring that you can experiment and build without financial strain. By making informed choices about
models, processing times, and operational limits, you can maximize your development budget.
Cost Breakdown: Agent Evaluation
Runs
Accelerate Learning: Communities &
Courses
To further accelerate your learning and connect with the AI agent community, consider these valuable resources:

[Link]: Offers specialized courses like "AI Agents in LangGraph" to deepen your understanding and
practical skills. GitHub: A treasure trove of open-source projects, including CrewAI Examples and LangGraph
Templates, for hands-on learning. Discord & Reddit: Join vibrant communities like the CrewAI Community and
LangChain Forums to ask questions, share insights, and collaborate. Moveworks Dev Portal: Explore their Agent
Automation Playground for practical application and experimentation in enterprise settings.
Evaluation, Metrics &
Pitfalls of AI Agents
Discover how to effectively evaluate AI agents through key
performance metrics and avoid common pitfalls. This presentation
guides future-focused professionals and entrepreneurs on
leveraging AI agents to build impactful startups, deploy MVPs, and
scale strategically.
The Future of Work is Agentic
Agentic Systems Entrepreneurial Opportunity

• Self-planning capabilities You don’t need to hard-code every logic step. Agents
• Seamless tool integrations learn and improve dynamically, enabling rapid
• Memory and learning from feedback innovation and adaptability in business processes.

• Dynamic reasoning machines


6 Real-World Agentic Startup
Models
Sector Agent Idea Benefit

AgriTech Crop advisory agent Improves crop yield


(weather + soil aware) with localized guidance

EdTech Exam AI buddy Personalized coaching


(adaptive tutor) at scale

HR & Talent Resume screener + Accelerates hiring with


Interview question unbiased evaluation
generator

Finance Investment planner DIY wealth


agent (risk modeling) management with
smart risk
HealthTech Patient Q&A and report AI nurse or counselor
explainer support

Retail Purchase advisor bot Boosts conversions with


personalized advice
AI Agent Deployment Blueprint
1 Start Small 2 Test & Score
One agent, one task using LangGraph or CrewAI Rigorous testing with human-in-the-loop fallback;
with GPT-4o or Claude API integration. use tools like Galileo and PromptLayer for scoring.

3 Add Guardrails 4 Launch with UI


Prevent infinite loops and offensive outputs to Deploy on platforms like Gradio, Streamlit,
ensure safe, effective behavior. Telegram, or WhatsApp for user engagement.
Common Pitfalls to Avoid
Trying to Do Everything at
Once
Build a vertical MVP focusing on core functionality before expanding.

Unbounded API
Use
Control costs by leveraging GPT-3.5 and model switching strategies.

Ignoring Safety &


Feedback
Implement approval workflows and feedback logging for continuous
improvement.

Agent Doesn’t
Scale
Use auto-scaling technologies like Kubernetes or serverless
architectures.
Innovator’s 90-Day AI Agent Path
0–30 Days: Learn Core
1 Tools
Hands-on projects with LangChain, LLM APIs, and Jupyter notebooks for solid
fundamentals.

30–60 Days: Build


Agents
2 Create 1–2 agents that use tools and memory, such as a resume
reviewer or investment planner.

60–90 Days: MVP


Launch
3 Develop a full-stack MVP with UI and feedback
systems; deploy a functional prototype.
Evaluation Metrics for AI Agents
Task Completion Rate Tool Success Response
Ratio Coherence
Measures how effectively agents Evaluates the accuracy and Assesses the logical consistency
achieve assigned goals end-to-end. reliability of integrated tool calls and relevance of agent outputs in
within workflows. context.
Final Thoughts & Vision
AI agents represent the next frontier of automation, combining
reasoning and adaptive learning. They free humans for higher-level
creative and strategic work, fostering equitable digital assistants
accessible to millions. Whether you’re an entrepreneur, developer,
or domain expert, the time to innovate with AI agents is now.
Empower your startup ideas and scale thoughtfully to lead in this
transformative era.
The Ethics, Safety &
Governance of AI
Agents
This chapter focuses on deploying AI agents responsibly, covering
ethical frameworks, guardrails, transparency, and bias detection. It
provides recommendations for enterprise-level safe deployment,
critical for professionals and startups planning real-world AI agent
by Sadanand
deployment.

Rajpurohit
Core Guardrails Every Agent Needs
Guardrails are protective mechanisms that restrict unsafe or unintended behavior in AI agents. Based on Agent SDK best
practices, these layers ensure safe and reliable operation.

Safety Classifier Blocks prompt injections or unsafe "Ignore previous instructions…" is


inputs flagged and blocked

PII Filter Removes personal info from output Prevents sharing of names, locations,
IDs
Moderation API Uses OpenAI or Anthropic’s Detects toxicity, hate speech, etc.
moderation filters

Tool Safeguards Classifies external tools by risk Only allow refunds < $1000 without
supervisor

Relevance Filter Ensures outputs stay on-topic Flags "What’s the height of Mount
Everest?" in an HR agent conversation

Human-in-Loop Lets agents hand off tasks they can’t Escalates complex cases like large
handle payment approvals
Evaluation, Intervention & Global Regulation
Even with good prompts, agents can behave unpredictably. Ongoing evaluation with tools like Galileo, built-in feedback loops, and intervention triggers are crucial for learning from failures and managing
high-risk decisions.
Evaluation & Feedback Global AI Ethics & Regulation

• Ongoing evaluation with tools like Galileo Staying aligned with national and international standards is essential for responsible AI deployment.
• Feedback loops to log and learn from failure cases
• Intervention triggers based on error thresholds UNESCO

AI Ethics Recommendations

EU AI Act

Tiered risk-based regulation

USA

NIST AI Risk Management Framework


Emerging Trends and the
Future of Agentic AI
This presentation explores the next phase of AI agents. Expect
autonomous evolving systems, marketplaces, and AI-powered
workplaces.
by Sadanand
Rajpurohit
The Rise of Autonomous,
Self-Improving Agents
LLM Bot
Basic prompt-driven interaction with limited memory.

Multi-Agent
Workflow
Collaborative agents with shared tasks and feedback
loops.
Reflexive Self-
Learner
Continuous learning with reflection and self-
improvement.
Next-Gen Agent Types and
Applications
Reflexive Agents Environment Memory-Enhanced Self-Learning
Agents Agents Agents
QA automation, medical Robotics and smart IoT Adaptive systems for
triage using self- systems manipulating Personal assistants research and market
reasoning. the real world. remembering user predictions.
context.
Tool-Augmented
Agents
Business copilots integrating external APIs seamlessly.
Key Shifts in Agent
Design
Long-Term Inter-Agent
Autonomy
Agents work continuously Communication
maintaining personas and Agents collaborate,

goals. delegate, and coordinate


tasks dynamically.

Natural Language
APIs
AI executes commands via natural language instead of REST.
AI Agent Marketplaces
and Platforms
Agent App Stores Leading Platforms
Plug-and-play vertical AI Moveworks, CrewAI,
agents like Legal and LangGraph, and OpenAI's
FinTech AI. OpenAgents.

Low-Code Builders
Enable non-experts to create custom personalized agents.
The Future Workplace Powered by AI
Agents
Agent-First Teams 24/7 Operations Scaled Knowledge
Work
Teams combining AI copilots and Continuous, automated workflows Automate routine tasks and
humans for task automation. covering global time zones. decision support at large scale.
Challenges and Risks in
Agentic AI
Memory Limits Evaluation Regulation Security
Current LLMs Testing complex Privacy, ethics, and Sandboxing agents
struggle with long- autonomous agents misuse need careful is critical to prevent
term recall. remains difficult. governance. harm.
Summary and Next
Steps
Evolving Agentic AI
From chatbots to self-learning, multi-agent systems.

Market and Career


Growth
New startups and jobs in AI agent orchestration and safety.

Join the Agentic Age


Prepare to lead with AI copilots and personalized agents.
Building Financial Research AI
Agents with LangGraph and
GPT-4o
This presentation walks software developers and AI
engineers through implementing a basic skill evaluation AI
agent named Fred using LangGraph and GPT-4o. Fred is
designed as a financial research assistant that can answer
complex investment questions by decomposing queries,
conducting web research, synthesizing insights, and
delivering verdicts. By exploring Fred's architecture,
workflow, evaluation methods, and real metrics dashboards,
by Sadanand
developers will gain hands-on guidance to build, test, and
Rajpurohit
monitor effective AI agents. This chapter exemplifies
integrating cutting-edge LLMs into practical portfolio
Defining the Agent Use Case:
Financial Research Assistant
Agent Objective Core Tasks

Fred answers investor The agent researches


questions by breaking using web tools like
down complex queries Tavily, synthesizes
such as “Should I findings reliably, and
invest in Tesla given delivers a well-
the EV market in justified, judgment-
2024?”
Key into
Benefits backed response to
manageable
The approachsub- the
promotes thorough user.
analysis,
questions.
transparency in reasoning, and alignment with user
queries, building trust in AI-driven investment advice.
Agent Architecture and
Workflow
Planner Agent & Replanner Judge

Breaks down the primary query Executes online research via Leverages GPT-4o for
into focused smaller steps to tools like Tavily and dynamically evaluation, scoring responses
address critical aspects such as adjusts plans by adding or on context relevance, clarity,
market demand and competitor removing steps based on and adherence to the original
analysis. intermediate findings. question.
Tools, Packages & Setup for
Building Fred
Required Packages API Integration Environment Setup

• langgraph==0.2.56 Acquire necessary API keys Ensure dependencies are


• langchain- from OpenAI for GPT models installed correctly and
• community==0.3.9
langchain-openai and from Tavily for web search environment variables securely

• tavily-python functionality to enable real-time configured to start

• promptquality data access. development and


experimentation smoothly.
Detailed Execution Flow of
Agent Fred
Planning Step

Decomposes the input query into research sub-questions


such as analyzing EV demand, Tesla's finances, and
competitors.
Execution

Uses Tavily for web-based insights, capturing


intermediate outputs to refine results continuously.

Replanning

Identifies gaps or uncertainties and updates the plan by


adding or removing steps before finalizing.

Judge Evaluation

GPT-4o assesses the output on multiple dimensions:


context adherence, relevance, and clarity of the answer.
Performance Metrics and
Monitoring with Galileo AI Studio
Context Adherence Tracks how well the agent
stays on the original topic.

Token Usage Measures cost efficiency


through token consumption
monitoring.
Latency Records the time taken to
generate the final answer.

Number of Replans Indicates complexity of the


problem or inefficiencies in
planning steps.
Overall Score An aggregate metric
reflecting comprehensive
agent performance across
multiple runs.
Outcomes and Applications
ofEfficiency
Agent Gains
Fred
Fred reduced financial analysis time by 70%, substantially accelerating
decision-making.

Cost Optimization

Operating cost per query dropped to about $0.0025 thanks to efficient


token use and process design.

Expert-Level Quality

Reached high context adherence scores (0.84+), enabling reliable and


trustworthy financial judgments.

Extensible Template

Fred’s design translates well to customer support bots, market


research analysts, student advisors, and lead scoring engines.
Key Takeaways and
Next Steps
Stateful & Reflective
AI
Effective AI agents maintain state, iteratively reassess, and
evaluate outputs for continuous improvement.

Modular & Scalable


Design
Breaking down complex queries and isolating components aids
debugging and scaling agent functionality.

Monitoring &
Evaluation
Combine tools like Galileo dashboards and GPT-based judges
to measure and tune agent behaviors.

Build from Use


Cases
Rooting projects in real-world scenarios accelerates learning
and develops portfolio-ready AI innovations.

You might also like