IBM AI Ethics Board
AI agents:
Opportunities, risks, and mitigations
Attributions
With gratitude to,
AI Ethics Board workstream’s co-chairs and executive
sponsors: Christina Montgomery, Francesca Rossi, Kush
Varshney, Manish Bhide, and Rob Parkin
Contributors: Saishruthi Swaminathan, Ambrish Rawat, Chan
Nasseb, Christopher Noessel, Daniel Karl I. Weidele, David
Piorkowski , Heloísa Candello, Jamie VanDodick, John
Richards, Matt Bellio, Michael Hind, Michael Muller, Mihaela
Bornea, Milena Pribic, Phaedra Boinodris, Rogerio Abreu de
Paula, Sara E Berger, and Simon Rogers
Acknowledgments
Co-chairs, executive sponsors, and contributors would like to
acknowledge the following members for their review and
feedback: Alina Glaubitz, Amit Dhurandhar, Christopher Hay,
Katherine Fick, Kevin Black, Manish Goyal, Maryam Ashoori,
Monica Patel, Pin-Yu Chen, Ryan Hagemann, and Wouter
Oosterbosch
2 AI agents: Opportunities, risks, and mitigations | March 2025
Table of contents
04 12
Introduction Mitigation and Governance
05
Benefits
06
Risks, Challenges, and
Societal Impacts
3 AI agents: Opportunities, risks, and mitigations | March 2025
Introduction
Artificial intelligence (AI) is revolutionizing the way people IBM is a hybrid cloud and AI company and by using the strength
interact with the world around them. As AI technology continues of our research, product, and consulting teams, along with
to evolve, it unlocks new opportunities and helps solves complex external partners and the AI Ethics Board, IBM helps bring the
problems in diverse fields such as finance, healthcare, power of AI agents to our clients. Some examples include
environment, education, and sports. We are now entering an age Software Engineering Agent from IBM Research®, AI
of AI where businesses are adopting AI agents to transform their Integration services from IBM Consulting®, Granite BeeAI
processes, maximize business impact across functions, and Agent Framework, IBM watsonx OrchestrateTM and [Link]®
accelerate time to value. An AI agent is a software entity that from IBM Technology.
employs AI techniques and has agency to act in its environment
based on set goals, which means it can decide which actions to With rapid advancements in AI technology, AI agents are
perform and has the ability to execute them. AI agents have becoming more powerful and sophisticated, raising questions
existed for a long time, starting from rule-based AI agents to about their ethical ramifications and societal impact. This paper
recent large language model (LLM) agents, which are AI agents describes IBM’s current point of view on the ethics and
that employ large language models and demonstrate great governance of AI agents. It is version one, and future iterations
potential across several sectors. Agentic AI systems are will expand on various aspects of IBM’s ethics and governance
software systems that leverage AI agents (together with other approach concerning its AI agents.
components like tools, planners, memory, and datasets), pursue
goals, and can operate autonomously. Though AI agents and
agentic AI systems are software, they can be used to control
hardware.
Note: In this paper, AI agent and agentic AI system terms are
used interchangeably.
AI agents: Opportunities, risks, and mitigations | March 2025
4
Benefits
AI agents can significantly enhance how effectively humans Improved efficiency and productivity
perform tasks and achieve business outcomes. These benefits
include: AI agents can operate continuously, manage multiple tasks
simultaneously, and tackle complex challenges. This can help
Augmenting human intelligence accelerate delivery speed, boost productivity, and improve the
efficiency of business operations. For example, IBM is
AI agents can integrate into workflows to help reduce the expanding their partnership with Salesforce to provide pre-built
amount of time that is spent completing tasks and augmenting AI agents that will be available to customers around the clock,
human performance. For example, the IBM AI Agent SWE-1.0 supporting sales and services. Leveraging watsonx Orchestrate,
system consists of a localization agent and an editing agent that IBM will create AI agents for Agentforce, Salesforce’s suite of
integrate into GitHub workflows to help developers by reducing autonomous agents, to help businesses improve productivity,
the amount of time that developers spend finding bugs, while maintaining security.
developing and testing fixes to software defects.
Enhanced decision-making and quality of responses
Automation
AI agents can be connected with multiple external resources,
AI agents can automate routine or time-consuming tasks to tools, and other agents to help enhance their decision-making
enable increased focus on innovation and strategic work. For and quality of their [Link] can also provide responses
example, IBM’s AI Digital Assistant AskHR uses AI agents to that are comprehensive and personalized to the user resulting
automate common HR processes such as employee support and in better customer experience. For example, IBM Consulting
onboarding. The AI agents are built on watsonx Orchestrate and worked with a global life sciences company to bring a series of
powered by generative AI. With this automation, IBM's AskHR AI agents together to accelerate generation of technical
now handles 94% of employee queries and resolves around documentation that is fact-based and traceable.
10.1 million interactions per year, allowing IBM’s HR team to
focus more on important work such as strategic planning.
5 AI agents: Opportunities, risks, and mitigations | March 2025
Risks, Challenges, and
Societal Impacts
Like all rapidly advancing technologies, AI agents can create These characteristics, as well as the range of autonomy levels
risks along with benefits. Various legal and regulatory actions that AI agents and agentic AI system may have, lead to several
may be associated with some of these risks, as well as risks, challenges, and societal impacts, as shown in the tables
reputational and operational consequences. In general, risks below. We build upon the risks and challenges identified for
raise sociotechnical questions and should be addressed and foundation models and thus only include in the tables below
mitigated through sociotechnical methods, including software risks, challenges, and societal impacts related to AI agents
tools, risk assessment processes, AI ethics frameworks, that are new or that amplify those in the foundation model
governance mechanisms, multistakeholder consultations, risks table.
standards, and regulation.
Note: In the tables below, AI agent is used for describing both
Approach AI agent and agentic AI system.
AI agents can perform three types of actions:
• Take actions that impact the world (physical or
digital.
• Consult resources and use tools.
• Decide which process to choose in the selection of
resources/tools/other AI agents and select them.
AI agents can perform these actions autonomously, which
means they can perform them without continuous human
oversight.
Because of their nature, AI agents and agentic AI systems can
have the following characteristics:
• Opaqueness due to limited visibility into how AI agents
operate, including their inner workings and interactions.
• Open-endedness in selecting resources/tools/other AI
agents to execute actions. This may add to the possibility
of executing unexpected actions.
• Complexity emerges as a consequence of open-endedness
and compounds with scaling open-endedness.
• Non-reversibility as a consequence of taking actions that
could impact the world.
6 AI agents: Opportunities, risks, and mitigations | March 2025
Risks
Group Risk Risk Indicator with Reason
Value Misaligned actions: AI agents can take actions that are not aligned with Amplified
Alignment relevant human values, ethical considerations, guidelines, and policies.
Misaligned actions can occur in different ways such as: Reason:
• Applying learned goals inappropriately to new or unforeseen situations. • Autonomy of AI agents to take
• Using AI agents for a purpose/goals that are beyond their intended use. actions
• Selecting resources or tools in a biased way.
• Using deceptive tactics to achieve the goal by developing the capacity for
scheming based on the instructions given within a specific context.
• Compromising on AI agent values to work with another AI agent or tool to
accomplish the task.
Fairness Discriminatory actions: AI agents can take actions where one group of humans Amplifed
is unfairly advantaged over another due to the decisions of the model. This may
be caused by AI agents’ biases in the actions that impact the world, in the Reason:
resources consulted, and in the resource selection process. For example, an AI • Autonomy of AI agents to
agent can generate code that can be biased. take actions
o Take actions that impact
the world
o Consulting biased
resources
o Biased resource selection
process
Data bias: Specific actions taken by the AI agent such as modifying a dataset or New
a database can introduce bias in the resource that gets used by others or by
itself to take actions. Reason:
• AI agents taking actions
that impact the world. Here,
bias is due to a specific
action
• Open-endedness
Misplaced Over- or under-reliance: Reliance, that is the willingness to accept an AI agent Amplified
Trust behavior, depends on how much a user trusts that agent and what they are
using it for. Over-reliance occurs when a user puts too much trust in an AI agent, Reason:
accepting an AI agent’s behavior even when it is likely undesired. Under-
reliance is the opposite, where the user doesn’t trust the AI agent but should. • AI Agents actions make the
trust evaluation harder
Increasing autonomy (to take action, select, and consult resources/tools) of AI
agents and the possibility of opaqueness and open-endedness increase the
variability and visibility of agent behavior leading to difficulty in calibrating trust
and possibly contributing to both over- and under-reliance.
7 AI agents: Opportunities, risks, and mitigations | March 2025
Group Risk Risk Indicator with Reason
Computation Redundant actions: AI agents can execute actions that are not needed for New
Inefficiency achieving the goal. These actions could result in wasting computation
resources, reduce AI agent’s efficiency in achieving the goal, and lead to Reason:
potentially harmful outcomes. • AI agents can take
actions
In an extreme case, AI agents could enter a cycle of executing the same
actions repeatedly without any progress. This could happen because of
unexpected conditions in the environment, the AI agent’s failure to reflect on
its action, AI agent reasoning and planning errors or the AI agent’s lack of
knowledge about the problem. This could prevent the AI agent from achieving
the goal and exhaust computational resources.
Robustness Attack on AI agents external resources: Attackers intentionally create New
vulnerabilities or exploit existing vulnerabilities in external resources (tools/
database/applications/services/other agents) that AI agents rely on to Reason
execute their intended actions or to achieve their goals.
• AI agents can take
Compromised external resources could impact the AI agent’s performance in actions
different ways such as: • Open-endedness and
complexities due to
• Manipulating AI agents to pursue a different goal. Example: Change AI open-endedness
agents’ goal to add positive reviews to a product of attacker’s
• Access to more
preference when the original user goal is to add queries about the
resources
product.
• Manipulating AI agents to execute undesired actions. Example: Trick AI
agents to download malware.
• Capturing and relaying interactions between AI agents to malicious
actors.
• Getting AI agents to share personal or confidential information.
Unauthorized use: If attackers can gain access to the AI agent and its Amplified
components, they can perform actions that can have different levels of
harm depending on the agent’s capabilities and information it has access Reason
to. Additionally, attackers may execute actions that lead to system
• AI agents can take
degradation, such as exhausting available resources and impairing
actions
performance. Examples:
• AI agents are more
• Using stored personal information to mimic identity or impersonate capable
with an intent to deceive. • Personalization feature
of AI agents
• Manipulating AI agent’s behavior via feedback to the AI agent or
corrupting its memory to change its behavior.
• Manipulating the problem description or the goal to get the AI agent to
behave badly or run harmful commands.
8 AI agents: Opportunities, risks, and mitigations | March 2025
Group Risk Risk Indicator with Reason
Exploit trust mismatch: Attackers could initiate injection attacks to bypass Amplified
the trust boundary which is a distinct point or conceptual line where the level
Reason:
of trust in a system, application, or network changes. This could lead to
mismatched (expected vs. realized) trust boundaries and could result in
• Open-endedness
unintended tool use, excessive agency, and privilege escalation. Also,
• Complexity
background execution in multi-agent environments increases the risk of
covert channels if input/output validation is weak.
New
Function-calling hallucination: AI agents could make mistakes when
generating function calls (calls to tools to execute actions). Those function Reason
calls could result in incorrect, unnecessary or harmful actions. Examples:
Generating wrong functions or wrong parameters for the functions. • AI agents can consult
resources and tools
• AI agents can take actions
Privacy and Sharing IP/PI/confidential information with user: AI agents with Amplified
IP unrestricted access to resources or databases or tools could potentially store
and share PI/IP/confidential information with system users when performing Reason:
their actions. • Multi-component nature
with the ability to perform
actions
Sharing IP/PI/confidential information with tools: AI agents with New
unrestricted access to resources or databases or tools could potentially store
and share PI/IP/confidential information with other tools or agents when Reason:
performing their actions. • Multi-component nature
with the ability to perform
actions
Explainability Unexplainable and untraceable actions: Explanations, lineage and trace Amplified
and information, and source attribution for AI agent actions might be difficult,
Transparency Reason:
imprecise, or unobtainable.
• Difficult to trace cause and
effect or influence of
different components
including LLMs across final
actions
Lack of transparency: Lack of transparency is due to insufficient Amplified
documentation of the AI agent design, development, evaluation process,
absence of insights into inner workings of the AI agent, and interaction with Reason:
other agents/tools/resources.
• Rely on other documents
available for other tools/
agents
9 AI agents: Opportunities, risks, and mitigations | March 2025
Challenges
Challenge Challenge Indicator with Reason
Evaluation: Challenge in evaluating AI agents’ performance/accuracy Amplified
because of the complexity of the system and open-endedness.
Reason:
• Complexity of the agentic
systems
Mitigation and maintenance: Challenge in knowing where in the system Amplified
something is going wrong and how to fix it or what might need maintenance
protocols. Reason:
• Complexity and open-
endedness of the agentic
systems
Reproducibility: Challenge in reproducing the agent’s behavior or output New
because of unavailability or changes to the tools or resources used for
executing the actions. Reason:
• Open-endedness
Accountability: Challenge in assigning responsibility for an action taken by Amplified
an agentic AI system.
Reason:
• Complexity & open-endedness.
Components can be from
different vendors
Compliance: C hallenge in determining regulatory compliance as AI agents are Amplified
complex and there may not be enough information to understand whether
the whole agentic AI system is compliant. Reason:
• Open-endedness
• Complexity
• Lack of transparency
10 AI agents: Opportunities, risks, and mitigations | March 2025
Societal Impacts
Societal Impact Sociatal Impact indicator
Impact on human dignity: If human workers perceive AI agents as being Amplified
better at doing their jobs than they are, they could experience a decline in their
self-worth and wellbeing.
Impact on human agency: The autonomous nature of AI agents in Amplified
performing tasks or taking actions could affect the individuals’ ability to
engage in critical thinking, make choices and act independently.
Impact on jobs: Widespread adoption of AI agents to perform complex tasks Amplified
might lead to widespread automation of roles and could lead to job
displacement.
Impact on environment: Complexity of the tasks and possibility of AI agents Amplified
performing redundant actions could led to computational inefficiencies and
add to the environmental impact.
11 AI agents: Opportunities, risks, and mitigations | March 2025
Mitigation and
Governance
As AI accelerates business transformation, trust has become
imperative to navigating a competitive business environment.
IBM’s commitment to trust is embodied in our Principles for
Trust and Transparency and Pillars of Trustworthy AI and this
paper highlights IBM’s holistic approach to culture,
processes, and tools and how IBM research, product, and
consulting teams come together with the AI Ethics Board to
build responsible agentic AI solutions.
AI Ethics Board
IBM has established a culture that supports the responsible
development, deployment, and use of AI. At the center of our
organizational governance is the multidisciplinary AI Ethics
Board which is responsible for the governance and decision-
making process for AI ethics policies and practices. The AI
Ethics Board has been in place for over five years. Among its
many activities, the AI Ethics Board works with focal points
within each IBM business units to assess AI use cases and
align these with IBM’s core values.
12 AI agents: Opportunities, risks, and mitigations | March 2025
Integrated Governance
Integrated Governance Program (IGP) is a unified approach to IBM Consulting’s AI Strategy and Governance offering
responsibility and compliance. By creating a holistic, end-to-end empowers organizations to harness the transformative
view of data and models built and used, IGP scales governance potential of responsibly curated AI – from traditional to agentic
workflows around data, privacy, and AI without disrupting
–rooted in their enterprise data to advance and reshape their
innovation and business processes. IGP has empowered IBM to
business strategies. We advise clients to ‘do the right AI’,
deploy uniform internal data standards, enabling data
unlocking ROI from AI investments, and ‘do AI right’, enabling
transparency and trustworthy AI development, while also
accountability for AI models they build and buy, so they can
allowing for innovation at scale.
scale responsibly while de-risking their investments. This
holistic framllowing highlights practices implemented and
recommended by the offering to help businesses mitigate
Products and offerings agentic AI risks:
IBM [Link]® enables organizations to drive • We enable observability to understand what actions the
responsible, transparent and explainable AI. It provides agent took with which inputs to determine attributability
complete lifecycle governance capabilities across AI, including for the corresponding outputs.
agentic AI, ffrom use case request to deployment, including • We examine the entire trace to determine if the AI agent
initial risk assessment and risk evaluation to help identify risks has progressed towards the overarching goal with each
early in the process. It has retrieval-augmented generation (RAG) action taken, to debug the flow, and identify bottlenecks
and agentic AI evaluation metrics such as faithfulness, context in system design that repeatedly surface redundant or
relevance, and answer similarity to help confirm if AI agents are unnecessary actions.
acting appropriately. IBM [Link] will have • We build ontologies to improve accuracy, reliability,
additional metrics designed to monitor and improve agent provide data lineage and provenance.
performance. It will also have capabilities for detecting tool • We perform functional testing which involves rigorously
calling’s hallucinations, managing risks, aiding regulatory testing each component of the agentic system in isolation
compliance, and capturing metadata and other details about AI (i.e., tools, agents), interactions between the
agents in a factsheet. components, the ability to select appropriate tools with
correct parameters and values, and guardrails to mitigate
vulnerabilities.
IBM [Link] simplifies, unifies, and optimizes the agent
lifecycle management (known as AgentOps), providing • We configure AI agents to improve computation
transparency, traceability, and flexibility to discover, manage, efficiency by detecting when AI agents enter possible
monitor and optimize AI agents. infinite loop scenarios by tracking, per task, elapsed time,
total tokens consumed, or number of unsuccessful
iterations.
IBM watsonx Orchestrate puts AI to work by helping users build,
deploy and manage powerful AI agents that automate tasks with • We create hyper-focused AI agents and provide them
generative AI. Orchestrate will provide transparency into AI only fit-for-purpose tools necessary to complete their
agent reasoning on tool selection and/or agent collaboration so a tasks while making sure execution is always done using
user can understand how a task or workflow has been the authorization context of the accessing user.
completed. • We define guardrails at the model level to detect and
mitigate HAP content, jailbreaking, prompt injection
IBM Guardium® AI Security helps customers to continuously attempts, unauthorized sensitive information disclosure,
monitor the controls of their Generative AI models in production and hallucinations.
and enables a secure and responsible deployment.
13 AI agents: Opportunities, risks, and mitigations | March 2025
Models Techniques and methods
The Granite Guardian models are a robust suite of safeguards • Techniques like forcing humans to take more time to think
designed to detect risks in both prompts and responses. In RAG (time-based de-anchoring strategy) help achieve optimal
use-cases, the guardian models assess context relevance, human-AI collaboration and reduce anchoring bias, where
groundedness, and answer relevance. It also has detectors for humans rely blindly on the AI decision. Another technique
function calling hallucination within the agentic workflows. This is based on the value-based AI-human collaborative
includes assessing the validity of function calls and detecting framework which introduces friction by nudging humans
fabricated information, particularly during query translation. with decision recommendations, when needed, in human-
AI interaction when the human is the final decision maker.
• Methods such as Multi-level Explanation for Generative
Human-in-the-loop and human oversight Language Models and Contrastive Explanations for Large
Language Models help in explainability and source
[Link] like adversarial collaboration, where
Human oversight and review can help identify risks and correct
AI scrutinizes the basis for human’s decision instead of
errors. Human validation and feedback help ensure that the
offering alternaxqte recommendation, help address the
actions taken by the AI agents are accurate, relevant, and
impact of automation and downgraded agency on human
aligned. As a part of the IBM use case AI ethics assessment
dignity.
process, question about human-in-the-loop is considered and
appropriate human oversight is implemented. IBM's Augmenting • Attack Atlas, an intuitive and organized taxonomy of single-
Human Intelligence POV outlines sample use cases, Key turn input attack vectors, provides the community with a
Performance Indicators and best practices for enhancing human unified starting point in the rapidly growing field of
intelligence with AI and empowering individuals to navigate a generative AI security and red-teaming.
competitive business environment in partnership with AI.
Education Tools and benchmarks
IBM provides education on AI ethics and governance across • Model tuning approaches such as those employed by the
teams, clients, and community. IBM SkillsBuild, IBM Technology IBM Alignment Studio help with mitigating value alignment
YouTube channel, IBM Developer watsonx courses, and IBM AI .
risks.
Academy are some resources through which IBM educates
people about AI ethics and governance • The IBM open source toolkit AI Fairness 360 helps
mitigate fairness risks.
• ITBench, a set of benchmarks, offer AI practitioners a way
The following highlights cutting-edge research and tools from
to measure how effective the agents they’re building are at
IBM to help users throughout the AI agent lifecycle and build
solving real problems and how their agents compare to
responsible agentic AI solutions.
others on tasks that businesses conduct every day.
• Carbon for AI, an open source design system from IBM,
leverages an interactive AI icon to promote explainability
and a clearer understanding of AI capabilities.
14 AI agents: Opportunities, risks, and mitigations | March 2025
© Copyright IBM Corporation 2025
IBM Corporation
New Orchard Road
Armonk, NY 10504
Produced in the
United States of America
March 2025
IBM, the IBM logo, watsonx Orchestrate, [Link],
[Link], IBM Guardium AI Security, IBM Research and IBM Consulting
care trademarks or registered trademarks of International Business
Machines Corporation, in the United States and/or other countries. Other
product and service names might be trademarks of IBM or other
companies. A current list of IBM trademarks is available on [Link]/
trademark.
This document is current as of the initial date of publication and may
be changed by IBM at any time. Not all offerings are available in every
country in which IBM operates.
THE INFORMATION IN THIS DOCUMENT IS PROVIDED “AS IS” WITHOUT
ANY WARRANTY, EXPRESS OR IMPLIED, INCLUDING WITHOUT ANY
WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR
PURPOSE AND ANY WARRANTY OR CONDITION OF NON-
INFRINGEMENT.
IBM products are warranted according to the terms and conditions of the
agreements under which they are provided.