1.
Welcome to OCI Data Science
1. Course overview
Oracle Cloud Infrastructure (OCI) Data
Science – Course Introduction
Overview
Data Science is the art and science of extracting valuable insights from data to solve real-
world and business problems.
This course focuses on using Oracle Cloud Infrastructure (OCI) for building, training,
deploying, and managing Machine Learning (ML) solutions.
Ideal for individuals seeking to upskill or reskill for the growing demand in Data
Science and AI.
Course Contributors
The course features multiple experts, including:
o Wes Prichard
o John Peach
o John Stanesby
o JR Gauthier
o Lyudmil Pelov
o Praveen Patil
o Hemant Gahankari
Developed by a large team of professionals across Oracle Cloud Infrastructure.
Intended Audience
Primary Audience:
o Data Scientists
Also Suitable For:
o Machine Learning (ML) Engineers
o Artificial Intelligence (AI) Engineers
Prerequisites
Participants should:
Be proficient in Python for machine learning.
Have general knowledge of open-source ML and data science libraries.
Ideally possess 1+ years of experience in Data Science, ML, or AI roles.
Have hands-on experience using OCI (recommended).
Course Objectives
Gain proficiency in using OCI Data Science and related services.
Learn to identify OCI services for implementing ML solutions.
Apply cloud and ML best practices.
Understand how to build, train, deploy, and manage ML models in OCI.
Use additional OCI Data & AI services to create end-to-end ML solutions.
Course Structure (5 Major Modules)
1. Introduction to OCI Data Science
o Overview of OCI Data Science and setting up tenancy.
2. Workspace & Environment Setup
o Configuring your OCI Data Science workspace for project development.
3. Machine Learning Lifecycle
o Understanding all steps supported by OCI Data Science throughout the ML
lifecycle.
4. MLOps Practices
o Scaling, monitoring, and automating ML workflows using OCI tools.
5. Related OCI Services
o Exploring other OCI cloud services useful for building comprehensive data
science solutions.
💡 Recommendation: Follow modules in order, as later ones build on earlier concepts.
Hands-On Learning
End-to-End Lab:
o Uses an Employee Attrition use case to predict the likelihood of employees
leaving the organization.
Demo Lessons:
o Recorded demonstrations throughout modules illustrate core concepts.
Requirements:
o Oracle Cloud Account (Free trial available at [Link])
o Optional GitHub access for OCI Data Science AI Samples repository.
Support and Community
Ask Your Instructor Form:
o Submit questions for personalized help from experts.
OU Community:
o Connect with fellow learners and subject matter experts.
o Discuss topics, ask questions, and share experiences.
Tips for Success
Take notes on topics based on your existing knowledge.
Use transcripts to follow along.
Schedule regular breaks and stay active.
Sign up for a free Oracle Cloud account and get hands-on practice.
Complete all skill checks, exam prep, and practice exams.
Stay consistent and review regularly for certification readiness.
Feedback and Continuous Improvement
Oracle continuously updates and refines training content based on learner feedback.
Rate the course and provide feedback on what’s helpful or needs improvement.
Your input helps enhance the learning experience for everyone.
Summary
This course equips learners with the knowledge and practical experience to:
Navigate OCI Data Science tools and services.
Build and deploy machine learning models in OCI.
Follow best practices for MLOps and cloud-based data science.
Prepare confidently for the OCI Data Science certification exam.
2. Expert Tips Intro
OCI Data Science Professional – Welcome Message
Instructor Introduction
Speaker: Hemant Gahankari
Role: Senior Principal Training Lead, Oracle University
Message Overview
Welcome to the OCI Data Science Professional certification course.
The course is designed to help learners gain hands-on skills to become proficient
data science professionals using Oracle Cloud Infrastructure (OCI).
Core Responsibilities of a Data Scientist / ML Engineer
In day-to-day work, data professionals typically:
1. Collect and Prepare Data – Gather, clean, and preprocess datasets.
2. Build and Train Models – Develop and tune machine learning models.
3. Evaluate Models – Assess performance using appropriate metrics.
4. Deploy and Scale Models – Implement trained models into production
environments.
5. Automate ML Pipelines – Streamline processes for repeatable, efficient
workflows.
Using OCI for Data Science
The OCI Data Science and AI Services enable efficient execution of all the
above tasks within a unified cloud platform.
These services provide:
o Scalable compute power
o Integrated data management
o Built-in tools for model training, evaluation, and deployment
o MLOps capabilities for automation and monitoring
Course Content Highlights
The course includes expert tip videos demonstrating:
o Key OCI Data Science and AI Service features
o Real-world best practices
o Simple and efficient ways to build end-to-end data science solutions
Designed to make the learning practical, engaging, and hands-on.
Closing Note
Learners are encouraged to explore, practice, and apply what they learn.
The expert sessions aim to make you confident in using OCI tools effectively.
“Hope you will find these videos useful.” – Hemant Gahankari
2. Introduction and Configuration
1. Data Science: Introduction
Module 1: Introduction and Configuration
Lesson 1 – Introduction to Oracle Cloud Infrastructure
(OCI) Data Science Service
Instructor: Wes Pritchard
Role: Senior Principal Product Manager, Data Science & AI Services, Oracle
1. Historical Background of Data Science
Period Key Contributor Contribution
Introduced Ockham’s Razor — simpler solutions are
1300s William Ockham preferable. Principle applies to ML by favoring simpler
models.
Advocated that more data leads to better accuracy —
1700s Tobias Mayer
considered one of the first data scientists.
Coined the term Machine Learning; created a self-
1952 Arthur Samuel (IBM)
learning checkers program.
Predicted the rise of empirical data analysis through
1962 John W. Tukey electronic computing — precursor to modern data
science.
Defeated chess champion Garry Kasparov; highlighted
1997 IBM’s Deep Blue
the power of computation and AI.
DJ Patil (LinkedIn) & Jeff Coined the term “Data Science” for extracting business
2008
Hammerbacher (Facebook) value from data.
Coined “The Great Resignation”, showcasing how
2021 Anthony Klotz
data science can be used to study employment trends.
2. Data Science Today
Data science now drives real-world business problem solving, e.g., predicting
employee attrition.
The course includes a hands-on lab using an Employee Attrition use case.
o You will build a predictive ML model yourself using OCI Data Science.
3. Oracle’s Approach to Data Science and AI
a. The Foundation: Data
Modern businesses handle diverse data types:
o Structured: Databases, business apps.
o Unstructured: Text, images, videos, audio, sensor data, customer interactions.
Goal: Use all data to gain insights, enhance experiences, and optimize operations.
b. Oracle AI Portfolio
Oracle AI encompasses multiple cloud services designed to help organizations leverage all their
data effectively.
Architecture Overview:
1. Data Layer: Foundation for all AI and ML activities.
2. AI & ML Services Layer:
o Machine Learning Services:
Used by data scientists to build, train, deploy, and manage ML models.
Utilize open-source frameworks within OCI Data Science.
Supported by OCI Data Labeling for supervised learning.
o AI Services:
Contain pre-built ML models for specific use cases.
Some are pre-trained, others can be trained with customer data.
Accessible via simple API calls — no infrastructure management
required.
3. Applications Layer:
o Where AI is consumed — apps, analytics, or business processes.
c. Supporting Services
Data & Graph Analytics
Data Integration and Management
Business Intelligence
Underlying OCI Infrastructure (Compute, Storage, Networking)
4. Oracle Cloud Infrastructure (OCI) Data Science
Definition
A cloud service designed to help data scientists:
Build, train, deploy, and manage ML models efficiently.
Support the full ML lifecycle using Python and open-source tools.
Operate within a JupyterLab notebook environment.
5. Core Principles of OCI Data Science
1. Accelerate Individual Productivity
o Built for Python & open-source workflows.
o Eliminates manual setup of environments.
o Provides on-demand compute power (CPU/GPU) without infrastructure
management.
o Includes Oracle’s Accelerated Data Science (ADS) SDK for automating
common ML tasks.
2. Enable Collaboration
o Shared projects and assets for teamwork.
o Promotes reproducibility, traceability, and auditability.
o Reduces duplication across data science teams.
3. Enterprise-Grade Reliability
o Integrated with OCI security, IAM, and compliance.
o Fully managed infrastructure (maintenance, patching, upgrades).
o Focus remains on solving business problems, not managing servers.
6. Key Concepts and Terminology
Term Description
A collaborative workspace to organize and document data science assets
Project
(notebooks, models, datasets).
An interactive JupyterLab environment with pre-installed open-source
Notebook Session
libraries for coding, training, and experimentation.
Open-source package & environment manager for Python. Used to
Conda
install and manage dependencies easily in OCI Data Science.
Oracle’s Python library for automating tasks like data connection,
Accelerated Data
visualization, AutoML training, evaluation, and model explainability.
Science (ADS) SDK
Provides simple access to model catalog and OCI services.
Mathematical representation of data and business logic. Created in
Model
notebook sessions and stored in the Model Catalog.
Central repository for storing, tracking, and sharing ML models with full
Model Catalog metadata and provenance details. Enables team collaboration and
reproducibility.
Publishes a trained model as an HTTP endpoint on managed
Model Deployment
infrastructure for real-time predictions.
Define and execute repeatable ML tasks on OCI-managed compute
environments.
Term Description
Browser-based interface to manage all features of OCI Data Science
OCI Console
(used throughout the course).
REST API / SDKs /
Alternative access methods:
CLI
APIs: For programmatic interaction.
SDKs: Available for Python, Java, JS, .NET, Go, Ruby.
CLI: Command-line management with full functionality. |
| Regions | Globally distributed OCI data centers offering secure, high-performance
environments. OCI Data Science is available in commercial, government, and dedicated
regions. |
7. Summary
OCI Data Science empowers data scientists to manage the entire ML lifecycle within the
cloud.
It supports open-source Python workflows, provides managed infrastructure, and
promotes collaboration.
The next lesson focuses on provisioning and configuring OCI Data Science
environments to begin practical implementation.
2. ADS SDK
OCI Data Science Professional – ADS SDK Module Notes
Instructor: John Peach
Role: Data Scientist, Oracle Cloud Infrastructure (OCI) Data Science Service Team
Overview
The Accelerated Data Science (ADS) SDK is a powerful, end-to-end toolkit built by data
scientists for data scientists, designed to streamline the machine learning (ML) lifecycle on
Oracle Cloud Infrastructure.
Goals of ADS SDK:
Integrate OCI services into the data scientist workflow.
Simplify common ML tasks like EDA, feature engineering, model training, tuning,
and deployment.
Enhance productivity with automation, explainability, and reproducibility.
Versions and Access
Versions of ADS:
1. Public Version – downloadable from GitHub or PyPI.
2. OCI-Integrated Version – pre-installed in OCI Data Science notebooks; includes
AutoML and ML Explainability features.
Accessing ADS:
Available in Conda environments on OCI Data Science service.
Installable via pip install ads or directly from GitHub.
Key Features of ADS
1. Data Connectivity
ADS supports connections to multiple data sources, enabling seamless access to data regardless
of location.
Supported Data Sources:
Local Storage (block storage within notebooks)
Object Storage via OCI protocol + pandas integration (APE Spec protocol)
Oracle Databases (via DB Secret Keeper and OCI Vault)
Autonomous Database (ADB) integration
Third-party Clouds: AWS S3, Google Cloud, Azure, Dropbox, etc.
NoSQL Databases: Using Dataset Factory Class
Big Data Service (BDS): Connects directly to HDFS
Web Data: Import via HTTP/HTTPS directly into DataFrames
2. Data Visualization and EDA
Understanding data is a crucial step in any ML pipeline.
ADS Visualization Tools:
Smart Plotting: Automatic default plots for different data types.
Feature Types: Provides reusable visualizations based on measurement type.
Summary Statistics: Generate feature summaries and correlation heatmaps.
Custom Visualization: Create reusable plots across multiple projects.
3. Feature Engineering
Improving data quality leads to better models.
Feature Engineering with ADS:
Uses the ADS Dataset class (wrapper around pandas DataFrame).
Provides automatic feature transformation and suggestions.
Handles:
o Categorical encoding
o Null value imputation
o Transformation recommendations
Supports both manual and automated feature engineering.
4. Model Training
ADS offers flexible options for building models.
Training Methods:
AutoML: Fully automates model selection, tuning, and evaluation.
ADSTuner: Performs hyperparameter optimization.
Model Packaging: Automatically creates model artifacts for deployment.
Model Catalog Integration: Easily push trained models to production.
5. Model Evaluation
Evaluating performance is essential to model development.
ADS Evaluator:
Compares multiple models using consistent metrics.
Supports binary, multinomial classification, and regression.
Automatically generates appropriate evaluation metrics and charts.
Reduces the need for manual chart recreation.
6. Model Interpretability and Explainability
Building trust in models through transparency.
Interpretability Tools:
Model-Agnostic: Works with any ML model.
Local Explainability: Understand predictions for specific observations.
Global Explainability: Understand overall model behavior.
Partial Dependence (PDP) and Accumulated Local Effects (ALE) plots.
What-if Analysis: Test how input changes affect predictions.
These tools help verify that the model is learning the right relationships from data.
7. Model Deployment
Easily move models from notebooks to production.
Deployment Features:
ADS Model Framework: Simplifies model deployment with few commands.
Supports:
o Oracle AutoML
o PyTorch
o Scikit-learn
o TensorFlow
o Generic Models
Integrates with OCI Logging Service for:
o Prediction logs
o Access logs
These logs help monitor model usage and performance in production.
Summary
The ADS SDK supports the entire ML lifecycle — from data access to deployment.
It simplifies every step with automation, reproducibility, and explainability.
Key Takeaways:
Integrates seamlessly with OCI and third-party services.
Streamlines EDA, feature engineering, and model training.
Automates hyperparameter tuning and evaluation.
Enhances interpretability and simplifies deployment.
ADS enables data scientists to move from experimentation to production efficiently and
securely.
3. Tenancy Configuration Basics
OCI Data Science Professional – Tenancy Configuration
Basics
Instructor: Jon Stanesby
Lesson Focus: Understanding how to configure your OCI tenancy for Data
Science — including compartments, user groups, dynamic groups, and policies.
1. Overview
Tenancy configuration in Oracle Cloud Infrastructure (OCI) is essential for organizing
and securing access to Data Science resources.
This setup ensures that users, groups, and services have the right permissions
within the right compartments.
How Data Science Components work together:
Assign User to appropriate groups. -> Create dynamic group for data science
resources. -> Create policies that grant access to resources in a compartment.
2. Key Concepts
a. Compartments
Definition: Logical containers for organizing OCI resources.
Purpose: Control access to cloud resources by grouping them logically.
Access: Only user groups with explicit permissions can access resources in a
compartment.
Steps to Create a Compartment:
1. Go to Identity → Compartments in the OCI Console.
2. Click Create Compartment.
3. Enter a name and description, optionally add tags.
4. Click Create Compartment and note the OCID (useful later).
b. User Groups
Definition: Collections of users granted access to OCI resources.
Purpose: Simplifies permission management for multiple users.
Steps to Create User Groups:
1. Create Users → Go to Identity → Users → Create User.
o Provide a username, description, and optional email.
2. Create Group → Go to Identity → Groups → Create Group.
o Add name and description.
3. Add Users to Group → Click Add User to Group, select user(s), and confirm.
c. Dynamic Groups
Definition: Groups of resources (not users) that match defined rules.
Examples of resources:
o Data Science Notebook Sessions
o Model Deployments
o Job Runs
Key Points:
Membership changes dynamically as matching resources are created/deleted.
These resources act as principal actors, capable of making API calls per
assigned policies.
Enables secure resource-to-service interactions (e.g., notebook session
accessing Object Storage).
Steps to Create Dynamic Group:
1. Go to Identity → Dynamic Groups → Create Dynamic Group.
2. Add name and description.
3. Define matching rules, replacing the compartment OCID with your Data Science
compartment’s OCID.
Example Matching Rules:
Include all Notebook Sessions, Model Deployments, and Job Runs within a
compartment.
3. Policies
a. Definition
Policies define what principals (users or resources) can do within a compartment.
They are the backbone of access control in OCI.
Syntax:
Allow group <group-name> to <verb> <resource-type> in compartment
<compartment-name>
b. Key Components
Component Meaning
Group Name The user group or dynamic group name
Verb The level of access granted
Resource Type The OCI resource or family (e.g., data-science-family)
Compartment Name The specific compartment in which access is granted
c. Verbs (Access Levels)
Verb Description Access Level
🔹 Least
Inspect List resources (no metadata access)
Permissive
Read Inspect + get resource metadata
Read + work with resource (update, but not
Use
create/delete)
🔹 Most
Manage Full access (create, update, delete)
Permissive
d. Resource Types
Policies can target individual resources (e.g., data-science-models)
Or aggregate resource types like data-science-family (includes related
resources such as models, jobs, notebook sessions, etc.)
4. Required Data Science Policies
a. Granting Access to Data Science Resources
1. Allow user group to manage all Data Science resources:
2. Allow group <user-group> to manage data-science-family in compartment
<compartment-name>
3. Allow dynamic group (resources) to manage Data Science resources:
4. Allow dynamic-group <dynamic-group> to manage data-science-family in
compartment <compartment-name>
b. Access to Metrics and Logging
Purpose Policy Statement
Allow users to read Allow group <user-group> to read metrics in
metrics compartment <compartment-name>
Purpose Policy Statement
Allow dynamic groups to Allow dynamic-group <dynamic-group> to use log-
use log content content in compartment <compartment-name>
Allow users to manage Allow group <user-group> to manage log-groups in
log groups compartment <compartment-name>
Allow users to use log Allow group <user-group> to use log-content in
content compartment <compartment-name>
c. Networking-Related Policies
(Required if using custom networking in Data Science)
Allow service datascience to use virtual-network-family in compartment
<compartment-name>
Allow group <user-group> to use virtual-network-family in compartment
<compartment-name>
Allow dynamic-group <dynamic-group> to use virtual-network-family in compartment
<compartment-name>
d. Additional Useful Policies
(For related OCI services, e.g., Object Storage)
Allow group <user-group> to manage object-family in compartment <compartment-
name>
Allow dynamic-group <dynamic-group> to manage object-family in compartment
<compartment-name>
These policies enable object storage integration (e.g., for reading/writing data
from notebooks or deployed models).
5. Summary
In this lesson, we covered the core tenancy configuration components and their
relationships:
Compartments: Logical boundaries for organizing resources.
User Groups: Control user access and permissions.
Dynamic Groups: Automatically manage resource-based access.
Policies: Define who can do what, where, and how.
We explored:
The three-step setup for each (Compartment → Group → Policy).
The required policies for Data Science.
Networking and logging policies for advanced configurations.
Optional policies for integrating with related OCI services.
Takeaway:
Proper tenancy configuration ensures secure, scalable, and organized
management of OCI Data Science environments.
4. Configure a Tenancy with OCI Resource
Manager
OCI Data Science Professional – Lesson: Configuring Tenancy with OCI Resource
Manager
Instructor Introduction
Speaker: John Stanesby
Topic: Configuring tenancy using Oracle Cloud Infrastructure (OCI) Resource
Manager
Lesson Overview
This lesson demonstrates how to use Oracle Resource Manager (ORM) to configure a
tenancy for OCI Data Science.
Rather than setting up resources manually, learners can deploy a preconfigured Data
Science Service template to automate the process.
Key Concepts
1. What is Oracle Resource Manager (ORM)?
A managed service in OCI that helps automate resource provisioning using
Terraform.
Enables repeatable, version-controlled infrastructure deployment.
2. Using the Data Science Service Template
The Data Science Service template is available within Resource Manager.
It automatically creates all required identity and access management (IAM)
components for a basic data-science setup.
Template Configuration Details
When you run the Data Science Service template, it creates:
A. User Group
A group with a name you define.
Used to assign and manage permissions for your human users.
B. Dynamic Group
A group with a name you define, including matching rules for the following
resource types:
o datasciencenotebooksession
o datasciencemodeldeployment
o datasciencejobrun
C. Policy
A policy with statements granting the following permissions:
1. Allow user group to manage data-science-family resources in the
compartment.
2. Allow dynamic group to manage data-science-family resources in the
compartment.
3. Allow user group to read metrics.
4. Allow dynamic group to use log content.
Steps to Run the Resource Manager Stack
1. Create a Stack
o Go to Resource Manager → Stacks → Create Stack.
2. Select Template
o Choose Template as your origin, then Service → Data Science.
o Click Select Template.
3. Choose Compartment
o Select the compartment where Data Science resources will be created.
o Click Next.
4. Configure Variables (Optional)
o You can modify additional configuration variables.
o Click Next to proceed.
5. Run the Stack
o Choose to run apply immediately.
o Click Create and wait for the job to finish.
6. Add Users
o After the stack runs successfully, add your users to the newly created user
group.
Alternative Configuration Option
Instead of using the prebuilt template, you can manually deploy using your own
Terraform script.
Oracle provides an official Terraform example available at a public GitHub
repository.
Summary
The Data Science Service template in Resource Manager automates tenancy
configuration.
It creates essential user groups, dynamic groups, and policies for data-
science workloads.
The setup process involves creating and running a stack, then adding users.
Advanced users can use Terraform scripts from the public GitHub repo for
custom setups.
5. Networking for Data Science
OCI Data Science Professional – Lesson: Networking for Data Science
Instructor Introduction
Speaker: John Stanesby
Topic: Networking concepts for Oracle Cloud Infrastructure (OCI) Data Science
Lesson Overview
This lesson provides a high-level introduction to essential OCI networking
components and how they relate to Data Science workloads.
It explains how to configure network connectivity for data science resources using either
default networking or custom networking.
⚠️ Note: The session does not cover in-depth networking topics; it focuses on core
concepts relevant to data science.
1. Key Cloud Networking Components
Virtual Cloud Network (VCN)
A virtual private network created within Oracle data centers.
Acts as the foundational network structure for cloud resources.
Subnets
Subdivisions of a VCN used to group resources.
Each subnet:
o Contains Virtual Network Interface Cards (VNICs) attached to instances.
o Shares the same route tables, security lists, and DHCP options.
Virtual Network Interface Card (VNIC)
Defines how an instance connects to internal and external endpoints.
Determines the IP address and network connectivity of a resource.
2. Virtual Routers and Gateways
Dynamic Routing Gateway (DRG)
Enables private network traffic between:
o Your VCN and on-premises network (via VPN or FastConnect).
o Different VCNs across regions.
Provides private communication without using public IP addresses.
Network Address Translation (NAT) Gateway
Allows outbound internet access for private resources.
Keeps those resources hidden from incoming public connections.
Service Gateway
Enables private traffic between your VCN and Oracle Services Network.
Example use case:
o A database in a private subnet can back up data to Object Storage
without public IPs or internet access.
3. Data Science Workloads and Networking
Workload Types
Notebook Sessions
Jobs / Job Runs
Model Deployments
These are collectively referred to as data science workloads.
External Asset Access
Workloads often need to access:
Code files
Data sources
Libraries or dependencies
Secrets (e.g., credentials)
Logs
These assets might reside:
On the public internet, or
Inside a private network (e.g., enterprise Git servers or databases).
4. Networking Patterns in OCI Data Science
A. Default Networking
Fastest and simplest option to get started.
Automatically attaches a workload to a pre-configured service-managed
subnet via a secondary VNIC.
Provides:
o Egress to the public internet via a NAT gateway.
o Access to OCI services via a Service Gateway.
No need to manually create VCNs or define networking policies.
Recommended if you only need:
o Internet access
o Access to OCI-managed services
B. Custom Networking
Suitable for advanced configurations and private network access.
You provide your own existing subnet for data science workloads.
When the workload is created:
o OCI Data Science connects it to your selected subnet via secondary
VNIC.
Enables:
o Secure access to private enterprise assets
o Full control over routing, security, and policies
Use custom networking if workloads must access:
On-prem databases
Private Git repositories
Other internal resources
📌 Note: Custom networking requires coordination with your network administrator
and additional IAM policies, as discussed in the Tenancy Configuration lesson.
5. Setting Up Networking for Data Science (Using VCN Wizard)
Steps:
1. Navigate to Networking → Virtual Cloud Networks.
2. Click Start VCN Wizard.
3. Select Create VCN with Internet Connectivity.
4. Enter a VCN name.
5. Click Next, then Create.
6. Wait for the VCN and associated resources to be created.
7. Click View Virtual Cloud Network to confirm successful setup.
⚠️ If you already configured your tenancy using OCI Resource Manager, this VCN is
automatically created and you don’t need to repeat these steps.
6. Summary
Networking Components Covered:
o VCN, Subnets, VNICs, DRG, NAT Gateway, Service Gateway.
Two Networking Options for Data Science:
1. Default Networking – for quick setup and public/OCI service access.
2. Custom Networking – for private network access and advanced control.
VCN Wizard provides a simple way to set up networking manually if not already
configured.
✅ Key Takeaway:
OCI provides flexible networking options to securely connect data science workloads to
the resources and data they need — whether over the public internet or within private
enterprise networks.
6. Authenticate to OCI APIs
OCI Data Science Professional – Lesson: Authenticate to OCI APIs
Instructor Introduction
Speaker: Jon Stanesby
Topic: Authentication methods for interacting with OCI APIs in the Data Science
Service
Lesson Overview
This lesson explains how to authenticate to Oracle Cloud Infrastructure (OCI) APIs
when working with Data Science resources such as:
Notebook Sessions
Jobs
Model Deployments
You’ll learn about different authentication options, including Resource Principals and
OCI configuration file with API keys, and how they are used across the ADS SDK,
OCI Python SDK, and OCI CLI.
⚠️ Note: This lesson focuses only on authentication (verifying identity), not
authorization (permissions and access control), which was covered in Lesson 2:
Tenancy Configuration.
1. Understanding Authentication in OCI
When working in OCI Data Science, your workloads may need to interact with other
OCI services through the OCI REST APIs.
For example:
Reading or writing data to Object Storage
Creating and running Data Flow applications
Accessing logs, secrets, or model repositories
To perform these tasks, your code must authenticate with OCI — meaning OCI must
recognize the identity performing the API calls.
2. Common Interfaces for API Interaction
Data scientists typically interact with OCI APIs using one of the following:
ADS SDK (Accelerated Data Science SDK)
OCI Python SDK
OCI Command Line Interface (CLI)
Each interface supports different authentication methods that are covered below.
3. What Are Resource Principals?
A Resource Principal is a special feature in OCI Identity and Access Management
(IAM) that allows cloud resources themselves (not just users) to act as authenticated
principals.
Key Points:
Each resource (like a notebook session or job run) has its own unique identity.
It authenticates using automatically managed certificates — no need for
manual credential handling.
These certificates are:
o Created and assigned automatically
o Securely stored and rotated by OCI
This removes the need to manually upload API keys or configuration files to workloads.
4. How Resource Principals Work in Data Science
The Data Science service enables workloads (like notebook sessions and job runs)
to use their own resource principal for authentication.
Benefits:
Provides secure authentication without exposing credentials.
Simplifies authentication for automated workloads (like job runs) that don’t have
an interactive interface.
Reduces operational complexity — no need to manage .oci/config or .pem files
manually.
If a resource principal is not explicitly used, the SDK or CLI will fall back to using the
OCI configuration file and API key approach.
5. Resource Principal Token Behavior
A resource principal token is cached for 15 minutes.
If you make changes to IAM policies or dynamic groups, those updates will
take effect only after the token cache expires.
You can continue development while waiting, but access permissions will not
refresh instantly.
💡 Tip: If authentication changes seem delayed, wait ~15 minutes before retesting.
6. Authenticating with Resource Principals
You can authenticate using resource principals in:
ADS SDK
OCI Python SDK
OCI CLI
Each interface uses slightly different syntax for setting up resource principal
authentication.
It’s recommended to refer to official SDK documentation or pause the lecture at the
relevant code examples to note down the commands.
7. Alternative Authentication: Configuration File & API Key
If you prefer not to use resource principals, you can authenticate as your personal IAM
user using a configuration file and API key.
Steps:
1. Upload your OCI configuration file (usually located at ~/.oci/config) into the
notebook session’s OCI directory.
2. For the profile defined in the config file, also upload or create the required .pem
key files.
3. Alternatively, use the api_keys notebook to generate and configure keys
automatically.
To Launch the api_keys Notebook:
Open JupyterLab Launcher → Click Notebook Examples → Select api_keys
notebook.
Follow the guided steps to create and set up your OCI API keys directly within
your Data Science environment.
8. Comparison: Resource Principals vs. API Key Authentication
Aspect Resource Principals Config File + API Key
Security High (no manual keys) Depends on secure key storage
Setup Automatic Manual (upload config + keys)
Automated jobs, production Personal use, testing, custom SDK
Best For
workloads scripts
Token
15 min cache Persistent
Lifespan
Aspect Resource Principals Config File + API Key
Ease of Use Simplifies automation Requires manual setup in notebook
9. Lesson Summary
Authentication is required to interact with OCI APIs securely.
You can authenticate via Resource Principals (recommended) or via
Configuration File + API Key.
Resource Principals:
o Are secure, automated, and ideal for jobs and model deployments.
o Use auto-managed certificates — no need to handle credentials manually.
Config + API Key:
o Suitable for user-level authentication, testing, or local SDK use.
Tokens refresh every 15 minutes, so policy or group changes take time to apply.
The api_keys notebook simplifies creating and managing config-based
credentials inside JupyterLab.
✅ Key Takeaway:
Use Resource Principals as the default authentication method for OCI Data Science
workloads — they are secure, fully managed, and ideal for production automation. The
configuration file + API key method remains available for personal authentication or
custom scripting.
3. Workspace Design and Setup
1. Projects
Module 2: Workspace Design and Setup
Lesson Title: Project
Instructor: Jon Stanesby
Summary
This lesson introduces projects, the central component of an OCI Data Science
workspace. Projects act as collaborative environments where data scientists organize
their work around specific use cases or business questions. The lesson explains how to
create, view, edit, and delete projects using both the OCI Console and the ADS
SDK.
Key Points
1. Definition of a Data Science Project
A project is a collaborative workspace for teams of data scientists.
It helps organize work around a particular use case or business question.
All data science resources (e.g., notebook sessions, models) are created
within a project.
Projects can be created, named, described, and tagged by data scientists.
2. Creating Projects
Two main methods:
1. From the OCI Console UI
o Log in with the necessary policies.
o Navigate: Menu → Analytics & AI → Data Science → Create Project.
o Choose a compartment to add the project to.
o Optionally add:
Unique name (auto-generated if omitted).
Short description (helps others understand the project’s purpose).
Tags for easy tracking and organization.
o To view the project after creation, select “View Detail Page” and click
Create.
2. From the ADS SDK
o Use the ProjectCatalog object and the create_project() method.
o Specify the compartment ID.
o The example uses a preset environment variable to match the
compartment of the notebook session.
3. Managing Projects
Viewing Projects
All projects appear on the Project List page.
Each project displays metadata such as:
o Display name & description
o OCID
o Creation date/time and creator
o Tags
Editing Projects
Editable fields: display name, description, and tags.
Edits can be made through any OCI interface.
Deleting Projects
A project must be empty before deletion (delete all associated resources first).
In the Console:
o Locate the project in the list or detail page.
o Click Delete.
o Type the exact project name to confirm.
Deleted projects show a “Deleting” status and can later be filtered out using
state filters (e.g., “Active”).
4. Demonstration Recap
Create a project: Enter a name, description, and tags → Click Create.
Edit a project: Change name or description, manage tags via the detail view.
Delete a project: Confirm deletion with the exact name → project status
becomes “Deleted”.
Conclusion
In this lesson, we learned:
What a data science project is and why it’s central to OCI Data Science.
How to create projects using both the Console and ADS SDK.
How to view, edit, and delete projects effectively.
2. Notebook Sessions
🧱 Module 2 – Workspace Design and Setup
Lesson 2 – Notebook Sessions
Instructor: John Stanesby
🧱 Overview
Notebook Sessions in OCI Data Science provide a fully managed JupyterLab
interface for building, training, and managing machine learning models.
They abstract away infrastructure complexity — meaning you don’t need to
handle compute or storage provisioning, patching, or lifecycle management
yourself.
⚙️ Key Features
Feature Description
Managed OCI automatically provisions and manages compute,
Infrastructure storage, and updates.
Choose from AMD, Intel, or NVIDIA shapes for performance
CPU & GPU Support
needs.
Persistent Storage Block storage retains data, notebooks, and environments.
Change shapes (scale up/down) during the notebook
Scalable Compute
session lifecycle.
Activate, deactivate, or delete notebook sessions easily via
Lifecycle Actions
Console or SDK.
🧱 Creating a Notebook Session (Console)
Prerequisite: You must have a project already created.
Steps:
1. Navigate to your Project Details page.
2. Click Create Notebook Session.
3. Select:
o Compartment
o Optional name
o Compute shape (CPU/GPU)
o Block storage size (50 GB – 10,240 GB)
o Networking option:
Default networking – service-managed VCN
Custom networking – choose your own VCN & subnet
4. Add tags (optional).
5. Click Create.
⏳ It may take a few minutes to start up. Once active, the Open button will be
enabled.
🔧 Managing Notebook Sessions
📋 Viewing
View all notebook sessions on the Notebook Sessions List page.
Metadata includes:
o Display name & OCID
o Creator and creation time
o Compute shape
o Storage size
o VCN and subnet
o Tags
✏️ Editing
When active, only the display name can be edited.
Tag editing is also possible through the details view.
🔁 Activation & Deactivation
Action Description
Starts up compute resources. You can update shape, storage, or
Activate
network configuration during activation.
Shuts down compute instance to save cost, retaining block storage
Deactivate
for data and files. Boot volume data is deleted.
Starts a new compute instance and reattaches the same block
Reactivate
volume, restoring your previous environment.
💡 Tip: Use deactivation when you’re not actively working — it saves compute
costs without losing work stored on the block volume.
🗑️ Deleting a Notebook Session
1. From the Details or List page, click Delete.
2. Confirm deletion by typing the exact display name.
3. OCI terminates the compute instance and destroys block storage.
4. The notebook moves from Deleting → Deleted status.
⚠️ Important:
Any files or code not backed up or committed before deletion are permanently
lost.
📈 Viewing Metrics
You can view metrics from the Details page:
CPU Utilization
Memory Usage
Network In/Out Traffic
Metrics help monitor resource usage during training or large data
operations.
🧱 Lesson Summary
Topic Description
Definition Fully managed JupyterLab environment for ML development.
Creation Through OCI Console with flexible compute and network setup.
Management Includes viewing, editing, activating, deactivating, and deleting.
Storage Persistent block storage for notebooks, data, and environments.
Metrics Monitor CPU, memory, and network traffic for optimization.
✅ Key Takeaways
Notebook sessions simplify model development by handling infrastructure
automatically.
Persistent storage allows safe deactivation without data loss.
Scaling compute shapes up/down is supported mid-lifecycle.
Always back up work before deleting a session — deletion destroys
storage permanently.
3. How to Work with JupyterLab
Module 2: Workspace Design and Setup
Lesson Title: How to Work with JupyterLab
Instructor: Jon Stanesby
Summary
This lesson covers JupyterLab, the next-generation web-based interface used in OCI
Data Science notebook sessions. It explains how data scientists interact with
notebooks, files, terminals, and environments through JupyterLab, highlights its features
and structure, and walks through how to use it effectively for developing and running
code.
1. Introduction to JupyterLab
JupyterLab is a web-based user interface and the main environment for
notebook sessions.
It supports interactive computing, integrates multiple document types, and
handles a wide range of file formats:
o Images, CSV, JSON, Markdown, PDF, Vega, and Vega Lite.
Data scientists use JupyterLab because it is familiar, intuitive, and makes the
OCI Data Science service easy to adopt.
2. Differences Between Open-Source JupyterLab and OCI Interface
Although the layout and menus are similar, the OCI version includes three key
additions:
1. Launcher Access – quick shortcuts to notebooks, terminals, console,
Environment Explorer, and notebook examples.
2. Environment Explorer – a GUI for managing and searching Conda
environments (covered in the next lesson).
3. GitHub Extension – enables version control and Git integration within notebook
sessions (covered later in Lesson 9).
3. Core Components of the JupyterLab Interface
A. Menu Bar
Located at the top of JupyterLab.
Contains top-level menus with keyboard shortcuts for major actions.
B. Launcher
Provides quick access to:
o New notebooks
o Consoles
o Text editors
o Terminals
o Environment Explorer
o Notebook examples
New items can be added via commands or extensions.
C. Left Sidebar
Contains several useful tabs:
1. File Browser – Navigate directories, open, add, or delete folders/files.
2. Running Terminals and Kernels – View and shut down active sessions.
3. Git Extension – Manage version control.
4. Commands Panel – Access and search all available commands.
5. Property Inspector – View notebook properties.
6. Open Tabs List – See all open files and tabs.
7. Table of Contents – Automatically generated from Markdown headings.
8. Extension Manager – View installed JupyterLab extensions.
To hide the sidebar: click an icon twice.
To reopen: use the plus (+) icon or select File → New Launcher.
D. Main Work Area
Central workspace for documents and activities.
Allows tabbed and resizable panels.
The active tab has a blue border.
Supports multiple views for live document editing.
E. Code Consoles and Kernels
Code consoles serve as interactive scratch pads.
Kernel-backed documents let you run code interactively.
Kernel management is handled from the Kernel Menu.
F. Command Palette
Provides a keyboard-driven interface to search and execute commands quickly.
G. Right Sidebar
Contains the Property Inspector for notebooks and user interface
customization.
4. Top Chrome Bar Overview
Located above JupyterLab:
Oracle Logo → Returns to the OCI Cloud Console home page.
Notebook Session Name → Navigates to the session details page in OCI
Console.
Session Remaining Time → Shows automatic logout timer (default: 1 hour
inactivity, extendable up to 24 hours).
Help Icon → Links to OCI documentation.
Sign Out → Logs out of the notebook session.
5. Using the Launcher
Access through File Browser toolbar (+ icon) or File → New Launcher.
The launcher is divided into:
o Left Side:
Environment Explorer → Discover and manage Conda
environments.
Notebook Examples → Tutorials and documentation for common
data science use cases.
o Right Side:
Create new notebooks, consoles, terminals, or text/Markdown
files.
6. Creating and Running a Notebook
1. From the launcher, select Python 3 kernel → Create Notebook.
2. A new notebook includes useful preloaded tips such as:
o Checking internet connectivity.
o Common ADS imports and environment variables.
3. You can rename the notebook via right-click → Rename.
4. Run code cells using:
o The triangle icon,
o Run menu → Run Selected Cell, or
o Keyboard shortcut: Shift + Enter.
5. While code runs:
o A star (✱) indicates the cell is executing.
o A number appears once it completes (execution order).
6. Change cell type using the dropdown (e.g., Code → Markdown).
7. Reorder cells by dragging them up or down.
7. Editing and Kernel Options
From the Edit menu, you can:
o Merge or split selected cells.
o Change the kernel (via kernel name dropdown in top-right).
8. Viewing and Running Example Notebooks
Use Launcher → Notebook Examples to open preloaded demos (e.g., Binary
Classification – Attrition).
You can run all cells or interrupt the kernel to stop execution.
Use Variable Inspector to view variables currently loaded in memory.
To view side-by-side:
o Drag tabs until screen sections highlight blue.
o To revert to full screen, drag back to the title bar.
9. Viewing Options and Settings
From the View menu, you can:
o Show line numbers.
o Collapse or expand code and output cells.
Themes and other preferences can be adjusted in Settings.
10. File Management
Create new files by right-clicking in File Explorer.
Upload files via drag-and-drop (e.g., CSV files).
File type determines view mode (e.g., CSVs display tabular data).
11. Using the Terminal
Open from Launcher → Terminal.
Supports standard Linux commands (ls, etc.).
Preinstalled tools include:
o odsc conda CLI
o git CLI
o oci CLI
12. Help and Documentation
Access via the Help Menu:
o Reference materials
o FAQs
o Official documentation
Conclusion
In this lesson, we:
Explored JupyterLab’s interface and features in the OCI Data Science
environment.
Learned how to navigate menus, panels, and the launcher.
Practiced creating and running notebooks and terminals.
Covered key functions such as variable inspection, kernel management, and
file handling.
JupyterLab provides a flexible, interactive, and integrated environment for all stages
of data science development in OCI.
4. Conda Environments: Overview
Conda Environments: Overview
Instructor: John Peach
Module Objective:
Understand what Conda environments are, their benefits, and how they are managed in
Oracle Cloud Infrastructure (OCI) Data Science Service.
1. What is a Conda Environment?
Conda is an open-source package management and environment system.
It allows bundling of:
o Python interpreter
o Python modules and external libraries
o Other required programs
o Into a single, isolated environment.
2. Benefits of Using Conda Environments
1. Install Only What You Need:
o Add or update packages selectively without affecting others.
2. Isolation of Configurations:
o Prevents conflicts between different projects or models.
o Example: One Conda for Computer Vision (TensorFlow, OpenCV), another
for Linear Regression.
3. Easy Switching Between Environments:
o Move between different setups effortlessly.
4. Shareable Configurations:
o Teams can share pre-configured environments instead of recreating
setups.
5. Consistency Across Workflows:
o The same Conda used in notebooks, jobs, and model deployments
ensures consistent results.
6. Reproducibility:
o Track exact software versions to reproduce experiments or debug model
performance changes.
3. Managing Conda Environments with Environment Explorer
Environment Explorer Overview
A graphical interface (GUI) within OCI Data Science.
Used to view, manage, and explore Conda environments.
Offers:
o Card view (for details)
o List view (for multiple environments)
o Search and filters (e.g., show only GPU-compatible or active Condas)
4. Types of Conda Environments in OCI
1. Data Science Conda Environments
Managed and curated by Oracle Data Science Service team.
Built for specific needs:
o Frameworks (e.g., TensorFlow, PyTorch)
o Industry-specific (e.g., Healthcare)
o General-purpose (e.g., Exploratory Data Analysis)
Contain pre-installed, relevant libraries and frameworks.
2. Published Conda Environments
Created and managed by users.
Stored in Object Storage Buckets for:
o Team sharing
o Cross-session use
o Reproducible model training and deployment
Example: Custom Conda with specific versions of scikit-learn and pandas.
3. Installed Conda Environments
Active environments installed in a specific notebook session.
Required for notebook execution.
Can be either Data Science Conda or Published Conda.
Persist on the block volume, meaning they remain available after notebook
sessions are deactivated and reactivated.
5. Summary
Conda environments act as containers for your project’s software stack.
They enhance reproducibility, collaboration, and isolation.
The Environment Explorer in OCI helps you manage and filter environments
efficiently.
Data Science, Published, and Installed Conda environments each serve
unique roles in development and deployment workflows.
5. Data Science Conda Environments
Data Science Conda Environments
Instructor: John Peach
Module Objective:
Understand what Data Science Service Conda Environments are, how they are
structured, their naming conventions, and the key types available for different data
science use cases in Oracle Cloud Infrastructure (OCI).
1. Introduction
Conda environments can be difficult to build from scratch.
Oracle Cloud Infrastructure (OCI) Data Science Service provides pre-built
Conda environments, designed on open-source software and easily
customizable.
These environments are accessible through JupyterLab → Launcher →
Environment Explorer.
Each OCI Conda Environment Includes:
OCI Python SDK
ADS (Accelerated Data Science) SDK
Note: The public ADS version does not include Oracle Labs AutoML and Model
Explainability.
These features are only available in specific Conda environments (e.g., General
Machine Learning for CPU/GPU).
2. Types of Conda Environments
A. Application-Based Conda Environments
Built around a specific software or framework:
Examples:
o ONNX
o PyPGX
o PySpark
o Intel Extension for Sciket-Learn
o Pytorch
o Rapids
o TensorFlow
B. Use Case-Based Conda Environments
Built for specific data science tasks or domains:
Examples:
o Computer Vision
o Data Exploration and Manipulation
o Financial Services
o General Machine Learning
o Natural Language Processing
o Neurophysiology
o Oracle Database
Each pack includes optimized “best of breed” libraries for its focus area.
3. Conda Environment Families and Naming Conventions
Environment Families
Grouped by:
Python Version
Architecture (CPU/GPU)
Example:
Natural Language Processing for CPU (Python 3.7, Version 2)
Natural Language Processing for CPU (Python 3.7, Version 1)
→ Both have the same Python version & architecture, but different software
builds.
General Machine Learning Example:
General Machine Learning for CPU (Python 3.6, Version 1)
General Machine Learning for CPU (Python 3.7)
Naming Convention:
1. Application-Based
<Software> <Version> for <Architecture> on Python <Version>
Example:
PyTorch 1.10 for GPU on Python 3.7
TensorFlow 2.7 for CPU on Python 3.7 (v1)
2. Use Case-Based
<Task> for <Architecture> on Python <Version>
Example:
Data Exploration and Manipulation for CPU on Python 3.7
4. Popular Data Science Conda Environments
🧱 Computer Vision
Focus: Image & video processing, object detection, facial recognition, tracking,
etc.
Includes libraries:
o scikit-image
o Pillow
o PyTorch
o OpenCV
Use cases:
o Object detection/tracking
o Image stitching & compression
o Eye tracking & facial recognition
📊 Data Exploration & Manipulation
Focus: Exploratory Data Analysis (EDA), visualization, and data ingestion.
Includes:
o Oracle- ADS
o Kafka-Python
o pandas, pandaparallel, dask
o Matplotlib, Seaborn, Plotly, Bokeh
Use Cases:
o Data set ingestion,Processing and visualization
o Stream consumption from Oracle Cloud Infrastructure Streaming
⚙️ General Machine Learning
Focus: Generic ML tasks & model explainability.
Includes:
o xgboost, lightgbm, Keras, TensorFlow
o Oracle AutoML
o Oracle MLX (Model Explainability Library)
o Oracle-ads
o Scikit-learn
o TensorFlow
Supports:
o Data manipulation
o Supervised ML
o Generic Machine Learning
o AutoML for model optimization
o Machine Learning explainability (MLX)
Comes with multiple notebook examples.
💬 Natural Language Processing (NLP)
Focus: Text analysis & language modeling.
Includes:
o Oracle-ads, eli5, Lime,nltk, keybert, transformers, pytorch-lightning,
simpletransformers
Supports tasks like:
o Text extraction
o Key phrase extraction
o Parts-of-speech tagging
o Deep learning–based text processing
🔁 ONNX (Open Neural Network Exchange)
Focus: Model portability and interoperability.
Features:
o Saves models in a common format for any framework.
o Supports conversion from scikit-learn, TensorFlow, etc.
o Runs via ONNX Runtime (framework-independent).
Used for:
o Model conversion & inference
o Portability between ML Frameworks
o Transfer models between ML frameworks
o ONNX Runtime library allows you to run the models on different platforms
o Generating workflow graphs
o Model deployment base environment in OCI
Includes:
o Onnx
o Onnxconverter-common
o Onnxmltools
o Onnxruntime
o Oracle-ads
🗄 Oracle Database
Focus: Working with on-premise and Autonomous Databases (ATP/ADW).
Includes:
o ipython-sql
o mysql-connetor-python
o oracle-ads
o SQLAlchemy
Enables:
o ETL jobs
o Database queries using the ADS Connetor, SQLAlchemy , and python-sql
o Perform analytics in the database without having to move the data to the
notebook
o Batch transformations
o Direct SQL queries within notebooks
Use ipython-sql for running SQL commands in cells, and ADS Connector for
easy database integration.
🔥 PyTorch
Focus: Deep learning, computer vision, and NLP.
Use cases:
o Computer vision, NLP, and General Ml
o Deep neural networks and algorithms for deep learning
o Tensor computing with strong acceleration on GPUS
Includes:
o daal4py (Intel optimization)
o oneAPI Data Analytics Library
o category-encoders
o Pandas
o Scikit-learn
o Oracle-ads
Benefits:
o GPU acceleration
o Efficient CPU-based deep learning
o Strong support for neural networks
⚡ PySpark
Focus: Distributed data processing using Apache Spark.
Includes:
o sparksql-magic
o oracle-ads
o oraclejdk
o scikit-learn
o pyspark
o MLlib for machine learning
Enables:
o Writing code in notebooks
o Python-based API for Apache Spark that contains MLlib library for
machine learning
o Develop and test your spark application in the notebook and run it on Data
Flow.
o Running jobs on OCI Data Flow (Spark)
🧱 TensorFlow
Focus: Building and deploying deep neural networks.
Use Cases:
o Machine learning
o Deep neural networks
o Flexible architecture runs on CPUs, GPUs, and TPUs
Includes:
o TensorFlow
o TensorBoard (for visualization)
o Oracle ADS
o Pandas
o Scikit-learn
o Oracle-ads
o Category-encoders
Supports:
o Image recognition
o NLP
o RNNs and other ML applications
5
. Summary
D
ata Science Conda Environments are curated by Oracle for reliability and
convenience.
They are categorized by use case, Python version, and architecture
(CPU/GPU).
They provide:
o Pre-configured tools for faster setup
o Easy customization
o Portability and reproducibility across OCI services
Popular environments include:
o Computer Vision
o Data Exploration and Manipulation
o General Machine Learning
o NLP
o ONNX
o Oracle Database
o PyTorch
o PySpark
o TensorFlow
6. Manage Conda Environments
🧱 Module: Manage Conda Environments
Instructor: John Peach
Role: Data Scientist, OCI Data Science Service Team
Overview
This lesson explains how to manage Conda environments in Oracle Cloud
Infrastructure (OCI) Data Science using the odsc command-line tool.
It provides more control than the Environment Explorer GUI and allows advanced
management tasks such as browsing, installing, cloning, modifying, publishing,
and deleting environments.
🔍 Recap: What Are Conda Environments?
A collection of software packages bundled together.
OCI provides:
o Managed Conda packs by Oracle’s Data Science Service team.
o Custom Conda packs published by users.
Management can be done:
o Through Environment Explorer (GUI), or
o Through odsc CLI for more control.
⚙️ Key Functionalities of odsc CLI
1. Browse Conda Environments
View details (name, description, libraries, slug).
Command:
odsc conda list
o Shows all OCI Data Science Conda environments.
o Add --local option to list only installed environments.
o Use --override to view published Conda environments (from Object
Storage).
2. Search Conda Environments
No direct search command.
Combine with Unix tools like grep:
odsc conda list | grep -e ‘^[ ]*name:’ -e ‘^ [ ]*slug:’
o Filters specific details (e.g., environment names or slugs).
Can also be used with tools like awk or perl for advanced searches.
3. Install Conda Environments
Install Data Science or published Conda environments.
Command:
odsc conda install --slug <slug_name>
To install from Object Storage (published):
odsc conda install --slug <slug_name> --override
4. Clone Conda Environments
Duplicate an environment safely before major modifications.
Command:
odsc conda clone --fromenv SOURCE_SLUG --env CONDA_NAME
o Automatically creates a new slug for the cloned environment.
5. Modify Conda Environments
Done using standard conda commands (not odsc).
Steps:
o You must activate the conda environment in your terminal
o conda activate /home/datascience/conda/<SLUG>/
o Now you can change/modify the environment
o For Example to upgrade ADS:
o python3 -m pip install oracle-ads --upgrade
o Changes apply only to the activated environment.
6. Publish Conda Environments
Share customized environments via Object Storage (for others, co-workers,
jobs, or model deployments).
Steps:
1. Create a bucket in Object Storage.
2. Initialize configuration:
3. odsc conda init --bucket_namespace <namespace> --bucket_name
<bucket_name>
4. Publish environment:
5. odsc conda publish --slug <slug_name>
7. Delete Conda Environments
Command:
odsc conda delete --slug <slug_name>
o Frees up storage by removing unused environments.
8. Create Custom Conda Environments
Build from a YAML manifest file.
Command:
odsc conda create --file [Link]
o Default includes base packages from /opt/[Link].
o Use --empty to skip installing base packages.
🧱 Summary
odsc CLI enables full management of Conda environments.
Key operations: browse, search, install, clone, modify, publish, delete, and
create.
Provides greater flexibility compared to the GUI Environment Explorer.
Essential for custom workflows, reproducible research, and collaboration.
7. Demo: Manage Conda Environments
🧱 Module: Demo – Manage Conda Environments
Instructor: John Peach
Role: Data Scientist, OCI Data Science Service
Overview
In this demo, you learn how to manage Conda environments using the odsc
command-line tool within Oracle Cloud Infrastructure (OCI) Data Science.
The demo covers browsing, searching, installing, cloning, modifying, publishing,
deleting, and creating custom Conda environments from a YAML file.
⚙️ ODSC Command Overview
Primary command: odsc
Subcommand for Conda operations: odsc conda
Key available options include:
list, init, show-configuration, publish, install, clone, delete, create, etc.
🔍 Browsing Conda Environments
Lists available Data Science Service–managed Conda environments.
Command:
odsc conda list
o Returns a YAML file with information such as:
Name of the Conda pack
Slug (unique identifier)
Type (dataScience = managed by the service)
To view installed environments:
odsc conda list --local
o Type will show as local.
To view published environments (from Object Storage):
odsc conda list --override
🔎 Searching Conda Environments
Since list output is in YAML, command-line tools are used for filtering:
odsc conda list | grep -e ‘^[ ]*name:’ -e ‘^[ ]*slug:’
grep allows pattern matching (using regular expressions).
Filters keywords like name or slug to make results more readable.
Example:
odsc conda list | grep data_exploration
o Filters only environments related to data exploration.
📦 Installing Conda Environments
To install a Data Science–managed Conda pack:
odsc conda install --slug <slug_name>
You will be prompted to select a version number.
After installation, verify using:
odsc conda list --local | grep <slug_name>
o Confirms successful installation.
🔁 Cloning Conda Environments
Used to duplicate an existing environment before making modifications.
Command:
odsc conda clone --fromenv <source_slug> --env "<new_env_name>"
System automatically assigns a new slug to the cloned environment.
Example:
odsc conda clone --fromenv data_exploration-py37-cpu-v3 --env
"My_Data_Exploration"
After cloning, list installed environments to confirm:
odsc conda list --local | grep -e ‘^[ ]*name:’ -e ‘^[ ]*slug:’
🧱 Modifying Conda Environments
Modifications are done using conda, not odsc.
Steps:
conda activate /home/datascience/conda/<slug_name>
python3 -m pip install pendulum
conda deactivate
This approach installs new packages or upgrades existing ones in the selected
environment.
☁️ Publishing Conda Environments
Publishes Conda packs to Object Storage for sharing or reuse.
Steps:
1. Create a bucket in OCI Object Storage (e.g., published-conda-environments).
2. Note the bucket name and namespace.
3. Initialize configuration:
4. odsc conda init --bucket_namespace <namespace> --bucket_name
<bucket_name>
5. Publish the Conda pack:
6. odsc conda publish --slug <slug_name>
7. Verify by listing published Condas:
8. odsc conda list --override | grep -e ‘^[ ]*name:’ -e ‘^[ ]*slug:’
The bucket will now contain a folder structure:
Conda Environments/
CPU/
My_Data_Exploration/
v1/
<conda-files>
🧱 Deleting Conda Environments
Remove unwanted Conda packs to free space.
Command:
odsc conda delete --slug <slug_name>
Prompts for confirmation before deletion.
🧱 Creating Custom Conda Environments (YAML File)
You can build Conda environments from scratch using a YAML manifest.
YAML defines:
o Channels
o Dependencies
o Other required packages
Command:
odsc conda create --file <my_env.yaml>
By default, installs base dependencies from /opt/[Link].
To skip these:
odsc conda create --file <my_env.yaml> --empty
🧱 Summary
In this demo, you learned how to:
Use the odsc CLI to browse, search, install, clone, modify, publish, delete,
and create Conda environments.
Combine CLI commands with tools like grep for efficient filtering.
Manage environment lifecycles directly from the OCI Data Science notebook
terminal.
8. OCI Vault: Introduction
Overview
This module explains why data scientists should use the OCI Vault service to securely
manage credentials, keys, and secrets — instead of storing them directly in code or
configuration files.
Key Concepts & Purpose
OCI Vault is a centralized, Oracle-managed service that securely stores
encryption keys and secrets (like passwords, tokens, API keys).
It helps data scientists securely connect to external databases, APIs, and OCI
services without exposing credentials in notebooks or scripts.
Integration: Works seamlessly with the OCI SDK, CLI, and API clients, and
integrates with other OCI services such as Object Storage, Block Storage, and
more.
Vault Components
1. Vaults – Logical containers for keys and secrets.
o Created in a compartment.
o Two types:
Virtual Private Vault:
Dedicated, isolated hardware partition.
Stores up to 1,000 key versions.
Can be backed up to Object Storage.
Supports disaster recovery and cross-region replication.
Default (Shared) Vault:
Shared with other Oracle customers.
Lower cost, but no backups.
Charged only for stored keys and secrets.
2. Keys – Logical entities representing cryptographic material used for encryption
and digital signatures.
o Supported algorithms: AES, RSA, ECDSA.
o AES: Symmetric (same key for encryption & decryption).
o RSA/ECDSA: Asymmetric (public/private key pairs).
Types of Keys:
o Master Encryption Key: Created or imported by the user.
o Data Encryption Key: Dynamically generated using a master key (used
for encrypting actual data).
o Wrapping Keys: Used for securely sharing or transferring encryption
keys.
→ Uses envelope encryption — master keys encrypt data keys, and data keys encrypt
data.
3. Secrets – Sensitive credentials like passwords, tokens, Usernames, SSH keys,
Auth tokens or API keys.
o Stored securely in Vault and retrieved when needed.
o Reduces risk of credential leaks from notebooks or scripts.
o Each secret has a unique OCID and supports versioning for rotation.
Key Rotation
Every key and secret can be rotated periodically.
Rotating keys generates new versions automatically or allows importing new
material.
Old key versions can still decrypt data but cannot encrypt new data.
Benefit: Reduces security risk if a key is ever compromised.
Best Practices
Always store credentials in OCI Vault — never in code or config files.
Use master keys from your Vault for encryption of storage or data resources.
Rotate keys and secrets periodically to limit exposure.
Use OCIDs in code to retrieve secrets at runtime via the SDK, CLI, or API.
Control access to secrets using IAM policies instead of notebook-level
permissions.
Summary
In this module, you learned:
What the OCI Vault service is and why it’s important.
The two types of vaults and their characteristics.
The types of keys and the role of key rotation.
How secrets help you keep credentials secure and out of your code.
How Vault integration enhances security and compliance for Data Science
workflows.
9. Using OCI Vault in OCI Data Science
Overview
This module explains how to use OCI Vault for managing encryption, keys, and
secrets in your Data Science workflows. You’ll learn the difference between Oracle-
managed and Customer-managed keys, and how to use both the OCI SDK and the
ADS SDK to store and retrieve secrets securely.
1. Encryption in OCI
OCI uses encryption everywhere (data at rest and in transit).
You’ll often be asked whether to use:
o Oracle Managed Keys – Handled automatically by OCI.
o Customer Managed Keys – Keys created and managed by you in your
own Vault.
🔹 Oracle Managed Keys
OCI handles key creation, encryption, and decryption.
Used when provisioning resources like Object Storage, Block Volume, OKE
clusters, etc.
Data is always encrypted — this cannot be disabled.
🔹 Customer Managed Keys
Stored in your own Vault.
Used when stricter security or compliance rules apply (e.g., Security Zone
compartments).
You manage key rotation, lifecycle, and access control.
The key may be imported into the Vault or generated by the Vault service.
The master key in your Vault creates data encryption keys that perform the
actual encryption.
2. Setting Up Customer Managed Keys
1. Choose Customer Managed Key when creating a resource.
2. Locate your Vault.
3. Select a master key to use for generating data encryption keys.
➡️ This allows you to control encryption behavior and key rotation independently from
Oracle.
3. Working with Secrets in Python
There are two main approaches:
A. Using the OCI SDK
The OCI SDK is a general-purpose API for working with Vaults, keys, and secrets.
Storing a Secret
1. Create a credentials dictionary (e.g., for database connection):
2. credentials = {
3. "database": "ADB",
4. "username": "admin",
5. "password": "mypassword"
6. }
7. Convert it to JSON → Base64 encode it.
(Use helper function to handle conversion.)
8. Create a Base64SecretContentDetails object containing the encoded data.
9. Create a SecretDetails object that includes:
o Compartment OCID
o Vault ID
o Key ID
o Secret name & description
o Encoded content
10. Use the VaultsClient class:
o Load the OCI config file.
o Create a VaultsClient instance.
o Call:
o create_secret_and_wait_for_state(secret_details, state='ACTIVE')
Retrieving a Secret
1. Use the SecretsClient class with the same OCI config.
2. Call:
3. get_secret_bundle(secret_ocid)
4. Access the secret from:
5. [Link].secret_bundle_content.content
6. Decode Base64 → JSON → Python dictionary.
✅ Result: You now have your secret data securely retrieved as a Python dictionary.
B. Using the ADS SDK (Simplified for Data Scientists)
The Accelerated Data Science (ADS) SDK provides specialized SecretKeeper
classes designed for common data science use cases.
These make secret management much easier than using the low-level OCI SDK.
Available SecretKeeper Classes
SecretKeeper Class Purpose
For Oracle Autonomous Database credentials (can also store
ADBSecretKeeper
wallet files).
BDSSecretKeeper For OCI Big Data Service (HDFS, Hive, etc.).
MySQLDBSecretKeeper For Oracle MySQL Database credentials.
AuthTokenSecretKeeper For authentication/access tokens (e.g., Streaming, GitHub).
Encode The Secret:
# Encode the secret.
def dict_to_secret(dictionary):
return base64.b64encode([Link](dictionary).encode('ascii')).decode('ascii')
secret_content_details = Base64SecretContentDetails
content_type = [Link].CONTENT_TYPE_BASE64
stage=[Link].Base64SecretContentDetails.STAGE_CURRENT
content=dict_to_secret((credentials))
# Bundle the secret and metadata about it.
secrets_details = CreateSecretDetails(
compartment_id=compartment_id,
description="Data Science service test secret"
secret_content = "secret_content_details"
secret_name = "Database creds"
vault_id = vault_id
key_id = key_id
)
The data is covered into JSON format, and then encoded into base64 string to
be stored in the Vault .
The contents of the secret are stored in a Base64Secert ContentDetails object
Store Secret in the Vault:
The VaultsClient class takes a configuration object and establishes a connection to the
Vault service
# Store secret and wait for the secret to become active
config = from_file([Link]([Link]("~"), ".oci", "config"), "DEFAULT")
vaults_client_composite = VaultsClientCompositeOperations(VaultsClient(config))
secret = vaults_client_composite.create_secret_and_wait_for_state(
create_secret_details=secret_details,
wait_for_states=[[Link].LIFECYCLE_STATE_ACTIVE]).data
Retrieve the Secret from the Vault:
The SecretsClient class takes a configuration object.
The get_secret_bundle method takes the secret’s OCID and returns a
Response object.
Its data attribute return the SecretBundle object. This has an attribute
secret_bundle_content that has the object
Base64SecretBundleContentDetails and the content attribute of this object
has acutal secret.
# Retrieve the secret bundle and extract the content
def secret_to_dict(wallet):
return [Link](base64.b64decode(wallet).decode('ascii'))
secret_bundle = SecretsClient(config).get_secret_bundle(secret_id)
secret_content =
secret_to_dict(secret_bundle.data.secret_bundle_content.content)
Using OCI Vault with ADS :
ADS provides a set of classes to make it easier to store and retrieve secrets in OCI
Vaults when using Oracle Autonomous Database, BDS, OCI MySQL Database, or Auth
tokens.
ADBSecretKeeper : Store and retrieve credentials for autonomous database. It
optionally has support for the database wallet.
BDSSecretKeeper : Stores and retrieves credentials for the OCI Big Data
Service.
MySQLDBSecretKeeper : Stores and retrieves credentials to Oracle MySQL
Database.
AuthTokenSecretKeeper : Stores and retrieves Auth Token or Access Token
string. This could be an Auth Token to use to connect to streaming, Github, and
so on.
MySQLDBSecretKeeper:
Understands what is needed to connect to MySQL
Store / retrieves the secrets for MySQL
Works with the ADS Database connection tool.
# Use ADS to store secret for MySQL
from [Link] import MySQLDBSecretKeeper
mysql_keeper = MySQLDBSecretKeeper(vault_id=vault_id, key_id=key_id, credentials)
mysqldb_keeper.save (name="mysql_employee", description="My DB credentials")
# Use ADS to get a secret for MySQL
with MySQLDBSecretKeeper.load_secret('[Link].oc1..<unique_ID>') as mysqldb_secret :
print(mysqldb_secret['user_name'])
4. Advantages of Using ADS SDK
Tailored for data science workflows.
Integrates directly with Autonomous DB, MySQL, Big Data Service, etc.
Automatically handles encoding, storing, and retrieving secrets.
Cleaner, shorter, and safer code compared to OCI SDK.
5. Summary
OCI encrypts all data by default using Oracle-managed or Customer-managed
keys.
Customer-managed keys offer better control and compliance (stored in your
Vault).
Secrets (like database credentials) should never be stored in code.
Use OCI SDK for full control, or ADS SDK for a data-scientist-friendly workflow.
SecretKeeper classes simplify saving and retrieving secrets securely.
10. Code Repositories (Git)
🧱 Module Overview
Instructor: John Peach (Data Scientist, OCI Data Science Service Team)
Topic: How version control systems (Git) integrate with OCI Data Science to manage
source code, Jupyter notebooks, and related resources.
🧱 1. Version Control Systems (VCS) Basics
Also known as Source Code Management (SCM) systems.
Allow tracking, managing, and reverting to different versions of files — including
code, notebooks, data, and reports.
Originally designed for software development, now essential for data science
workflows.
🔹 Examples of VCS:
CVS, Subversion, Perforce, Mercurial, Bazaar, CodeCommit —
but Git is by far the most popular and is integrated in OCI Data Science.
🗂️ 2. What is a Repository (Repo)?
A repo is like a filing cabinet for your project — it holds all files, code, and
tracked changes.
Each project (analysis, model, report) has its own repo.
Git allows:
o Version control (track and revert changes)
o Collaboration (merge changes from multiple users)
o Branching (work on different versions in parallel)
o Archiving (keep historical versions)
👥 3. Centralized vs Distributed Version Control
Type Description Examples Advantages
Single main server stores all
Centralized CVS, Simpler setup, controlled
versions; developers commit
VCS Subversion workflow.
to it.
Work offline, no single
Distributed Each user has a full copy of Git, Mercurial,
point of failure, flexible
VCS the repo. Bazaar
branching.
➡️ Git is distributed — but teams often use a hybrid model with a central peer (like
GitHub or OCI Code Repo).
💻 4. Git for Data Science
Tracks code + notebooks + reports.
Enables collaboration, experimentation, and rollback.
Fast (most operations are local).
Fault tolerant — even if the central repo is down, local copies remain.
🧱 5. Git in OCI Data Science (JupyterLab Integration)
OCI’s JupyterLab Git Extension provides a GUI for Git inside notebook
sessions.
Accessible via sidebar icon or top Git menu.
You can:
Create / clone repos
Stage and commit changes
Push / pull to/from remote repos
View diffs between versions
Supports integrations with:
OCI Code Repository
GitHub
GitLab
Bitbucket
Custom Git servers
🧱 6. Key Git Terminology
Term Meaning
Commit Snapshot of your project at a point in time (with SHA ID).
Repository Directory tracking all project versions.
Working Area The current local folder where files are being edited.
Staging Marking files to include in the next commit.
🔁 7. Git Workflow Summary
1. Modify files in the working area.
2. Stage selected files (git add).
3. Commit changes with a message (git commit -m "message").
4. Push to remote repo (git push).
5. Pull updates from remote (git pull).
Frequent commits = safer rollback and better traceability.
☁️ 8. OCI Code Repository
Acts as a Git-based, centralized peer hosted inside OCI.
Integrated with OCI IAM (Identity & Access Management).
Each repo gets an OCID, visible in the OCI Console.
Supports:
o Commit, branch, clone, delete operations
o Viewing commits and repo size
o Integration with external repos (e.g. GitHub, GitLab)
o Replication of GitHub repos into OCI via secure Vault-stored credentials
🔐 9. Connecting to External Repos
Use SSH keys or HTTPS tokens (stored in OCI Vault as secrets).
GitHub connection steps:
1. Install and configure Git locally.
2. Create a GitHub account and generate SSH keys.
3. Add the public key to GitHub.
4. Create or clone a repo.
5. Work locally → commit → push to remote.
🧱 10. Common Git Commands
Command Description
git init Create a new repo
git clone <url> Copy remote repo to local
git add <files> Stage files
git commit -m "msg" Commit staged files
git push Send commits to remote
git pull Get + merge updates from remote
git fetch Download changes (no merge yet)
git remote Manage remote connections
🧱 11. Summary
Version control = essential for collaboration and reproducibility.
Git = most widely used distributed VCS.
OCI Data Science integrates Git directly into JupyterLab.
OCI Code Repository offers secure, IAM-integrated Git hosting.
External services (GitHub, GitLab) can connect securely via OCI Vault.
Basic Git workflow: init → add → commit → push → pull.
11. Demo: Code Repositories (Git)
🧱 Module: Demo — Code Repositories (Git)
Instructor: John Peach, Data Scientist (OCI Data Science Team)
Objective:
Demonstrate how to:
1. Configure Git in a JupyterLab notebook session
2. Generate and use SSH keys for authentication
3. Create and link local and remote (GitHub) repositories
4. Commit, push, and sync notebooks between OCI and GitHub
🧱 1. Git Configuration
Git is pre-installed in OCI Data Science Notebook Sessions.
Configure username and email for commit identity:
git config --global [Link] "Your Name"
git config --global [Link] "you@[Link]"
git config --list # verify settings
These details appear in commit history to identify contributors.
📂 2. Create a Local Repository
1. In JupyterLab → open File Browser.
2. Create a new folder (e.g., demo).
3. Go to the Git sidebar → click Initialize Repository.
4. You now have a local Git repository with staging and history tools visible.
🔐 3. Generate and Add SSH Keys (for GitHub Access)
GitHub requires authentication for pushing changes.
Steps:
1. Generate a new key pair:
2. ssh-keygen -t rsa -b 4096 -C your_email@[Link]
3. list keys : ls .ssh/id_<ssh-key>
4. Two files are created:
o Private key → keep secret.
o Public key → share safely.
5. Start the SSH agent:
6. eval "$(ssh-agent -s)"
7. Add the key to the agent:
8. ssh-add -k ~/.ssh/<private key>
9. Copy the public key and add it to GitHub:
GitHub → Settings > SSH and GPG keys > New SSH Key → Paste → Save.
☁️ 4. Create and Link Remote GitHub Repository
1. On GitHub → click New Repository → Name it (e.g., demo).
2. Choose SSH clone URL (e.g., git@[Link]:user/[Link]).
3. In JupyterLab Git panel → Add Remote Repository → Paste URL.
4. Verify connection:
5. git remote -v
→ should list origin for fetch and push.
🧱 5. First Commit & Push
1. Create a new notebook (File > New Notebook).
2. It appears as Untracked in Git tab.
3. Right-click → Track → file becomes staged.
4. Add commit message → Initial Commit → Commit.
5. Push to GitHub:
6. git push --set-upstream origin master
7. Afterwards, use Push to Remote for future commits.
✅ Check GitHub → the notebook and commit message appear.
🧱 6. Updating and Syncing Changes
1. Edit notebook → Save.
2. Git tab shows modified file.
3. Stage → Commit → Push with a message (e.g., “some math”).
4. Refresh GitHub → updated file and commit are visible.
🧱 7. Summary
In this demo, you learned how to:
Configure Git user identity
Generate and manage SSH keys
Create and initialize a local Git repo
Create a matching GitHub repo
Link both via SSH
Stage, commit, and push notebook files
Maintain versioned synchronization between OCI Data Science and GitHub
4. Machine Learning Lifecycle
1. ML Lifecycle: Overview
🧱 Module: Machine Learning Lifecycle — Overview
Instructor: Wes Prichard, Senior Principal Product Manager (Data Science & AI
Services)
🎯 Objective
Understand the six-step lifecycle of building, deploying, and managing machine
learning (ML) models within OCI Data Science — from accessing data to monitoring
models in production.
🔄 Simplified ML Lifecycle (6 Steps)
Each ML project begins with a business problem and proceeds through these stages:
1. Data Access
2. Data Exploration & Preparation
3. Modeling (Building & Training)
4. Model Validation (Evaluation)
5. Model Deployment
6. Model Monitoring (and Refresh/Retirement)
⚙️ The process is iterative, not linear — data scientists refine multiple steps until
model performance meets business goals.
1⃣ Data Access
Purpose: Gather relevant data for the business problem.
Data Sources:
o Enterprise systems → data lakes, relational/non-relational databases
o OCI Object Storage, Data Lakehouse, Data Catalog
o External/public datasets, APIs, sensors, surveys, web scraping
OCI Tip: Store working data within the notebook session for fast access.
2⃣ Data Exploration & Preparation
Goal: Cleanse, transform, and understand data before modeling.
Tasks:
o Identify missing, corrupt, or duplicate data → fix/remove
o Detect and handle outliers
o Check for bias or imbalance
o Perform feature analysis (distributions, correlations, summary stats)
o Visualize data relationships (pair plots, histograms, box plots, etc.)
Feature Engineering:
o Create new features (e.g., time-of-day from timestamps)
o Convert categorical → binary (one-hot encoding)
o Normalize or scale numerical features
If data lacks labels: Use OCI Data Labeling Service to annotate datasets.
3⃣ Modeling (Building & Training)
Choose ML type:
o Supervised Learning: labeled data (classification, regression)
o Unsupervised Learning: unlabeled data (clustering, segmentation)
Process:
o Split data into training and testing sets.
o Train multiple model candidates using different algorithms.
o Experiment with feature subsets to optimize performance and cost.
Goal: Find the best-performing model for the defined objective.
4⃣ Model Validation (Evaluation)
Purpose: Assess how well the trained model performs on unseen data.
Key Evaluation Metrics:
o Classification: accuracy, precision, recall, F1-score, confusion matrix
o Regression: RMSE, MAE, R² (coefficient of determination)
o Unsupervised: cluster cohesion and separation
Metric Selection: Should align with the business goal (e.g., precision > accuracy
in fraud detection or rare disease prediction).
5⃣ Model Deployment
Goal: Make the trained model available for use.
Deployment Modes:
o Batch inference: scheduled predictions (e.g., daily churn scoring)
o Real-time inference: on-demand predictions (e.g., fraud detection)
Considerations:
o Response time (latency requirements)
o Number of requests
o Data volume
OCI Tip: Deploy models via OCI Data Science and integrate with MLOps
pipelines for automated workflows.
6⃣ Model Monitoring & Refresh
Purpose: Ensure models remain accurate and reliable after deployment.
Types of Monitoring:
1. Drift/Statistical Monitoring:
Detect data drift or concept drift (distribution changes)
Compare training vs. live data distributions
2. Operational (Ops) Monitoring:
Track latency, throughput, CPU/memory usage, reliability
Set up logs & alerts for incident analysis
When to Retrain:
o When prediction accuracy degrades
o When live data deviates significantly from training data
Outcome: Retrain → redeploy → repeat lifecycle as needed.
🔁 Key Takeaways
The ML lifecycle is cyclical and collaborative (data scientists + ML engineers +
DevOps).
Each step builds upon the previous, but feedback loops are frequent.
OCI Data Science provides managed tools for each stage — from data access
and labeling to deployment and monitoring.
2. Access Data
🧱 Lesson Title: Access Data
Instructor: Himanshu Raj — Data Scientist & Senior Training Lead (AI/ML), Oracle
Module: Machine Learning Lifecycle — Step 1: Access Data
🎯 Objective
Understand:
Why data is needed in machine learning
How data is collected
What data sources are supported in Oracle Cloud Infrastructure (OCI) Data
Science
How to connect and access data through the Accelerated Data Science (ADS)
SDK and console
🧱 1. Importance of Data
Everything we do—digitally or non-digitally—generates information.
This information is the foundation for insights and decision-making in data
science.
Data enables:
o Hypothesis-driven research
o Data-driven insights
o Problem-solving through modeling
🗝️ Without data, we cannot train models, test hypotheses, or make reliable business
conclusions.
⚙️ 2. Types of Data (by Source and Nature)
Type Description Examples
Collected over time from scheduled or daily
Batch Data Backups, data migrations
operations
Streaming
Real-time messages or logs IoT devices, user events
Data
Application Logs, event tracking, API
Generated by app events or APIs
Data calls
All this data must be brought into OCI Data Science for preprocessing and model
training.
☁️ 3. How to Access Data in OCI Data Science
You can access data either via:
Console / UI (simple uploads)
Command Line / ADS SDK (programmatic access via Python)
OCI Data Science supports multiple data sources:
🔹 a. Oracle Object Storage
Main and most common data source.
Use ADS to load data from Object Storage into a DataFrame.
Access via:
o API Key authentication
ads.set_auth(auth="api_key", profile="DEFAULT")
bucket_name = <bukect_name>
file_name= <file_name>
namespace = <namespace>
df = pd.read_csv(f"oci://{bucket_name}@{namespace}/{file_name}",
storage_options={"signer": default_signer()})
o Resource Principal (used in serverless functions)
ads.set_auth(auth="resource_principal", profile="DEFAULT")
bucket_name = <bukect_name>
file_name= <file_name>
namespace = <namespace>
df = pd.read_csv(f"oci://{bucket_name}@{namespace}/{file_name}",
storage_options={"signer": default_signer()})
Example functions:
from [Link] import DatasetFactory
[Link]("oci://<bucket_name>@<namespace>/<path_to_file>")
Use set_auth() to toggle between resource principal and key pair authentication.
🔹 b. Local Storage
Access local files using standard Pandas methods:
import pandas as pd
df = pd.read_csv("local_path/[Link]")
🔹 c. Oracle Autonomous Databases (ATP/ADW)
Supported through ads.read_sql(), which is 15× faster than pandas.read_sql()
because it bypasses ORM.
With Wallet File:
from ads import read_sql
read_sql("SELECT * FROM TABLE", connection_parameters)
Without Wallet File: (ADS ≥ 2.5.6)
connection_parameters = {
"user": "admin",
"password": "mypassword",
"host": "hostname",
"port": 1521,
"service_name": "servicename"
}
⚠️ Use bind variables to prevent SQL injection.
Performance depends on network latency; can be optimized using indexes and
efficient SQL.
🔹 d. MySQL
Same as Oracle Autonomous DB, but set engine as MySQL.
Available in ADS version 2.5.6+.
ads.to_sql(df, "table_name", engine="mysql")
🔹 e. Amazon S3
Supports public and private buckets.
Use Pandas with ADS storage_options for credentials:
df = pd.read_csv("s3://bucket/[Link]", storage_options={"key": "...", "secret":
"..."})
🔹 f. HTTP / HTTPS Endpoints
Access data via direct URL:
df = pd.read_csv("[Link]
🔹 g. DatasetBrowser
Built-in ADS utility to explore reference datasets from:
o Seaborn
o Scikit-learn
o GitHub, etc.
Functions:
from [Link].dataset_browser import DatasetBrowser
[Link]() # view all available datasets
[Link]("iris") # load a specific dataset
🔹 h. PyArrow (OCI File System Integration)
Used for big data access and processing.
The OCI FS library enables file-system-like operations for OCI File Systems.
🧱 4. Data Types (Semantic Detection in ADS)
ADS automatically detects data types when loading datasets:
Data Type Description Example
Categorical Labeled groups without numeric
Eye color, shirt size
(Qualitative) meaning
Ordinal Ordered categories with intrinsic ranking Education level
Height,
Continuous Measurable quantitative data
temperature
Datetime Temporal data Timestamp, date
You can inspect data types using:
dataset.feature_types
dataset.show_in_notebook()
📦 5. Supported Sources & Formats
✅ Supported Formats: CSV, JSON, Parquet, ORC, XLSX, etc.
❌ Unsupported Formats: TXT, DOC, PDF, raw images, lists, tuples, etc.
➡️ For DOC/PDF, ADS provides a text extraction module to convert them into plain
text.
🧱 Key Takeaways
Data is the first and most critical step of the ML lifecycle.
OCI Data Science + ADS SDK provides seamless ways to access data from
cloud, databases, and local sources.
Always ensure:
o Correct authentication method (API key / resource principal)
o Efficient queries & secure SQL
o Awareness of data types and formats
Use DatasetBrowser and PyArrow for efficient exploration and handling of large
or reference datasets.
3. Data Preprocessing
🧱 Lesson Title: Data Preprocessing
Instructor: Himanshu Raj — Data Scientist & Senior Training Lead (AI/ML), Oracle
Module: Machine Learning Lifecycle — Step 2: Data Exploration and Preparation
🎯 Objective
Learn:
Why data preprocessing is needed
Common data cleaning and transformation steps
How to use OCI Data Science’s ADS tools for preprocessing
How to split data for training, testing, and validation
🧱 1. What Is Data Preprocessing and Why It’s Needed
Preprocessing : Clean, Impute, Engineer, and normalize features.
Real-world data is imperfect — it can contain:
o Missing values
o Errors
o Duplicates
o Outliers
o Inconsistent formats
Before model training, data must be cleaned and standardized to ensure
accurate results.
Preprocessing is often the largest and most time-consuming part of the ML
lifecycle.
Preprocessing of data involves various steps :
o Combining and Cleaning Data
o Data Imputation
o Dummy Variables
o Outlier detection
o Feature Scaling
o Feature Engineering
o Feature Selection
o Feature Extraction
🔹 Data transformations and manipulations prepare raw data for meaningful analysis
and modeling.
⚙️ 2. Common Preprocessing Operations
a. Combining and Cleaning Data
Data often comes from multiple sources and needs merging.
Perform row/column operations such as:
o Append / Delete / Filter
o Join / Concatenate (vertically or horizontally)
Maintain consistency:
o Use proper formats, units, and naming conventions
o Remove duplicates and incomplete rows
ADS datasets support all operations that can be performed on a Pandas DataFrame.
b. Data Imputation (Handling Missing Values)
Missing data can occur due to human error, bad sensors, or transmission
failures.
Approaches:
o Deletion: Remove incomplete rows (not recommended).
o Imputation: Replace missing values with:
Mean / Median / Mode
Mode is preferred for categorical features.
ADS supports built-in methods for automatic imputation.
c. Encoding Categorical Data
Categorical variables must be converted into numbers before modeling.
Method Description Use Case
Label Good for nominal categories
Converts each category into a number.
Encoding (no order).
One-Hot Creates binary (dummy) columns for Best for ordinal or
Encoding each category. unordered data.
Example (Pandas):
For label encoding : From [Link].label_encoder import DataFrrameLabelEncoder.
pd.get_dummies(df['Category'])
Or use ADS fit_transform() to encode all categorical columns at once.
d. Outlier Detection
Outliers are data points that deviate significantly from others.
They can be:
o Errors or
o Valid but rare observations
Detection methods:
o Visualization: Scatterplots, Boxplots
o Statistical analysis: Deviation from normal distribution
Supervised methods require labeled data (time-consuming).
Unsupervised methods assume outliers are few and distinct from normal
samples.
e. Feature Scaling
Brings all features to a comparable scale, essential for algorithms using
Euclidean distances (e.g., regression, clustering).
Method Description Formula
(x - min) / (max -
Normalization (Min-Max) Scales values between 0–1.
min)
Standardization (Z- Centers around mean 0 with unit
(x - μ) / σ
score) variance.
Ensures that all features contribute equally to model learning.
f. Dimensionality Reduction
High-dimensional data is computationally expensive and prone to overfitting.
Two main approaches:
1. Feature Selection — Choose the most relevant features
o Variance Thresholds
o Correlation Thresholds
o Genetic Algorithm
2. Feature Extraction — Create new features from existing ones (e.g.,
PCA).
o Principal Component Analysis
o AutoEncoders
o Linear Discriminant Analysis
Reduces complexity while preserving essential information.
g. Text Data Preprocessing
Text requires specialized steps such as:
o Tokenization
o Removing stop words
o POS tagging
o Stemming / Lemmatization
o Vectorization
ADS provides built-in utilities for text transformation.
🧱 3. ADS Data Transformation Tools
🔹 1. suggest_recommendations()
Detects issues in data (e.g., missing values, imbalance, correlations).
Recommends actions (like imputation, dropping columns, etc.).
User can apply suggestions manually or accept all via dropdown.
Transformed data is retrieved with get_transformed_dataset().
🔹 2. auto_transform()
Applies all recommended transformations automatically.
Handles:
o Missing values
o Strongly correlated columns
o Imbalanced classes (upsampling/downsampling)
o Non-predictive columns (like primary keys)
Returns a cleaned, optimized dataset.
⚙️ fix_imbalance=True (default) ensures class balance automatically.
🔹 3. visualize_transforms()
Shows a visual flow diagram of all transformations applied.
Displays detected correlations, imbalance, and feature summaries.
Only visualizes automated transformations, not custom ones.
📊 Example Workflow
1. Apply suggest_recommendations() → View detected issues.
2. Accept or modify transformations.
3. Use auto_transform() → Automatically optimize dataset.
4. Visualize results via visualize_transforms() → Flowchart of preprocessing steps.
Example dataset: Employee Attrition Dataset — demonstrates type discovery,
imbalance correction, and transformation summary.
🔀 4. Splitting Data
Before Inputting data to ML algorithm, we have to split data into data into train, test, and
split.
train, test = transformed_ds.train_test_split()
Splitting ensures models generalize well on unseen data.
Split Type Purpose Default Ratio
Training Set Used to train model 80%
Testing Set Evaluate performance 10%
Validation Set Fine-tune model 10%
For large datasets, 80–90% for training is fine.
For smaller datasets, use 60–70% training to retain test representativeness.
Example in ADS:
This examples sets split to 70%, 15%, 15%
data_split = transformed_ds.train_validation_test_split(
test_size=0.15,
validation_size=0.15
)
train, validation, test = data_split
print(data_split)
🧱 Key Takeaways
Data preprocessing ensures clean, consistent, and reliable input for models.
Major steps include:
o Cleaning, merging, imputing
o Encoding, scaling, detecting outliers
o Reducing dimensionality
o Splitting data effectively
ADS provides automated transformation tools to simplify and accelerate the
workflow.
Use auto_transform and visualize_transforms to optimize preprocessing with
minimal manual effort.
4. Demo: Data Preprocessing
Lesson Title: Demo: Data Preprocessing
Instructor: Himanshu Raj – Data Scientist & Senior Training Lead (AI/ML), Oracle
Summary
This demo provides a hands-on walkthrough of data preprocessing using Oracle
Cloud Infrastructure’s Accelerated Data Science (ADS) SDK. It showcases how to
access datasets, apply automated transformations, encode categorical data, handle
class imbalance through upsampling, and finally split the data for training and testing.
Key Steps Demonstrated
1. Dataset Overview
Dataset used: Employee Attrition Dataset
Size: 1,470 rows
Features: 36 total
o 22 ordinal
o 11 categorical
o 3 constant
Contains demographic, job satisfaction, compensation, and performance-related
data.
Imbalance: Fewer employees leave compared to those who stay.
2. Loading Data
Imported required libraries including Accelerated Data Science (ADS) and
pandas.
Dataset loaded from Object Storage using:
[Link]()
Defined bucket and namespace parameters.
Set target feature as attrition.
3. Suggest Recommendations
The suggest_recommendations() tool analyzes the dataset and provides:
o Detected data issues
o Suggested fixes (e.g., correlations, imbalance)
o Ready-to-use code snippets for corrections
4. Auto Transform
The auto_transform() method automatically applies all recommended
transformations.
Optimizations include:
o Missing value imputation
o Noise reduction
o Dropping strongly correlated columns
o Handling class imbalance (upsampling/downsampling)
o Removing non-predictive columns (e.g., primary keys)
Greatly reduces manual preprocessing time.
5. Visualize Transforms
visualize_transforms() displays the sequence of transformations performed.
Helps compare results with and without auto_transform.
Provides a clear understanding of data changes applied automatically.
6. Encoding Categorical Data
Example: Encoding the job_function feature.
Used ADS’s built-in Label Encoder:
from [Link].label_encoder import LabelEncoder
Converts categorical values into numeric form.
7. Handling Class Imbalance
Used upsampling via:
from [Link] import upsample
Balanced the dataset by repeating underrepresented class samples.
Verified balance using value counts before and after upsampling.
8. Splitting the Data
After preprocessing, the dataset was split into:
o 80% training
o 10% testing
o 10% validation
Default ratio in ADS; adjustable based on dataset size.
Conclusion
The demo showcased the power of ADS preprocessing tools:
o Fast, automated, and customizable transformations
o Built-in encoders and samplers
o Visual and interactive analysis
Users can further explore the ADS documentation, GitHub labs, and sample
notebooks for deeper understanding.
5. Data Visualization
Lesson Title: Data Visualization
Instructor: Jon Stanesby – Senior Principal Instructor, Oracle University
Summary
This lesson explains the importance of Data Visualization (DV) in the data science
lifecycle and demonstrates how Oracle’s Accelerated Data Science (ADS) SDK
provides automated and customizable visualization tools. It highlights how DV simplifies
exploratory data analysis, enables better insights, and improves decision-making
through clear and flexible visual representations.
Key Concepts
1. Importance of Data Visualization
DV is a core part of Exploratory Data Analysis (EDA) — used to discover
insights and relationships early in the process.
Makes data understandable through visual representations — charts, plots,
maps, and dashboards.
Helps both technical and non-technical users interpret data easily.
Supports decision-making and storytelling with data.
2. Characteristics of a Good Visualization Tool
Connects to multiple data sources, whether on-premises or in the cloud.
AI/ML-powered analytics make it easier for non-technical users.
Pre-built connectors simplify data integration and blending.
Enables collaboration — shareable across the organization.
Offers flexibility: manual control or automated visualization.
Features like drag-and-drop and auto-layout adjustments enhance ease of
use.
3. Data Visualization in ADS
ADS Smart Visualization automatically detects data types and chooses the
most suitable plots.
Supports both automatic and custom visualizations using any plotting library.
Generates visual insights like:
o Summary statistics
o Distribution charts
o Correlation maps
o Anomaly detection (e.g., missing values, high cardinality)
4. ADS Automatic Visualization Methods
Method Description
Calculates and displays correlation matrices. Uses different
corr
methods for each data type pair.
Displays summarized information about the dataset (type, feature
show_in_notebook
overview, correlations, sample preview).
Automatically determines the best plot type for given variables
plot
(e.g., histograms, violin plots, heatmaps).
Creates visualizations for individual or multiple features. Supports
feature_plot
both univariate and multivariate plots.
5. Correlation Methods in ADS
Data Type Correlation
Description Range
Combination Method
Continuous– Linear correlation between two
Pearson -1 to 1
Continuous continuous variables
Continuous– Correlation Ratio Measures curvilinear relationships
0 to 1
Categorical (η) across categories
Categorical– Association strength between two
Cramer’s V 0 to 1
Categorical nominal variables
6. The show_in_notebook() Function
Provides a comprehensive data overview:
o Dataset type (regression, binary, multiclass)
o Row/column count
o Feature types and distributions
o Correlation map and header preview
Uses smart sampling — statistically significant subset (95% confidence level,
1% interval) to improve performance.
7. The plot() Function
Automatically selects plot type based on data:
o Categorical data: Bar chart
o Continuous data: Histogram
o Categorical vs Continuous: Violin plot
o Continuous vs Continuous: Scatter plot or Gaussian heatmap
Offers flexible axis control (x, y parameters) for exploring relationships.
8. Feature Type System in ADS
Separates data representation from data meaning.
Extends Pandas DataFrames with metadata, validation, and visualization
features.
Allows data scientists to create custom feature types with specialized plots.
Provides warnings/validation to ensure data quality.
feature_plot() Method:
feature_plot() creates custom visualization
Create univariate plots that are customized to the feature type.
Result plots across different feature types.
series.feature_plot() creates a single plot for that feature
df.feature_plot() creates a collection of plots for all the features in the dataframe.
Multiple inheritance allows you to reuse plots from other features.
9. Custom Visualizations
Users can override default ADS plotting using:
[Link](<custom_plot_function>)
Integrates seamlessly with other libraries like:
o Seaborn → Pair plots showing pairwise relationships
o Matplotlib → Custom charts (e.g., geographic earthquake plots)
Enables complete flexibility for personalized data storytelling.
10. Lesson Takeaways
Data Visualization is essential for insight generation and data storytelling.
ADS simplifies and accelerates visualization through automation and smart
charting.
Users can easily switch between automatic and manual plotting methods.
Supports both exploratory and presentation-ready visualizations.
✅ In short:
Oracle’s ADS provides a powerful, intelligent, and flexible visualization framework
that combines automation with customization — allowing data scientists to move
smoothly from exploration to insight to presentation.
6. Model Training
📘 Key Concepts
🔹 What is Model Training?
Model training builds a mathematical representation of the relationships
between:
o Features → Target (supervised learning)
o Features ↔ Features (unsupervised learning)
The output of this process is a model artifact, which captures these learned
patterns.
Determines the best algorithm for model training; considers tradeoffs in terms of
compute, storage, complexity, performance, explainability, and so on
🔹 Core Components in Training
1. Score Function:
Evaluates how well the model fits the data (e.g., accuracy, log-likelihood).
2. Loss Function (Cost Function):
Measures the difference between predictions and actual values.
o The goal is to minimize loss (e.g., MSE, cross-entropy).
o The graph example shows:
Green dots → True values
Black line → Predictions
Red arrows → Loss values
3. Update Function:
Updates model parameters iteratively (e.g., via gradient descent).
🧱 Open Source + Oracle Ecosystem
OCI Data Science integrates both Oracle proprietary and open-source frameworks.
Examples:
o scikit-learn, TensorFlow, PyTorch, XGBoost, LightGBM, SpaCy,
MXNet, Keras (open source)
o Oracle AutoML, ADS (Accelerated Data Science SDK) (Oracle tools)
You can:
Use built-in conda environments in OCI Data Science.
Install custom libraries via the terminal if needed.
Train models through:
o Jupyter Notebooks
o Conda Environments using ADS/MLX/AutoML tools
o Jobs (batch training) – covered later in Module 4.
🧱 Takeaways
Model training = learning the mathematical mapping between data and target.
OCI Data Science provides flexibility: use open-source libraries, Oracle’s own
AutoML, or both.
You can train models interactively (in notebooks) or in production (as jobs).
7. Expert Tips: Training a ML model on OCI
📘 1. Training ML Models with Jobs
You can train models easily on OCI Data Science by creating Jobs using ADS
(Accelerated Data Science SDK).
The Job defines:
o The resources (compute, memory, etc.)
o The Job Run specifies the training code and output storage.
Training code can be provided in:
o Python scripts, or
o YAML configuration files.
Source code can be hosted on GitHub, while results/artifacts can be stored in
OCI Object Storage.
⚙️ 2. Distributed Training
OCI supports distributed training to handle:
o Large datasets
o Compute-intensive workloads
Distributed training allows parallelized model training across multiple nodes —
improving speed without losing accuracy.
Supported frameworks include:
o Dask
o PyTorch Distributed
o Horovod
o TensorFlow Distributed
These can be implemented using ADS utilities within OCI Data Science.
You can choose to use Docker containers or GitHub repositories for
implementation.
🤖 3. AutoML and AutoMLX
OCI provides the AutoMLX package (included in the automlx conda
environment).
AutoMLX automatically:
o Selects the best algorithm for your data,
o Performs hyperparameter tuning, and
o Generates a ready-to-deploy model pipeline.
The process can be initialized via the INIT() function.
It supports parallel processing via the task or local engine.
📚 4. Additional Recommendations
Explore:
o ADS documentation
o AutoMLX documentation
o OCI Data Science Docs for distributed training frameworks.
Review release notes regularly for new features.
Share your experiments and projects in the Oracle University (OU)
Community.
🧱 Key Takeaways
Use Jobs in OCI to automate and scale training runs.
Perform distributed training for large or compute-heavy datasets.
Leverage AutoMLX for automated model selection and tuning.
Combine ADS + GitHub + Object Storage for a complete ML workflow on OCI.
8. Oracle AutoML: Introduction
📘 Overview
In this lesson, Jon Stanesby introduces Oracle AutoML — a part of the Accelerated
Data Science (ADS) SDK in OCI Data Science.
AutoML (Automated Machine Learning) automates the model training, selection,
feature tuning, and evaluation process — enabling faster, more efficient machine
learning model development with minimal manual intervention.
⚙️ 1. What Is AutoML?
AutoML (Automated Machine Learning) automates:
o Model selection
o Hyperparameter tuning
o Feature selection
o Evaluation and optimization
It reduces manual effort and time while maintaining accuracy.
Useful since model optimization rarely happens on the first iteration —
multiple experiments are usually required.
🧱 2. Common AutoML Approaches
AutoML systems differ in how they optimize model performance and configuration.
Here are the main approaches mentioned:
a. Bayesian Optimization
Uses a probabilistic model to explore hyperparameter performance.
Example: AutoSklearn, which applies Random Forest-based Sequential
Model-Based Optimization.
Uses meta-learning to identify similar previously-optimized datasets to guide
optimization.
b. Recommender System Approach
Keeps records of best configurations from previous datasets.
For a new dataset, recommends new configurations based on similarity and
Probabilistic Matrix Factorization (PMF).
c. Genetic Evolutionary Algorithms
Example: TPOT (Tree-based Pipeline Optimization Tool).
Uses evolutionary search to optimize pipelines built around scikit-learn.
⚡ 3. Oracle AutoML’s Feed-Forward Approach
Oracle AutoML uses a non-iterative feed-forward method:
Predicts relative performance of algorithms before training.
Uses meta-learned proxy models to make fast decisions about which pipelines
to build.
Builds and tunes only the best candidate models, improving efficiency.
🔹 Advantages:
Shorter runtime
Avoids the cold start problem
(using meta-learning trained on diverse datasets)
Predictive efficiency — evaluates algorithm potential without full training.
🧱 4. Key Benefits of Oracle AutoML
No-code or low-code automation for model development.
Automates:
o Algorithm selection
o Hyperparameter tuning
o Feature selection
o Adaptive sampling
Boosts productivity by saving compute time and reducing manual tweaking.
Delivers the best-performing model within a given time and budget.
5 .Oracle AutoML Workflow
Selects a model from a large number of viable candidate models
Tunes the hyperparameters for each model
Selects predictive features to speed up the pipeline and reduce overfitting
Ensures the model trained is generalized and works for unseen data
🔄 6. Oracle AutoML Pipline
1. Algorithm Selection → Choose top-performing algorithms.
2. Adaptive Sampling → Efficiently determine ideal dataset size.
3. Feature Selection → Remove irrelevant or noisy features.
4. Model Tuning → Optimize hyperparameters for final model accuracy.
🔍 7. Step-by-Step Details
a. Algorithm Selection
Identifies algorithms that yield the maximum predictive score.
Uses meta-learning trained on many datasets.
Ranks algorithms by expected performance before training.
b. Adaptive Sampling
Starts from small subsets → increases gradually.
Evaluates performance convergence to find minimal data sample size.
Detects unbalanced datasets and adjusts automatically.
Reduces training cost and time.
c. Feature Selection
Selects highly predictive features.
Removes features that:
o Have too many missing or constant values,
o Are uncorrelated with the target,
o Have high cardinality (too many unique values).
Ranks features using multiple techniques → identifies optimal subset using
meta-learning.
d. Model Tuning
Tunes hyperparameters for selected algorithms efficiently.
Avoids exhaustive grid search.
Example: For Decision Tree, tunes:
o Max tree depth
o Minimum split percentage
Supports:
o n_jobs → Degree of parallelism (default = -1, all cores).
o log_level → Control verbosity.
🧱 8. Evaluation & Output
Produces summaries of:
o Training data info
o Pipeline details
o Model trials
Note: Adaptive sampling is skipped for datasets with < 1000 records.
Allows:
o Custom model lists (model_list)
o Custom score metrics
Binary: roc_auc
Multiclass: recall_macro
Regression: neg_mean_squared_error
o Time budget in seconds
o Minimum feature list (to preserve essential features)
💡 9. Advantages of ADS AutoML
Shorter model training time
No-code workflow automation
Meta-learning-driven algorithm selection
Efficient adaptive sampling and tuning
Flexible control (time budget, metrics, features)
High-quality models optimized automatically
🧱 Key Takeaways
Oracle AutoML automates every major step of the ML workflow.
Uses feed-forward meta-learning for faster and more accurate model selection.
Greatly enhances productivity and reduces training time.
Provides interpretable results and flexibility to customize pipelines.
9. Demo: Oracle AutoML
🧱 Demo: Building a Classifier using Oracle AutoMLx
🎯 Objective
To build a binary classification model using Oracle AutoMLx on the Census Income
dataset (from the UCI Machine Learning Repository) — predicting whether a person
earns more than $50K/year.
🧱 Concept Recap
Machine Learning model building typically involves:
1. Data preprocessing – cleaning, imputing, feature engineering, normalization.
2. Model selection – choosing the best algorithm for the dataset.
3. Hyperparameter tuning – optimizing algorithm parameters for performance.
These tasks are time-consuming and dataset-specific.
Oracle AutoMLx automates this entire process through a Python API, reducing manual
effort and accelerating model development.
⚙️ Environment Setup
Conda Environment:
o Oracle AutoML and Model Explanation for Python 3.8 (v2.0)
Libraries Used:
gzip, pandas, numpy, matplotlib, seaborn, scikit-learn, and automlx.
from automl import AutoML
from automl import init
Notebook setup: includes %matplotlib inline and autoreload magics.
📊 Dataset: Census Income
Downloaded using:
from [Link] import fetch_openml
data = fetch_openml(name='Census-Income', as_frame=True)
Target variable: Income (>50K or <=50K)
Dataset has mixed numerical and categorical columns.
Some numeric columns (like age, hours-per-week) are mislabeled as categorical
— fixed by converting to int.
🧱 Data Preprocessing
1. Fix mislabeled data types
Convert age and hours-per-week from category → int.
2. Handle missing values
AutoMLx automatically:
o Drops features with too many missing values
o Imputes remaining ones based on feature type
3. Train-Test Split
4. from sklearn.model_selection import train_test_split
5. X_train, X_test, y_train, y_test = train_test_split(X, y, train_size=0.7)
6. Binary Encoding of Target
o >50K → 1
o <=50K → 0
⚡ Initializing AutoMLx
init(engine='local', n_jobs=2)
Engine: local (Python multiprocessing)
Creates a parallel AutoML engine.
🧱 AutoMLx Pipeline Structure
[Link](task='classification')
🔸 Steps inside the pipeline:
1. Preprocessing – Cleans, imputes, normalizes data automatically.
2. Algorithm Selection – Chooses best algorithms (e.g. XGBoost, LightGBM,
RandomForest, etc.)
3. Adaptive Sampling – Dynamically selects representative subsets of data.
4. Feature Selection – Finds optimal subset of features.
5. Hyperparameter Tuning – Optimizes parameters of selected models.
🧱 Model Training
automl_pipeline = [Link](task='classification')
automl_pipeline.fit(X_train, y_train)
AutoMLx performs cross-validation (cv=5).
Automatically disables algorithms not suitable for large datasets (e.g., SVC with
>10k samples).
Selected Model: LGBMClassifier (Light Gradient Boosting Machine)
Final AUC ROC score: 0.91
📈 Evaluation
Metric: roc_auc
(Area under the ROC curve)
Confusion Matrix Visualization:
from [Link] import confusion_matrix
[Link](cm_normalized, annot=True)
o Rows = actual values
o Columns = predicted values
o Shows model’s classification performance.
🔧 Advanced AutoMLx Features
1. Model Control
Limit AutoMLx to specific algorithms:
model_list = ['LogisticRegression']
automl_pipeline = [Link](model_list=model_list)
2. Custom Validation Set
Use your own validation split:
automl_pipeline.fit(X_train, y_train, X_val=X_valid, y_val=y_valid)
3. Tune Multiple Models
Optimize top-N algorithms:
automl_pipeline = [Link](top_n=2)
4. Change Scoring Metric
Default = neg_log_loss
Can switch to accuracy, f1, etc.
automl_pipeline = [Link](score_metric='accuracy')
5. User-Defined Scoring Function
Create a custom scoring metric:
from [Link] import make_scorer, f1_score
f1_scorer = make_scorer(f1_score)
6. Time Budget
Restrict optimization time:
automl_pipeline = [Link](time_budget=10)
7. Minimum Feature List
Force inclusion of specific features:
automl_pipeline = [Link](minimum_features=['fnlwgt', 'native-country'])
🧱 Summary Output
AutoMLx logs:
Selected algorithm (e.g., LGBMClassifier)
Selected hyperparameters
Selected features (e.g., age, education-num, marital-status)
Performance metrics (AUC, accuracy, etc.)
Use:
automl_pipeline.print_summary()
✅ Final Results
Model: LightGBM Classifier
AUC ROC: 0.91
Key Features: age, education-num, workclass, marital-status, etc.
Dropped Features: fnlwgt, native-country (optional reinclusion via
minimum_features)
🧱 Key Takeaways
Oracle AutoMLx automates all major ML pipeline steps: preprocessing →
selection → tuning → evaluation.
Provides flexibility with manual overrides (time, models, metrics, features).
Accelerates the process of building and tuning high-performance models.
Achieved strong performance (AUC ≈ 0.91) with minimal manual tuning.
10. Hyperparameter Tuning: ADSTuner
🧱 Lesson: Hyperparameter Tuning – ADSTuner
🎯 Objective
To understand how to use ADSTuner — the Oracle Accelerated Data Science (ADS)
SDK’s tool for automated hyperparameter tuning — to optimize machine learning
models.
⚙️ What Are Hyperparameters?
Hyperparameters are configuration settings that control the learning process
of a machine learning algorithm.
Unlike model parameters (which are learned from data), hyperparameters are
set before training.
🔸 Examples:
Number of trees in a Random Forest
Learning rate in a Gradient Boosting model
Kernel type in an SVM
🔍 What Is Hyperparameter Tuning?
It is the process of:
1. Selecting a set of hyperparameter values
2. Training and evaluating a model for each set
3. Choosing the best-performing combination
This process is iterative and can be computationally expensive, which is why tools like
ADSTuner automate it efficiently.
⚡ ADSTuner Overview
ADSTuner is part of the Oracle Accelerated Data Science (ADS) SDK, designed for
the Oracle Cloud Infrastructure (OCI) Data Science service.
It provides:
Multiple search strategies for hyperparameter optimization
Support for user-defined search spaces
Compatibility with any ML library (e.g., scikit-learn, XGBoost, LightGBM)
🧱 How ADSTuner Works
1. Initialization
You instantiate an ADSTuner object by referencing:
The model to tune
Optional parameters such as:
o Number of cross-validation folds
o The search strategy to use
from [Link] import ADSTuner
tuner = ADSTuner(model=my_model, cv=3, strategy="perfunctory")
2. Search Strategies
ADSTuner can search through hyperparameter spaces in different ways:
🔹 Perfunctory Search
Focuses on key hyperparameters only
Covers a small search space
Ideal for quick tests or early-stage tuning
Goal: Reduce computational cost and quickly gauge model quality
🔹 Detailed Search
Covers a larger, more exhaustive space
Tunes more hyperparameters
Used after identifying the best model type
Goal: Fine-tune model performance
🔹 Custom Search
Define your own search space as a Python dictionary
Useful if you already have intuition about which ranges or values work best
search_space = {
"n_estimators": [100, 200, 300],
"max_depth": [3, 5, 7],
"learning_rate": [0.01, 0.05, 0.1]
tuner = ADSTuner(model=my_model, strategy=search_space)
3. Running the Tuning Process
Use the .tune() method to start the search:
tuning_results = [Link](X, y)
X = features
y = target variable
You can specify stopping criteria (e.g., number of trials or maximum time) using the
exit_criterion parameter.
4. Stopping Criteria (exit_criterion)
ADSTuner will stop once one of these conditions is met:
Maximum number of trials
Time budget reached
Convergence achieved
tuning_results = [Link](X, y, exit_criterion={"MAX_TRIALS": 50})
5. Modifying the Search Space
You can:
Add or remove hyperparameters
Modify the range of numeric parameters
Adjust categorical or continuous values anytime before re-running tuning
📊 Output: Tuning Report
After running, ADSTuner generates a comprehensive report containing:
All hyperparameter trial results
Best-performing combinations
Summary statistics
Visuals comparing different parameter sets and performance scores
tuning_results.show_best(n=5)
🧱 Cross-Validation Support
ADSTuner integrates cross-validation (CV) into its search:
Ensures model robustness
Reduces overfitting risk
Uses CV results to select the best hyperparameters
✅ Key Takeaways
Concept Summary
ADSTuner
Automates hyperparameter tuning within OCI Data Science
Purpose
Supports Any ML library (scikit-learn, XGBoost, etc.)
Perfunctory (quick), Detailed (comprehensive), Custom (user-
Strategies
defined)
Cross-Validation Built-in support for CV folds
Output Tuning report with trials, best performer, and statistics
Exit Criteria Time, trials, or performance threshold
🧱 Example Workflow
from [Link] import ADSTuner
from [Link] import RandomForestClassifier
# Step 1: Define model
model = RandomForestClassifier()
# Step 2: Initialize ADSTuner
tuner = ADSTuner(model=model, cv=3, strategy="perfunctory")
# Step 3: Run tuning
results = [Link](X_train, y_train, exit_criterion={"MAX_TRIALS": 30})
# Step 4: Review best hyperparameters
results.show_best(n=3)
11. Model Evaluation
🧱 Lesson: Model Evaluation
🎯 Objective
To understand the importance of evaluating ML models, the benefits of evaluation,
and how to use ADS Evaluators in Oracle Cloud Infrastructure (OCI) Data Science for
performance analysis and benchmarking.
⚙️ 1. What Is Model Evaluation?
Model evaluation is performed after model training to determine how well a
model performs on unseen data.
It helps measure accuracy, precision, recall, and other performance metrics.
Evaluation compares the predicted outputs against the true labels using a
validation dataset.
🔸 Purpose:
To quantify the predictive power of the model and ensure it generalizes well to new
data.
🧱 2. Why Model Evaluation Matters
✅ Benefits
Benefit Description
Benchmarking Compare different models or algorithms on standardized metrics.
Identify issues like high accuracy but low precision (e.g., class
Pitfall Detection
imbalance).
Trade-off Understand where models perform well or poorly (e.g., one model
Analysis performs better in clear weather, another in poor conditions).
🧱 3. The Role of ADS Evaluator
Oracle’s Accelerated Data Science (ADS) SDK provides an Evaluation module to
simplify performance measurement and visualization.
🔹 Key Classes:
ADSEvaluator – Computes metrics and generates visual charts.
ADSModel – Wraps trained models for standardized evaluation.
🧱 Supported Evaluator Types
Type Description Output Example
Binary Two-class problems (e.g.,
Spam detection
Classification Yes/No, 0/1)
Multiclass More than two discrete Sentiment
Classification classes (Positive/Neutral/Negative)
Continuous numeric
Regression House prices, temperature
prediction
📊 4. Using ADS Evaluator
Example Setup
from [Link] import ADSEvaluator
from [Link].sklearn_model import SklearnModel
from sklearn.linear_model import LogisticRegression
from [Link] import RandomForestClassifier
# Convert fitted estimator into ADS model
lr_model = SklearnModel(LogisticRegression().fit(X_train,
y_train)).prepare(inference_conda_env="generalml_p38")
rf_model = SklearnModel(RandomForestClassifier().fit(X_train,
y_train)).prepare(inference_conda_env="generalml_p38")
# Create Evaluator
evaluator = ADSEvaluator([lr_model, rf_model], X_test, y_test)
# Display metrics
[Link]
evaluator.show_in_notebook(perfect=True)
📘 5. Evaluating Different ML Tasks
🧱 A. Binary Classification
Examples: Fraud detection, churn prediction, disease diagnosis
Common Metrics:
Accuracy
Precision / Recall / F1 Score
Hamming Loss
ROC AUC
Log Loss
Confusion Matrix
Visual Charts in ADS:
Lift & Gain Charts
Precision-Recall Curve
Normalized Confusion Matrix
Binary Classification Metrics Example :
Data has to be split into a testing and training set with the features in X_train and
X_test and the responses in y_train and y_test.
To generate metrics and charts using ADSEvaluator.
lr_clf = LogisticRegression(random_state=0, solver='lbfgs',
multi_class='multinomial').fit(X_train, y_train)
rf_clf = RandomForestClassifier(n_estimators=10).fit(X_train, y_train)
from [Link] import ADSModel
bin_lr_model = ADSModel.from_estimator(lr_clf, classes=[0,1])
bin_rf_model = ADSModel.from_estimator(rf_clf, classes=[0,1])
from [Link] import ADSEvaluator
from [Link] import MLData
evaluator = ADSEvaluator(test, models=[bin_lr_model, bin_rf_model],
training_data=train)
📌 Tip:
perfect=True plots a perfect classifier line in lift/gain charts for comparison.
To show all of the metrics in a table:
[Link]
To show all of the charts :
Evaluator.show_in_notebook(perfect=True)
To add a custom metrics :
Evaluator.add_metrics
🧱 B. Multiclass Classification
Examples: Sentiment analysis, species classification, image labeling
Common Metrics:
Accuracy
Hamming Loss
F1 Score (micro, macro, weighted)
Precision & Recall (micro, macro, weighted)
ROC AUC (per class)
Key Points:
When using ADSEvaluator, specify class levels (e.g., 0, 1, 2) in the classes
argument.
Supports multiple models for side-by-side metric comparison.
Charts include:
o Multiclass ROC Curve
o Precision-by-Label
o F1-by-Label
o Multiclass Precision-Recall Curve
Example :
Change the number of classes in from_estimator function.
lr_clf = LogisticRegression(random_state=0, solver='lbfgs',
multi_class='multinomial').fit(X_train, y_train)
rf_clf = RandomForestClassifier(n_estimators=10).fit(X_train, y_train)
from [Link] import ADSModel
bin_lr_model = ADSModel.from_estimator(lr_clf, classes=[0,1,2])
bin_rf_model = ADSModel.from_estimator(rf_clf, classes=[0,1,2])
from [Link] import ADSEvaluator
from [Link] import MLData
evaluator = ADSEvaluator(test, models=[bin_lr_model, bin_rf_model],
training_data=train)
🧱 C. Regression Models
Examples: Predicting house prices, sales, or temperature
Common Metrics:
R² Score
Explained Variance Score
Mean Squared Error (MSE)
Root Mean Squared Error (RMSE)
Mean Absolute Error (MAE)
Mean Residuals
Visual Charts:
Observed vs Predicted
Residuals QQ Plot (should form a straight line for good models)
Residuals vs Predicted
Residuals vs Observed
[Link]
evaluator.show_in_notebook()
📈 6. Interpreting Results
Compare metrics between models (e.g., LogisticRegression vs RandomForest).
Review charts for bias patterns or prediction errors.
Use visual diagnostics (like lift and residual plots) to refine or retrain models.
✅ 7. Summary
Concept Key Insight
Purpose Assess and compare ML model performance
Concept Key Insight
Tool ADSEvaluator and ADSModel classes from ADS SDK
Evaluator Types Binary, Multiclass, Regression
Key Metrics Accuracy, F1, Precision, Recall, R², MSE
Visualizations Confusion Matrix, ROC, Lift & Gain, Residual Plots
Customization Add custom metrics and enable perfect classifier comparison
🧱 Key Takeaway
Model evaluation ensures that your ML model is accurate, reliable, and generalizable.
The ADS Evaluator in OCI Data Science simplifies this process by providing ready-to-
use metrics, charts, and comparisons for multiple models — all within your
JupyterLab notebook.
12. Expert Tips: ADS Evaluators
🧱 Lesson: Expert Tips — ADS Evaluators
🎯 Objective
To demonstrate how to use ADS Evaluators in OCI Data Science to easily calculate
and visualize model performance metrics for different ML tasks — binary, multiclass,
and regression.
👨🏫 Instructor
Hemant Gahankari
Senior Principal Training Lead, Oracle University
⚙️ 1. What Are ADS Evaluators?
ADS Evaluators are built-in tools in the Accelerated Data Science (ADS) SDK that:
Simplify model performance evaluation
Automatically generate key metrics and visual charts
Support multiple model types
🧱 Three Types of ADSEvaluators
Evaluator Type Use Case Example
Binary Classifier Two-class problems Spam detection (Yes/No)
Multinomial Classifier Multi-class problems Handwritten digit recognition
Regression Evaluator Continuous output Predicting house prices
🧱 2. Why Use ADS Evaluators
ADS Evaluators provide a streamlined and unified interface for:
Computing standard metrics (accuracy, recall, F1, etc.)
Generating confusion matrices and charts
Comparing multiple models simultaneously
Evaluating both training and testing datasets
🧱 3. Practical Example: Binary Classification
▶️ Step-by-Step Workflow
1. Import required modules
2. Generate or load a dataset
3. Split data into training and testing sets
4. Train multiple models (e.g., Logistic Regression, Random Forest)
5. Wrap models with ADSModel
6. Create an ADSEvaluator object
7. Display metrics and charts
🧱💻 Code Example
# Step 1: Import libraries
from [Link] import ADSEvaluator
from [Link].sklearn_model import SklearnModel
from [Link] import make_classification
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from [Link] import RandomForestClassifier
# Step 2: Create a binary classification dataset
X, y = make_classification(n_samples=1000, n_features=10, random_state=42)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2)
# Step 3: Train models
lr = LogisticRegression().fit(X_train, y_train)
rf = RandomForestClassifier().fit(X_train, y_train)
# Step 4: Wrap with ADSModel
lr_model = SklearnModel(lr).prepare(inference_conda_env="generalml_p38")
rf_model = SklearnModel(rf).prepare(inference_conda_env="generalml_p38")
# Step 5: Create ADSEvaluator
evaluator = ADSEvaluator([lr_model, rf_model], X_test, y_test)
# Step 6: Print metrics
[Link]
# Step 7: Plot charts
evaluator.show_in_notebook()
📊 4. Output and Interpretation
Metrics
Automatically prints side-by-side comparison of metrics for both models.
Shows results for both training and testing datasets.
Example metrics:
o Accuracy
o Precision
o Recall
o F1 Score
o ROC AUC
Visual Charts
Confusion Matrix
ROC Curve
Precision-Recall Curve
Lift and Gain Charts
evaluator.show_in_notebook(perfect=True)
📌 Setting perfect=True overlays a perfect model line on performance charts for
comparison.
🧱 5. Key Advantages
Feature Benefit
Unified Interface Works across multiple algorithms and problem types
Automation Computes all metrics and visualizations in one call
Model Easily compare different models (e.g., Logistic Regression vs
Comparison Random Forest)
ADS Integration Fully integrated within OCI Data Science notebooks
Visualization Instant insight via interactive charts
✅ 6. Summary
Concept Description
Purpose Simplify model evaluation using ADS SDK
Concept Description
Evaluator Types Binary, Multinomial, Regression
Main Classes ADSModel, ADSEvaluator
Outputs Metrics + Visual charts
Function Calls .metrics → text summary, .show_in_notebook() → charts
🧱 Example: Quick Recap
# Quick 3-line evaluation
from [Link] import ADSEvaluator
evaluator = ADSEvaluator([lr_model, rf_model], X_test, y_test)
evaluator.show_in_notebook(perfect=True)
💡 Expert Tip
“Evaluators make the job of calculating and plotting metrics very simple.
Try them with different datasets and explore!”
— Hemant Gahankari, Oracle University
13. Model Explanations : Global Explainer
🧱 Lesson: Model Explanations – Global Explainer
👨🏫 Instructor
Himanshu Raj
Data Scientist & Senior Training Lead, Oracle University
🎯 Objective
Understand model explainability and its importance in the model validation phase of
the machine learning lifecycle.
Focus on global explanation techniques used in Oracle Cloud Infrastructure (OCI)
Data Science.
🔍 1. What is Model Explainability and Interpretability?
Concept Definition
The ability to explain why a machine learning model made a particular
Explainability
prediction.
Interpretability How easily a human can understand those explanations.
👉 Together, they help demystify model behavior — especially for complex models like
XGBoost, Neural Networks, etc.
⚙️ 2. Why Explainability Matters
Complex ML models act as black boxes.
Lack of transparency can limit adoption and trust.
Explainability builds confidence and aids in:
o Debugging models
o Complying with regulations
o Communicating insights to stakeholders
🧱 3. Types of Model Explanations
Type Description Scope
Global
Describes the model’s overall behavior. Entire model
Explanation
Local Describes why the model made a specific
Single prediction
Explanation prediction.
What-If Analyzes how changing input features affects Hypothetical
Explanation predictions. scenario
This lesson focuses on Global Explanation only.
🧱 4. Global Explanation Techniques in ADS (Model-Agnostic)
All global explainers in OCI Data Science are model-agnostic, meaning:
They treat the model as a black box — relying only on inputs and outputs, not internal
model weights.
Technique Purpose Visualization
Feature Permutation Measures each feature’s impact on Box Plot, Bar Chart,
Importance prediction accuracy. Scatter Plot
Feature Dependence Shows how model predictions change with Line, Bar, Heatmap,
(PDP & ICE) feature values. Violin
Similar to PDP, but isolates the true effect
Accumulated Local
of each feature (handles correlated Line or Bar plots
Effects (ALE)
features better).
🔍 5. Technique 1: Feature Permutation Importance
🧱 Concept
Measures how much the model’s error increases when a feature’s values are
shuffled (destroying its relationship with the target).
A higher increase in error ⇒ more important feature.
🧱 Steps
1. Compute baseline prediction error (using F1 for classification, R² for
regression).
2. Randomly shuffle values of one feature.
3. Recalculate prediction error.
4. Compare with baseline.
📈 If the model depends heavily on that feature → prediction error rises sharply.
📊 Visualizations
Plot Type Description
Bar Chart Shows average feature importance (longer bar = higher importance).
Box Plot Shows distribution of importance across runs.
Scatter Plot Displays feature importance per iteration.
Interpretation:
Features higher on the chart = more impact.
Bars show mean ± standard deviation of importance.
🔍 6. Technique 2: Feature Dependence Explanations (PDP & ICE)
🧱 Concept
Shows how model predictions change as a single feature’s value varies.
⚙️ Process
1. Select a feature to analyze.
2. Take multiple values from its distribution.
3. Replace the feature in all rows with each selected value.
4. Generate predictions → compare differences.
📘 Two Methods
Method Description Output
PDP (Partial Dependence Averages predictions across all Average trend line
Plot) samples for each feature value. or bar chart
ICE (Individual Plots individual predictions for each Multiple lines (one
Conditional Expectation) sample. per instance)
📊 Visual Examples
Categorical Feature (PDP):
Bar chart showing average predicted value for each category (e.g., Sex →
Female vs Male).
Numerical Feature (PDP):
Line plot of feature value (x-axis) vs average prediction (y-axis).
Two Features (PDP):
Heatmap showing average prediction intensity.
ICE Plot:
Shows individual sample prediction lines (continuous or categorical). Median line
shows trend.
🔍 7. Technique 3: Accumulated Local Effects (ALE)
🧱 Concept
Improves upon PDP by isolating the true effect of a feature — even when features are
correlated.
⚙️ How It Works
1. Divide the feature’s range into intervals.
2. For each interval:
o Compute change in prediction when the feature increases/decreases
within that range.
o Average over all samples in that interval.
3. Plot the cumulative differences.
🧱 Advantages
Handles correlated features better than PDP.
Produces more realistic feature effect curves.
📊 Visualization
Numerical Feature: Line plot showing effect relative to average prediction.
Categorical Feature: Vertical bar chart showing effect difference per category.
📈 8. Example Context (as per lesson)
Dataset: Titanic
Model: XGBoost Classifier
Built using ADS AutoML provider
Used to visualize global feature importance and dependence plots
✅ 9. Summary
Concept Description
Explainability Explains why a model made a decision
Interpretability How well humans can understand that reasoning
Global
Feature Importance, PDP/ICE, ALE
Explainers
Visual, automatic, model-agnostic, and easy to integrate in OCI
ADS Advantage
notebooks
💡 Instructor’s Key Takeaway
“Global explanations help you understand how your model behaves overall —
which features drive predictions and how strongly they influence results.”
— Himanshu Raj
14. Model Explanations: Local Explainer
🧱 Lesson: Model Explanations – Local and What-If Explainers
👨🏫 Instructor
Himanshu Raj
Data Scientist & Senior Training Lead, Oracle University
🎯 Lesson Objective
Understand how to interpret individual predictions made by ML models using Local
Explainability and What-If Analysis tools in Oracle Cloud Infrastructure (OCI) Data
Science.
🔍 1. Overview of Model Explainability Techniques
Type Description Focus
Global
Explains the model’s overall behavior. Entire model
Explanation
Explains why a model made a specific
Local Explanation Single observation
prediction.
What-If Shows how changing feature values affects Hypothetical
Explanation predictions. changes
This lesson focuses on Local and What-If explainers.
🧱 2. Local Explainability (LIME in OCI ADS)
OCI Data Science provides an enhanced version of LIME —
Local Interpretable Model-Agnostic Explanations.
⚙️ Key Idea
Even though a model’s global behavior can be complex, its local behavior (around a
single prediction) is often simpler and can be approximated using an interpretable
surrogate model, such as a linear model.
🧱 How LIME Works
1. Start with a trained model.
2. Select a specific sample (observation) to explain.
3. Generate random samples around it (local neighborhood).
4. Use the complex model to predict for each of these local samples.
5. Fit a simple surrogate model (e.g., linear regression) to these local predictions.
6. Use this surrogate model to interpret which features influenced the prediction.
🧱 3. Components of the ADS LIME Explainer
ADS LIME consists of three main sections:
Model, Explainer, and Explanations.
🧱 A. Model Section
Left column:
o Shows details about the ML model.
o Displays:
True label or value
Model’s predicted label/value
Prediction probabilities (for classification) or predicted values (for
regression)
Right column:
o Shows the sample being explained.
o For tabular data: displays features and their corresponding values.
o For text data: shows the input text itself.
⚙️ B. Explainer Section
Left column: configuration of the LIME explainer, including:
o Algorithm used (e.g., LIME)
o Type of surrogate model (e.g., linear)
o Number of generated samples (e.g., 5,000)
o Whether continuous features were discretized
Right column:
o Legend showing how to interpret colors and bar directions in the
explanation.
📊 C. Explanations Section
Displays the actual local explanation results.
🔸 For Classification
A local explanation can be generated for each class label.
In binary classification, one class’s explanation mirrors the other.
In multiclass, each class gets a separate row showing how features contribute
toward or against that class.
🔸 For Regression
Shows how each feature increases or decreases the predicted target value.
📈 Feature Importance Visualization
Shown as horizontal bar charts:
o Bars ordered by relative importance.
o Longer bars = greater influence.
Positive values (right side): feature increases prediction.
Negative values (left/gray side): feature decreases prediction.
⚖️ Explanation Quality
Evaluates how well the surrogate model mimics the original model.
Component Description
Sample Distance Shows how generated samples are distributed around the
Distribution explained sample (locality).
Indicate how accurately the surrogate model approximates the
Evaluation Metrics
black-box model (using regression or classification metrics).
🔍 4. What-If Explainer
🎯 Purpose
Helps understand how changes in input features impact the model’s predicted
outcome.
It includes two main exploration techniques:
Method Description
Interactively change feature values for a single observation and see
Explore Sample
how predictions change.
Explore
Examine model predictions over feature distributions (1D or 2D).
Predictions
🧱 A. Explore Sample
Opens a graphical interface for one observation.
User can:
o Adjust feature values manually.
o Click Run Inference to recompute predictions.
The interface shows both:
o Original feature values.
o Updated feature values.
Allows observing how small changes affect predictions.
📈 B. Explore Predictions
Analyzes how prediction values vary across feature distributions.
Case Description Visualization
One Plots relationship between a single feature (x-axis) and Line or scatter
Feature prediction (y-axis). plot
Two Uses both features as x and y axes. Color scale
Heatmap
Features represents predicted value (target).
Example:
x = Age, y = CRIM (crime rate)
Color intensity = predicted target value
✅ 5. Summary
Concept Description
Local Explainability Explains why a specific prediction was made by approximating
(LIME) model behavior locally.
What-If Explainer Tests how changes in input values affect predictions.
ADS LIME
Model, Explainer, Explanations
Components
Evaluation Metrics Measure surrogate model accuracy and sample locality.
Horizontal bar charts, sample distance plots, prediction
Visualization Tools
heatmaps.
💡 Instructor’s Key Takeaway
“While global explanations describe what the model has learned overall,
local and what-if explainers reveal why a model made a particular prediction
and how changes in input features would alter that outcome.”
— Himanshu Raj
15. Expert Tips: Explainers
🧱 Topic: Expert Tips — Explainers
Instructor: Hemant Gahankari (Senior Principal Training Lead, Oracle University)
Key Concepts:
The Explainer objects are part of Oracle AutoMLx (AutoML for Python).
To use them, you need to have the automlx_p28_cpu conda environment
active.
Explainers help understand how models make predictions — providing both
local and global interpretability.
⚙️ Steps to Use Explainer Objects
1. Import and Initialize AutoMLx
2. import automlx
3. [Link]()
4. Train an Estimator
5. from automlx import AutoML
6. model = AutoML()
7. [Link](X_train, y_train)
8. Obtain the Explainer Object
9. explainer = [Link]()
10. Call Explainability Methods
o Local Explanation: Explains one specific prediction.
o Global Explanation: Explains model behavior across all data.
11. explainer.local_explanation(sample=X_sample)
12. explainer.global_explanation()
📚 Extra Notes
Depending on your data type, AutoMLx automatically chooses:
o TabularExplainer for structured/tabular data.
o TextExplainer for text data.
These Explainers can visualize feature importance, partial dependence, and
individual feature impact.
Documentation reference: ML Explainer Interface inside AutoMLx.
16. Model Catalog: Overview
🧱 Lesson: Model Catalog — Overview
Instructor: Jon Stanesby
Course: OCI Data Science
Purpose: To understand how models are stored, tracked, versioned, and deployed
within the OCI Model Catalog.
🧱 What Is the Model Catalog?
The Model Catalog provides:
A centralized, immutable repository for storing and managing ML models.
Enables model provenance, reproducibility, and auditability.
Ensures all models can be traced back to their exact training artifacts.
🔒 Immutability: Once saved, models cannot be modified — to make changes, a
new version must be created.
📦 Model Artifact Components
Each model artifact (a .zip file) contains:
1. [Link] → Python script that loads the model and defines the inference logic.
2. [Link] → Defines the Conda environment and dependencies for
deployment.
3. [Link] (optional) → Contains introspection tests to verify the model
artifact.
4. [Link] → Lists dependencies for running validation.
5. [Link] → Provides setup and saving instructions.
6. Model file(s) (like .pkl, .onnx, etc.)
🧱 Important: Files above the level of [Link] in the directory are ignored during
deployment — keep everything at or below that level.
⚙️ Key Files Explained
🧱 [Link]
Contains:
o load_model() → Loads serialized model into memory.
o predict() → Defines inference endpoint logic.
Can include helper functions (e.g., for feature transformations).
📄 [Link]
Specifies:
Conda environment to use (training or custom)
Environment slug, path, and type
Supported Python versions: 3.6 and 3.7
Example structure:
inference_env_slug: "generalml_p38_cpu_v1"
inference_env_type: "data_science"
inference_env_path: "mybucket@namespace/envs/generalml_p38_cpu_v1"
python_version: 3.7
📊 Metadata & Documentation
Each model includes four documentation types:
1. Input/Output Schema
o Describes expected features and payload format.
o Defines contract between client and model API.
2. Model Provenance
o Auto-extracted from Git if available.
o Tracks:
Training code, environment, and data
Compute resources and configurations
3. Model Introspection Tests
o Optional pre-save checks verifying operational health.
o Generates a local test_json_output.json.
4. Model Taxonomy
o Describes:
Use case (e.g., regression, binary classification)
Framework (e.g., TensorFlow, scikit-learn)
Algorithm type and hyperparameters (JSON)
Optional custom metadata fields
🧱 Custom metadata allows adding key-value pairs, with optional category and
description fields.
⚠️ Combined metadata file size limit: 32 KB.
🗂️ Artifact Size Limits
Console upload: ≤ 100 MB
ADS SDK / CLI upload: ≤ 20 GB
🔐 Access & Policies
Like all OCI resources, IAM policies are required.
Policies control:
o Model catalog management
o Model deployment
o Object Storage access
Example policy:
Allow group DataScientists to manage data-science-model-family in compartment
<compartment_name>
🧱 Key Benefits
✅ Centralized model storage
✅ Full model version control
✅ Provenance and reproducibility
✅ Integration with ADS SDK and OCI Console
✅ Seamless deployment to OCI Data Science endpoints
17. Model Serialization
🧱 Lesson: Model Serialization
Instructor: Jon Stanesby
Course: Oracle Cloud Infrastructure (OCI) Data Science
Focus: How to serialize, save, and manage models within the OCI Model Catalog.
🧱 What Is Model Serialization?
Serialization is the process of converting an object (e.g., a trained ML model) into a
storable or transmittable format — and deserialization is the reverse process.
Also known as marshaling, it allows models to be:
Saved to disk
Transferred between systems
Reloaded later for inference or retraining
🗂️ Common Serialization Formats
Format Use Case / Description
JSON Human-readable; ideal for configuration or structured data
XML Hierarchical data storage
Common for large numerical arrays and neural networks (e.g., Keras
HDF5
models)
Pickle
Python’s native binary format for serializing objects
(.pkl)
Joblib Efficient for NumPy arrays and large scikit-learn models
🧱 ADS Model Serialization
The Accelerated Data Science (ADS) SDK supports multiple ML frameworks:
Framework ADS Serialization Class
scikit-learn SklearnModel
TensorFlow TensorFlowModel
PyTorch PyTorchModel
XGBoost XGBoostModel
Generic Models GenericModel
⚙️ It’s not possible to have a specific serializer for every framework — use
GenericModel for unsupported ones.
💾 Saving Models to the Model Catalog
The save() method:
1. Packages model artifacts ([Link], [Link], model files, etc.)
2. Reloads the latest versions of these files from disk.
3. Optionally runs introspection tests (if ignore_introspection=False).
4. Uploads artifacts to the Model Catalog.
5. Returns the Model OCID.
6. Introspect() can be called after [Link]()
If issues are detected during introspection, ADS provides remediation suggestions.
🧱 Preparing a Generic Model
When working with a custom or unsupported framework:
Use prepare_generic_model() to wrap it into an ADS model object
The GenericModel class works with any unsupported model framework that has
a .predict() method.
The verify() method simlulates a model deployment by calling the load_model()
and predict() methods in [Link] file
With the .verify() method, you can debug your [Link] file without deploying any
models.
The .save() method deploys a model artifact to the model catalog.
The .deploy() method deploys a model to a REST endpoint.
You have to serialize your method in the GenericModel class.
🔧 Ways to Save and Manage Models
You can use:
1. ADS SDK (in Python)
2. OCI Python SDK
3. OCI Console (UI)
Most data scientists prefer ADS SDK, as it automates the artifact generation and
introspection.
🧱 Model Management Operations
After saving to the Model Catalog, you can:
Operation Description
View / Edit Change metadata (name, description, tags, taxonomy).
Move models between compartments (e.g., from “Development” to
Move
“Production”).
Activate / Toggle model usability in deployments. Inactive models cannot be
Deactivate deployed but remain available.
Delete Permanently remove models (soft-deleted for 30 days).
Tagging Apply defined or free-form tags for organization.
🕒 Deleted models stay in the list for 30 days and can be filtered via the state filter.
📋 Metadata & Provenance Views in OCI Console
Provenance View: Shows training source, notebook session, and Git details.
Taxonomy View: Displays description, algorithm, framework, and custom
metadata.
Schema View: Displays input/output schema definitions (read-only).
Introspection Tests: Lists test results (Success, Failed, Not Tested).
✅ Always ensure all introspection tests pass before saving a model.
🧱 CLI / SDK Capabilities
Using CLI or Python SDK, you can:
Create
Update
List
Delete
…models within the Model Catalog programmatically.
🧱 Key Takeaways
✅ Serialization enables storing and reusing ML models
✅ OCI Model Catalog keeps artifacts immutable and versioned
✅ Supports both framework-specific and generic models
✅ Enables full lifecycle management — save → validate → deploy
✅ Integrates seamlessly with ADS SDK, OCI Console, and CLI
18. Model Deployment
🧱 OCI Data Science – Model Deployment
👨🏫 Instructor
Himanshu Raj – Senior Training Lead, AIML at Oracle
🔹 Overview
After training and evaluating models, the best candidates are stored in the Model
Catalog.
Model Deployment enables you to serve predictions using those models — either for
batch or real-time consumption.
⚙️ Model Deployment Flow
1. Client Application sends API calls to a deployed model endpoint.
2. The deployed model is hosted in OCI Data Science.
3. The service manages compute, environment, and scaling automatically.
Two types of prediction consumption:
🕒 Batch consumption: Scheduled (hourly, daily, etc.)
⚡ Real-time consumption: Triggered instantly (e.g., fraud detection)
🧱 Model Deployment Architecture
Component Description
Load Balancer Distributes traffic across multiple model servers.
VM Instance Pool Hosts model server, conda env, and model artifact.
Model Artifact Contains the model file and prediction code ([Link]).
Conda Environment Includes all third-party dependencies (e.g., NumPy, XGBoost).
Logs Emit logs to OCI Logging for monitoring & debugging.
🛠️ Creating a Model Deployment
1⃣ From Console
Steps:
1. Enter deployment name
2. Select model from catalog
3. Choose compute shape and instances
4. Configure logging service
5. Set load balancer bandwidth
💡 Bandwidth Tip:
If payload = 1024 KB, requests = 120/sec
→ Bandwidth = 1024 × 120 × 8 / 1024 × 1.2 = 1152 Mbps
2⃣ From ADS SDK
[Link](deployment_properties)
or define properties directly using the .deploy() method.
3⃣ From OCI CLI
oci data-science model-deployment create --config-file [Link]
Optionally include log_config.json for access & prediction logs.
🚀 Invoking Model Deployments
Send HTTP requests with feature vectors → get predictions in response.
Can use:
o OCI CLI
o OCI Python SDK
o OCI Java SDK
Payload limit: 10 MB
Timeout: 60 seconds
If latency critical → use streaming inference
Must use Base64 encoding
🔧 Managing Model Deployments
From the OCI Console or via SDK/CLI, you can:
Operation Description
View/Edit Check OCID, compute, logs, etc.
Invoke Call the predict endpoint.
Update Change model, name, VM shape, or instances.
Deactivate / Reactivate Stop or restart deployments.
Delete Remove deployment (metadata preserved until deletion).
💡 When inactive, all compute billing stops but metadata is retained.
📊 Monitoring Model Deployments
🔸 Using OCI Logging
Access Logs: Capture all HTTP requests.
Predict Logs: Capture logs from [Link].
🔸 Using OCI Monitoring
Built-in metrics:
o CPU utilization
o Memory utilization
o Network utilization
o Request count
o Latency
o Bandwidth
You can:
o View metrics in Metrics Explorer
o Create Alarms for thresholds
✅ Summary
Model Deployment allows real-time or batch inference via HTTP endpoints.
Components: Load Balancer, VM Pool, Model Artifact, Conda Env, Logs.
Create via Console, ADS SDK, or CLI.
Monitor via Logging and Monitoring services.
Deactivation halts billing, reactivation restores endpoint access.
19. Demo: Model Deployment
🧱 OCI Data Science – Demo: Model Deployment
👨🏫 Instructor
Himanshu Raj – Senior Training Lead, AIML at Oracle
🎯 Objective
Hands-on demo showing how to create, configure, and manage a model deployment
in the OCI Data Science project environment.
🧱 Steps to Create a Model Deployment
1⃣ Navigate to Model Deployments
In your OCI Data Science Project, click on your project (e.g., test-ds).
On the left sidebar, select Model Deployments.
Click on Create Model Deployment.
2⃣ Basic Setup
Compartment: Ensure you’re in the correct compartment (e.g., OCI Data
Science Compartment).
Name: Enter a unique name (up to 255 characters).
o If not provided, OCI generates one automatically.
o Example: test-model-deploy
Description: Optional. Example: This is to test model deployment.
3⃣ Select Model
Choose an active model from the Model Catalog.
o Example: RF Classifier
Click Select → Submit.
4⃣ Configure Compute
Compute Shape: Choose the compute configuration for deployment.
o Example: 1 OCPU, 15 GB memory
Number of Instances:
o Determines scalability.
o Example: 2 instances → handles more concurrent requests by distributing
load.
5⃣ Enable Logging (Optional)
Click Select under Logging.
Two types of logs:
o Access Logs: Capture request details.
o Predict Logs: Capture stdout and stderr outputs from prediction code.
Select log names → click Next / Summary.
6⃣ Advanced Options – Load Balancing
Define load balancer bandwidth (Mbps).
Formula:
Bandwidth = (Payload Size in KB × Requests/sec × 8 / 1024) × 1.2
Example:
o Payload = 1024 KB
o Requests = 120/sec
o → Bandwidth = 1152 Mbps
For demo: used 10 Mbps.
7⃣ Create Deployment
Click Create.
Wait for the deployment to initialize.
Status changes to Active once ready.
🔍 Post-Deployment Details
📄 General Information
Displays:
o Deployment name, OCID, description, compartment.
o Associated model, owner, and tags.
📊 Metrics Dashboard
Shows:
o Success Rate
o Request Count
o CPU Utilization
o Memory Utilization
o Network Utilization
🧱 Logs & Work Requests
Two log categories available:
o Predict Logs
o Access Logs
Work Requests: Track creation status (e.g., Succeeded – 100% Complete).
🚀 Invoking the Model Deployment
Once active, you can invoke the model endpoint via:
Interface Command/Usage
HTTP Endpoint Use provided link for REST API call.
OCI CLI Use sample CLI command shown in console.
Python SDK Invoke through OCI Python SDK scripts.
Java SDK Use Java SDK to call endpoint.
🔧 Managing the Deployment
You can:
Deactivate / Reactivate deployment (pauses or resumes instances).
Delete deployment when no longer needed.
o Frees up resources and stops billing.
✅ Summary
Demonstrated end-to-end model deployment process in OCI.
Covered setup, configuration, compute selection, logging, and load balancing.
Showed how to monitor metrics, view logs, and invoke predictions.
Illustrated deactivation/reactivation for efficient resource management.
20. Demo: Model Deployment using Tensor Flow
🧱 OCI Data Science – Demo: Model Deployment using TensorFlow Model Class
👨🏫 Instructor
Himanshu Raj – Senior Training Lead, AIML at Oracle
🎯 Objective
Demonstrate end-to-end model deployment in Oracle Cloud Infrastructure (OCI)
Data Science using the Accelerated Data Science (ADS) library and TensorFlow
model class.
🧱 Overview
ADS provides framework-specific model classes (TensorFlow, PyTorch, Scikit-learn,
etc.) that help you register, prepare, verify, and deploy models into OCI Data Science
with minimal code.
⚙️ Step 1: Setup & Authentication
🧱 Imported Libraries
ads → main interface to OCI Data Science
logging → manage log output
os → interact with OS paths
pandas → tabular data handling
tempfile → manage temporary directories
tensorflow → build & train ML models
tensorflow_datasets → load built-in datasets (e.g., Fashion-MNIST)
warnings → suppress warnings
🔐 Authentication
Configured using Resource Principal authentication (recommended for OCI notebook
sessions).
🧱 Step 2: Dataset – Fashion-MNIST
Training set: 60,000 images
Test set: 10,000 images
Each image: 28×28 grayscale, labeled with 10 classes (e.g., shirts, shoes).
Visualized dataset using matplotlib.
🏗️ Step 3: Build and Train TensorFlow Model
Model Architecture
Layer Type Description
1 Flatten Converts 2D image into 1D vector
2 Dense(128, ReLU) Fully-connected layer
Layer Type Description
3 Dropout Prevents overfitting
4 Dense(10, Softmax) Output layer – 10 classes
Training Details
Optimizer: Adam
Loss: Sparse Categorical Cross-Entropy
Metric: Accuracy
Dataset scaled to [0, 1]
Used first 10,000 samples to reduce compute time
Example result: loss = 0.7899, accuracy = 0.7235
🧱 Step 4: Create Model Serialization Object
TensorFlowModel() constructor wraps the trained model and creates an ADS
model object.
This object provides helper methods to:
o Prepare artifacts
o Verify model
o Save to catalog
o Deploy
o Predict
from [Link].tensorflow_model import TensorFlowModel
tf_model = TensorFlowModel(estimator=model, artifact_dir="/tmp/artifacts")
🧱 Step 5: Prepare Model Artifacts
Files Generated by .prepare()
File Description
input_schema.json Defines input feature types and structure
model.h5 Serialized TensorFlow model
output_schema.json Defines output format
[Link] Runtime environment & conda setup
[Link] Contains load_model() and predict() functions
Default model format: HDF5 (.h5)
The .prepare() step also captures metadata like code provenance and model
parameters.
🧱 Step 6: Review Metadata
After preparation, the TensorFlow model object includes:
runtime → environment name, conda pack, Python version
model_provenance → training data and source code info
_input / _output → schema details
metadata_custom / metadata_taxonomy → key-value metadata for
classification, framework, and use case
🧱 Step 7: Verify the Model
Use .verify() method to test [Link] without deploying.
Ensures:
o load_model() loads correctly
o predict() works as expected
tf_model.verify()
✅ Verification successful → speeds up debugging and avoids failed deployments.
🗂️ Step 8: Save Model to Model Catalog
Use .save() method to register model in OCI Model Catalog.
Returns Model OCID.
The model appears in console under Models → Active Status.
🚀 Step 9: Deploy Model
Use .deploy() to create model deployment.
You can specify:
o display_name
o description
o instance_type
o instance_count
o bandwidth
o logging groups
deployment = tf_model.deploy(display_name="demo-tf-model")
Progress tracked using .summary_status()
Once Active, the model becomes available as HTTPS endpoint.
🤖 Step 10: Invoke Predictions
Two modes:
1. Local – via model’s .predict() method before deployment
2. Deployed – via same .predict() method, which now sends requests to the live
endpoint
🧱 Step 11: Cleanup
Always remove resources after testing:
1. Delete deployment using .delete_deployment()
2. Delete model from catalog
3. Delete local artifact directory
This ensures no unnecessary billing and keeps workspace clean.
✅ Summary
Step Action Outcome
1 Import & Authenticate Setup OCI + ADS
2 Load Dataset Fashion-MNIST
3 Train Model TensorFlow Sequential
4 Create Model Object ADS TensorFlowModel
5 Prepare Artifacts 5 core files generated
6 Verify Local functional test
7 Save Register to Model Catalog
8 Deploy Create HTTPS endpoint
9 Predict Serve predictions
10 Cleanup Delete deployment + model
🧱 Key Takeaways
ADS TensorFlowModel class abstracts all deployment complexity.
Verify before deploy to save time and cost.
Model Catalog acts as the central repository for reproducibility.
Deployment cleanup is mandatory to avoid billing.
21. LLM Training & LangChain Integration
🧱 Lesson: Large Language Model (LLM) Training & LangChain Integration
Course: OCI Data Science
Topic: Training and integrating large language models using OCI Data Science Jobs
and ADS
🔹 Overview
OCI Data Science Jobs provides fully managed infrastructure for training large
language models (LLMs) at scale.
It supports both:
o Full-parameter fine-tuning
o Parameter-efficient fine-tuning
Using the Accelerated Data Science (ADS) library, you can start training jobs
directly from GitHub repositories — without modifying the source code.
🔹 Steps for Fine-Tuning a Model
1. Access Pre-trained Model
o Obtain the model from Meta or Hugging Face.
2. Define the Training Job
o Use the ADS Python API to define the training configuration.
3. Create and Start Job Run
o Launch the job run via API.
o Stream the job run outputs in real-time.
4. Job Run Workflow
o Sets up the Conda environment and installs dependencies.
o Fetches source code from GitHub and checks out the specified commit.
o Runs the training script with defined arguments.
o Downloads model and dataset automatically.
o Saves outputs and checkpoints to OCI Object Storage when training
completes.
🔹 Infrastructure Handling
No need to manually define:
o Number of nodes
o Number of GPUs
ADS automatically configures compute resources based on:
o Replica count
o VM shape specified
🔹 Post-training Output
Fine-tuning results (checkpoints) are saved to your OCI Object Storage
bucket.
These outputs can be used for deployment or further evaluation.
🔹 Integration with LangChain
OCI Generative AI Service supports:
o Text generation
o Summarization
o Embedding models
These models can be integrated with LangChain through ADS.
Authentication
By default, ADS uses the authentication method configured with:
ads.set_auth()
Optionally, you can specify authentication explicitly using the auth keyword (e.g.,
resource principal).
🧱 Key Takeaways
OCI Data Science + ADS = streamlined, scalable, end-to-end LLM fine-tuning
and integration workflow.
Fine-tune pre-trained models without modifying code.
Automatic infrastructure management.
Seamless integration with LangChain and OCI Generative AI for downstream
NLP tasks.
22. Demo: Deploy LangChain based RAG to OCI
Data Science
🤖 Lesson: Deploying LangChain-based RAG to OCI Data Science
Course: OCI Data Science
Topic: Deploying Retrieval-Augmented Generation (RAG) Applications
🔹 Overview
This demo demonstrates how to deploy a LangChain-based Retrieval-Augmented
Generation (RAG) application to OCI Data Science using the ADS library and OCI
Generative AI models.
The process covers:
Building embeddings and retrievers
Creating a LangChain retrieval QA pipeline
Preparing and deploying the model as an OCI Data Science model
🔹 Steps in the Demo
1. Import Dependencies
Import necessary classes from ADS and LangChain libraries.
2. Authenticate
Use Resource Principal authentication for secure, OCI-native access.
from [Link] import AuthType
auth = AuthType.RESOURCE_PRINCIPAL
3. Create Embeddings and Model
Use:
o GenerativeAIEmbeddings class to create vector embeddings
o GenerativeAI class to create the LLM model
from [Link] import GenerativeAIEmbeddings, GenerativeAI
embedding = GenerativeAIEmbeddings()
llm = GenerativeAI()
4. Load and Process Documents
Create a text loader to load documents.
Split the document into manageable chunks for retrieval.
from langchain.document_loaders import TextLoader
loader = TextLoader("docs/ai_foundations.txt")
documents = loader.load_and_split()
5. Create Vector Store
Build a vector database using the previously generated embeddings and
documents.
from [Link] import FAISS
vector_store = FAISS.from_documents(documents, embedding)
6. Create Retriever and Chain
Initialize a retriever from the vector store.
Create a Retrieval QA Chain using:
o The retriever
o The LLM model
from [Link] import RetrievalQA
retriever = vector_store.as_retriever()
chain = RetrievalQA.from_chain_type(llm=llm, retriever=retriever)
7. Prepare Model for Deployment
Create a temporary directory for storing model artifacts.
Use the ChainDeployment class to package the LangChain model.
Prepare artifacts using:
chain_deployment.prepare()
This generates files like:
o [Link]
o [Link]
o input_schema.json
o output_schema.json
8. Verify Model
Test the prepared model locally before deployment.
[Link](prompt="What is AI Foundations course?")
✅ The response is successfully generated → model verified.
9. Save and Deploy
Save model to the OCI Model Catalog.
Deploy the model using:
[Link](display_name="LangChain_RAG_Model")
Deployment typically takes 10–15 minutes.
Once active, the model is accessible via HTTPS endpoint or SDK calls.
10. Test the Deployed Model
Use the predict() method to query the deployed model.
response = [Link]("Who are instructors for AI Foundation?")
print(response)
✅ Output:
Instructors for the Oracle Cloud Infrastructure AI Foundations course
are Hemant Gahankari, Himansha Raj, and Nick Commisso.
🔹 Outcome
Successfully deployed a LangChain RAG model to OCI Data Science.
Verified real-time question-answering capability using retrieval-based
contextual data.
Model integrates seamlessly with other LLM applications through OCI
endpoints.
🧱 Key Takeaways
OCI Data Science + LangChain enables production-grade deployment of RAG
apps.
Use ADS model classes for easy preparation, verification, saving, and
deployment.
Embedding-based retrieval enhances accuracy and context relevance in
responses.
End-to-end deployment requires minimal configuration thanks to ADS
automation.
23. Demo: OCI Data Science Operators
🧱 Lesson Summary — OCI Data Science Operators
Concept:
OCI Data Science Operators are low-code, prebuilt AI/ML tools that simplify tasks
like:
⏳ Forecasting — The Forecasting Operator leverages historical time series
data to generate accurate forecasts for future trends.(e.g., weather forecasts)
⚠️ Anomaly Detection — The Anomaly Detection Operator is a low-code tool
for integrating Anomaly Detection into any enterprise application.(Credit card
fraud)
🔒 PII Detection — The PII Operator aims to detect and redact Personally
Identifiable Information (PII )in datasets by combining pattern match and
machine learning solution.
Features:
Run via CLI or notebook session
Executed inside Conda environments (preconfigured)
Each operator comes with YAML templates for configuration
Process shown in demo:
1. Install operator Conda environment (e.g., Forecasting)
2. Initialize operator:
3. ads opctl init -t forecasting -o my-forecast
4. Configure [Link]
5. Activate environment:
6. conda activate forecast
7. Run operator:
8. ads opctl run -f my-forecast/[Link]
9. Check results:
Generated in /results → contains [Link], [Link], etc.
🧱 Lab / GitHub Resources
You can find all official OCI Data Science Operator labs and YAML examples on
GitHub here:
🔗 OCI Data Science Samples Repository:
👉 [Link]
Inside that repo, look for folders like:
/operators/forecasting/
/operators/anomaly-detection/
/operators/pii/
Each directory contains:
Example notebooks (.ipynb)
YAML configuration files
Quick start command examples (exactly like the demo)
24. Demo: OCI AI Quick Actions
🚀 Lesson Summary — Demo: OCI AI Quick Actions
🧱 What It Is
AI Quick Actions is a new low-code interface in OCI Data Science that lets you:
Quickly deploy pre-trained Large Language Models (LLMs)
Fine-tune them with your own dataset
Evaluate model performance — all from the OCI Data Science Notebook UI
It’s designed for developers who want to use models like Mistral 7B, Meta LLaMA, or
Cohere models without deep infrastructure setup.
🧱 Demo Workflow
1. Open Notebook Session
o In OCI Data Science, launch your Notebook Session (JupyterLab
interface).
2. Click “AI Quick Actions”
o Found on the top toolbar of your notebook.
o Opens the AI Quick Actions window.
3. Tabs Overview
o Models Tab: Shows available out-of-the-box LLMs (e.g., Mistral 7B
Instruct).
o Deployments Tab: Manage your deployed models.
o Evaluations Tab: Assess models using custom datasets.
4. Deploy an LLM
o Click “Create Deployment”
o Choose:
Model: e.g. Mistral 7B Instruct v0.02
Compute Shape: [Link].8NR.1
Log Group (optional)
o Then click Deploy
5. Use & Test
o Once active:
You’ll get a REST endpoint
You can test prompts directly in the notebook (e.g., “Tell us about
Las Vegas”)
Adjust generation parameters (temperature, max tokens, etc.)
🧱 Related Labs & Resources
You can find OCI AI Quick Actions examples and documentation here:
🔗 OCI Data Science AI Quick Actions Labs / Samples
👉 [Link]
actions
Inside this directory, you’ll find:
🧱 Example notebooks (e.g., deploy_llm_quick_actions.ipynb)
⚙️ Steps to configure policies for Quick Actions
💡 Example of testing the model endpoint with Python
🧱 Official Documentation
For setup (policies, permissions, prerequisites):
📘 OCI Data Science — AI Quick Actions Documentation
5. MLOps Practices
1. MLOps Architecture
🧱 Lesson: MLOps Architecture
Instructor: Lyudmil Pelov — Product Manager, Oracle Cloud Infrastructure (OCI) Data
Science & AI Services
Course: OCI Data Science Professional
Topic: MLOps (Machine Learning Operations) on Oracle Cloud
🔍 What is MLOps?
MLOps (Machine Learning Operations) applies DevOps principles to machine
learning systems.
It standardizes, automates, and governs the entire ML lifecycle, from data preparation
to deployment and retraining.
Goal:
To make ML development and deployment more efficient, consistent, and scalable —
just like how DevOps transformed software delivery.
⚙️ Core Concepts
DevOps Practice MLOps Equivalent Description
Continuous Integration Model & Data Incorporates new datasets and model
(CI) Validation versions
Continuous Model Release to Automates safe rollout of trained
Deployment (CD) Production models
Continuous Training Retrains models automatically when
Unique to MLOps
(CT) new data arrives
🔄 Why Continuous Training Matters
Unlike traditional software, data constantly changes.
This causes data drift — when the statistical properties of input data evolve, reducing
model accuracy over time.
To prevent drift degradation, models must be:
Continuously monitored,
Retrained on fresh data,
Validated before redeployment.
🧱 Maturity Levels in MLOps Automation
Level Description Example
1⃣ Manual Jupyter-based experimentation; data OCI Data Science
(Experimental) prep, & training done manually Notebook
2⃣ Automated ML Automated training/validation ADS pipelines or OCI
Pipeline triggered by new data DevOps triggers
Level Description Example
OCI DevOps + Model
3⃣ Full CI/CD Fully automated model retraining,
Catalog + Deployment
Pipeline testing, and redeployment
Pipelines
🏗️ MLOps Architecture in OCI
1. Data Ingestion
o Business or sensor data flows into OCI Object Storage or data lake.
2. Model Development
o Data scientists use OCI Data Science Notebook sessions for
exploration and training.
3. CI/CD Integration
o New data or notebook changes trigger OCI DevOps Build Pipelines.
4. Model Catalog
o The trained model is versioned and stored in the OCI Model Catalog.
5. Deployment Pipeline
o The OCI DevOps Deploy Pipeline takes the validated model, deploys it
to an endpoint for testing.
6. Testing & Approval
o The model is validated internally; if approved, it’s promoted to
production.
7. Monitoring
o OCI monitors model metrics and production data.
If performance drops, it triggers retraining — closing the continuous
learning loop.
🧱 End-to-End Flow Summary
Data → Notebook → CI/CD Pipeline → Model Catalog → Test Deployment → Approval
→ Production → Monitoring → Retrain
🧱 Key OCI Services Used
OCI Data Science – Model development (Jupyter, ADS SDK)
OCI DevOps – CI/CD pipelines for model automation
OCI Model Catalog – Versioning and governance
OCI Monitoring / Logging – Performance tracking
OCI Object Storage – Data and model artifact storage
2. Data Science Jobs
🧱 Lesson: Data Science Jobs
Instructor: Lyudmil Pelov — Product Manager, Oracle Data Science & AI Services
Course: Oracle Cloud Infrastructure (OCI) Data Science Professional
Topic: Jobs Service — Automating and Scaling MLOps Tasks
⚙️ 1. What Are OCI Data Science Jobs?
The OCI Data Science Jobs service allows you to run repeatable, automated
machine learning and data tasks on fully managed infrastructure — only when
needed.
It is a key MLOps enabler because it:
Automates parts of the ML lifecycle (data prep, model training, evaluation,
inference).
Reduces cost by provisioning compute only for the duration of the job.
Eliminates manual setup and maintenance of batch processing systems.
🧱 2. Key Concepts
Concept Description
Template defining what to run — includes the code (artifact), compute shape,
Job
environment, and configurations.
A specific execution instance of a Job. You can override parameters or
Job Run
environment variables for each run.
Each Job can have multiple runs, e.g., to test different model hyperparameters or
input data.
🧱 3. Components of a Job
1. Job Artifact – The code or script or instruction to execute (Python, Bash,
ZIP/TAR file).
o Immutable once uploaded.
o Required
o Defines the entry point for execution (JOB_RUN_ENTRYPOINT).
2. Compute Shape – Determines CPU/GPU and memory resources.
o Can be edited between runs.
3. Environment Variables & CLI Arguments –Parameters that can be customize
each job run.
4. Logging & Storage – Define log groups and block storage size.
5. VCN (Virtual Cloud Network) – Optional; enables access to internal or secure
resources.
6. Max Runtime – Up to 30 days per job run.
🚀 4. The Job Lifecycle
A job run is the actual processor that executes the instructions in the artifact and follows
the parameters set in a job and a job run.
Create Job → Configure Artifact & Compute → Start Job Run → Monitor & Log →
Complete or Cancel → Deprovision
Create Job :
o Name
o Artifact
o Environment Variables
o Command Line Arguments
o Compute : CPU, GPU
o Logging
o Block Storage
o Max Run Time
o VCN
Run Job :
o Name
o Environment Variables
o Command line Arguments
o Logging
o Max run time
Moniter + Log :
o Compute Metrics
o Service Logging
o Custome Logging
End :
o Finish/Cancel
o Deprovisioning
o Events
🧱 5. Supported Artifact Types
Artifact Type Description Use Case
Python / Bash
Single-file artifact Quick one-off tasks
Script
Full project, can include YAML runtime
ZIP / TAR Archive Complex ML pipelines
configuration
Defines environment, dependencies, and Reproducibility and
Runtime YAML
variables control
OCI provides Python preinstalled and supports Conda environments for dependency
isolation.
🔐 6. Integration and Access
Jobs can securely access OCI resources (Object Storage, ADW, Databases).
Jobs can also integrate with on-prem or third-party systems via VCN or OCI
Vault credentials.
Supported interfaces include:
o OCI Console (UI)
o OCI CLI
o SDKs (Python, Java, Go, Ruby, JavaScript)
o Terraform
o CI/CD pipelines (e.g., Bitbucket, GitHub, Jenkins)
⚡ 7. Batch Inference Modes
Type Description Example Frequency
Processes full dataset
Regular Batch Daily model scoring Moderate
periodically
Processes smaller data Fraud detection every
Mini Batch High
slices more frequently few minutes
Distributed Splits massive datasets into Parallel model Long-running,
Batch parallel jobs training or analytics high-scale
Each approach balances speed, resource usage, and data volume.
📈 8. Scaling Resources
You can scale up or down:
o Compute shapes (CPU/GPU cores, memory)
o Block storage size
Scaling is available for Jobs and Notebook Sessions within OCI Data Science.
🧱 9. MLOps Use Cases for Jobs
✅ Data preprocessing pipelines
✅ Automated model training and evaluation
✅ Batch predictions (inference)
✅ Data validation or transformation
✅ Periodic retraining or scoring workflows
🧱 10. Summary
OCI Data Science Jobs provides:
Fully managed, on-demand compute for ML workflows.
A repeatable, secure, and cost-optimized way to automate tasks.
Integration with OCI and third-party ecosystems.
Flexibility for regular, mini, and distributed batch pipelines.
In essence:
“Jobs automate the heavy lifting of MLOps — from one-off scripts to full-scale pipelines
— while you pay only for what you use.”
3. Demo: Create Artifacts
🧱 Concept Recap: What’s a Job Artifact?
A job artifact in OCI Data Science is basically a Python script (or a zip file of a Python
project) that contains the code to execute inside an OCI Job.
It defines what happens when your job runs — e.g. data processing, model training,
batch inference, etc.
⚙️ Example: Simple Python Job Artifact
You can create a file named job_artifact.py:
from datatime import datetime
import os
import argparse
NAME = [Link]("NAME", "UNDIFINED")
parser = [Link]()
parser.add_argument("-g","--greeting", required=False, default="Hello")
args = parser.parse_args()
print(f'Job Run {[Link]().strftime("%Y-%m-%d %H:%M:%S")}')
print(f"{[Link]}, Your Environment Variable has value of : {NAME}")
print("Job Done.")
💻 Optional: Add OCI SDK for Resource Principal Authentication
If you want to use OCI SDK (for example, to access Object Storage), you can extend it
like this:
import argparse
import oci
import os
# Resource Principal
TENACY_OCID, dsc = None, None
COMPARTMENT_OCID = [Link]("PROJECT_COMPARTMENT_OCID","UNDEFINED")
RP = [Link]("OCI_RESOURCE_PRINCIPAL_VERSION","UNDEFINED")
if not RP == "UNDEFINED":
# LOCAL RUN
config = [Link].from_file("~/.oci/config","BIGDATA")
dsc = oci.data_science.DataScienceClient(config=config)
TENACY_OCID = config["tenancy"]
else:
# JOB RUN
singer = [Link].get_resource_principals_signer()
dsc = oci.data_science.DataScienceClient(config={}, signer=singer)
TENACY_OCID = singer.tenancy_id
# You 2 weeks ago * - adding base job wiht RP example
# Command Line Arguments
parser = [Link]()
parser.add_argument("-g", "--greeting", required=False, default="Hello")
args = parser.parse_args()
# print
print(
f'{[Link]} {[Link]("Name","Unknown")} in tenancy OCID
{TENACY_OCID}!'
)
# OCI SDK Clinet
if not COMPARTMENT_OCID or COMPARTMENT_OCID == "UNDEFINED":
shapes = dsc.list_job_shapes(compartment_id=COMPARTMENT_OCID)
print([Link][0])
else:
print("No PROJECT_COMPARTMENT_OCID set!")
print("Job Done")
🧱 Step-by-Step Lab (to simulate the demo)
1. Create the file locally
2. nano job_artifact.py
Paste the code above.
3. Test it locally
4. export NAME="Haroon"
5. python job_artifact.py -g "Welcome"
✅ Output:
Job started at: 2025-10-06 10:05:33
Welcome, Haroon!
Job finished at: 2025-10-06 10:05:34
6. Upload as a Job Artifact in OCI Console
o Go to OCI → Data Science → Jobs → Create Job
o Under Job Artifact, upload your job_artifact.py
o Set Environment Variable → NAME=Haroon
o Add Command-line argument → -g Welcome
o Choose a compute shape
o Click Run
4. Demo: Create and Manage Jobs
🧱 LAB: Create and Manage OCI Data Science Jobs
🧱 Objective
You’ll learn how to create a Data Science Job, upload your Python artifact, and
configure compute, logging, and networking.
🧱 Prerequisites
Access to an Oracle Cloud tenancy
A Data Science Project already created
A Python artifact file (like job_artifact.py from the previous lab)
🧱 Step-by-Step Instructions
1⃣ Navigate to the Data Science Service
Log in to the OCI Console
Click the ☰ (hamburger menu) → Analytics & AI → Machine Learning →
Data Science
2⃣ Create a Project (if not created yet)
Click “Create Project”
Enter a name and description (e.g. DataScience_Jobs_Demo)
Click Create
3⃣ Create a Job
In your project dashboard, select Jobs in the left sidebar
Click Create Job
Fill out the job details:
Setting Description
Name Optional — e.g. GreetingJob
Description e.g. “Simple Python job artifact demo”
Compartment Leave default or choose one
Artifact Upload your Python file job_artifact.py
Max upload size 100 MB via Console (use SDK for larger uploads)
4⃣ Configure Environment and Arguments
Under Environment Variables, add:
NAME = Haroon
Under Command-line Arguments, add:
-g Hey
This matches your artifact’s parameters (--greeting or -g and environment variable
NAME).
5⃣ Set Runtime and Compute
Max runtime (minutes): 100
Compute Shape: Choose one of:
o Fast Launch (pre-warmed): quicker startup
o Custom Configuration: allows GPU/Intel shapes (e.g., VM.GPU3.1)
6⃣ Enable Logging (Recommended)
Under Logging, click Select
Choose or create a Log Group
Select “Automatically create log for every job run”
Click Select
Logs will go to OCI Logging Service.
7⃣ Configure Storage
Select Block Storage Size (e.g. 50 GB)
This storage is automatically attached to your job run — adjust size if processing
large datasets.
8⃣ Configure Networking
Choose Default Network if you don’t have custom VCN requirements.
(This provides basic internet and OCI service access.)
Advanced users can select Custom Network with their own VCN/Subnet.
9⃣ Create the Job
Review your settings
Click Create
✅ Result:
Your job will be created — but not yet running.
This is a template describing the infrastructure and artifact.
🔟 Manage Your Job
After creation, you can:
View General Info, Job Artifact, Logging, Infrastructure Config
Edit job name, description, compute shape, or storage
Download the artifact
Move the job to another compartment
Add tags
Delete the job
To run the job, you’ll create a Job Run — that’s the next step in the lesson series.
🧱 Example Summary Configuration
Setting Value
Job Name GreetingJob
Artifact job_artifact.py
Env Var NAME=Haroon
Argument -g Hey
Max Runtime 100 mins
Compute Shape VM.Standard2.1 (Fast Launch)
Logging Enabled (auto-create)
Storage 50 GB
Network Default Network
5. Demo: Start and Manage a Job Run
🧱 LAB: Start and Manage a Job Run
🧱 Objective
Learn how to start, monitor, clone, and cancel OCI Data Science Job Runs.
🧱 Prerequisites
You have already created a Data Science Project
You have a Job created with a valid Python artifact (from the “Create and
Manage Jobs” lab)
🧱 Step-by-Step Instructions
1⃣ Go to Your Job
Open OCI Console → Data Science Service
Navigate to your Project → Jobs
Select the job you created earlier (e.g. GreetingJob)
2⃣ Start a Job Run
Click Start Job Run
You’ll see a configuration screen before the run starts.
3⃣ Configure Job Run Options
You can override parameters from the original job if you want:
Setting Example
Logging Keep default or choose another log group
Environment Variable NAME = Haroon Khan
Command-line Argument -g Hello again!
Max Runtime 60 minutes
After reviewing, click Start.
4⃣ Watch the Job Run Lifecycle
The job will progress through several states:
State Description
Accepted OCI acknowledges your job submission
Provisioning Infrastructure (VM + storage) is being created
Running Your Python artifact is executing
Succeeded / Failed Job completed successfully or with an error
Cancelled Job manually stopped
⏳ The provisioning and running phases are billed — metering stops automatically
when the run ends or is canceled.
5⃣ Monitor the Run
You can monitor in the Jobs → Job Runs table
View columns such as:
o Status
o Lifecycle Detail
o Created By
o Start Time
If a run fails, click it → Logs → check OCI Logging for error output (e.g. Python
exception, missing variable, etc.)
6⃣ Run Multiple Jobs in Parallel
You can start another job run while one is executing — for example, to test different
model parameters or hyperparameters.
Click Start Job Run again
Provide new environment variables or arguments
(e.g. different greeting or learning rate for ML model)
✅ Multiple job runs can execute simultaneously.
7⃣ Clone a Job Run
If you want to re-run a previous job with the same parameters:
Select the previous Job Run
Click Clone
Change only what’s needed (like updating the greeting or timestamp)
Click Clone Job Run
This saves time by keeping all prior environment variables and arguments intact.
8⃣ Cancel a Running Job
If you made a mistake or need to stop execution:
While the job is Provisioning or Running, click Cancel
Confirm cancellation when prompted
Once canceled:
Infrastructure is destroyed
Billing stops immediately
9⃣ Delete or Rename Jobs
You can also:
Edit a job’s name or tags
Delete unused jobs (to keep your workspace clean)
Download artifacts for local modification or debugging
🧱 Example Lifecycle Summary
Action Description
Start Job Run Launches new job execution
Clone Job Run Duplicates settings from past runs
Action Description
Cancel Job Run Stops infrastructure + billing
View Logs Debug failed or completed runs
Parallel Runs Run multiple jobs simultaneously
🧱 Pro Tip
Use job runs to automate batch tasks or model retraining by scheduling them
through the OCI CLI, Python SDK, or DevOps CI/CD pipelines.
6. Demo: Scaling
Demo: Scaling – OCI Data Science
Presenter: Lyudmil Pelov, Product Manager – Oracle Data Science & AI Services
Overview:
This demo explains how to scale up or down jobs and notebook sessions in Oracle
Cloud Data Science Service to optimize CPU, memory, and storage usage.
🔹 Scaling Jobs
If a job shows high CPU, memory, or storage utilization, you can edit the job
to change its compute shape or storage size.
Steps:
1. Select the job → Click Edit.
2. Under Change Shape, select a new shape:
Fast launch shape
Standard shape (e.g., VM.Standard2.1, VM.Standard3.4, etc.)
GPU shapes for model training tasks.
3. Optionally, increase block storage size.
4. Click Save Changes and start a new job run — it will use the updated
configuration immediately.
🔹 Scaling Notebooks
You can monitor metrics such as CPU and memory utilization for your notebook
instance.
If performance is low, scale up your notebook shape.
Steps to scale a notebook:
1. Deactivate the notebook first.
o Click Deactivate → Confirm → Notebook stops billing but keeps block
storage data safe.
2. Once inactive, click Activate again.
o During activation, choose a new compute shape (e.g., Intel, Flex, or
GPU) and optionally increase storage.
3. Click Activate to restart the notebook with new resources.
Note:
Deactivating a notebook stops billing but preserves data in block storage. When
reactivated, all previous files remain intact.
✅ Key Benefits
Dynamically adjust resources for performance optimization.
Cost-efficient: Pay only for active sessions or job runtimes.
Flexible compute options for both CPU- and GPU-intensive tasks.
Safe scaling: Data persistence ensured across shape changes.
7. Jobs Monitoring and Logging
Demo / Module: Jobs Monitoring and Logging – OCI Data Science
Presenter: Lyudmil Pelov, Product Manager – Oracle Data Science and AI Services
📘 Overview
This module focuses on monitoring and logging in OCI Data Science Jobs — the final
step in the job lifecycle before the infrastructure is deprovisioned.
It explains how to track job performance, metrics, logs, and events for effective
troubleshooting and optimization.
🔹 Monitoring
Purpose:
Monitoring helps check the health, capacity, and performance of cloud resources in
real time.
Components:
1. Metrics – Continuously emitted data points that measure:
o CPU and GPU utilization
o Memory usage
o Network bytes in/out
o Disk utilization
2. Alarms – Passive monitoring that triggers when metrics cross thresholds (e.g.,
CPU > 80%).
o Sends notifications via Slack, SMS, or Email through OCI Notifications
Service.
Use Cases:
Identify resource bottlenecks.
Scale up compute or storage when workloads increase.
Debug performance issues.
Track and maintain system health.
🔹 Logging
Purpose: Capture and record information about job execution and artifacts for
debugging and auditing.
Types of Logs:
1. Service Logs
o Automatically emitted by job runs to the OCI Logging Service.
o Capture both standard output (stdout) and standard error (stderr)
streams.
o Requires job run’s resource principal to have logging permissions.
o Recommended to enable for all jobs.
2. Custom Logs
o Defined by the user for specific contexts or outputs.
o User specifies where logs are stored.
o Multiple job runs can share the same log or use individual logs.
Automatic Logging:
You can enable automatic log creation, letting the Data Science service create
logs within your log group automatically.
Even if a job or job run is deleted, logs persist and must be managed manually.
🔹 Event Service
Purpose:
Detect and respond to changes in resources (like job or job run lifecycle events).
How It Works:
Events represent Create, Read, Update, Delete (CRUD) operations.
Users can create rules to monitor specific events and trigger actions.
Actions include:
o Notifications (email, Slack, etc.)
o Oracle Functions
o Streaming
Multiple actions can be tied to a single rule.
OCI guarantees at least one delivery for each action.
✅ Key Takeaways
Monitoring and logging provide visibility, diagnostics, and automation for Data
Science jobs.
Metrics & Alarms → Detect and respond to performance issues.
Service & Custom Logs → Capture outputs for debugging and auditing.
Event Service → Automate actions based on resource state changes.
Logs and events remain accessible even after job deletion for analysis and
record-keeping.
8. Data Science Pipeline
Lesson: Data Science Pipeline
Instructor: Hemant Gahankari, Senior Principal Training Lead – Oracle University
📘 Overview
This lesson introduces Data Science Pipelines in Oracle Cloud Infrastructure (OCI)
Data Science Service — a powerful feature that allows users to build and automate
end-to-end machine learning workflows composed of multiple tasks (called steps).
Pipelines help orchestrate complex data science processes such as data
preprocessing, model training, evaluation, and deployment in a structured,
reusable, and automated manner.
🔹 What Is a Data Science Pipeline?
A Pipeline is a workflow made up of one or more steps.
Each step performs a specific task — e.g.,
o Step 1: Data preprocessing
o Step 2: Model training (could include multiple models)
o Step 3: Model evaluation
o Step 4: Model deployment
Steps can run in sequence (with dependencies) or in parallel to improve
efficiency.
Pipelines allow integration of different environments and programming
languages within one workflow (e.g., Python for preprocessing, Java for model
training).
🔹 Pipeline Configuration
Each Pipeline and Pipeline Step has its own configuration:
Pipeline-Level Configuration
Compute Shape – Defines the processing power and memory.
Block Storage – Determines storage capacity.
Environment Variables – Used to pass data or parameters (e.g., dataset paths).
Logging Settings – Enable or define log destinations.
Maximum Runtime – Defines time limits for execution.
Default configurations apply to all steps unless overridden.
Step-Level Configuration
Can override pipeline-level defaults.
Each step can be implemented as either:
1. A Script (Python, Bash, or Java file — single or zipped).
2. An OCI Data Science Job (referenced using its OCID).
Steps can access OCI resources (Object Storage, Database, etc.) if proper IAM
policies and VCN configurations are set.
🔹 Pipeline Lifecycle
1. Creating – Pipeline is being set up.
2. Active – Pipeline is ready for execution.
3. Pipeline Run – Each execution instance of a pipeline; multiple runs can be
created.
4. Deletion – Pipeline can be deleted once no longer needed.
🔹 Demonstration Scenario
In the demo setup:
Step 1 – Data Preprocessing
Reads dataset from OCI Object Storage.
Performs operations like:
o Dropping unnecessary columns
o Label encoding
o Scaling features
o Splitting into train/test datasets
Saves processed data back to Object Storage.
Updates Pipeline Variables with file locations for downstream steps.
Step 2 – Model Training
Trains multiple models using algorithms such as:
o Linear Regression
o Random Forest
o XGBoost
Stores all trained models in the OCI Model Catalog.
Step 3 – Model Evaluation and Deployment
Retrieves all trained models from the Model Catalog.
Evaluates and selects the best-performing model.
Deploys the best model for inference.
✅ Key Takeaways
A Pipeline automates and connects the full machine learning workflow.
Steps can be independent or dependent, and run sequentially or in parallel.
Configuration flexibility allows different environments and compute settings per
step.
Integration with Object Storage, Model Catalog, and Jobs enables seamless
ML automation.
Pipelines are essential for scalability, reproducibility, and efficiency in modern
data science projects.
9. Demo: Data Science Pipeline
☁️ Demo: Data Science Pipelines — OCI Data Science
🎯 Objective
To understand how to create and run a Data Science Pipeline in OCI to automate an
end-to-end machine learning workflow — from data preprocessing to model
deployment.
🧱 What Is a Data Science Pipeline?
A pipeline is an automated workflow that chains multiple steps of a machine learning
lifecycle — data processing, model training, evaluation, and deployment — into a single
executable sequence.
🗂️ Setup
Use the Oracle Samples GitHub repository →
🔗 oci-data-science-ai-samples
Navigate to:
pipelines/samples/employee-attrition/
The folder contains multiple .zip files — each representing a step in the pipeline:
o [Link] → Data preprocessing
o [Link] → Train Linear Regression model
o [Link] → Train Random Forest model
o [Link] → Train XGBoost model
o evaluate_deploy.zip → Evaluate models & deploy best one
🏗️ Creating the Pipeline
1. Create a Data Science Project
In OCI Console → Data Science → Create new Project
Use the Pipeline section inside your project.
2. Create the Pipeline
Click Create Pipeline
Give it a name: e.g., Employee_Attrition_Pipeline
Add description (optional)
Configure:
o Environment variable:
data_location = oci://<your-bucket-name>@namespace/pipeline-temp-
bucket
o Compute Shape: VM.Standard2.2
o Block Volume Size: 50 GB
o Logging Configuration: Select an existing log group
⚙️ Defining Pipeline Steps
Step Description Depends On Artifact Entry Point
Step 1 Data Preprocessing — [Link] [Link]
Step
Train Linear Regression Step 1 [Link] [Link]
2a
Step
Train Random Forest Step 1 [Link] [Link]
2b
Step
Train XGBoost Step 1 [Link] [Link]
2c
Evaluate & Deploy Best Steps 2a, 2b,
Step 3 evaluate_deploy.zip evaluate_deploy.py
Model 2c
All steps are “Build by Script” type.
Steps 2a, 2b, and 2c run in parallel (to train multiple models simultaneously).
▶️ Running the Pipeline
Click Start Pipeline Run
Name the run (e.g., Run-11)
Add environment variable:
CONDA_ENV_SLUG → select default service environment
Check logs under configured log group
Observe status transitions:
o Waiting → Accepted → In Progress → Succeeded
📊 Results
AUC Scores (example):
o Linear Regression → 0.85 ✅ (Best model)
o Random Forest → 0.81
o XGBoost → 0.837
The best-performing model (Linear Regression) is automatically deployed.
You can verify deployment in:
o Models tab → shows the selected model and deployment status
🧱 Key Takeaways
OCI Data Science Pipelines automate model training, evaluation, and
deployment.
You can reuse step artifacts and integrate with OCI Logging & Object Storage.
Supports parallel model training and end-to-end reproducibility.
Reduces manual effort and ensures consistency across ML workflows.
10. Model Deployment: Autoscaling
🎯 Objective
To understand Autoscaling in OCI Data Science model deployments — a feature
that automatically adjusts compute resources to balance performance, availability,
and cost-efficiency.
⚙️ What Is Autoscaling?
Autoscaling allows a deployed model to automatically scale up or down the number
of compute instances based on real-time demand.
It helps maintain optimal performance while minimizing unnecessary cost —
especially useful for unpredictable workloads.
🚀 Why Autoscaling Matters
When deploying models, choosing the right compute shape and instance count can
be difficult.
Autoscaling solves this by dynamically managing resources based on usage thresholds.
🧱 Key Benefits of Autoscaling
# Benefit Description
Dynamic Resource Automatically increases or decreases compute instances
1⃣
Adjustment based on demand.
2⃣ Cost Efficiency You only pay for resources actually used — no idle cost.
Works with load balancers to reroute traffic to healthy
3⃣ Enhanced Availability
instances if one fails.
Supports user-defined NQL expressions to control when
4⃣ Customizable Triggers
scaling happens.
Load Balancer Load balancer bandwidth scales automatically to match
5⃣
Compatibility traffic.
Prevents too-frequent scaling by pausing actions for a set
6⃣ Cooldown Periods
duration after a scale event.
🧱 How Autoscaling Works
🔹 Metric-Based Autoscaling
The only supported autoscaling type in model deployments.
Triggered when a metric (like CPU or memory usage) meets/exceeds a defined
threshold.
Metrics are collected by the Monitoring Service and aggregated over time.
Example:
If CPU utilization > 80% for 3 consecutive intervals → scale up instance count.
⏱️ Autoscaling Event Flow
1. Metrics (e.g., CPU%) are monitored.
2. If threshold is exceeded for several intervals → scaling event triggered.
3. A cooldown period begins → no further scaling during this time.
4. When cooldown ends → system reevaluates metrics and adjusts if needed.
🧱 Autoscaling Policy Types
Type Description
Predefined Choose from built-in metrics like CPU Utilization or Memory
Metric Utilization.
Use Monitoring Query Language (NQL) expressions for advanced
Custom Metric
control — combine multiple metrics, aggregations, and logical
(NQL)
conditions (AND/OR).
🧱 Setting Up Autoscaling
You can configure autoscaling:
During model deployment creation
Or for an existing deployment
Methods:
OCI Console
OCI CLI
OCI Data Science API
🧱 Required IAM Policy
Add this to your tenancy:
Allow service datascience to read metrics in tenancy
where [Link] = 'oci_datascience_modeldeploy'
This allows autoscaling to access model deployment metrics.
🧱 Custom Metric Example
You can create an NQL expression using any available model deployment metrics, such
as:
CPUUtilization
MemoryUtilization
ResponseTime
RequestsPerSecond
StatusCode, etc.
Example Query:
CpuUtilization[1m].mean() > 75 or MemoryUtilization[1m].mean() > 80
This query scales up if either CPU > 75% or memory > 80% in the last 1 minute.
🔄 Deployment States and Updates
Deployment
Allowed Updates
State
Can modify autoscaling policy independently (not combined with
Active
other changes).
All options (including autoscaling configuration) can be modified
Inactive
together.
📊 Monitoring Metrics
Metrics for model deployments are automatically emitted under namespace:
oci_datascience_modeldeploy
Common Metric Dimensions:
resourceId
statusCode
statusFamily
instanceId
result
networkType
No need to manually enable monitoring — it’s automatically active for every
deployment.
💡 Summary
Feature Description
Purpose Automatically scale model deployment resources.
Scaling Type Metric-based (CPU/Memory or custom).
Key Benefit Performance stability with cost efficiency.
Integrations Load balancer + OCI Monitoring.
Control NQL expressions for custom triggers.
11. Expert Tips: Pipelines
☁️ Expert Tips: Pipelines
(OCI Data Science Course — by Hemant Gahankari, Senior Principal Training
Lead, Oracle University)
🎯 Objective
To understand the Pipelines feature in OCI Data Science Service and how it
automates end-to-end machine learning workflows.
⚙️ What Are Pipelines?
Pipelines are a recently introduced feature in OCI Data Science that allow you to
automate complete ML workflows.
A pipeline consists of multiple steps that can be executed either:
Sequentially (one after another), or
In parallel (simultaneously).
🧱 Typical Steps in a Pipeline
Examples of workflow stages include:
1. Data Extraction — retrieving data from sources.
2. Data Validation — ensuring data quality and correctness.
3. Data Preparation / Preprocessing — cleaning and transforming data.
4. Model Training — applying ML algorithms.
5. Model Evaluation & Deployment — assessing and deploying the trained
model.
🚀 Key Purpose
Automate repetitive machine learning tasks.
Ensure consistency across model training and evaluation runs.
Support complex workflows with dependencies and multiple environments.
🌟 Instructor’s Recommendation
Hemant emphasizes exploring OCI Pipelines through:
Documentation
Hands-on experimentation
This will help users fully leverage the feature to streamline and scale their
machine learning projects.
🧱 Summary Table
Feature Description
Service Oracle Cloud Infrastructure (OCI) Data Science
Feature Description
Feature Name Pipelines
Function Automate and manage multi-step ML workflows
Execution Mode Sequential or Parallel
Benefits Automation, efficiency, scalability, and reproducibility
Recommended By Hemant Gahankari (Oracle University)
6. Related OCI Services
1. Spark Applications, Data Flow, and Data
Science
🧱 Demo Notes: Spark Applications, Data Flow, and Data Science
👨🏫 Instructor
Jean-Rene Gauthier, Product Manager – OCI Data Science
🚀 1. Introduction to OCI Data Flow
Oracle Cloud Infrastructure (OCI) Data Flow
A fully managed, serverless Apache Spark service for big data processing
and machine learning at scale.
Enables you to run Spark applications (PySpark, SQL, Java, Scala) without
provisioning or managing infrastructure.
Key use cases:
Data aggregation & transformation
Feature engineering
Data cleaning
Model training (via MLlib or other libraries)
Why Spark?
Popular engine for scalable, distributed data processing.
Used for AI/ML workloads due to its high-speed parallel processing.
⚙️ 2. Data Flow Core Components
Component Description
Library Central repository of all Spark applications in Data Flow.
Reusable Spark app template (code + dependencies + parameters +
Application
runtime).
Execution instance of a Data Flow application — includes output, logs,
Run
stats, etc.
Logs Automatically stored in OCI Object Storage for debugging and auditing.
🔑 3. Key Capabilities
✅ Connect to Spark data sources and launch jobs in seconds.
✅ Manage all Spark apps from one interface.
✅ Create reusable Spark applications in any Spark language.
✅ Secure, isolated environment (no cluster sharing).
✅ Data encrypted in transit and at rest.
✅ Integrated with OCI IAM for secure access.
✅ Supports Object Storage and other connectors.
🔒 4. Security Model
Each run executes on behalf of the user who launched it.
Uses IAM policies for authorization.
Jobs run in isolated pools (no shared clusters).
Logs → Object Storage (securely stored).
Data encrypted end-to-end.
🧱 5. Spark + Data Flow for AI/ML Workloads
Spark provides:
ETL (Extract, Transform, Load) for large datasets
Feature engineering
Scalable model training
Spark includes MLlib, offering:
Common ML algorithms (classification, regression, clustering)
Training on distributed dataframes
Ability to integrate with PySpark or Spark SQL
⚙️ 6. Spark Architecture Overview
Component Description
Driver Central coordinator that controls execution.
Cluster Manager Manages cluster resources and launches applications.
Executors Distributed worker nodes running tasks.
SparkContext Entry point for Spark functions.
When configuring resources:
Choose driver/executor shape and number of executors.
The more data or shorter processing time → more OCPUs needed.
o Example: 500 GB in 10 hrs → ~5 executor OCPUs.
🧱 7. Integration with OCI Data Science
Data Flow integrates with OCI Data Science Notebooks (JupyterLab interface).
Uses Accelerated Data Science (ADS) library.
You can:
o Submit Spark jobs from notebooks
o Fetch logs
o Manage Data Flow runs
o Sync PySpark scripts
o Add custom Python libraries
📋 8. Prerequisites for Using Data Flow with Data Science
You need:
1. A Data Science Project + Notebook Session
2. An Object Storage bucket for logs & data
3. A PySpark / Spark SQL / Java application uploaded
4. Correct IAM policies for access permissions
💡 9. Best Practices for PySpark Development
1. Use sample data during development (not full dataset).
2. Break code into cells for iterative testing.
3. Convert Spark DF → Pandas DF for visualization.
4. Use Matplotlib or similar for analysis.
5. Remove unnecessary cells before converting notebook → script:
6. jupyter nbconvert --to script [Link]
7. Test locally:
8. spark-submit [Link]
9. Before uploading:
o Replace sample data paths with full dataset.
o Remove external library calls (like pandas/sklearn).
o Or include them in Data Flow environment manually.
📘 10. Summary
Data Flow = Serverless Spark for large-scale batch & ML jobs.
No infrastructure management – on-demand compute.
Ideal for ETL, feature engineering, model training.
Integrated with OCI Data Science notebooks and ADS SDK.
Secure, scalable, and cost-effective Spark job execution.
2. Oracle Open Data
🌍 Lesson Notes: Oracle Open Data
👨🏫 Instructor
Wes Prichard, Senior Principal Product Manager – Oracle Data Science & AI Services
🧱 1. Overview
Oracle Open Data is a public data repository by Oracle for Research that provides
free access to curated datasets for researchers, scientists, and developers worldwide.
Goal: To make large, high-quality datasets easily available for research and
AI/ML applications.
Access:
🔓 No login or payment required.
🌐 Website: [Link]
🗂️ 2. What It Offers
Oracle Open Data includes datasets from multiple scientific and technical domains:
Domain Example Data Sources / Types
🛰️ Geospatial
Satellite imagery from GOES, MODIS, Landsat
Data
🧱 Life Sciences Chemical compounds, protein sequences, CT scans, genomic data
Datasets for training models — text corpora, images, audio,
🤖 AI & ML
annotated files
💡 3. Why It’s Valuable
Key Benefits:
✅ Trusted & Curated Data — All datasets come from reputable institutions like
NASA, Stanford, and DeepMind.
✅ Easy Access & Navigation — User-friendly platform for searching and
downloading.
✅ Regular Updates — Continuously refreshed with the latest research data.
✅ Ready to Use — Datasets include:
Documentation
Code samples
Usage examples for reproducibility
🧱 4. How to Use It
1. Visit [Link]
2. Click “Explore Repository”
3. Browse datasets by category (Geospatial, Life Sciences, AI/ML, etc.)
4. Download data directly for analysis or model training.
📘 5. Summary
Oracle Open Data = Free, curated, research-grade datasets from trusted
global sources.
Supports AI, ML, and scientific research across multiple domains.
Simple to use, no account needed.
Accessible at: [Link]
3. OCI Data Labeling
🏷️ Lesson Notes: OCI Data Labeling
👨🏫 Instructor
Praveen Patil, Principal Product Manager – Data Science & AI Services, Oracle
🧱 1. What is Data Labeling?
Data labeling is the process of identifying and annotating properties (labels) for
image or text data.
It’s an essential step in preparing data for AI and machine learning (ML) training.
Example: Labeling 100 images of tigers enables an AI system to identify tigers in
new, unlabeled images.
Labeled data → Used to train supervised ML models.
🗂️ 2. Key Components
Term Description
A collection of data records (images, text files, or documents) and their
Data Set
associated labels.
Data
An individual data item (e.g., one image or one text file).
Record
Labels Assigned categories or properties (e.g., “tiger,” “signature”).
These datasets are interoperable across OCI Data Science and AI Services —
supporting supervised learning workflows.
⚙️ 3. OCI Data Labeling Service
A fully managed OCI service for quickly labeling raw data with minimal setup.
Steps:
1. Load your data
2. Label it using built-in UIs or APIs
3. Export labeled datasets for use in OCI Vision, Data Science, or other AI
services
Purpose: Enables seamless movement of labeled data across OCI for model training
and deployment.
👥 4. Who Uses It
Persona Role & Use
Build labeled datasets, use prebuilt UIs, export data for model
Data Scientists
training in OCI Data Science.
Developers & Label data to fine-tune AI models (e.g., image recognition) before
Engineers deployment.
🏭 5. Industry Applications
Industry Use Case
🏪 Retail & E-Commerce Analyze customer behavior, recommend products
🏥 Healthcare Label medical scans to detect anomalies
🏛️ Government Categorize documents, automate processes
🎬 Media & Entertainment Flag inappropriate content, analyze sentiment
⚙️ Manufacturing Detect defects or classify parts
🔁 6. Role in the AI/ML Lifecycle
Data labeling occurs early in the ML pipeline — it’s foundational for quality
model training.
Poor labeling → inaccurate models → wrong predictions.
ML development is iterative — labeling may need updates to handle:
o Missing or underrepresented classes
o Model drift (changes in real-world data)
Key takeaway: High-quality labeled data = high-quality model.
💡 7. Current Capabilities
✅ Simple and fast labeling for images, text, and documents
✅ Interactive UI or API-based labeling
✅ Easy export to OCI Vision, Data Science, and other AI services
✅ Seamless integration for building and retraining models
📘 8. Summary
Data labeling = foundational step for supervised learning.
OCI Data Labeling Service simplifies annotation, management, and export.
Used by data scientists and developers across multiple industries.
Enables iterative, accurate ML model training with high-quality labeled data.