0% found this document useful (0 votes)
2 views208 pages

Notes Compressed

The OCI Data Science course provides a comprehensive introduction to using Oracle Cloud Infrastructure for machine learning, aimed at data scientists and ML engineers. It covers essential skills such as building, training, deploying, and managing ML models, with hands-on labs and expert contributions. Participants will learn best practices and prepare for the OCI Data Science certification, while also gaining insights into data science's historical context and modern applications.

Uploaded by

arnav2005gupta
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views208 pages

Notes Compressed

The OCI Data Science course provides a comprehensive introduction to using Oracle Cloud Infrastructure for machine learning, aimed at data scientists and ML engineers. It covers essential skills such as building, training, deploying, and managing ML models, with hands-on labs and expert contributions. Participants will learn best practices and prepare for the OCI Data Science certification, while also gaining insights into data science's historical context and modern applications.

Uploaded by

arnav2005gupta
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

1.

Welcome to OCI Data Science


1. Course overview

Oracle Cloud Infrastructure (OCI) Data


Science – Course Introduction
Overview
 Data Science is the art and science of extracting valuable insights from data to solve real-
world and business problems.
 This course focuses on using Oracle Cloud Infrastructure (OCI) for building, training,
deploying, and managing Machine Learning (ML) solutions.
 Ideal for individuals seeking to upskill or reskill for the growing demand in Data
Science and AI.

Course Contributors
 The course features multiple experts, including:
o Wes Prichard
o John Peach
o John Stanesby
o JR Gauthier
o Lyudmil Pelov
o Praveen Patil
o Hemant Gahankari
 Developed by a large team of professionals across Oracle Cloud Infrastructure.

Intended Audience
 Primary Audience:
o Data Scientists
 Also Suitable For:
o Machine Learning (ML) Engineers
o Artificial Intelligence (AI) Engineers
Prerequisites
Participants should:

 Be proficient in Python for machine learning.


 Have general knowledge of open-source ML and data science libraries.
 Ideally possess 1+ years of experience in Data Science, ML, or AI roles.
 Have hands-on experience using OCI (recommended).

Course Objectives
 Gain proficiency in using OCI Data Science and related services.
 Learn to identify OCI services for implementing ML solutions.
 Apply cloud and ML best practices.
 Understand how to build, train, deploy, and manage ML models in OCI.
 Use additional OCI Data & AI services to create end-to-end ML solutions.

Course Structure (5 Major Modules)


1. Introduction to OCI Data Science
o Overview of OCI Data Science and setting up tenancy.
2. Workspace & Environment Setup
o Configuring your OCI Data Science workspace for project development.
3. Machine Learning Lifecycle
o Understanding all steps supported by OCI Data Science throughout the ML
lifecycle.
4. MLOps Practices
o Scaling, monitoring, and automating ML workflows using OCI tools.
5. Related OCI Services
o Exploring other OCI cloud services useful for building comprehensive data
science solutions.

💡 Recommendation: Follow modules in order, as later ones build on earlier concepts.

Hands-On Learning
 End-to-End Lab:
o Uses an Employee Attrition use case to predict the likelihood of employees
leaving the organization.
 Demo Lessons:
o Recorded demonstrations throughout modules illustrate core concepts.
 Requirements:
o Oracle Cloud Account (Free trial available at [Link])
o Optional GitHub access for OCI Data Science AI Samples repository.

Support and Community


 Ask Your Instructor Form:
o Submit questions for personalized help from experts.
 OU Community:
o Connect with fellow learners and subject matter experts.
o Discuss topics, ask questions, and share experiences.

Tips for Success


 Take notes on topics based on your existing knowledge.
 Use transcripts to follow along.
 Schedule regular breaks and stay active.
 Sign up for a free Oracle Cloud account and get hands-on practice.
 Complete all skill checks, exam prep, and practice exams.
 Stay consistent and review regularly for certification readiness.

Feedback and Continuous Improvement


 Oracle continuously updates and refines training content based on learner feedback.
 Rate the course and provide feedback on what’s helpful or needs improvement.
 Your input helps enhance the learning experience for everyone.

Summary
This course equips learners with the knowledge and practical experience to:
 Navigate OCI Data Science tools and services.
 Build and deploy machine learning models in OCI.
 Follow best practices for MLOps and cloud-based data science.
 Prepare confidently for the OCI Data Science certification exam.

2. Expert Tips Intro

OCI Data Science Professional – Welcome Message

Instructor Introduction

 Speaker: Hemant Gahankari


Role: Senior Principal Training Lead, Oracle University

Message Overview

 Welcome to the OCI Data Science Professional certification course.

 The course is designed to help learners gain hands-on skills to become proficient
data science professionals using Oracle Cloud Infrastructure (OCI).

Core Responsibilities of a Data Scientist / ML Engineer

In day-to-day work, data professionals typically:


1. Collect and Prepare Data – Gather, clean, and preprocess datasets.
2. Build and Train Models – Develop and tune machine learning models.

3. Evaluate Models – Assess performance using appropriate metrics.

4. Deploy and Scale Models – Implement trained models into production


environments.
5. Automate ML Pipelines – Streamline processes for repeatable, efficient
workflows.

Using OCI for Data Science


 The OCI Data Science and AI Services enable efficient execution of all the
above tasks within a unified cloud platform.

 These services provide:


o Scalable compute power

o Integrated data management

o Built-in tools for model training, evaluation, and deployment

o MLOps capabilities for automation and monitoring

Course Content Highlights

 The course includes expert tip videos demonstrating:

o Key OCI Data Science and AI Service features

o Real-world best practices

o Simple and efficient ways to build end-to-end data science solutions


 Designed to make the learning practical, engaging, and hands-on.

Closing Note

 Learners are encouraged to explore, practice, and apply what they learn.

 The expert sessions aim to make you confident in using OCI tools effectively.
 “Hope you will find these videos useful.” – Hemant Gahankari

2. Introduction and Configuration


1. Data Science: Introduction

Module 1: Introduction and Configuration


Lesson 1 – Introduction to Oracle Cloud Infrastructure
(OCI) Data Science Service
Instructor: Wes Pritchard
Role: Senior Principal Product Manager, Data Science & AI Services, Oracle

1. Historical Background of Data Science


Period Key Contributor Contribution
Introduced Ockham’s Razor — simpler solutions are
1300s William Ockham preferable. Principle applies to ML by favoring simpler
models.
Advocated that more data leads to better accuracy —
1700s Tobias Mayer
considered one of the first data scientists.
Coined the term Machine Learning; created a self-
1952 Arthur Samuel (IBM)
learning checkers program.
Predicted the rise of empirical data analysis through
1962 John W. Tukey electronic computing — precursor to modern data
science.
Defeated chess champion Garry Kasparov; highlighted
1997 IBM’s Deep Blue
the power of computation and AI.
DJ Patil (LinkedIn) & Jeff Coined the term “Data Science” for extracting business
2008
Hammerbacher (Facebook) value from data.
Coined “The Great Resignation”, showcasing how
2021 Anthony Klotz
data science can be used to study employment trends.

2. Data Science Today


 Data science now drives real-world business problem solving, e.g., predicting
employee attrition.
 The course includes a hands-on lab using an Employee Attrition use case.
o You will build a predictive ML model yourself using OCI Data Science.

3. Oracle’s Approach to Data Science and AI


a. The Foundation: Data
 Modern businesses handle diverse data types:
o Structured: Databases, business apps.
o Unstructured: Text, images, videos, audio, sensor data, customer interactions.
 Goal: Use all data to gain insights, enhance experiences, and optimize operations.

b. Oracle AI Portfolio

Oracle AI encompasses multiple cloud services designed to help organizations leverage all their
data effectively.

Architecture Overview:

1. Data Layer: Foundation for all AI and ML activities.


2. AI & ML Services Layer:
o Machine Learning Services:
 Used by data scientists to build, train, deploy, and manage ML models.
 Utilize open-source frameworks within OCI Data Science.
 Supported by OCI Data Labeling for supervised learning.
o AI Services:
 Contain pre-built ML models for specific use cases.
 Some are pre-trained, others can be trained with customer data.
 Accessible via simple API calls — no infrastructure management
required.
3. Applications Layer:
o Where AI is consumed — apps, analytics, or business processes.

c. Supporting Services

 Data & Graph Analytics


 Data Integration and Management
 Business Intelligence
 Underlying OCI Infrastructure (Compute, Storage, Networking)

4. Oracle Cloud Infrastructure (OCI) Data Science


Definition

A cloud service designed to help data scientists:

 Build, train, deploy, and manage ML models efficiently.


 Support the full ML lifecycle using Python and open-source tools.
 Operate within a JupyterLab notebook environment.
5. Core Principles of OCI Data Science
1. Accelerate Individual Productivity
o Built for Python & open-source workflows.
o Eliminates manual setup of environments.
o Provides on-demand compute power (CPU/GPU) without infrastructure
management.
o Includes Oracle’s Accelerated Data Science (ADS) SDK for automating
common ML tasks.
2. Enable Collaboration
o Shared projects and assets for teamwork.
o Promotes reproducibility, traceability, and auditability.
o Reduces duplication across data science teams.
3. Enterprise-Grade Reliability
o Integrated with OCI security, IAM, and compliance.
o Fully managed infrastructure (maintenance, patching, upgrades).
o Focus remains on solving business problems, not managing servers.

6. Key Concepts and Terminology


Term Description
A collaborative workspace to organize and document data science assets
Project
(notebooks, models, datasets).
An interactive JupyterLab environment with pre-installed open-source
Notebook Session
libraries for coding, training, and experimentation.
Open-source package & environment manager for Python. Used to
Conda
install and manage dependencies easily in OCI Data Science.
Oracle’s Python library for automating tasks like data connection,
Accelerated Data
visualization, AutoML training, evaluation, and model explainability.
Science (ADS) SDK
Provides simple access to model catalog and OCI services.
Mathematical representation of data and business logic. Created in
Model
notebook sessions and stored in the Model Catalog.
Central repository for storing, tracking, and sharing ML models with full
Model Catalog metadata and provenance details. Enables team collaboration and
reproducibility.
Publishes a trained model as an HTTP endpoint on managed
Model Deployment
infrastructure for real-time predictions.
Define and execute repeatable ML tasks on OCI-managed compute
environments.
Term Description
Browser-based interface to manage all features of OCI Data Science
OCI Console
(used throughout the course).
REST API / SDKs /
Alternative access methods:
CLI

 APIs: For programmatic interaction.


 SDKs: Available for Python, Java, JS, .NET, Go, Ruby.
 CLI: Command-line management with full functionality. |
| Regions | Globally distributed OCI data centers offering secure, high-performance
environments. OCI Data Science is available in commercial, government, and dedicated
regions. |

7. Summary
 OCI Data Science empowers data scientists to manage the entire ML lifecycle within the
cloud.
 It supports open-source Python workflows, provides managed infrastructure, and
promotes collaboration.
 The next lesson focuses on provisioning and configuring OCI Data Science
environments to begin practical implementation.

2. ADS SDK

OCI Data Science Professional – ADS SDK Module Notes


Instructor: John Peach

Role: Data Scientist, Oracle Cloud Infrastructure (OCI) Data Science Service Team

Overview
The Accelerated Data Science (ADS) SDK is a powerful, end-to-end toolkit built by data
scientists for data scientists, designed to streamline the machine learning (ML) lifecycle on
Oracle Cloud Infrastructure.
Goals of ADS SDK:

 Integrate OCI services into the data scientist workflow.


 Simplify common ML tasks like EDA, feature engineering, model training, tuning,
and deployment.
 Enhance productivity with automation, explainability, and reproducibility.

Versions and Access


Versions of ADS:

1. Public Version – downloadable from GitHub or PyPI.


2. OCI-Integrated Version – pre-installed in OCI Data Science notebooks; includes
AutoML and ML Explainability features.

Accessing ADS:

 Available in Conda environments on OCI Data Science service.


 Installable via pip install ads or directly from GitHub.

Key Features of ADS


1. Data Connectivity

ADS supports connections to multiple data sources, enabling seamless access to data regardless
of location.

Supported Data Sources:

 Local Storage (block storage within notebooks)


 Object Storage via OCI protocol + pandas integration (APE Spec protocol)
 Oracle Databases (via DB Secret Keeper and OCI Vault)
 Autonomous Database (ADB) integration
 Third-party Clouds: AWS S3, Google Cloud, Azure, Dropbox, etc.
 NoSQL Databases: Using Dataset Factory Class
 Big Data Service (BDS): Connects directly to HDFS
 Web Data: Import via HTTP/HTTPS directly into DataFrames

2. Data Visualization and EDA


Understanding data is a crucial step in any ML pipeline.

ADS Visualization Tools:

 Smart Plotting: Automatic default plots for different data types.


 Feature Types: Provides reusable visualizations based on measurement type.
 Summary Statistics: Generate feature summaries and correlation heatmaps.
 Custom Visualization: Create reusable plots across multiple projects.

3. Feature Engineering

Improving data quality leads to better models.

Feature Engineering with ADS:

 Uses the ADS Dataset class (wrapper around pandas DataFrame).


 Provides automatic feature transformation and suggestions.
 Handles:
o Categorical encoding
o Null value imputation
o Transformation recommendations
 Supports both manual and automated feature engineering.

4. Model Training

ADS offers flexible options for building models.

Training Methods:

 AutoML: Fully automates model selection, tuning, and evaluation.


 ADSTuner: Performs hyperparameter optimization.
 Model Packaging: Automatically creates model artifacts for deployment.
 Model Catalog Integration: Easily push trained models to production.

5. Model Evaluation

Evaluating performance is essential to model development.

ADS Evaluator:
 Compares multiple models using consistent metrics.
 Supports binary, multinomial classification, and regression.
 Automatically generates appropriate evaluation metrics and charts.
 Reduces the need for manual chart recreation.

6. Model Interpretability and Explainability

Building trust in models through transparency.

Interpretability Tools:

 Model-Agnostic: Works with any ML model.


 Local Explainability: Understand predictions for specific observations.
 Global Explainability: Understand overall model behavior.
 Partial Dependence (PDP) and Accumulated Local Effects (ALE) plots.
 What-if Analysis: Test how input changes affect predictions.

These tools help verify that the model is learning the right relationships from data.

7. Model Deployment

Easily move models from notebooks to production.

Deployment Features:

 ADS Model Framework: Simplifies model deployment with few commands.


 Supports:
o Oracle AutoML
o PyTorch
o Scikit-learn
o TensorFlow
o Generic Models
 Integrates with OCI Logging Service for:
o Prediction logs
o Access logs

These logs help monitor model usage and performance in production.

Summary
The ADS SDK supports the entire ML lifecycle — from data access to deployment.
It simplifies every step with automation, reproducibility, and explainability.

Key Takeaways:

 Integrates seamlessly with OCI and third-party services.


 Streamlines EDA, feature engineering, and model training.
 Automates hyperparameter tuning and evaluation.
 Enhances interpretability and simplifies deployment.

ADS enables data scientists to move from experimentation to production efficiently and
securely.

3. Tenancy Configuration Basics

OCI Data Science Professional – Tenancy Configuration


Basics
Instructor: Jon Stanesby

Lesson Focus: Understanding how to configure your OCI tenancy for Data
Science — including compartments, user groups, dynamic groups, and policies.

1. Overview

Tenancy configuration in Oracle Cloud Infrastructure (OCI) is essential for organizing


and securing access to Data Science resources.
This setup ensures that users, groups, and services have the right permissions
within the right compartments.

How Data Science Components work together:


Assign User to appropriate groups. -> Create dynamic group for data science
resources. -> Create policies that grant access to resources in a compartment.

2. Key Concepts

a. Compartments

 Definition: Logical containers for organizing OCI resources.


 Purpose: Control access to cloud resources by grouping them logically.

 Access: Only user groups with explicit permissions can access resources in a
compartment.
Steps to Create a Compartment:

1. Go to Identity → Compartments in the OCI Console.

2. Click Create Compartment.

3. Enter a name and description, optionally add tags.

4. Click Create Compartment and note the OCID (useful later).

b. User Groups

 Definition: Collections of users granted access to OCI resources.

 Purpose: Simplifies permission management for multiple users.

Steps to Create User Groups:

1. Create Users → Go to Identity → Users → Create User.

o Provide a username, description, and optional email.


2. Create Group → Go to Identity → Groups → Create Group.

o Add name and description.


3. Add Users to Group → Click Add User to Group, select user(s), and confirm.

c. Dynamic Groups

 Definition: Groups of resources (not users) that match defined rules.

 Examples of resources:

o Data Science Notebook Sessions

o Model Deployments

o Job Runs
Key Points:

 Membership changes dynamically as matching resources are created/deleted.


 These resources act as principal actors, capable of making API calls per
assigned policies.

 Enables secure resource-to-service interactions (e.g., notebook session


accessing Object Storage).
Steps to Create Dynamic Group:

1. Go to Identity → Dynamic Groups → Create Dynamic Group.

2. Add name and description.

3. Define matching rules, replacing the compartment OCID with your Data Science
compartment’s OCID.
Example Matching Rules:

 Include all Notebook Sessions, Model Deployments, and Job Runs within a
compartment.

3. Policies

a. Definition

Policies define what principals (users or resources) can do within a compartment.


They are the backbone of access control in OCI.
Syntax:

Allow group <group-name> to <verb> <resource-type> in compartment


<compartment-name>

b. Key Components

Component Meaning

Group Name The user group or dynamic group name

Verb The level of access granted

Resource Type The OCI resource or family (e.g., data-science-family)

Compartment Name The specific compartment in which access is granted


c. Verbs (Access Levels)

Verb Description Access Level

🔹 Least
Inspect List resources (no metadata access)
Permissive

Read Inspect + get resource metadata

Read + work with resource (update, but not


Use
create/delete)

🔹 Most
Manage Full access (create, update, delete)
Permissive

d. Resource Types

 Policies can target individual resources (e.g., data-science-models)

 Or aggregate resource types like data-science-family (includes related


resources such as models, jobs, notebook sessions, etc.)

4. Required Data Science Policies

a. Granting Access to Data Science Resources

1. Allow user group to manage all Data Science resources:

2. Allow group <user-group> to manage data-science-family in compartment


<compartment-name>
3. Allow dynamic group (resources) to manage Data Science resources:

4. Allow dynamic-group <dynamic-group> to manage data-science-family in


compartment <compartment-name>

b. Access to Metrics and Logging

Purpose Policy Statement

Allow users to read Allow group <user-group> to read metrics in


metrics compartment <compartment-name>
Purpose Policy Statement

Allow dynamic groups to Allow dynamic-group <dynamic-group> to use log-


use log content content in compartment <compartment-name>

Allow users to manage Allow group <user-group> to manage log-groups in


log groups compartment <compartment-name>

Allow users to use log Allow group <user-group> to use log-content in


content compartment <compartment-name>

c. Networking-Related Policies
(Required if using custom networking in Data Science)

Allow service datascience to use virtual-network-family in compartment


<compartment-name>

Allow group <user-group> to use virtual-network-family in compartment


<compartment-name>

Allow dynamic-group <dynamic-group> to use virtual-network-family in compartment


<compartment-name>

d. Additional Useful Policies

(For related OCI services, e.g., Object Storage)

Allow group <user-group> to manage object-family in compartment <compartment-


name>
Allow dynamic-group <dynamic-group> to manage object-family in compartment
<compartment-name>
These policies enable object storage integration (e.g., for reading/writing data
from notebooks or deployed models).

5. Summary

In this lesson, we covered the core tenancy configuration components and their
relationships:
 Compartments: Logical boundaries for organizing resources.
 User Groups: Control user access and permissions.

 Dynamic Groups: Automatically manage resource-based access.

 Policies: Define who can do what, where, and how.

We explored:
 The three-step setup for each (Compartment → Group → Policy).

 The required policies for Data Science.

 Networking and logging policies for advanced configurations.

 Optional policies for integrating with related OCI services.


Takeaway:
Proper tenancy configuration ensures secure, scalable, and organized
management of OCI Data Science environments.

4. Configure a Tenancy with OCI Resource


Manager

OCI Data Science Professional – Lesson: Configuring Tenancy with OCI Resource
Manager

Instructor Introduction

 Speaker: John Stanesby

 Topic: Configuring tenancy using Oracle Cloud Infrastructure (OCI) Resource


Manager

Lesson Overview

This lesson demonstrates how to use Oracle Resource Manager (ORM) to configure a
tenancy for OCI Data Science.
Rather than setting up resources manually, learners can deploy a preconfigured Data
Science Service template to automate the process.

Key Concepts
1. What is Oracle Resource Manager (ORM)?

 A managed service in OCI that helps automate resource provisioning using


Terraform.

 Enables repeatable, version-controlled infrastructure deployment.


2. Using the Data Science Service Template

 The Data Science Service template is available within Resource Manager.

 It automatically creates all required identity and access management (IAM)


components for a basic data-science setup.

Template Configuration Details

When you run the Data Science Service template, it creates:

A. User Group

 A group with a name you define.

 Used to assign and manage permissions for your human users.


B. Dynamic Group

 A group with a name you define, including matching rules for the following
resource types:

o datasciencenotebooksession

o datasciencemodeldeployment

o datasciencejobrun
C. Policy

 A policy with statements granting the following permissions:


1. Allow user group to manage data-science-family resources in the
compartment.
2. Allow dynamic group to manage data-science-family resources in the
compartment.
3. Allow user group to read metrics.
4. Allow dynamic group to use log content.
Steps to Run the Resource Manager Stack

1. Create a Stack

o Go to Resource Manager → Stacks → Create Stack.

2. Select Template

o Choose Template as your origin, then Service → Data Science.

o Click Select Template.

3. Choose Compartment

o Select the compartment where Data Science resources will be created.


o Click Next.

4. Configure Variables (Optional)

o You can modify additional configuration variables.


o Click Next to proceed.

5. Run the Stack

o Choose to run apply immediately.

o Click Create and wait for the job to finish.

6. Add Users

o After the stack runs successfully, add your users to the newly created user
group.

Alternative Configuration Option

 Instead of using the prebuilt template, you can manually deploy using your own
Terraform script.

 Oracle provides an official Terraform example available at a public GitHub


repository.

Summary

 The Data Science Service template in Resource Manager automates tenancy


configuration.
 It creates essential user groups, dynamic groups, and policies for data-
science workloads.
 The setup process involves creating and running a stack, then adding users.
 Advanced users can use Terraform scripts from the public GitHub repo for
custom setups.

5. Networking for Data Science

OCI Data Science Professional – Lesson: Networking for Data Science

Instructor Introduction

 Speaker: John Stanesby

 Topic: Networking concepts for Oracle Cloud Infrastructure (OCI) Data Science

Lesson Overview

This lesson provides a high-level introduction to essential OCI networking


components and how they relate to Data Science workloads.
It explains how to configure network connectivity for data science resources using either
default networking or custom networking.

⚠️ Note: The session does not cover in-depth networking topics; it focuses on core
concepts relevant to data science.

1. Key Cloud Networking Components

Virtual Cloud Network (VCN)

 A virtual private network created within Oracle data centers.

 Acts as the foundational network structure for cloud resources.


Subnets

 Subdivisions of a VCN used to group resources.

 Each subnet:
o Contains Virtual Network Interface Cards (VNICs) attached to instances.

o Shares the same route tables, security lists, and DHCP options.

Virtual Network Interface Card (VNIC)

 Defines how an instance connects to internal and external endpoints.

 Determines the IP address and network connectivity of a resource.

2. Virtual Routers and Gateways

Dynamic Routing Gateway (DRG)


 Enables private network traffic between:

o Your VCN and on-premises network (via VPN or FastConnect).

o Different VCNs across regions.

 Provides private communication without using public IP addresses.

Network Address Translation (NAT) Gateway

 Allows outbound internet access for private resources.

 Keeps those resources hidden from incoming public connections.

Service Gateway

 Enables private traffic between your VCN and Oracle Services Network.

 Example use case:


o A database in a private subnet can back up data to Object Storage
without public IPs or internet access.

3. Data Science Workloads and Networking

Workload Types

 Notebook Sessions

 Jobs / Job Runs

 Model Deployments

These are collectively referred to as data science workloads.


External Asset Access

Workloads often need to access:


 Code files

 Data sources

 Libraries or dependencies

 Secrets (e.g., credentials)

 Logs

These assets might reside:


 On the public internet, or

 Inside a private network (e.g., enterprise Git servers or databases).

4. Networking Patterns in OCI Data Science

A. Default Networking

 Fastest and simplest option to get started.


 Automatically attaches a workload to a pre-configured service-managed
subnet via a secondary VNIC.

 Provides:
o Egress to the public internet via a NAT gateway.
o Access to OCI services via a Service Gateway.

 No need to manually create VCNs or define networking policies.

 Recommended if you only need:

o Internet access

o Access to OCI-managed services

B. Custom Networking

 Suitable for advanced configurations and private network access.

 You provide your own existing subnet for data science workloads.
 When the workload is created:
o OCI Data Science connects it to your selected subnet via secondary
VNIC.

 Enables:
o Secure access to private enterprise assets

o Full control over routing, security, and policies

Use custom networking if workloads must access:

 On-prem databases

 Private Git repositories

 Other internal resources

📌 Note: Custom networking requires coordination with your network administrator


and additional IAM policies, as discussed in the Tenancy Configuration lesson.

5. Setting Up Networking for Data Science (Using VCN Wizard)

Steps:

1. Navigate to Networking → Virtual Cloud Networks.

2. Click Start VCN Wizard.

3. Select Create VCN with Internet Connectivity.

4. Enter a VCN name.


5. Click Next, then Create.

6. Wait for the VCN and associated resources to be created.


7. Click View Virtual Cloud Network to confirm successful setup.

⚠️ If you already configured your tenancy using OCI Resource Manager, this VCN is
automatically created and you don’t need to repeat these steps.

6. Summary

 Networking Components Covered:

o VCN, Subnets, VNICs, DRG, NAT Gateway, Service Gateway.


 Two Networking Options for Data Science:

1. Default Networking – for quick setup and public/OCI service access.

2. Custom Networking – for private network access and advanced control.

 VCN Wizard provides a simple way to set up networking manually if not already
configured.

✅ Key Takeaway:
OCI provides flexible networking options to securely connect data science workloads to
the resources and data they need — whether over the public internet or within private
enterprise networks.

6. Authenticate to OCI APIs

OCI Data Science Professional – Lesson: Authenticate to OCI APIs


Instructor Introduction

 Speaker: Jon Stanesby

 Topic: Authentication methods for interacting with OCI APIs in the Data Science
Service

Lesson Overview

This lesson explains how to authenticate to Oracle Cloud Infrastructure (OCI) APIs
when working with Data Science resources such as:

 Notebook Sessions

 Jobs

 Model Deployments
You’ll learn about different authentication options, including Resource Principals and
OCI configuration file with API keys, and how they are used across the ADS SDK,
OCI Python SDK, and OCI CLI.
⚠️ Note: This lesson focuses only on authentication (verifying identity), not
authorization (permissions and access control), which was covered in Lesson 2:
Tenancy Configuration.

1. Understanding Authentication in OCI

When working in OCI Data Science, your workloads may need to interact with other
OCI services through the OCI REST APIs.
For example:
 Reading or writing data to Object Storage
 Creating and running Data Flow applications

 Accessing logs, secrets, or model repositories

To perform these tasks, your code must authenticate with OCI — meaning OCI must
recognize the identity performing the API calls.

2. Common Interfaces for API Interaction

Data scientists typically interact with OCI APIs using one of the following:
 ADS SDK (Accelerated Data Science SDK)

 OCI Python SDK

 OCI Command Line Interface (CLI)

Each interface supports different authentication methods that are covered below.

3. What Are Resource Principals?

A Resource Principal is a special feature in OCI Identity and Access Management


(IAM) that allows cloud resources themselves (not just users) to act as authenticated
principals.

Key Points:

 Each resource (like a notebook session or job run) has its own unique identity.

 It authenticates using automatically managed certificates — no need for


manual credential handling.

 These certificates are:


o Created and assigned automatically

o Securely stored and rotated by OCI

This removes the need to manually upload API keys or configuration files to workloads.

4. How Resource Principals Work in Data Science

The Data Science service enables workloads (like notebook sessions and job runs)
to use their own resource principal for authentication.

Benefits:

 Provides secure authentication without exposing credentials.

 Simplifies authentication for automated workloads (like job runs) that don’t have
an interactive interface.

 Reduces operational complexity — no need to manage .oci/config or .pem files


manually.
If a resource principal is not explicitly used, the SDK or CLI will fall back to using the
OCI configuration file and API key approach.

5. Resource Principal Token Behavior

 A resource principal token is cached for 15 minutes.


 If you make changes to IAM policies or dynamic groups, those updates will
take effect only after the token cache expires.

 You can continue development while waiting, but access permissions will not
refresh instantly.

💡 Tip: If authentication changes seem delayed, wait ~15 minutes before retesting.

6. Authenticating with Resource Principals

You can authenticate using resource principals in:


 ADS SDK

 OCI Python SDK

 OCI CLI
Each interface uses slightly different syntax for setting up resource principal
authentication.
It’s recommended to refer to official SDK documentation or pause the lecture at the
relevant code examples to note down the commands.

7. Alternative Authentication: Configuration File & API Key

If you prefer not to use resource principals, you can authenticate as your personal IAM
user using a configuration file and API key.

Steps:

1. Upload your OCI configuration file (usually located at ~/.oci/config) into the
notebook session’s OCI directory.

2. For the profile defined in the config file, also upload or create the required .pem
key files.

3. Alternatively, use the api_keys notebook to generate and configure keys


automatically.
To Launch the api_keys Notebook:

 Open JupyterLab Launcher → Click Notebook Examples → Select api_keys


notebook.

 Follow the guided steps to create and set up your OCI API keys directly within
your Data Science environment.

8. Comparison: Resource Principals vs. API Key Authentication

Aspect Resource Principals Config File + API Key

Security High (no manual keys) Depends on secure key storage

Setup Automatic Manual (upload config + keys)

Automated jobs, production Personal use, testing, custom SDK


Best For
workloads scripts

Token
15 min cache Persistent
Lifespan
Aspect Resource Principals Config File + API Key

Ease of Use Simplifies automation Requires manual setup in notebook

9. Lesson Summary

 Authentication is required to interact with OCI APIs securely.

 You can authenticate via Resource Principals (recommended) or via


Configuration File + API Key.
 Resource Principals:

o Are secure, automated, and ideal for jobs and model deployments.

o Use auto-managed certificates — no need to handle credentials manually.


 Config + API Key:

o Suitable for user-level authentication, testing, or local SDK use.


 Tokens refresh every 15 minutes, so policy or group changes take time to apply.

 The api_keys notebook simplifies creating and managing config-based


credentials inside JupyterLab.

✅ Key Takeaway:
Use Resource Principals as the default authentication method for OCI Data Science
workloads — they are secure, fully managed, and ideal for production automation. The
configuration file + API key method remains available for personal authentication or
custom scripting.

3. Workspace Design and Setup


1. Projects

Module 2: Workspace Design and Setup


Lesson Title: Project

Instructor: Jon Stanesby

Summary

This lesson introduces projects, the central component of an OCI Data Science
workspace. Projects act as collaborative environments where data scientists organize
their work around specific use cases or business questions. The lesson explains how to
create, view, edit, and delete projects using both the OCI Console and the ADS
SDK.

Key Points

1. Definition of a Data Science Project

 A project is a collaborative workspace for teams of data scientists.

 It helps organize work around a particular use case or business question.

 All data science resources (e.g., notebook sessions, models) are created
within a project.
 Projects can be created, named, described, and tagged by data scientists.

2. Creating Projects

Two main methods:


1. From the OCI Console UI

o Log in with the necessary policies.


o Navigate: Menu → Analytics & AI → Data Science → Create Project.

o Choose a compartment to add the project to.

o Optionally add:
 Unique name (auto-generated if omitted).

 Short description (helps others understand the project’s purpose).


 Tags for easy tracking and organization.
o To view the project after creation, select “View Detail Page” and click
Create.

2. From the ADS SDK


o Use the ProjectCatalog object and the create_project() method.

o Specify the compartment ID.

o The example uses a preset environment variable to match the


compartment of the notebook session.

3. Managing Projects

Viewing Projects

 All projects appear on the Project List page.

 Each project displays metadata such as:


o Display name & description
o OCID

o Creation date/time and creator

o Tags

Editing Projects

 Editable fields: display name, description, and tags.

 Edits can be made through any OCI interface.


Deleting Projects

 A project must be empty before deletion (delete all associated resources first).

 In the Console:

o Locate the project in the list or detail page.


o Click Delete.

o Type the exact project name to confirm.

 Deleted projects show a “Deleting” status and can later be filtered out using
state filters (e.g., “Active”).
4. Demonstration Recap

 Create a project: Enter a name, description, and tags → Click Create.

 Edit a project: Change name or description, manage tags via the detail view.

 Delete a project: Confirm deletion with the exact name → project status
becomes “Deleted”.

Conclusion

In this lesson, we learned:

 What a data science project is and why it’s central to OCI Data Science.

 How to create projects using both the Console and ADS SDK.

 How to view, edit, and delete projects effectively.

2. Notebook Sessions

🧱 Module 2 – Workspace Design and Setup

Lesson 2 – Notebook Sessions

Instructor: John Stanesby

🧱 Overview

Notebook Sessions in OCI Data Science provide a fully managed JupyterLab


interface for building, training, and managing machine learning models.

They abstract away infrastructure complexity — meaning you don’t need to


handle compute or storage provisioning, patching, or lifecycle management
yourself.

⚙️ Key Features
Feature Description

Managed OCI automatically provisions and manages compute,


Infrastructure storage, and updates.

Choose from AMD, Intel, or NVIDIA shapes for performance


CPU & GPU Support
needs.

Persistent Storage Block storage retains data, notebooks, and environments.

Change shapes (scale up/down) during the notebook


Scalable Compute
session lifecycle.

Activate, deactivate, or delete notebook sessions easily via


Lifecycle Actions
Console or SDK.

🧱 Creating a Notebook Session (Console)

Prerequisite: You must have a project already created.


Steps:

1. Navigate to your Project Details page.

2. Click Create Notebook Session.

3. Select:

o Compartment

o Optional name
o Compute shape (CPU/GPU)

o Block storage size (50 GB – 10,240 GB)

o Networking option:

 Default networking – service-managed VCN

 Custom networking – choose your own VCN & subnet

4. Add tags (optional).

5. Click Create.

⏳ It may take a few minutes to start up. Once active, the Open button will be
enabled.
🔧 Managing Notebook Sessions

📋 Viewing

 View all notebook sessions on the Notebook Sessions List page.

 Metadata includes:

o Display name & OCID


o Creator and creation time

o Compute shape

o Storage size

o VCN and subnet

o Tags

✏️ Editing

 When active, only the display name can be edited.

 Tag editing is also possible through the details view.

🔁 Activation & Deactivation

Action Description

Starts up compute resources. You can update shape, storage, or


Activate
network configuration during activation.

Shuts down compute instance to save cost, retaining block storage


Deactivate
for data and files. Boot volume data is deleted.

Starts a new compute instance and reattaches the same block


Reactivate
volume, restoring your previous environment.

💡 Tip: Use deactivation when you’re not actively working — it saves compute
costs without losing work stored on the block volume.
🗑️ Deleting a Notebook Session

1. From the Details or List page, click Delete.

2. Confirm deletion by typing the exact display name.

3. OCI terminates the compute instance and destroys block storage.

4. The notebook moves from Deleting → Deleted status.

⚠️ Important:
Any files or code not backed up or committed before deletion are permanently
lost.

📈 Viewing Metrics

You can view metrics from the Details page:

 CPU Utilization

 Memory Usage

 Network In/Out Traffic


Metrics help monitor resource usage during training or large data
operations.

🧱 Lesson Summary

Topic Description

Definition Fully managed JupyterLab environment for ML development.

Creation Through OCI Console with flexible compute and network setup.

Management Includes viewing, editing, activating, deactivating, and deleting.

Storage Persistent block storage for notebooks, data, and environments.

Metrics Monitor CPU, memory, and network traffic for optimization.

✅ Key Takeaways
 Notebook sessions simplify model development by handling infrastructure
automatically.

 Persistent storage allows safe deactivation without data loss.


 Scaling compute shapes up/down is supported mid-lifecycle.

 Always back up work before deleting a session — deletion destroys


storage permanently.

3. How to Work with JupyterLab

Module 2: Workspace Design and Setup

Lesson Title: How to Work with JupyterLab

Instructor: Jon Stanesby

Summary

This lesson covers JupyterLab, the next-generation web-based interface used in OCI
Data Science notebook sessions. It explains how data scientists interact with
notebooks, files, terminals, and environments through JupyterLab, highlights its features
and structure, and walks through how to use it effectively for developing and running
code.

1. Introduction to JupyterLab

 JupyterLab is a web-based user interface and the main environment for


notebook sessions.
 It supports interactive computing, integrates multiple document types, and
handles a wide range of file formats:
o Images, CSV, JSON, Markdown, PDF, Vega, and Vega Lite.

 Data scientists use JupyterLab because it is familiar, intuitive, and makes the
OCI Data Science service easy to adopt.

2. Differences Between Open-Source JupyterLab and OCI Interface


Although the layout and menus are similar, the OCI version includes three key
additions:
1. Launcher Access – quick shortcuts to notebooks, terminals, console,
Environment Explorer, and notebook examples.
2. Environment Explorer – a GUI for managing and searching Conda
environments (covered in the next lesson).

3. GitHub Extension – enables version control and Git integration within notebook
sessions (covered later in Lesson 9).

3. Core Components of the JupyterLab Interface

A. Menu Bar

 Located at the top of JupyterLab.


 Contains top-level menus with keyboard shortcuts for major actions.

B. Launcher

 Provides quick access to:

o New notebooks

o Consoles

o Text editors
o Terminals

o Environment Explorer
o Notebook examples

 New items can be added via commands or extensions.


C. Left Sidebar

Contains several useful tabs:


1. File Browser – Navigate directories, open, add, or delete folders/files.

2. Running Terminals and Kernels – View and shut down active sessions.

3. Git Extension – Manage version control.

4. Commands Panel – Access and search all available commands.


5. Property Inspector – View notebook properties.

6. Open Tabs List – See all open files and tabs.

7. Table of Contents – Automatically generated from Markdown headings.

8. Extension Manager – View installed JupyterLab extensions.

To hide the sidebar: click an icon twice.


To reopen: use the plus (+) icon or select File → New Launcher.

D. Main Work Area

 Central workspace for documents and activities.


 Allows tabbed and resizable panels.

 The active tab has a blue border.

 Supports multiple views for live document editing.

E. Code Consoles and Kernels

 Code consoles serve as interactive scratch pads.

 Kernel-backed documents let you run code interactively.

 Kernel management is handled from the Kernel Menu.

F. Command Palette

 Provides a keyboard-driven interface to search and execute commands quickly.

G. Right Sidebar
 Contains the Property Inspector for notebooks and user interface
customization.

4. Top Chrome Bar Overview

Located above JupyterLab:


 Oracle Logo → Returns to the OCI Cloud Console home page.

 Notebook Session Name → Navigates to the session details page in OCI


Console.
 Session Remaining Time → Shows automatic logout timer (default: 1 hour
inactivity, extendable up to 24 hours).
 Help Icon → Links to OCI documentation.

 Sign Out → Logs out of the notebook session.

5. Using the Launcher

 Access through File Browser toolbar (+ icon) or File → New Launcher.

 The launcher is divided into:


o Left Side:

 Environment Explorer → Discover and manage Conda


environments.

 Notebook Examples → Tutorials and documentation for common


data science use cases.
o Right Side:

 Create new notebooks, consoles, terminals, or text/Markdown


files.

6. Creating and Running a Notebook

1. From the launcher, select Python 3 kernel → Create Notebook.

2. A new notebook includes useful preloaded tips such as:

o Checking internet connectivity.

o Common ADS imports and environment variables.


3. You can rename the notebook via right-click → Rename.

4. Run code cells using:

o The triangle icon,


o Run menu → Run Selected Cell, or

o Keyboard shortcut: Shift + Enter.

5. While code runs:

o A star (✱) indicates the cell is executing.

o A number appears once it completes (execution order).


6. Change cell type using the dropdown (e.g., Code → Markdown).

7. Reorder cells by dragging them up or down.

7. Editing and Kernel Options

 From the Edit menu, you can:

o Merge or split selected cells.

o Change the kernel (via kernel name dropdown in top-right).

8. Viewing and Running Example Notebooks

 Use Launcher → Notebook Examples to open preloaded demos (e.g., Binary


Classification – Attrition).
 You can run all cells or interrupt the kernel to stop execution.

 Use Variable Inspector to view variables currently loaded in memory.

 To view side-by-side:

o Drag tabs until screen sections highlight blue.

o To revert to full screen, drag back to the title bar.

9. Viewing Options and Settings


 From the View menu, you can:
o Show line numbers.

o Collapse or expand code and output cells.

 Themes and other preferences can be adjusted in Settings.

10. File Management

 Create new files by right-clicking in File Explorer.

 Upload files via drag-and-drop (e.g., CSV files).

 File type determines view mode (e.g., CSVs display tabular data).
11. Using the Terminal

 Open from Launcher → Terminal.

 Supports standard Linux commands (ls, etc.).

 Preinstalled tools include:

o odsc conda CLI

o git CLI

o oci CLI

12. Help and Documentation

 Access via the Help Menu:

o Reference materials

o FAQs

o Official documentation

Conclusion

In this lesson, we:


 Explored JupyterLab’s interface and features in the OCI Data Science
environment.
 Learned how to navigate menus, panels, and the launcher.

 Practiced creating and running notebooks and terminals.

 Covered key functions such as variable inspection, kernel management, and


file handling.

JupyterLab provides a flexible, interactive, and integrated environment for all stages
of data science development in OCI.
4. Conda Environments: Overview

Conda Environments: Overview

Instructor: John Peach

Module Objective:

Understand what Conda environments are, their benefits, and how they are managed in
Oracle Cloud Infrastructure (OCI) Data Science Service.

1. What is a Conda Environment?


 Conda is an open-source package management and environment system.

 It allows bundling of:

o Python interpreter

o Python modules and external libraries

o Other required programs

o Into a single, isolated environment.

2. Benefits of Using Conda Environments

1. Install Only What You Need:

o Add or update packages selectively without affecting others.


2. Isolation of Configurations:

o Prevents conflicts between different projects or models.

o Example: One Conda for Computer Vision (TensorFlow, OpenCV), another


for Linear Regression.
3. Easy Switching Between Environments:

o Move between different setups effortlessly.


4. Shareable Configurations:
o Teams can share pre-configured environments instead of recreating
setups.
5. Consistency Across Workflows:
o The same Conda used in notebooks, jobs, and model deployments
ensures consistent results.
6. Reproducibility:

o Track exact software versions to reproduce experiments or debug model


performance changes.

3. Managing Conda Environments with Environment Explorer

Environment Explorer Overview

 A graphical interface (GUI) within OCI Data Science.

 Used to view, manage, and explore Conda environments.

 Offers:
o Card view (for details)

o List view (for multiple environments)

o Search and filters (e.g., show only GPU-compatible or active Condas)

4. Types of Conda Environments in OCI

1. Data Science Conda Environments

 Managed and curated by Oracle Data Science Service team.

 Built for specific needs:

o Frameworks (e.g., TensorFlow, PyTorch)


o Industry-specific (e.g., Healthcare)

o General-purpose (e.g., Exploratory Data Analysis)

 Contain pre-installed, relevant libraries and frameworks.

2. Published Conda Environments


 Created and managed by users.

 Stored in Object Storage Buckets for:

o Team sharing

o Cross-session use

o Reproducible model training and deployment

 Example: Custom Conda with specific versions of scikit-learn and pandas.

3. Installed Conda Environments


 Active environments installed in a specific notebook session.

 Required for notebook execution.

 Can be either Data Science Conda or Published Conda.


 Persist on the block volume, meaning they remain available after notebook
sessions are deactivated and reactivated.

5. Summary

 Conda environments act as containers for your project’s software stack.

 They enhance reproducibility, collaboration, and isolation.

 The Environment Explorer in OCI helps you manage and filter environments
efficiently.
 Data Science, Published, and Installed Conda environments each serve
unique roles in development and deployment workflows.

5. Data Science Conda Environments

Data Science Conda Environments

Instructor: John Peach

Module Objective:
Understand what Data Science Service Conda Environments are, how they are
structured, their naming conventions, and the key types available for different data
science use cases in Oracle Cloud Infrastructure (OCI).

1. Introduction

 Conda environments can be difficult to build from scratch.

 Oracle Cloud Infrastructure (OCI) Data Science Service provides pre-built


Conda environments, designed on open-source software and easily
customizable.
 These environments are accessible through JupyterLab → Launcher →
Environment Explorer.

Each OCI Conda Environment Includes:

 OCI Python SDK


 ADS (Accelerated Data Science) SDK

Note: The public ADS version does not include Oracle Labs AutoML and Model
Explainability.
These features are only available in specific Conda environments (e.g., General
Machine Learning for CPU/GPU).

2. Types of Conda Environments

A. Application-Based Conda Environments


Built around a specific software or framework:

 Examples:

o ONNX

o PyPGX

o PySpark

o Intel Extension for Sciket-Learn


o Pytorch

o Rapids

o TensorFlow
B. Use Case-Based Conda Environments

Built for specific data science tasks or domains:

 Examples:

o Computer Vision

o Data Exploration and Manipulation

o Financial Services

o General Machine Learning

o Natural Language Processing


o Neurophysiology

o Oracle Database
Each pack includes optimized “best of breed” libraries for its focus area.

3. Conda Environment Families and Naming Conventions

Environment Families

Grouped by:
 Python Version

 Architecture (CPU/GPU)

Example:

 Natural Language Processing for CPU (Python 3.7, Version 2)


 Natural Language Processing for CPU (Python 3.7, Version 1)
→ Both have the same Python version & architecture, but different software
builds.

General Machine Learning Example:

 General Machine Learning for CPU (Python 3.6, Version 1)

 General Machine Learning for CPU (Python 3.7)

Naming Convention:
1. Application-Based
<Software> <Version> for <Architecture> on Python <Version>
Example:

 PyTorch 1.10 for GPU on Python 3.7


 TensorFlow 2.7 for CPU on Python 3.7 (v1)
2. Use Case-Based

<Task> for <Architecture> on Python <Version>


Example:

 Data Exploration and Manipulation for CPU on Python 3.7

4. Popular Data Science Conda Environments

🧱 Computer Vision

 Focus: Image & video processing, object detection, facial recognition, tracking,
etc.

 Includes libraries:
o scikit-image

o Pillow

o PyTorch

o OpenCV

 Use cases:

o Object detection/tracking

o Image stitching & compression

o Eye tracking & facial recognition

📊 Data Exploration & Manipulation

 Focus: Exploratory Data Analysis (EDA), visualization, and data ingestion.

 Includes:

o Oracle- ADS

o Kafka-Python
o pandas, pandaparallel, dask

o Matplotlib, Seaborn, Plotly, Bokeh

 Use Cases:

o Data set ingestion,Processing and visualization

o Stream consumption from Oracle Cloud Infrastructure Streaming

⚙️ General Machine Learning

 Focus: Generic ML tasks & model explainability.

 Includes:

o xgboost, lightgbm, Keras, TensorFlow

o Oracle AutoML

o Oracle MLX (Model Explainability Library)

o Oracle-ads

o Scikit-learn
o TensorFlow

 Supports:

o Data manipulation

o Supervised ML

o Generic Machine Learning

o AutoML for model optimization

o Machine Learning explainability (MLX)


 Comes with multiple notebook examples.

💬 Natural Language Processing (NLP)

 Focus: Text analysis & language modeling.

 Includes:
o Oracle-ads, eli5, Lime,nltk, keybert, transformers, pytorch-lightning,
simpletransformers

 Supports tasks like:


o Text extraction

o Key phrase extraction

o Parts-of-speech tagging

o Deep learning–based text processing

🔁 ONNX (Open Neural Network Exchange)

 Focus: Model portability and interoperability.

 Features:
o Saves models in a common format for any framework.

o Supports conversion from scikit-learn, TensorFlow, etc.

o Runs via ONNX Runtime (framework-independent).

 Used for:

o Model conversion & inference

o Portability between ML Frameworks

o Transfer models between ML frameworks

o ONNX Runtime library allows you to run the models on different platforms

o Generating workflow graphs


o Model deployment base environment in OCI

 Includes:

o Onnx
o Onnxconverter-common
o Onnxmltools
o Onnxruntime
o Oracle-ads
🗄 Oracle Database

 Focus: Working with on-premise and Autonomous Databases (ATP/ADW).

 Includes:

o ipython-sql

o mysql-connetor-python

o oracle-ads

o SQLAlchemy

 Enables:

o ETL jobs

o Database queries using the ADS Connetor, SQLAlchemy , and python-sql

o Perform analytics in the database without having to move the data to the
notebook

o Batch transformations

o Direct SQL queries within notebooks


 Use ipython-sql for running SQL commands in cells, and ADS Connector for
easy database integration.

🔥 PyTorch

 Focus: Deep learning, computer vision, and NLP.

 Use cases:

o Computer vision, NLP, and General Ml


o Deep neural networks and algorithms for deep learning
o Tensor computing with strong acceleration on GPUS

 Includes:

o daal4py (Intel optimization)

o oneAPI Data Analytics Library


o category-encoders

o Pandas

o Scikit-learn

o Oracle-ads

 Benefits:

o GPU acceleration

o Efficient CPU-based deep learning

o Strong support for neural networks

⚡ PySpark

 Focus: Distributed data processing using Apache Spark.

 Includes:

o sparksql-magic

o oracle-ads
o oraclejdk

o scikit-learn

o pyspark

o MLlib for machine learning

 Enables:

o Writing code in notebooks

o Python-based API for Apache Spark that contains MLlib library for
machine learning

o Develop and test your spark application in the notebook and run it on Data
Flow.
o Running jobs on OCI Data Flow (Spark)

🧱 TensorFlow
 Focus: Building and deploying deep neural networks.

 Use Cases:

o Machine learning
o Deep neural networks
o Flexible architecture runs on CPUs, GPUs, and TPUs

 Includes:

o TensorFlow

o TensorBoard (for visualization)

o Oracle ADS

o Pandas

o Scikit-learn

o Oracle-ads

o Category-encoders
 Supports:

o Image recognition

o NLP

o RNNs and other ML applications


5
. Summary

D
ata Science Conda Environments are curated by Oracle for reliability and
convenience.
 They are categorized by use case, Python version, and architecture
(CPU/GPU).

 They provide:

o Pre-configured tools for faster setup

o Easy customization

o Portability and reproducibility across OCI services

 Popular environments include:


o Computer Vision

o Data Exploration and Manipulation

o General Machine Learning

o NLP

o ONNX

o Oracle Database

o PyTorch

o PySpark
o TensorFlow

6. Manage Conda Environments

🧱 Module: Manage Conda Environments

Instructor: John Peach


Role: Data Scientist, OCI Data Science Service Team

Overview

This lesson explains how to manage Conda environments in Oracle Cloud


Infrastructure (OCI) Data Science using the odsc command-line tool.
It provides more control than the Environment Explorer GUI and allows advanced
management tasks such as browsing, installing, cloning, modifying, publishing,
and deleting environments.

🔍 Recap: What Are Conda Environments?

 A collection of software packages bundled together.

 OCI provides:

o Managed Conda packs by Oracle’s Data Science Service team.

o Custom Conda packs published by users.


 Management can be done:

o Through Environment Explorer (GUI), or

o Through odsc CLI for more control.

⚙️ Key Functionalities of odsc CLI

1. Browse Conda Environments

 View details (name, description, libraries, slug).

 Command:

 odsc conda list

o Shows all OCI Data Science Conda environments.

o Add --local option to list only installed environments.

o Use --override to view published Conda environments (from Object


Storage).

2. Search Conda Environments

 No direct search command.

 Combine with Unix tools like grep:

 odsc conda list | grep -e ‘^[ ]*name:’ -e ‘^ [ ]*slug:’

o Filters specific details (e.g., environment names or slugs).

 Can also be used with tools like awk or perl for advanced searches.

3. Install Conda Environments

 Install Data Science or published Conda environments.

 Command:
 odsc conda install --slug <slug_name>

 To install from Object Storage (published):

 odsc conda install --slug <slug_name> --override


4. Clone Conda Environments

 Duplicate an environment safely before major modifications.

 Command:

 odsc conda clone --fromenv SOURCE_SLUG --env CONDA_NAME

o Automatically creates a new slug for the cloned environment.

5. Modify Conda Environments


 Done using standard conda commands (not odsc).

 Steps:

o You must activate the conda environment in your terminal


o conda activate /home/datascience/conda/<SLUG>/
o Now you can change/modify the environment
o For Example to upgrade ADS:
o python3 -m pip install oracle-ads --upgrade

o Changes apply only to the activated environment.

6. Publish Conda Environments


 Share customized environments via Object Storage (for others, co-workers,
jobs, or model deployments).
Steps:

1. Create a bucket in Object Storage.

2. Initialize configuration:

3. odsc conda init --bucket_namespace <namespace> --bucket_name


<bucket_name>

4. Publish environment:

5. odsc conda publish --slug <slug_name>

7. Delete Conda Environments


 Command:

 odsc conda delete --slug <slug_name>

o Frees up storage by removing unused environments.

8. Create Custom Conda Environments

 Build from a YAML manifest file.

 Command:

 odsc conda create --file [Link]


o Default includes base packages from /opt/[Link].

o Use --empty to skip installing base packages.

🧱 Summary

 odsc CLI enables full management of Conda environments.

 Key operations: browse, search, install, clone, modify, publish, delete, and
create.

 Provides greater flexibility compared to the GUI Environment Explorer.

 Essential for custom workflows, reproducible research, and collaboration.

7. Demo: Manage Conda Environments

🧱 Module: Demo – Manage Conda Environments

Instructor: John Peach


Role: Data Scientist, OCI Data Science Service

Overview
In this demo, you learn how to manage Conda environments using the odsc
command-line tool within Oracle Cloud Infrastructure (OCI) Data Science.
The demo covers browsing, searching, installing, cloning, modifying, publishing,
deleting, and creating custom Conda environments from a YAML file.

⚙️ ODSC Command Overview

 Primary command: odsc

 Subcommand for Conda operations: odsc conda

 Key available options include:


list, init, show-configuration, publish, install, clone, delete, create, etc.

🔍 Browsing Conda Environments

 Lists available Data Science Service–managed Conda environments.


 Command:

 odsc conda list


o Returns a YAML file with information such as:

 Name of the Conda pack

 Slug (unique identifier)

 Type (dataScience = managed by the service)


 To view installed environments:

 odsc conda list --local

o Type will show as local.


 To view published environments (from Object Storage):

 odsc conda list --override

🔎 Searching Conda Environments

 Since list output is in YAML, command-line tools are used for filtering:

 odsc conda list | grep -e ‘^[ ]*name:’ -e ‘^[ ]*slug:’


 grep allows pattern matching (using regular expressions).

 Filters keywords like name or slug to make results more readable.

 Example:

 odsc conda list | grep data_exploration

o Filters only environments related to data exploration.

📦 Installing Conda Environments

 To install a Data Science–managed Conda pack:

 odsc conda install --slug <slug_name>

 You will be prompted to select a version number.

 After installation, verify using:

 odsc conda list --local | grep <slug_name>

o Confirms successful installation.

🔁 Cloning Conda Environments

 Used to duplicate an existing environment before making modifications.


 Command:

 odsc conda clone --fromenv <source_slug> --env "<new_env_name>"


 System automatically assigns a new slug to the cloned environment.

 Example:

 odsc conda clone --fromenv data_exploration-py37-cpu-v3 --env


"My_Data_Exploration"

 After cloning, list installed environments to confirm:

 odsc conda list --local | grep -e ‘^[ ]*name:’ -e ‘^[ ]*slug:’

🧱 Modifying Conda Environments

 Modifications are done using conda, not odsc.


 Steps:

 conda activate /home/datascience/conda/<slug_name>

 python3 -m pip install pendulum

 conda deactivate

 This approach installs new packages or upgrades existing ones in the selected
environment.

☁️ Publishing Conda Environments

 Publishes Conda packs to Object Storage for sharing or reuse.


Steps:

1. Create a bucket in OCI Object Storage (e.g., published-conda-environments).


2. Note the bucket name and namespace.

3. Initialize configuration:

4. odsc conda init --bucket_namespace <namespace> --bucket_name


<bucket_name>

5. Publish the Conda pack:

6. odsc conda publish --slug <slug_name>

7. Verify by listing published Condas:

8. odsc conda list --override | grep -e ‘^[ ]*name:’ -e ‘^[ ]*slug:’


 The bucket will now contain a folder structure:

 Conda Environments/

 CPU/

 My_Data_Exploration/

 v1/

 <conda-files>

🧱 Deleting Conda Environments


 Remove unwanted Conda packs to free space.
 Command:

 odsc conda delete --slug <slug_name>

 Prompts for confirmation before deletion.

🧱 Creating Custom Conda Environments (YAML File)

 You can build Conda environments from scratch using a YAML manifest.

 YAML defines:

o Channels

o Dependencies

o Other required packages


Command:

odsc conda create --file <my_env.yaml>

 By default, installs base dependencies from /opt/[Link].


 To skip these:

 odsc conda create --file <my_env.yaml> --empty

🧱 Summary

In this demo, you learned how to:


 Use the odsc CLI to browse, search, install, clone, modify, publish, delete,
and create Conda environments.

 Combine CLI commands with tools like grep for efficient filtering.

 Manage environment lifecycles directly from the OCI Data Science notebook
terminal.

8. OCI Vault: Introduction


Overview

This module explains why data scientists should use the OCI Vault service to securely
manage credentials, keys, and secrets — instead of storing them directly in code or
configuration files.

Key Concepts & Purpose

 OCI Vault is a centralized, Oracle-managed service that securely stores


encryption keys and secrets (like passwords, tokens, API keys).
 It helps data scientists securely connect to external databases, APIs, and OCI
services without exposing credentials in notebooks or scripts.
 Integration: Works seamlessly with the OCI SDK, CLI, and API clients, and
integrates with other OCI services such as Object Storage, Block Storage, and
more.

Vault Components

1. Vaults – Logical containers for keys and secrets.

o Created in a compartment.

o Two types:
 Virtual Private Vault:

 Dedicated, isolated hardware partition.


 Stores up to 1,000 key versions.

 Can be backed up to Object Storage.


 Supports disaster recovery and cross-region replication.
 Default (Shared) Vault:

 Shared with other Oracle customers.

 Lower cost, but no backups.

 Charged only for stored keys and secrets.


2. Keys – Logical entities representing cryptographic material used for encryption
and digital signatures.
o Supported algorithms: AES, RSA, ECDSA.

o AES: Symmetric (same key for encryption & decryption).

o RSA/ECDSA: Asymmetric (public/private key pairs).

Types of Keys:

o Master Encryption Key: Created or imported by the user.

o Data Encryption Key: Dynamically generated using a master key (used


for encrypting actual data).
o Wrapping Keys: Used for securely sharing or transferring encryption
keys.
→ Uses envelope encryption — master keys encrypt data keys, and data keys encrypt
data.
3. Secrets – Sensitive credentials like passwords, tokens, Usernames, SSH keys,
Auth tokens or API keys.

o Stored securely in Vault and retrieved when needed.

o Reduces risk of credential leaks from notebooks or scripts.


o Each secret has a unique OCID and supports versioning for rotation.

Key Rotation
 Every key and secret can be rotated periodically.

 Rotating keys generates new versions automatically or allows importing new


material.
 Old key versions can still decrypt data but cannot encrypt new data.

 Benefit: Reduces security risk if a key is ever compromised.

Best Practices

 Always store credentials in OCI Vault — never in code or config files.

 Use master keys from your Vault for encryption of storage or data resources.
 Rotate keys and secrets periodically to limit exposure.

 Use OCIDs in code to retrieve secrets at runtime via the SDK, CLI, or API.
 Control access to secrets using IAM policies instead of notebook-level
permissions.

Summary

In this module, you learned:


 What the OCI Vault service is and why it’s important.

 The two types of vaults and their characteristics.

 The types of keys and the role of key rotation.

 How secrets help you keep credentials secure and out of your code.

 How Vault integration enhances security and compliance for Data Science
workflows.

9. Using OCI Vault in OCI Data Science

Overview

This module explains how to use OCI Vault for managing encryption, keys, and
secrets in your Data Science workflows. You’ll learn the difference between Oracle-
managed and Customer-managed keys, and how to use both the OCI SDK and the
ADS SDK to store and retrieve secrets securely.

1. Encryption in OCI

 OCI uses encryption everywhere (data at rest and in transit).

 You’ll often be asked whether to use:


o Oracle Managed Keys – Handled automatically by OCI.

o Customer Managed Keys – Keys created and managed by you in your


own Vault.

🔹 Oracle Managed Keys

 OCI handles key creation, encryption, and decryption.


 Used when provisioning resources like Object Storage, Block Volume, OKE
clusters, etc.

 Data is always encrypted — this cannot be disabled.

🔹 Customer Managed Keys

 Stored in your own Vault.

 Used when stricter security or compliance rules apply (e.g., Security Zone
compartments).
 You manage key rotation, lifecycle, and access control.

 The key may be imported into the Vault or generated by the Vault service.

 The master key in your Vault creates data encryption keys that perform the
actual encryption.

2. Setting Up Customer Managed Keys

1. Choose Customer Managed Key when creating a resource.


2. Locate your Vault.

3. Select a master key to use for generating data encryption keys.

➡️ This allows you to control encryption behavior and key rotation independently from
Oracle.

3. Working with Secrets in Python


There are two main approaches:

A. Using the OCI SDK

The OCI SDK is a general-purpose API for working with Vaults, keys, and secrets.

Storing a Secret

1. Create a credentials dictionary (e.g., for database connection):

2. credentials = {

3. "database": "ADB",

4. "username": "admin",
5. "password": "mypassword"

6. }
7. Convert it to JSON → Base64 encode it.
(Use helper function to handle conversion.)
8. Create a Base64SecretContentDetails object containing the encoded data.

9. Create a SecretDetails object that includes:

o Compartment OCID

o Vault ID

o Key ID

o Secret name & description

o Encoded content
10. Use the VaultsClient class:

o Load the OCI config file.

o Create a VaultsClient instance.

o Call:

o create_secret_and_wait_for_state(secret_details, state='ACTIVE')
Retrieving a Secret

1. Use the SecretsClient class with the same OCI config.

2. Call:
3. get_secret_bundle(secret_ocid)

4. Access the secret from:

5. [Link].secret_bundle_content.content

6. Decode Base64 → JSON → Python dictionary.

✅ Result: You now have your secret data securely retrieved as a Python dictionary.

B. Using the ADS SDK (Simplified for Data Scientists)


The Accelerated Data Science (ADS) SDK provides specialized SecretKeeper
classes designed for common data science use cases.
These make secret management much easier than using the low-level OCI SDK.
Available SecretKeeper Classes

SecretKeeper Class Purpose

For Oracle Autonomous Database credentials (can also store


ADBSecretKeeper
wallet files).

BDSSecretKeeper For OCI Big Data Service (HDFS, Hive, etc.).

MySQLDBSecretKeeper For Oracle MySQL Database credentials.

AuthTokenSecretKeeper For authentication/access tokens (e.g., Streaming, GitHub).

Encode The Secret:


# Encode the secret.
def dict_to_secret(dictionary):
return base64.b64encode([Link](dictionary).encode('ascii')).decode('ascii')

secret_content_details = Base64SecretContentDetails
content_type = [Link].CONTENT_TYPE_BASE64
stage=[Link].Base64SecretContentDetails.STAGE_CURRENT
content=dict_to_secret((credentials))

# Bundle the secret and metadata about it.


secrets_details = CreateSecretDetails(
compartment_id=compartment_id,
description="Data Science service test secret"
secret_content = "secret_content_details"
secret_name = "Database creds"
vault_id = vault_id
key_id = key_id
)
 The data is covered into JSON format, and then encoded into base64 string to
be stored in the Vault .
 The contents of the secret are stored in a Base64Secert ContentDetails object

Store Secret in the Vault:


The VaultsClient class takes a configuration object and establishes a connection to the
Vault service
# Store secret and wait for the secret to become active
config = from_file([Link]([Link]("~"), ".oci", "config"), "DEFAULT")
vaults_client_composite = VaultsClientCompositeOperations(VaultsClient(config))
secret = vaults_client_composite.create_secret_and_wait_for_state(
create_secret_details=secret_details,
wait_for_states=[[Link].LIFECYCLE_STATE_ACTIVE]).data

Retrieve the Secret from the Vault:


 The SecretsClient class takes a configuration object.
 The get_secret_bundle method takes the secret’s OCID and returns a
Response object.
 Its data attribute return the SecretBundle object. This has an attribute
secret_bundle_content that has the object
Base64SecretBundleContentDetails and the content attribute of this object
has acutal secret.

# Retrieve the secret bundle and extract the content


def secret_to_dict(wallet):
return [Link](base64.b64decode(wallet).decode('ascii'))
secret_bundle = SecretsClient(config).get_secret_bundle(secret_id)
secret_content =
secret_to_dict(secret_bundle.data.secret_bundle_content.content)

Using OCI Vault with ADS :


ADS provides a set of classes to make it easier to store and retrieve secrets in OCI
Vaults when using Oracle Autonomous Database, BDS, OCI MySQL Database, or Auth
tokens.

 ADBSecretKeeper : Store and retrieve credentials for autonomous database. It


optionally has support for the database wallet.
 BDSSecretKeeper : Stores and retrieves credentials for the OCI Big Data
Service.
 MySQLDBSecretKeeper : Stores and retrieves credentials to Oracle MySQL
Database.
 AuthTokenSecretKeeper : Stores and retrieves Auth Token or Access Token
string. This could be an Auth Token to use to connect to streaming, Github, and
so on.

MySQLDBSecretKeeper:
 Understands what is needed to connect to MySQL
 Store / retrieves the secrets for MySQL
 Works with the ADS Database connection tool.
# Use ADS to store secret for MySQL
from [Link] import MySQLDBSecretKeeper
mysql_keeper = MySQLDBSecretKeeper(vault_id=vault_id, key_id=key_id, credentials)

mysqldb_keeper.save (name="mysql_employee", description="My DB credentials")

# Use ADS to get a secret for MySQL


with MySQLDBSecretKeeper.load_secret('[Link].oc1..<unique_ID>') as mysqldb_secret :
print(mysqldb_secret['user_name'])

4. Advantages of Using ADS SDK

 Tailored for data science workflows.

 Integrates directly with Autonomous DB, MySQL, Big Data Service, etc.

 Automatically handles encoding, storing, and retrieving secrets.

 Cleaner, shorter, and safer code compared to OCI SDK.

5. Summary

 OCI encrypts all data by default using Oracle-managed or Customer-managed


keys.

 Customer-managed keys offer better control and compliance (stored in your


Vault).
 Secrets (like database credentials) should never be stored in code.

 Use OCI SDK for full control, or ADS SDK for a data-scientist-friendly workflow.

 SecretKeeper classes simplify saving and retrieving secrets securely.

10. Code Repositories (Git)


🧱 Module Overview

Instructor: John Peach (Data Scientist, OCI Data Science Service Team)
Topic: How version control systems (Git) integrate with OCI Data Science to manage
source code, Jupyter notebooks, and related resources.

🧱 1. Version Control Systems (VCS) Basics

 Also known as Source Code Management (SCM) systems.

 Allow tracking, managing, and reverting to different versions of files — including


code, notebooks, data, and reports.

 Originally designed for software development, now essential for data science
workflows.

🔹 Examples of VCS:

CVS, Subversion, Perforce, Mercurial, Bazaar, CodeCommit —


but Git is by far the most popular and is integrated in OCI Data Science.

🗂️ 2. What is a Repository (Repo)?

 A repo is like a filing cabinet for your project — it holds all files, code, and
tracked changes.

 Each project (analysis, model, report) has its own repo.


 Git allows:
o Version control (track and revert changes)

o Collaboration (merge changes from multiple users)

o Branching (work on different versions in parallel)

o Archiving (keep historical versions)

👥 3. Centralized vs Distributed Version Control


Type Description Examples Advantages

Single main server stores all


Centralized CVS, Simpler setup, controlled
versions; developers commit
VCS Subversion workflow.
to it.

Work offline, no single


Distributed Each user has a full copy of Git, Mercurial,
point of failure, flexible
VCS the repo. Bazaar
branching.

➡️ Git is distributed — but teams often use a hybrid model with a central peer (like
GitHub or OCI Code Repo).

💻 4. Git for Data Science

 Tracks code + notebooks + reports.

 Enables collaboration, experimentation, and rollback.

 Fast (most operations are local).

 Fault tolerant — even if the central repo is down, local copies remain.

🧱 5. Git in OCI Data Science (JupyterLab Integration)

 OCI’s JupyterLab Git Extension provides a GUI for Git inside notebook
sessions.

 Accessible via sidebar icon or top Git menu.

You can:

 Create / clone repos

 Stage and commit changes


 Push / pull to/from remote repos

 View diffs between versions

Supports integrations with:


 OCI Code Repository

 GitHub
 GitLab

 Bitbucket

 Custom Git servers

🧱 6. Key Git Terminology

Term Meaning

Commit Snapshot of your project at a point in time (with SHA ID).

Repository Directory tracking all project versions.

Working Area The current local folder where files are being edited.

Staging Marking files to include in the next commit.

🔁 7. Git Workflow Summary

1. Modify files in the working area.

2. Stage selected files (git add).

3. Commit changes with a message (git commit -m "message").

4. Push to remote repo (git push).

5. Pull updates from remote (git pull).


Frequent commits = safer rollback and better traceability.

☁️ 8. OCI Code Repository

 Acts as a Git-based, centralized peer hosted inside OCI.

 Integrated with OCI IAM (Identity & Access Management).

 Each repo gets an OCID, visible in the OCI Console.

 Supports:

o Commit, branch, clone, delete operations

o Viewing commits and repo size


o Integration with external repos (e.g. GitHub, GitLab)

o Replication of GitHub repos into OCI via secure Vault-stored credentials

🔐 9. Connecting to External Repos

 Use SSH keys or HTTPS tokens (stored in OCI Vault as secrets).

 GitHub connection steps:

1. Install and configure Git locally.

2. Create a GitHub account and generate SSH keys.


3. Add the public key to GitHub.

4. Create or clone a repo.

5. Work locally → commit → push to remote.

🧱 10. Common Git Commands

Command Description

git init Create a new repo

git clone <url> Copy remote repo to local

git add <files> Stage files

git commit -m "msg" Commit staged files

git push Send commits to remote

git pull Get + merge updates from remote

git fetch Download changes (no merge yet)

git remote Manage remote connections

🧱 11. Summary

 Version control = essential for collaboration and reproducibility.


 Git = most widely used distributed VCS.
 OCI Data Science integrates Git directly into JupyterLab.

 OCI Code Repository offers secure, IAM-integrated Git hosting.

 External services (GitHub, GitLab) can connect securely via OCI Vault.

 Basic Git workflow: init → add → commit → push → pull.

11. Demo: Code Repositories (Git)

🧱 Module: Demo — Code Repositories (Git)

Instructor: John Peach, Data Scientist (OCI Data Science Team)

Objective:
Demonstrate how to:

1. Configure Git in a JupyterLab notebook session

2. Generate and use SSH keys for authentication

3. Create and link local and remote (GitHub) repositories


4. Commit, push, and sync notebooks between OCI and GitHub

🧱 1. Git Configuration

 Git is pre-installed in OCI Data Science Notebook Sessions.

 Configure username and email for commit identity:

 git config --global [Link] "Your Name"

 git config --global [Link] "you@[Link]"


 git config --list # verify settings

 These details appear in commit history to identify contributors.

📂 2. Create a Local Repository

1. In JupyterLab → open File Browser.


2. Create a new folder (e.g., demo).
3. Go to the Git sidebar → click Initialize Repository.

4. You now have a local Git repository with staging and history tools visible.

🔐 3. Generate and Add SSH Keys (for GitHub Access)

GitHub requires authentication for pushing changes.


Steps:

1. Generate a new key pair:

2. ssh-keygen -t rsa -b 4096 -C your_email@[Link]


3. list keys : ls .ssh/id_<ssh-key>

4. Two files are created:


o Private key → keep secret.

o Public key → share safely.

5. Start the SSH agent:

6. eval "$(ssh-agent -s)"

7. Add the key to the agent:

8. ssh-add -k ~/.ssh/<private key>


9. Copy the public key and add it to GitHub:
GitHub → Settings > SSH and GPG keys > New SSH Key → Paste → Save.

☁️ 4. Create and Link Remote GitHub Repository

1. On GitHub → click New Repository → Name it (e.g., demo).

2. Choose SSH clone URL (e.g., git@[Link]:user/[Link]).

3. In JupyterLab Git panel → Add Remote Repository → Paste URL.

4. Verify connection:
5. git remote -v
→ should list origin for fetch and push.
🧱 5. First Commit & Push

1. Create a new notebook (File > New Notebook).


2. It appears as Untracked in Git tab.

3. Right-click → Track → file becomes staged.

4. Add commit message → Initial Commit → Commit.

5. Push to GitHub:

6. git push --set-upstream origin master


7. Afterwards, use Push to Remote for future commits.

✅ Check GitHub → the notebook and commit message appear.

🧱 6. Updating and Syncing Changes

1. Edit notebook → Save.

2. Git tab shows modified file.


3. Stage → Commit → Push with a message (e.g., “some math”).

4. Refresh GitHub → updated file and commit are visible.

🧱 7. Summary

In this demo, you learned how to:


 Configure Git user identity

 Generate and manage SSH keys

 Create and initialize a local Git repo

 Create a matching GitHub repo

 Link both via SSH

 Stage, commit, and push notebook files

 Maintain versioned synchronization between OCI Data Science and GitHub


4. Machine Learning Lifecycle
1. ML Lifecycle: Overview

🧱 Module: Machine Learning Lifecycle — Overview

Instructor: Wes Prichard, Senior Principal Product Manager (Data Science & AI
Services)

🎯 Objective

Understand the six-step lifecycle of building, deploying, and managing machine


learning (ML) models within OCI Data Science — from accessing data to monitoring
models in production.

🔄 Simplified ML Lifecycle (6 Steps)

Each ML project begins with a business problem and proceeds through these stages:

1. Data Access

2. Data Exploration & Preparation

3. Modeling (Building & Training)

4. Model Validation (Evaluation)

5. Model Deployment

6. Model Monitoring (and Refresh/Retirement)

⚙️ The process is iterative, not linear — data scientists refine multiple steps until
model performance meets business goals.

1⃣ Data Access

 Purpose: Gather relevant data for the business problem.


 Data Sources:

o Enterprise systems → data lakes, relational/non-relational databases

o OCI Object Storage, Data Lakehouse, Data Catalog

o External/public datasets, APIs, sensors, surveys, web scraping


 OCI Tip: Store working data within the notebook session for fast access.

2⃣ Data Exploration & Preparation

 Goal: Cleanse, transform, and understand data before modeling.

 Tasks:

o Identify missing, corrupt, or duplicate data → fix/remove

o Detect and handle outliers

o Check for bias or imbalance


o Perform feature analysis (distributions, correlations, summary stats)

o Visualize data relationships (pair plots, histograms, box plots, etc.)


 Feature Engineering:

o Create new features (e.g., time-of-day from timestamps)

o Convert categorical → binary (one-hot encoding)

o Normalize or scale numerical features


 If data lacks labels: Use OCI Data Labeling Service to annotate datasets.

3⃣ Modeling (Building & Training)

 Choose ML type:

o Supervised Learning: labeled data (classification, regression)

o Unsupervised Learning: unlabeled data (clustering, segmentation)

 Process:

o Split data into training and testing sets.

o Train multiple model candidates using different algorithms.


o Experiment with feature subsets to optimize performance and cost.
 Goal: Find the best-performing model for the defined objective.

4⃣ Model Validation (Evaluation)

 Purpose: Assess how well the trained model performs on unseen data.

 Key Evaluation Metrics:

o Classification: accuracy, precision, recall, F1-score, confusion matrix

o Regression: RMSE, MAE, R² (coefficient of determination)

o Unsupervised: cluster cohesion and separation

 Metric Selection: Should align with the business goal (e.g., precision > accuracy
in fraud detection or rare disease prediction).

5⃣ Model Deployment

 Goal: Make the trained model available for use.

 Deployment Modes:

o Batch inference: scheduled predictions (e.g., daily churn scoring)


o Real-time inference: on-demand predictions (e.g., fraud detection)

 Considerations:

o Response time (latency requirements)

o Number of requests

o Data volume
 OCI Tip: Deploy models via OCI Data Science and integrate with MLOps
pipelines for automated workflows.

6⃣ Model Monitoring & Refresh

 Purpose: Ensure models remain accurate and reliable after deployment.

 Types of Monitoring:
1. Drift/Statistical Monitoring:

 Detect data drift or concept drift (distribution changes)

 Compare training vs. live data distributions


2. Operational (Ops) Monitoring:

 Track latency, throughput, CPU/memory usage, reliability

 Set up logs & alerts for incident analysis


 When to Retrain:

o When prediction accuracy degrades


o When live data deviates significantly from training data
 Outcome: Retrain → redeploy → repeat lifecycle as needed.

🔁 Key Takeaways

 The ML lifecycle is cyclical and collaborative (data scientists + ML engineers +


DevOps).

 Each step builds upon the previous, but feedback loops are frequent.
 OCI Data Science provides managed tools for each stage — from data access
and labeling to deployment and monitoring.

2. Access Data

🧱 Lesson Title: Access Data

Instructor: Himanshu Raj — Data Scientist & Senior Training Lead (AI/ML), Oracle
Module: Machine Learning Lifecycle — Step 1: Access Data

🎯 Objective

Understand:
 Why data is needed in machine learning
 How data is collected

 What data sources are supported in Oracle Cloud Infrastructure (OCI) Data
Science
 How to connect and access data through the Accelerated Data Science (ADS)
SDK and console

🧱 1. Importance of Data

 Everything we do—digitally or non-digitally—generates information.

 This information is the foundation for insights and decision-making in data


science.

 Data enables:
o Hypothesis-driven research

o Data-driven insights

o Problem-solving through modeling

🗝️ Without data, we cannot train models, test hypotheses, or make reliable business
conclusions.

⚙️ 2. Types of Data (by Source and Nature)

Type Description Examples

Collected over time from scheduled or daily


Batch Data Backups, data migrations
operations

Streaming
Real-time messages or logs IoT devices, user events
Data

Application Logs, event tracking, API


Generated by app events or APIs
Data calls

All this data must be brought into OCI Data Science for preprocessing and model
training.

☁️ 3. How to Access Data in OCI Data Science


You can access data either via:
 Console / UI (simple uploads)

 Command Line / ADS SDK (programmatic access via Python)

OCI Data Science supports multiple data sources:

🔹 a. Oracle Object Storage

 Main and most common data source.


 Use ADS to load data from Object Storage into a DataFrame.

 Access via:
o API Key authentication

ads.set_auth(auth="api_key", profile="DEFAULT")
bucket_name = <bukect_name>
file_name= <file_name>
namespace = <namespace>
df = pd.read_csv(f"oci://{bucket_name}@{namespace}/{file_name}",
storage_options={"signer": default_signer()})

o Resource Principal (used in serverless functions)

ads.set_auth(auth="resource_principal", profile="DEFAULT")
bucket_name = <bukect_name>
file_name= <file_name>
namespace = <namespace>
df = pd.read_csv(f"oci://{bucket_name}@{namespace}/{file_name}",
storage_options={"signer": default_signer()})

 Example functions:

from [Link] import DatasetFactory


[Link]("oci://<bucket_name>@<namespace>/<path_to_file>")

 Use set_auth() to toggle between resource principal and key pair authentication.
🔹 b. Local Storage

 Access local files using standard Pandas methods:

import pandas as pd
df = pd.read_csv("local_path/[Link]")

🔹 c. Oracle Autonomous Databases (ATP/ADW)

 Supported through ads.read_sql(), which is 15× faster than pandas.read_sql()


because it bypasses ORM.
 With Wallet File:

from ads import read_sql


read_sql("SELECT * FROM TABLE", connection_parameters)

 Without Wallet File: (ADS ≥ 2.5.6)

connection_parameters = {
"user": "admin",
"password": "mypassword",
"host": "hostname",
"port": 1521,
"service_name": "servicename"
}

 ⚠️ Use bind variables to prevent SQL injection.

 Performance depends on network latency; can be optimized using indexes and


efficient SQL.

🔹 d. MySQL

 Same as Oracle Autonomous DB, but set engine as MySQL.

 Available in ADS version 2.5.6+.

ads.to_sql(df, "table_name", engine="mysql")


🔹 e. Amazon S3

 Supports public and private buckets.

 Use Pandas with ADS storage_options for credentials:

df = pd.read_csv("s3://bucket/[Link]", storage_options={"key": "...", "secret":


"..."})

🔹 f. HTTP / HTTPS Endpoints

 Access data via direct URL:

df = pd.read_csv("[Link]

🔹 g. DatasetBrowser

 Built-in ADS utility to explore reference datasets from:

o Seaborn
o Scikit-learn

o GitHub, etc.

 Functions:

from [Link].dataset_browser import DatasetBrowser


[Link]() # view all available datasets
[Link]("iris") # load a specific dataset

🔹 h. PyArrow (OCI File System Integration)

 Used for big data access and processing.

 The OCI FS library enables file-system-like operations for OCI File Systems.
🧱 4. Data Types (Semantic Detection in ADS)

ADS automatically detects data types when loading datasets:

Data Type Description Example

Categorical Labeled groups without numeric


Eye color, shirt size
(Qualitative) meaning

Ordinal Ordered categories with intrinsic ranking Education level

Height,
Continuous Measurable quantitative data
temperature

Datetime Temporal data Timestamp, date

You can inspect data types using:

dataset.feature_types
dataset.show_in_notebook()

📦 5. Supported Sources & Formats

✅ Supported Formats: CSV, JSON, Parquet, ORC, XLSX, etc.


❌ Unsupported Formats: TXT, DOC, PDF, raw images, lists, tuples, etc.
➡️ For DOC/PDF, ADS provides a text extraction module to convert them into plain
text.

🧱 Key Takeaways

 Data is the first and most critical step of the ML lifecycle.

 OCI Data Science + ADS SDK provides seamless ways to access data from
cloud, databases, and local sources.

 Always ensure:

o Correct authentication method (API key / resource principal)

o Efficient queries & secure SQL


o Awareness of data types and formats
 Use DatasetBrowser and PyArrow for efficient exploration and handling of large
or reference datasets.

3. Data Preprocessing

🧱 Lesson Title: Data Preprocessing

Instructor: Himanshu Raj — Data Scientist & Senior Training Lead (AI/ML), Oracle
Module: Machine Learning Lifecycle — Step 2: Data Exploration and Preparation

🎯 Objective

Learn:
 Why data preprocessing is needed

 Common data cleaning and transformation steps

 How to use OCI Data Science’s ADS tools for preprocessing

 How to split data for training, testing, and validation

🧱 1. What Is Data Preprocessing and Why It’s Needed

Preprocessing : Clean, Impute, Engineer, and normalize features.

 Real-world data is imperfect — it can contain:

o Missing values

o Errors

o Duplicates
o Outliers

o Inconsistent formats
 Before model training, data must be cleaned and standardized to ensure
accurate results.
 Preprocessing is often the largest and most time-consuming part of the ML
lifecycle.

 Preprocessing of data involves various steps :


o Combining and Cleaning Data

o Data Imputation

o Dummy Variables

o Outlier detection

o Feature Scaling

o Feature Engineering

o Feature Selection

o Feature Extraction

🔹 Data transformations and manipulations prepare raw data for meaningful analysis
and modeling.

⚙️ 2. Common Preprocessing Operations

a. Combining and Cleaning Data

 Data often comes from multiple sources and needs merging.

 Perform row/column operations such as:

o Append / Delete / Filter

o Join / Concatenate (vertically or horizontally)

 Maintain consistency:
o Use proper formats, units, and naming conventions

o Remove duplicates and incomplete rows

ADS datasets support all operations that can be performed on a Pandas DataFrame.

b. Data Imputation (Handling Missing Values)

 Missing data can occur due to human error, bad sensors, or transmission
failures.
 Approaches:
o Deletion: Remove incomplete rows (not recommended).

o Imputation: Replace missing values with:

 Mean / Median / Mode

 Mode is preferred for categorical features.

ADS supports built-in methods for automatic imputation.

c. Encoding Categorical Data


 Categorical variables must be converted into numbers before modeling.

Method Description Use Case

Label Good for nominal categories


Converts each category into a number.
Encoding (no order).

One-Hot Creates binary (dummy) columns for Best for ordinal or


Encoding each category. unordered data.

Example (Pandas):

For label encoding : From [Link].label_encoder import DataFrrameLabelEncoder.

pd.get_dummies(df['Category'])

Or use ADS fit_transform() to encode all categorical columns at once.

d. Outlier Detection

 Outliers are data points that deviate significantly from others.

 They can be:


o Errors or

o Valid but rare observations

 Detection methods:
o Visualization: Scatterplots, Boxplots

o Statistical analysis: Deviation from normal distribution


 Supervised methods require labeled data (time-consuming).

 Unsupervised methods assume outliers are few and distinct from normal
samples.

e. Feature Scaling

 Brings all features to a comparable scale, essential for algorithms using


Euclidean distances (e.g., regression, clustering).

Method Description Formula

(x - min) / (max -
Normalization (Min-Max) Scales values between 0–1.
min)

Standardization (Z- Centers around mean 0 with unit


(x - μ) / σ
score) variance.

Ensures that all features contribute equally to model learning.

f. Dimensionality Reduction

 High-dimensional data is computationally expensive and prone to overfitting.

 Two main approaches:


1. Feature Selection — Choose the most relevant features

o Variance Thresholds
o Correlation Thresholds
o Genetic Algorithm
2. Feature Extraction — Create new features from existing ones (e.g.,
PCA).

o Principal Component Analysis


o AutoEncoders
o Linear Discriminant Analysis

Reduces complexity while preserving essential information.

g. Text Data Preprocessing

 Text requires specialized steps such as:


o Tokenization

o Removing stop words

o POS tagging

o Stemming / Lemmatization

o Vectorization

ADS provides built-in utilities for text transformation.

🧱 3. ADS Data Transformation Tools

🔹 1. suggest_recommendations()

 Detects issues in data (e.g., missing values, imbalance, correlations).

 Recommends actions (like imputation, dropping columns, etc.).


 User can apply suggestions manually or accept all via dropdown.

 Transformed data is retrieved with get_transformed_dataset().

🔹 2. auto_transform()

 Applies all recommended transformations automatically.

 Handles:

o Missing values

o Strongly correlated columns

o Imbalanced classes (upsampling/downsampling)

o Non-predictive columns (like primary keys)


 Returns a cleaned, optimized dataset.

⚙️ fix_imbalance=True (default) ensures class balance automatically.

🔹 3. visualize_transforms()

 Shows a visual flow diagram of all transformations applied.


 Displays detected correlations, imbalance, and feature summaries.
 Only visualizes automated transformations, not custom ones.

📊 Example Workflow

1. Apply suggest_recommendations() → View detected issues.

2. Accept or modify transformations.

3. Use auto_transform() → Automatically optimize dataset.

4. Visualize results via visualize_transforms() → Flowchart of preprocessing steps.

Example dataset: Employee Attrition Dataset — demonstrates type discovery,


imbalance correction, and transformation summary.

🔀 4. Splitting Data

Before Inputting data to ML algorithm, we have to split data into data into train, test, and
split.

train, test = transformed_ds.train_test_split()

Splitting ensures models generalize well on unseen data.

Split Type Purpose Default Ratio

Training Set Used to train model 80%

Testing Set Evaluate performance 10%

Validation Set Fine-tune model 10%

 For large datasets, 80–90% for training is fine.

 For smaller datasets, use 60–70% training to retain test representativeness.

 Example in ADS:

 This examples sets split to 70%, 15%, 15%

data_split = transformed_ds.train_validation_test_split(
test_size=0.15,
validation_size=0.15
)
train, validation, test = data_split
print(data_split)

🧱 Key Takeaways

 Data preprocessing ensures clean, consistent, and reliable input for models.

 Major steps include:

o Cleaning, merging, imputing

o Encoding, scaling, detecting outliers

o Reducing dimensionality

o Splitting data effectively


 ADS provides automated transformation tools to simplify and accelerate the
workflow.
 Use auto_transform and visualize_transforms to optimize preprocessing with
minimal manual effort.

4. Demo: Data Preprocessing

Lesson Title: Demo: Data Preprocessing

Instructor: Himanshu Raj – Data Scientist & Senior Training Lead (AI/ML), Oracle

Summary

This demo provides a hands-on walkthrough of data preprocessing using Oracle


Cloud Infrastructure’s Accelerated Data Science (ADS) SDK. It showcases how to
access datasets, apply automated transformations, encode categorical data, handle
class imbalance through upsampling, and finally split the data for training and testing.
Key Steps Demonstrated

1. Dataset Overview

 Dataset used: Employee Attrition Dataset

 Size: 1,470 rows

 Features: 36 total

o 22 ordinal

o 11 categorical

o 3 constant
 Contains demographic, job satisfaction, compensation, and performance-related
data.
 Imbalance: Fewer employees leave compared to those who stay.

2. Loading Data

 Imported required libraries including Accelerated Data Science (ADS) and


pandas.

 Dataset loaded from Object Storage using:

 [Link]()
 Defined bucket and namespace parameters.

 Set target feature as attrition.

3. Suggest Recommendations

 The suggest_recommendations() tool analyzes the dataset and provides:

o Detected data issues

o Suggested fixes (e.g., correlations, imbalance)

o Ready-to-use code snippets for corrections

4. Auto Transform
 The auto_transform() method automatically applies all recommended
transformations.

 Optimizations include:
o Missing value imputation

o Noise reduction

o Dropping strongly correlated columns

o Handling class imbalance (upsampling/downsampling)

o Removing non-predictive columns (e.g., primary keys)

 Greatly reduces manual preprocessing time.

5. Visualize Transforms

 visualize_transforms() displays the sequence of transformations performed.

 Helps compare results with and without auto_transform.

 Provides a clear understanding of data changes applied automatically.

6. Encoding Categorical Data

 Example: Encoding the job_function feature.

 Used ADS’s built-in Label Encoder:

 from [Link].label_encoder import LabelEncoder


 Converts categorical values into numeric form.

7. Handling Class Imbalance

 Used upsampling via:

 from [Link] import upsample

 Balanced the dataset by repeating underrepresented class samples.

 Verified balance using value counts before and after upsampling.


8. Splitting the Data

 After preprocessing, the dataset was split into:


o 80% training

o 10% testing

o 10% validation

 Default ratio in ADS; adjustable based on dataset size.

Conclusion
 The demo showcased the power of ADS preprocessing tools:

o Fast, automated, and customizable transformations

o Built-in encoders and samplers

o Visual and interactive analysis


 Users can further explore the ADS documentation, GitHub labs, and sample
notebooks for deeper understanding.

5. Data Visualization

Lesson Title: Data Visualization

Instructor: Jon Stanesby – Senior Principal Instructor, Oracle University

Summary

This lesson explains the importance of Data Visualization (DV) in the data science
lifecycle and demonstrates how Oracle’s Accelerated Data Science (ADS) SDK
provides automated and customizable visualization tools. It highlights how DV simplifies
exploratory data analysis, enables better insights, and improves decision-making
through clear and flexible visual representations.
Key Concepts

1. Importance of Data Visualization

 DV is a core part of Exploratory Data Analysis (EDA) — used to discover


insights and relationships early in the process.
 Makes data understandable through visual representations — charts, plots,
maps, and dashboards.

 Helps both technical and non-technical users interpret data easily.


 Supports decision-making and storytelling with data.

2. Characteristics of a Good Visualization Tool

 Connects to multiple data sources, whether on-premises or in the cloud.

 AI/ML-powered analytics make it easier for non-technical users.

 Pre-built connectors simplify data integration and blending.


 Enables collaboration — shareable across the organization.

 Offers flexibility: manual control or automated visualization.

 Features like drag-and-drop and auto-layout adjustments enhance ease of


use.

3. Data Visualization in ADS

 ADS Smart Visualization automatically detects data types and chooses the
most suitable plots.
 Supports both automatic and custom visualizations using any plotting library.

 Generates visual insights like:


o Summary statistics

o Distribution charts

o Correlation maps

o Anomaly detection (e.g., missing values, high cardinality)


4. ADS Automatic Visualization Methods

Method Description

Calculates and displays correlation matrices. Uses different


corr
methods for each data type pair.

Displays summarized information about the dataset (type, feature


show_in_notebook
overview, correlations, sample preview).

Automatically determines the best plot type for given variables


plot
(e.g., histograms, violin plots, heatmaps).

Creates visualizations for individual or multiple features. Supports


feature_plot
both univariate and multivariate plots.

5. Correlation Methods in ADS

Data Type Correlation


Description Range
Combination Method

Continuous– Linear correlation between two


Pearson -1 to 1
Continuous continuous variables

Continuous– Correlation Ratio Measures curvilinear relationships


0 to 1
Categorical (η) across categories

Categorical– Association strength between two


Cramer’s V 0 to 1
Categorical nominal variables

6. The show_in_notebook() Function

 Provides a comprehensive data overview:

o Dataset type (regression, binary, multiclass)

o Row/column count

o Feature types and distributions

o Correlation map and header preview


 Uses smart sampling — statistically significant subset (95% confidence level,
1% interval) to improve performance.
7. The plot() Function

 Automatically selects plot type based on data:


o Categorical data: Bar chart

o Continuous data: Histogram

o Categorical vs Continuous: Violin plot

o Continuous vs Continuous: Scatter plot or Gaussian heatmap

 Offers flexible axis control (x, y parameters) for exploring relationships.

8. Feature Type System in ADS

 Separates data representation from data meaning.

 Extends Pandas DataFrames with metadata, validation, and visualization


features.
 Allows data scientists to create custom feature types with specialized plots.

 Provides warnings/validation to ensure data quality.

 feature_plot() Method:

 feature_plot() creates custom visualization

 Create univariate plots that are customized to the feature type.


 Result plots across different feature types.
 series.feature_plot() creates a single plot for that feature

 df.feature_plot() creates a collection of plots for all the features in the dataframe.

 Multiple inheritance allows you to reuse plots from other features.

9. Custom Visualizations

 Users can override default ADS plotting using:

 [Link](<custom_plot_function>)

 Integrates seamlessly with other libraries like:


o Seaborn → Pair plots showing pairwise relationships

o Matplotlib → Custom charts (e.g., geographic earthquake plots)

 Enables complete flexibility for personalized data storytelling.

10. Lesson Takeaways

 Data Visualization is essential for insight generation and data storytelling.

 ADS simplifies and accelerates visualization through automation and smart


charting.
 Users can easily switch between automatic and manual plotting methods.

 Supports both exploratory and presentation-ready visualizations.

✅ In short:
Oracle’s ADS provides a powerful, intelligent, and flexible visualization framework
that combines automation with customization — allowing data scientists to move
smoothly from exploration to insight to presentation.

6. Model Training

📘 Key Concepts

🔹 What is Model Training?

 Model training builds a mathematical representation of the relationships


between:
o Features → Target (supervised learning)

o Features ↔ Features (unsupervised learning)

 The output of this process is a model artifact, which captures these learned
patterns.

 Determines the best algorithm for model training; considers tradeoffs in terms of
compute, storage, complexity, performance, explainability, and so on
🔹 Core Components in Training

1. Score Function:
Evaluates how well the model fits the data (e.g., accuracy, log-likelihood).
2. Loss Function (Cost Function):
Measures the difference between predictions and actual values.

o The goal is to minimize loss (e.g., MSE, cross-entropy).

o The graph example shows:

 Green dots → True values

 Black line → Predictions


 Red arrows → Loss values
3. Update Function:
Updates model parameters iteratively (e.g., via gradient descent).

🧱 Open Source + Oracle Ecosystem

OCI Data Science integrates both Oracle proprietary and open-source frameworks.

 Examples:
o scikit-learn, TensorFlow, PyTorch, XGBoost, LightGBM, SpaCy,
MXNet, Keras (open source)

o Oracle AutoML, ADS (Accelerated Data Science SDK) (Oracle tools)

You can:
 Use built-in conda environments in OCI Data Science.

 Install custom libraries via the terminal if needed.

 Train models through:


o Jupyter Notebooks

o Conda Environments using ADS/MLX/AutoML tools

o Jobs (batch training) – covered later in Module 4.

🧱 Takeaways
 Model training = learning the mathematical mapping between data and target.

 OCI Data Science provides flexibility: use open-source libraries, Oracle’s own
AutoML, or both.
 You can train models interactively (in notebooks) or in production (as jobs).

7. Expert Tips: Training a ML model on OCI

📘 1. Training ML Models with Jobs

 You can train models easily on OCI Data Science by creating Jobs using ADS
(Accelerated Data Science SDK).

 The Job defines:

o The resources (compute, memory, etc.)

o The Job Run specifies the training code and output storage.

 Training code can be provided in:


o Python scripts, or

o YAML configuration files.

 Source code can be hosted on GitHub, while results/artifacts can be stored in


OCI Object Storage.

⚙️ 2. Distributed Training

 OCI supports distributed training to handle:

o Large datasets

o Compute-intensive workloads

 Distributed training allows parallelized model training across multiple nodes —


improving speed without losing accuracy.

 Supported frameworks include:


o Dask
o PyTorch Distributed

o Horovod

o TensorFlow Distributed

 These can be implemented using ADS utilities within OCI Data Science.

 You can choose to use Docker containers or GitHub repositories for


implementation.

🤖 3. AutoML and AutoMLX

 OCI provides the AutoMLX package (included in the automlx conda


environment).
 AutoMLX automatically:

o Selects the best algorithm for your data,

o Performs hyperparameter tuning, and

o Generates a ready-to-deploy model pipeline.

 The process can be initialized via the INIT() function.


 It supports parallel processing via the task or local engine.

📚 4. Additional Recommendations

 Explore:
o ADS documentation

o AutoMLX documentation

o OCI Data Science Docs for distributed training frameworks.

 Review release notes regularly for new features.

 Share your experiments and projects in the Oracle University (OU)


Community.

🧱 Key Takeaways

 Use Jobs in OCI to automate and scale training runs.


 Perform distributed training for large or compute-heavy datasets.

 Leverage AutoMLX for automated model selection and tuning.

 Combine ADS + GitHub + Object Storage for a complete ML workflow on OCI.

8. Oracle AutoML: Introduction

📘 Overview

In this lesson, Jon Stanesby introduces Oracle AutoML — a part of the Accelerated
Data Science (ADS) SDK in OCI Data Science.
AutoML (Automated Machine Learning) automates the model training, selection,
feature tuning, and evaluation process — enabling faster, more efficient machine
learning model development with minimal manual intervention.

⚙️ 1. What Is AutoML?

 AutoML (Automated Machine Learning) automates:

o Model selection

o Hyperparameter tuning

o Feature selection
o Evaluation and optimization

 It reduces manual effort and time while maintaining accuracy.


 Useful since model optimization rarely happens on the first iteration —
multiple experiments are usually required.

🧱 2. Common AutoML Approaches

AutoML systems differ in how they optimize model performance and configuration.
Here are the main approaches mentioned:
a. Bayesian Optimization
 Uses a probabilistic model to explore hyperparameter performance.

 Example: AutoSklearn, which applies Random Forest-based Sequential


Model-Based Optimization.
 Uses meta-learning to identify similar previously-optimized datasets to guide
optimization.
b. Recommender System Approach

 Keeps records of best configurations from previous datasets.


 For a new dataset, recommends new configurations based on similarity and
Probabilistic Matrix Factorization (PMF).

c. Genetic Evolutionary Algorithms

 Example: TPOT (Tree-based Pipeline Optimization Tool).

 Uses evolutionary search to optimize pipelines built around scikit-learn.

⚡ 3. Oracle AutoML’s Feed-Forward Approach

Oracle AutoML uses a non-iterative feed-forward method:

 Predicts relative performance of algorithms before training.

 Uses meta-learned proxy models to make fast decisions about which pipelines
to build.
 Builds and tunes only the best candidate models, improving efficiency.

🔹 Advantages:

 Shorter runtime

 Avoids the cold start problem


(using meta-learning trained on diverse datasets)
 Predictive efficiency — evaluates algorithm potential without full training.

🧱 4. Key Benefits of Oracle AutoML

 No-code or low-code automation for model development.

 Automates:
o Algorithm selection

o Hyperparameter tuning

o Feature selection

o Adaptive sampling

 Boosts productivity by saving compute time and reducing manual tweaking.


 Delivers the best-performing model within a given time and budget.

5 .Oracle AutoML Workflow

 Selects a model from a large number of viable candidate models


 Tunes the hyperparameters for each model
 Selects predictive features to speed up the pipeline and reduce overfitting
 Ensures the model trained is generalized and works for unseen data

🔄 6. Oracle AutoML Pipline

1. Algorithm Selection → Choose top-performing algorithms.

2. Adaptive Sampling → Efficiently determine ideal dataset size.

3. Feature Selection → Remove irrelevant or noisy features.

4. Model Tuning → Optimize hyperparameters for final model accuracy.

🔍 7. Step-by-Step Details

a. Algorithm Selection

 Identifies algorithms that yield the maximum predictive score.

 Uses meta-learning trained on many datasets.

 Ranks algorithms by expected performance before training.


b. Adaptive Sampling

 Starts from small subsets → increases gradually.

 Evaluates performance convergence to find minimal data sample size.


 Detects unbalanced datasets and adjusts automatically.

 Reduces training cost and time.


c. Feature Selection

 Selects highly predictive features.

 Removes features that:

o Have too many missing or constant values,

o Are uncorrelated with the target,

o Have high cardinality (too many unique values).


 Ranks features using multiple techniques → identifies optimal subset using
meta-learning.
d. Model Tuning

 Tunes hyperparameters for selected algorithms efficiently.

 Avoids exhaustive grid search.


 Example: For Decision Tree, tunes:

o Max tree depth

o Minimum split percentage

 Supports:

o n_jobs → Degree of parallelism (default = -1, all cores).

o log_level → Control verbosity.

🧱 8. Evaluation & Output

 Produces summaries of:

o Training data info

o Pipeline details

o Model trials
 Note: Adaptive sampling is skipped for datasets with < 1000 records.

 Allows:
o Custom model lists (model_list)

o Custom score metrics


 Binary: roc_auc

 Multiclass: recall_macro

 Regression: neg_mean_squared_error
o Time budget in seconds

o Minimum feature list (to preserve essential features)

💡 9. Advantages of ADS AutoML

 Shorter model training time

 No-code workflow automation

 Meta-learning-driven algorithm selection

 Efficient adaptive sampling and tuning

 Flexible control (time budget, metrics, features)

 High-quality models optimized automatically

🧱 Key Takeaways

 Oracle AutoML automates every major step of the ML workflow.


 Uses feed-forward meta-learning for faster and more accurate model selection.

 Greatly enhances productivity and reduces training time.

 Provides interpretable results and flexibility to customize pipelines.

9. Demo: Oracle AutoML

🧱 Demo: Building a Classifier using Oracle AutoMLx

🎯 Objective
To build a binary classification model using Oracle AutoMLx on the Census Income
dataset (from the UCI Machine Learning Repository) — predicting whether a person
earns more than $50K/year.

🧱 Concept Recap

Machine Learning model building typically involves:


1. Data preprocessing – cleaning, imputing, feature engineering, normalization.

2. Model selection – choosing the best algorithm for the dataset.

3. Hyperparameter tuning – optimizing algorithm parameters for performance.

These tasks are time-consuming and dataset-specific.


Oracle AutoMLx automates this entire process through a Python API, reducing manual
effort and accelerating model development.

⚙️ Environment Setup

 Conda Environment:

o Oracle AutoML and Model Explanation for Python 3.8 (v2.0)


 Libraries Used:
gzip, pandas, numpy, matplotlib, seaborn, scikit-learn, and automlx.

from automl import AutoML

from automl import init


 Notebook setup: includes %matplotlib inline and autoreload magics.

📊 Dataset: Census Income

 Downloaded using:

 from [Link] import fetch_openml

 data = fetch_openml(name='Census-Income', as_frame=True)


 Target variable: Income (>50K or <=50K)

 Dataset has mixed numerical and categorical columns.


 Some numeric columns (like age, hours-per-week) are mislabeled as categorical
— fixed by converting to int.

🧱 Data Preprocessing

1. Fix mislabeled data types


Convert age and hours-per-week from category → int.
2. Handle missing values
AutoMLx automatically:

o Drops features with too many missing values

o Imputes remaining ones based on feature type


3. Train-Test Split

4. from sklearn.model_selection import train_test_split

5. X_train, X_test, y_train, y_test = train_test_split(X, y, train_size=0.7)


6. Binary Encoding of Target

o >50K → 1

o <=50K → 0

⚡ Initializing AutoMLx

init(engine='local', n_jobs=2)

 Engine: local (Python multiprocessing)

 Creates a parallel AutoML engine.

🧱 AutoMLx Pipeline Structure

[Link](task='classification')

🔸 Steps inside the pipeline:

1. Preprocessing – Cleans, imputes, normalizes data automatically.

2. Algorithm Selection – Chooses best algorithms (e.g. XGBoost, LightGBM,


RandomForest, etc.)
3. Adaptive Sampling – Dynamically selects representative subsets of data.

4. Feature Selection – Finds optimal subset of features.

5. Hyperparameter Tuning – Optimizes parameters of selected models.

🧱 Model Training

automl_pipeline = [Link](task='classification')

automl_pipeline.fit(X_train, y_train)

 AutoMLx performs cross-validation (cv=5).

 Automatically disables algorithms not suitable for large datasets (e.g., SVC with
>10k samples).
 Selected Model: LGBMClassifier (Light Gradient Boosting Machine)

 Final AUC ROC score: 0.91

📈 Evaluation

 Metric: roc_auc
(Area under the ROC curve)
 Confusion Matrix Visualization:

 from [Link] import confusion_matrix

 [Link](cm_normalized, annot=True)

o Rows = actual values

o Columns = predicted values

o Shows model’s classification performance.

🔧 Advanced AutoMLx Features

1. Model Control

Limit AutoMLx to specific algorithms:

model_list = ['LogisticRegression']
automl_pipeline = [Link](model_list=model_list)
2. Custom Validation Set

Use your own validation split:

automl_pipeline.fit(X_train, y_train, X_val=X_valid, y_val=y_valid)


3. Tune Multiple Models

Optimize top-N algorithms:

automl_pipeline = [Link](top_n=2)
4. Change Scoring Metric

Default = neg_log_loss
Can switch to accuracy, f1, etc.

automl_pipeline = [Link](score_metric='accuracy')
5. User-Defined Scoring Function

Create a custom scoring metric:

from [Link] import make_scorer, f1_score

f1_scorer = make_scorer(f1_score)
6. Time Budget

Restrict optimization time:

automl_pipeline = [Link](time_budget=10)
7. Minimum Feature List

Force inclusion of specific features:

automl_pipeline = [Link](minimum_features=['fnlwgt', 'native-country'])

🧱 Summary Output

AutoMLx logs:
 Selected algorithm (e.g., LGBMClassifier)

 Selected hyperparameters

 Selected features (e.g., age, education-num, marital-status)


 Performance metrics (AUC, accuracy, etc.)

Use:

automl_pipeline.print_summary()

✅ Final Results

 Model: LightGBM Classifier

 AUC ROC: 0.91

 Key Features: age, education-num, workclass, marital-status, etc.

 Dropped Features: fnlwgt, native-country (optional reinclusion via


minimum_features)

🧱 Key Takeaways

 Oracle AutoMLx automates all major ML pipeline steps: preprocessing →


selection → tuning → evaluation.

 Provides flexibility with manual overrides (time, models, metrics, features).

 Accelerates the process of building and tuning high-performance models.

 Achieved strong performance (AUC ≈ 0.91) with minimal manual tuning.

10. Hyperparameter Tuning: ADSTuner

🧱 Lesson: Hyperparameter Tuning – ADSTuner

🎯 Objective

To understand how to use ADSTuner — the Oracle Accelerated Data Science (ADS)
SDK’s tool for automated hyperparameter tuning — to optimize machine learning
models.
⚙️ What Are Hyperparameters?

 Hyperparameters are configuration settings that control the learning process


of a machine learning algorithm.

 Unlike model parameters (which are learned from data), hyperparameters are
set before training.

🔸 Examples:

 Number of trees in a Random Forest

 Learning rate in a Gradient Boosting model

 Kernel type in an SVM

🔍 What Is Hyperparameter Tuning?

It is the process of:

1. Selecting a set of hyperparameter values

2. Training and evaluating a model for each set

3. Choosing the best-performing combination


This process is iterative and can be computationally expensive, which is why tools like
ADSTuner automate it efficiently.

⚡ ADSTuner Overview

ADSTuner is part of the Oracle Accelerated Data Science (ADS) SDK, designed for
the Oracle Cloud Infrastructure (OCI) Data Science service.

It provides:
 Multiple search strategies for hyperparameter optimization

 Support for user-defined search spaces

 Compatibility with any ML library (e.g., scikit-learn, XGBoost, LightGBM)

🧱 How ADSTuner Works

1. Initialization
You instantiate an ADSTuner object by referencing:
 The model to tune

 Optional parameters such as:


o Number of cross-validation folds

o The search strategy to use

from [Link] import ADSTuner

tuner = ADSTuner(model=my_model, cv=3, strategy="perfunctory")

2. Search Strategies

ADSTuner can search through hyperparameter spaces in different ways:

🔹 Perfunctory Search

 Focuses on key hyperparameters only

 Covers a small search space

 Ideal for quick tests or early-stage tuning


 Goal: Reduce computational cost and quickly gauge model quality

🔹 Detailed Search

 Covers a larger, more exhaustive space

 Tunes more hyperparameters

 Used after identifying the best model type

 Goal: Fine-tune model performance

🔹 Custom Search

 Define your own search space as a Python dictionary

 Useful if you already have intuition about which ranges or values work best

search_space = {

"n_estimators": [100, 200, 300],

"max_depth": [3, 5, 7],


"learning_rate": [0.01, 0.05, 0.1]

tuner = ADSTuner(model=my_model, strategy=search_space)

3. Running the Tuning Process

Use the .tune() method to start the search:

tuning_results = [Link](X, y)
 X = features

 y = target variable
You can specify stopping criteria (e.g., number of trials or maximum time) using the
exit_criterion parameter.

4. Stopping Criteria (exit_criterion)

ADSTuner will stop once one of these conditions is met:

 Maximum number of trials

 Time budget reached

 Convergence achieved
tuning_results = [Link](X, y, exit_criterion={"MAX_TRIALS": 50})

5. Modifying the Search Space

You can:
 Add or remove hyperparameters

 Modify the range of numeric parameters

 Adjust categorical or continuous values anytime before re-running tuning

📊 Output: Tuning Report


After running, ADSTuner generates a comprehensive report containing:

 All hyperparameter trial results

 Best-performing combinations

 Summary statistics

 Visuals comparing different parameter sets and performance scores

tuning_results.show_best(n=5)

🧱 Cross-Validation Support

ADSTuner integrates cross-validation (CV) into its search:

 Ensures model robustness

 Reduces overfitting risk

 Uses CV results to select the best hyperparameters

✅ Key Takeaways

Concept Summary

ADSTuner
Automates hyperparameter tuning within OCI Data Science
Purpose

Supports Any ML library (scikit-learn, XGBoost, etc.)

Perfunctory (quick), Detailed (comprehensive), Custom (user-


Strategies
defined)

Cross-Validation Built-in support for CV folds

Output Tuning report with trials, best performer, and statistics

Exit Criteria Time, trials, or performance threshold

🧱 Example Workflow

from [Link] import ADSTuner


from [Link] import RandomForestClassifier
# Step 1: Define model
model = RandomForestClassifier()

# Step 2: Initialize ADSTuner


tuner = ADSTuner(model=model, cv=3, strategy="perfunctory")

# Step 3: Run tuning


results = [Link](X_train, y_train, exit_criterion={"MAX_TRIALS": 30})

# Step 4: Review best hyperparameters


results.show_best(n=3)

11. Model Evaluation

🧱 Lesson: Model Evaluation

🎯 Objective

To understand the importance of evaluating ML models, the benefits of evaluation,


and how to use ADS Evaluators in Oracle Cloud Infrastructure (OCI) Data Science for
performance analysis and benchmarking.

⚙️ 1. What Is Model Evaluation?

 Model evaluation is performed after model training to determine how well a


model performs on unseen data.

 It helps measure accuracy, precision, recall, and other performance metrics.

 Evaluation compares the predicted outputs against the true labels using a
validation dataset.

🔸 Purpose:

To quantify the predictive power of the model and ensure it generalizes well to new
data.
🧱 2. Why Model Evaluation Matters

✅ Benefits

Benefit Description

Benchmarking Compare different models or algorithms on standardized metrics.

Identify issues like high accuracy but low precision (e.g., class
Pitfall Detection
imbalance).

Trade-off Understand where models perform well or poorly (e.g., one model
Analysis performs better in clear weather, another in poor conditions).

🧱 3. The Role of ADS Evaluator

Oracle’s Accelerated Data Science (ADS) SDK provides an Evaluation module to


simplify performance measurement and visualization.

🔹 Key Classes:

 ADSEvaluator – Computes metrics and generates visual charts.

 ADSModel – Wraps trained models for standardized evaluation.

🧱 Supported Evaluator Types

Type Description Output Example

Binary Two-class problems (e.g.,


Spam detection
Classification Yes/No, 0/1)

Multiclass More than two discrete Sentiment


Classification classes (Positive/Neutral/Negative)

Continuous numeric
Regression House prices, temperature
prediction

📊 4. Using ADS Evaluator

Example Setup
from [Link] import ADSEvaluator
from [Link].sklearn_model import SklearnModel
from sklearn.linear_model import LogisticRegression
from [Link] import RandomForestClassifier

# Convert fitted estimator into ADS model


lr_model = SklearnModel(LogisticRegression().fit(X_train,
y_train)).prepare(inference_conda_env="generalml_p38")
rf_model = SklearnModel(RandomForestClassifier().fit(X_train,
y_train)).prepare(inference_conda_env="generalml_p38")

# Create Evaluator
evaluator = ADSEvaluator([lr_model, rf_model], X_test, y_test)

# Display metrics
[Link]
evaluator.show_in_notebook(perfect=True)

📘 5. Evaluating Different ML Tasks

🧱 A. Binary Classification

Examples: Fraud detection, churn prediction, disease diagnosis

Common Metrics:

 Accuracy

 Precision / Recall / F1 Score


 Hamming Loss

 ROC AUC

 Log Loss

 Confusion Matrix
Visual Charts in ADS:

 Lift & Gain Charts

 Precision-Recall Curve
 Normalized Confusion Matrix
Binary Classification Metrics Example :

 Data has to be split into a testing and training set with the features in X_train and
X_test and the responses in y_train and y_test.
 To generate metrics and charts using ADSEvaluator.
 lr_clf = LogisticRegression(random_state=0, solver='lbfgs',
multi_class='multinomial').fit(X_train, y_train)

 rf_clf = RandomForestClassifier(n_estimators=10).fit(X_train, y_train)

 from [Link] import ADSModel
 bin_lr_model = ADSModel.from_estimator(lr_clf, classes=[0,1])
 bin_rf_model = ADSModel.from_estimator(rf_clf, classes=[0,1])

 from [Link] import ADSEvaluator
 from [Link] import MLData

 evaluator = ADSEvaluator(test, models=[bin_lr_model, bin_rf_model],
training_data=train)

📌 Tip:
perfect=True plots a perfect classifier line in lift/gain charts for comparison.
To show all of the metrics in a table:

[Link]

To show all of the charts :

Evaluator.show_in_notebook(perfect=True)
To add a custom metrics :

Evaluator.add_metrics

🧱 B. Multiclass Classification

Examples: Sentiment analysis, species classification, image labeling


Common Metrics:

 Accuracy

 Hamming Loss

 F1 Score (micro, macro, weighted)

 Precision & Recall (micro, macro, weighted)

 ROC AUC (per class)


Key Points:

 When using ADSEvaluator, specify class levels (e.g., 0, 1, 2) in the classes


argument.

 Supports multiple models for side-by-side metric comparison.

 Charts include:

o Multiclass ROC Curve

o Precision-by-Label

o F1-by-Label

o Multiclass Precision-Recall Curve

Example :

Change the number of classes in from_estimator function.

 lr_clf = LogisticRegression(random_state=0, solver='lbfgs',


multi_class='multinomial').fit(X_train, y_train)

 rf_clf = RandomForestClassifier(n_estimators=10).fit(X_train, y_train)

 from [Link] import ADSModel
 bin_lr_model = ADSModel.from_estimator(lr_clf, classes=[0,1,2])
 bin_rf_model = ADSModel.from_estimator(rf_clf, classes=[0,1,2])

 from [Link] import ADSEvaluator
 from [Link] import MLData

 evaluator = ADSEvaluator(test, models=[bin_lr_model, bin_rf_model],
training_data=train)
🧱 C. Regression Models

Examples: Predicting house prices, sales, or temperature

Common Metrics:

 R² Score

 Explained Variance Score

 Mean Squared Error (MSE)

 Root Mean Squared Error (RMSE)

 Mean Absolute Error (MAE)

 Mean Residuals
Visual Charts:

 Observed vs Predicted

 Residuals QQ Plot (should form a straight line for good models)

 Residuals vs Predicted

 Residuals vs Observed

[Link]
evaluator.show_in_notebook()

📈 6. Interpreting Results

 Compare metrics between models (e.g., LogisticRegression vs RandomForest).

 Review charts for bias patterns or prediction errors.

 Use visual diagnostics (like lift and residual plots) to refine or retrain models.

✅ 7. Summary

Concept Key Insight

Purpose Assess and compare ML model performance


Concept Key Insight

Tool ADSEvaluator and ADSModel classes from ADS SDK

Evaluator Types Binary, Multiclass, Regression

Key Metrics Accuracy, F1, Precision, Recall, R², MSE

Visualizations Confusion Matrix, ROC, Lift & Gain, Residual Plots

Customization Add custom metrics and enable perfect classifier comparison

🧱 Key Takeaway

Model evaluation ensures that your ML model is accurate, reliable, and generalizable.
The ADS Evaluator in OCI Data Science simplifies this process by providing ready-to-
use metrics, charts, and comparisons for multiple models — all within your
JupyterLab notebook.

12. Expert Tips: ADS Evaluators

🧱 Lesson: Expert Tips — ADS Evaluators

🎯 Objective

To demonstrate how to use ADS Evaluators in OCI Data Science to easily calculate
and visualize model performance metrics for different ML tasks — binary, multiclass,
and regression.

👨🏫 Instructor

Hemant Gahankari
Senior Principal Training Lead, Oracle University
⚙️ 1. What Are ADS Evaluators?

ADS Evaluators are built-in tools in the Accelerated Data Science (ADS) SDK that:

 Simplify model performance evaluation

 Automatically generate key metrics and visual charts

 Support multiple model types

🧱 Three Types of ADSEvaluators

Evaluator Type Use Case Example

Binary Classifier Two-class problems Spam detection (Yes/No)

Multinomial Classifier Multi-class problems Handwritten digit recognition

Regression Evaluator Continuous output Predicting house prices

🧱 2. Why Use ADS Evaluators

ADS Evaluators provide a streamlined and unified interface for:

 Computing standard metrics (accuracy, recall, F1, etc.)

 Generating confusion matrices and charts

 Comparing multiple models simultaneously


 Evaluating both training and testing datasets

🧱 3. Practical Example: Binary Classification

▶️ Step-by-Step Workflow

1. Import required modules

2. Generate or load a dataset

3. Split data into training and testing sets

4. Train multiple models (e.g., Logistic Regression, Random Forest)

5. Wrap models with ADSModel

6. Create an ADSEvaluator object


7. Display metrics and charts

🧱💻 Code Example

# Step 1: Import libraries


from [Link] import ADSEvaluator
from [Link].sklearn_model import SklearnModel
from [Link] import make_classification
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from [Link] import RandomForestClassifier

# Step 2: Create a binary classification dataset


X, y = make_classification(n_samples=1000, n_features=10, random_state=42)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2)

# Step 3: Train models


lr = LogisticRegression().fit(X_train, y_train)
rf = RandomForestClassifier().fit(X_train, y_train)

# Step 4: Wrap with ADSModel


lr_model = SklearnModel(lr).prepare(inference_conda_env="generalml_p38")
rf_model = SklearnModel(rf).prepare(inference_conda_env="generalml_p38")

# Step 5: Create ADSEvaluator


evaluator = ADSEvaluator([lr_model, rf_model], X_test, y_test)

# Step 6: Print metrics


[Link]

# Step 7: Plot charts


evaluator.show_in_notebook()

📊 4. Output and Interpretation

Metrics

 Automatically prints side-by-side comparison of metrics for both models.

 Shows results for both training and testing datasets.

 Example metrics:
o Accuracy

o Precision

o Recall

o F1 Score

o ROC AUC
Visual Charts

 Confusion Matrix

 ROC Curve
 Precision-Recall Curve

 Lift and Gain Charts

evaluator.show_in_notebook(perfect=True)

📌 Setting perfect=True overlays a perfect model line on performance charts for


comparison.

🧱 5. Key Advantages

Feature Benefit

Unified Interface Works across multiple algorithms and problem types

Automation Computes all metrics and visualizations in one call

Model Easily compare different models (e.g., Logistic Regression vs


Comparison Random Forest)

ADS Integration Fully integrated within OCI Data Science notebooks

Visualization Instant insight via interactive charts

✅ 6. Summary

Concept Description

Purpose Simplify model evaluation using ADS SDK


Concept Description

Evaluator Types Binary, Multinomial, Regression

Main Classes ADSModel, ADSEvaluator

Outputs Metrics + Visual charts

Function Calls .metrics → text summary, .show_in_notebook() → charts

🧱 Example: Quick Recap

# Quick 3-line evaluation

from [Link] import ADSEvaluator

evaluator = ADSEvaluator([lr_model, rf_model], X_test, y_test)

evaluator.show_in_notebook(perfect=True)

💡 Expert Tip

“Evaluators make the job of calculating and plotting metrics very simple.
Try them with different datasets and explore!”
— Hemant Gahankari, Oracle University

13. Model Explanations : Global Explainer

🧱 Lesson: Model Explanations – Global Explainer

👨🏫 Instructor

Himanshu Raj
Data Scientist & Senior Training Lead, Oracle University

🎯 Objective

Understand model explainability and its importance in the model validation phase of
the machine learning lifecycle.
Focus on global explanation techniques used in Oracle Cloud Infrastructure (OCI)
Data Science.

🔍 1. What is Model Explainability and Interpretability?

Concept Definition

The ability to explain why a machine learning model made a particular


Explainability
prediction.

Interpretability How easily a human can understand those explanations.

👉 Together, they help demystify model behavior — especially for complex models like
XGBoost, Neural Networks, etc.

⚙️ 2. Why Explainability Matters

 Complex ML models act as black boxes.

 Lack of transparency can limit adoption and trust.

 Explainability builds confidence and aids in:


o Debugging models

o Complying with regulations

o Communicating insights to stakeholders

🧱 3. Types of Model Explanations

Type Description Scope

Global
Describes the model’s overall behavior. Entire model
Explanation

Local Describes why the model made a specific


Single prediction
Explanation prediction.

What-If Analyzes how changing input features affects Hypothetical


Explanation predictions. scenario

This lesson focuses on Global Explanation only.


🧱 4. Global Explanation Techniques in ADS (Model-Agnostic)

All global explainers in OCI Data Science are model-agnostic, meaning:

They treat the model as a black box — relying only on inputs and outputs, not internal
model weights.

Technique Purpose Visualization

Feature Permutation Measures each feature’s impact on Box Plot, Bar Chart,
Importance prediction accuracy. Scatter Plot

Feature Dependence Shows how model predictions change with Line, Bar, Heatmap,
(PDP & ICE) feature values. Violin

Similar to PDP, but isolates the true effect


Accumulated Local
of each feature (handles correlated Line or Bar plots
Effects (ALE)
features better).

🔍 5. Technique 1: Feature Permutation Importance

🧱 Concept

 Measures how much the model’s error increases when a feature’s values are
shuffled (destroying its relationship with the target).

 A higher increase in error ⇒ more important feature.

🧱 Steps

1. Compute baseline prediction error (using F1 for classification, R² for


regression).
2. Randomly shuffle values of one feature.

3. Recalculate prediction error.

4. Compare with baseline.

📈 If the model depends heavily on that feature → prediction error rises sharply.

📊 Visualizations
Plot Type Description

Bar Chart Shows average feature importance (longer bar = higher importance).

Box Plot Shows distribution of importance across runs.

Scatter Plot Displays feature importance per iteration.

Interpretation:

 Features higher on the chart = more impact.

 Bars show mean ± standard deviation of importance.

🔍 6. Technique 2: Feature Dependence Explanations (PDP & ICE)

🧱 Concept

Shows how model predictions change as a single feature’s value varies.

⚙️ Process

1. Select a feature to analyze.

2. Take multiple values from its distribution.

3. Replace the feature in all rows with each selected value.

4. Generate predictions → compare differences.

📘 Two Methods

Method Description Output

PDP (Partial Dependence Averages predictions across all Average trend line
Plot) samples for each feature value. or bar chart

ICE (Individual Plots individual predictions for each Multiple lines (one
Conditional Expectation) sample. per instance)

📊 Visual Examples

 Categorical Feature (PDP):


Bar chart showing average predicted value for each category (e.g., Sex →
Female vs Male).
 Numerical Feature (PDP):
Line plot of feature value (x-axis) vs average prediction (y-axis).
 Two Features (PDP):
Heatmap showing average prediction intensity.
 ICE Plot:
Shows individual sample prediction lines (continuous or categorical). Median line
shows trend.

🔍 7. Technique 3: Accumulated Local Effects (ALE)

🧱 Concept

Improves upon PDP by isolating the true effect of a feature — even when features are
correlated.

⚙️ How It Works

1. Divide the feature’s range into intervals.

2. For each interval:

o Compute change in prediction when the feature increases/decreases


within that range.

o Average over all samples in that interval.

3. Plot the cumulative differences.

🧱 Advantages

 Handles correlated features better than PDP.


 Produces more realistic feature effect curves.

📊 Visualization

 Numerical Feature: Line plot showing effect relative to average prediction.

 Categorical Feature: Vertical bar chart showing effect difference per category.

📈 8. Example Context (as per lesson)

 Dataset: Titanic
 Model: XGBoost Classifier

 Built using ADS AutoML provider

 Used to visualize global feature importance and dependence plots

✅ 9. Summary

Concept Description

Explainability Explains why a model made a decision

Interpretability How well humans can understand that reasoning

Global
Feature Importance, PDP/ICE, ALE
Explainers

Visual, automatic, model-agnostic, and easy to integrate in OCI


ADS Advantage
notebooks

💡 Instructor’s Key Takeaway

“Global explanations help you understand how your model behaves overall —
which features drive predictions and how strongly they influence results.”
— Himanshu Raj

14. Model Explanations: Local Explainer

🧱 Lesson: Model Explanations – Local and What-If Explainers

👨🏫 Instructor

Himanshu Raj
Data Scientist & Senior Training Lead, Oracle University

🎯 Lesson Objective
Understand how to interpret individual predictions made by ML models using Local
Explainability and What-If Analysis tools in Oracle Cloud Infrastructure (OCI) Data
Science.

🔍 1. Overview of Model Explainability Techniques

Type Description Focus

Global
Explains the model’s overall behavior. Entire model
Explanation

Explains why a model made a specific


Local Explanation Single observation
prediction.

What-If Shows how changing feature values affects Hypothetical


Explanation predictions. changes

This lesson focuses on Local and What-If explainers.

🧱 2. Local Explainability (LIME in OCI ADS)

OCI Data Science provides an enhanced version of LIME —


Local Interpretable Model-Agnostic Explanations.

⚙️ Key Idea

Even though a model’s global behavior can be complex, its local behavior (around a
single prediction) is often simpler and can be approximated using an interpretable
surrogate model, such as a linear model.

🧱 How LIME Works

1. Start with a trained model.

2. Select a specific sample (observation) to explain.

3. Generate random samples around it (local neighborhood).

4. Use the complex model to predict for each of these local samples.
5. Fit a simple surrogate model (e.g., linear regression) to these local predictions.

6. Use this surrogate model to interpret which features influenced the prediction.
🧱 3. Components of the ADS LIME Explainer

ADS LIME consists of three main sections:


Model, Explainer, and Explanations.

🧱 A. Model Section

 Left column:

o Shows details about the ML model.

o Displays:

 True label or value

 Model’s predicted label/value

 Prediction probabilities (for classification) or predicted values (for


regression)
 Right column:

o Shows the sample being explained.

o For tabular data: displays features and their corresponding values.

o For text data: shows the input text itself.

⚙️ B. Explainer Section

 Left column: configuration of the LIME explainer, including:

o Algorithm used (e.g., LIME)

o Type of surrogate model (e.g., linear)

o Number of generated samples (e.g., 5,000)

o Whether continuous features were discretized


 Right column:

o Legend showing how to interpret colors and bar directions in the


explanation.
📊 C. Explanations Section

Displays the actual local explanation results.

🔸 For Classification

 A local explanation can be generated for each class label.

 In binary classification, one class’s explanation mirrors the other.

 In multiclass, each class gets a separate row showing how features contribute
toward or against that class.

🔸 For Regression

 Shows how each feature increases or decreases the predicted target value.

📈 Feature Importance Visualization

 Shown as horizontal bar charts:

o Bars ordered by relative importance.


o Longer bars = greater influence.
 Positive values (right side): feature increases prediction.

 Negative values (left/gray side): feature decreases prediction.

⚖️ Explanation Quality

Evaluates how well the surrogate model mimics the original model.

Component Description

Sample Distance Shows how generated samples are distributed around the
Distribution explained sample (locality).

Indicate how accurately the surrogate model approximates the


Evaluation Metrics
black-box model (using regression or classification metrics).

🔍 4. What-If Explainer

🎯 Purpose
Helps understand how changes in input features impact the model’s predicted
outcome.

It includes two main exploration techniques:

Method Description

Interactively change feature values for a single observation and see


Explore Sample
how predictions change.

Explore
Examine model predictions over feature distributions (1D or 2D).
Predictions

🧱 A. Explore Sample

 Opens a graphical interface for one observation.

 User can:

o Adjust feature values manually.


o Click Run Inference to recompute predictions.

 The interface shows both:

o Original feature values.

o Updated feature values.

 Allows observing how small changes affect predictions.

📈 B. Explore Predictions

Analyzes how prediction values vary across feature distributions.

Case Description Visualization

One Plots relationship between a single feature (x-axis) and Line or scatter
Feature prediction (y-axis). plot

Two Uses both features as x and y axes. Color scale


Heatmap
Features represents predicted value (target).

Example:

 x = Age, y = CRIM (crime rate)


 Color intensity = predicted target value

✅ 5. Summary

Concept Description

Local Explainability Explains why a specific prediction was made by approximating


(LIME) model behavior locally.

What-If Explainer Tests how changes in input values affect predictions.

ADS LIME
Model, Explainer, Explanations
Components

Evaluation Metrics Measure surrogate model accuracy and sample locality.

Horizontal bar charts, sample distance plots, prediction


Visualization Tools
heatmaps.

💡 Instructor’s Key Takeaway

“While global explanations describe what the model has learned overall,
local and what-if explainers reveal why a model made a particular prediction
and how changes in input features would alter that outcome.”
— Himanshu Raj

15. Expert Tips: Explainers

🧱 Topic: Expert Tips — Explainers

Instructor: Hemant Gahankari (Senior Principal Training Lead, Oracle University)

Key Concepts:

 The Explainer objects are part of Oracle AutoMLx (AutoML for Python).
 To use them, you need to have the automlx_p28_cpu conda environment
active.
 Explainers help understand how models make predictions — providing both
local and global interpretability.

⚙️ Steps to Use Explainer Objects

1. Import and Initialize AutoMLx

2. import automlx

3. [Link]()
4. Train an Estimator

5. from automlx import AutoML

6. model = AutoML()

7. [Link](X_train, y_train)
8. Obtain the Explainer Object

9. explainer = [Link]()
10. Call Explainability Methods

o Local Explanation: Explains one specific prediction.

o Global Explanation: Explains model behavior across all data.

11. explainer.local_explanation(sample=X_sample)
12. explainer.global_explanation()

📚 Extra Notes

 Depending on your data type, AutoMLx automatically chooses:

o TabularExplainer for structured/tabular data.

o TextExplainer for text data.


 These Explainers can visualize feature importance, partial dependence, and
individual feature impact.

 Documentation reference: ML Explainer Interface inside AutoMLx.


16. Model Catalog: Overview

🧱 Lesson: Model Catalog — Overview

Instructor: Jon Stanesby


Course: OCI Data Science
Purpose: To understand how models are stored, tracked, versioned, and deployed
within the OCI Model Catalog.

🧱 What Is the Model Catalog?

The Model Catalog provides:

 A centralized, immutable repository for storing and managing ML models.


 Enables model provenance, reproducibility, and auditability.

 Ensures all models can be traced back to their exact training artifacts.

🔒 Immutability: Once saved, models cannot be modified — to make changes, a


new version must be created.

📦 Model Artifact Components

Each model artifact (a .zip file) contains:

1. [Link] → Python script that loads the model and defines the inference logic.
2. [Link] → Defines the Conda environment and dependencies for
deployment.
3. [Link] (optional) → Contains introspection tests to verify the model
artifact.
4. [Link] → Lists dependencies for running validation.
5. [Link] → Provides setup and saving instructions.

6. Model file(s) (like .pkl, .onnx, etc.)


🧱 Important: Files above the level of [Link] in the directory are ignored during
deployment — keep everything at or below that level.

⚙️ Key Files Explained

🧱 [Link]

 Contains:

o load_model() → Loads serialized model into memory.


o predict() → Defines inference endpoint logic.

 Can include helper functions (e.g., for feature transformations).

📄 [Link]

Specifies:

 Conda environment to use (training or custom)


 Environment slug, path, and type

 Supported Python versions: 3.6 and 3.7

Example structure:

inference_env_slug: "generalml_p38_cpu_v1"

inference_env_type: "data_science"

inference_env_path: "mybucket@namespace/envs/generalml_p38_cpu_v1"

python_version: 3.7

📊 Metadata & Documentation

Each model includes four documentation types:

1. Input/Output Schema

o Describes expected features and payload format.

o Defines contract between client and model API.


2. Model Provenance

o Auto-extracted from Git if available.


o Tracks:

 Training code, environment, and data

 Compute resources and configurations


3. Model Introspection Tests

o Optional pre-save checks verifying operational health.

o Generates a local test_json_output.json.


4. Model Taxonomy

o Describes:
 Use case (e.g., regression, binary classification)

 Framework (e.g., TensorFlow, scikit-learn)

 Algorithm type and hyperparameters (JSON)


 Optional custom metadata fields

🧱 Custom metadata allows adding key-value pairs, with optional category and
description fields.
⚠️ Combined metadata file size limit: 32 KB.

🗂️ Artifact Size Limits

 Console upload: ≤ 100 MB

 ADS SDK / CLI upload: ≤ 20 GB

🔐 Access & Policies

 Like all OCI resources, IAM policies are required.

 Policies control:

o Model catalog management


o Model deployment

o Object Storage access

Example policy:
Allow group DataScientists to manage data-science-model-family in compartment
<compartment_name>

🧱 Key Benefits

✅ Centralized model storage


✅ Full model version control
✅ Provenance and reproducibility
✅ Integration with ADS SDK and OCI Console
✅ Seamless deployment to OCI Data Science endpoints

17. Model Serialization

🧱 Lesson: Model Serialization

Instructor: Jon Stanesby


Course: Oracle Cloud Infrastructure (OCI) Data Science
Focus: How to serialize, save, and manage models within the OCI Model Catalog.

🧱 What Is Model Serialization?

Serialization is the process of converting an object (e.g., a trained ML model) into a


storable or transmittable format — and deserialization is the reverse process.

Also known as marshaling, it allows models to be:

 Saved to disk

 Transferred between systems

 Reloaded later for inference or retraining

🗂️ Common Serialization Formats


Format Use Case / Description

JSON Human-readable; ideal for configuration or structured data

XML Hierarchical data storage

Common for large numerical arrays and neural networks (e.g., Keras
HDF5
models)

Pickle
Python’s native binary format for serializing objects
(.pkl)

Joblib Efficient for NumPy arrays and large scikit-learn models

🧱 ADS Model Serialization

The Accelerated Data Science (ADS) SDK supports multiple ML frameworks:

Framework ADS Serialization Class

scikit-learn SklearnModel

TensorFlow TensorFlowModel

PyTorch PyTorchModel

XGBoost XGBoostModel

Generic Models GenericModel

⚙️ It’s not possible to have a specific serializer for every framework — use
GenericModel for unsupported ones.

💾 Saving Models to the Model Catalog

The save() method:

1. Packages model artifacts ([Link], [Link], model files, etc.)

2. Reloads the latest versions of these files from disk.


3. Optionally runs introspection tests (if ignore_introspection=False).

4. Uploads artifacts to the Model Catalog.


5. Returns the Model OCID.

6. Introspect() can be called after [Link]()


If issues are detected during introspection, ADS provides remediation suggestions.

🧱 Preparing a Generic Model

When working with a custom or unsupported framework:


 Use prepare_generic_model() to wrap it into an ADS model object

 The GenericModel class works with any unsupported model framework that has
a .predict() method.
 The verify() method simlulates a model deployment by calling the load_model()
and predict() methods in [Link] file

 With the .verify() method, you can debug your [Link] file without deploying any
models.

 The .save() method deploys a model artifact to the model catalog.


 The .deploy() method deploys a model to a REST endpoint.

 You have to serialize your method in the GenericModel class.

🔧 Ways to Save and Manage Models

You can use:


1. ADS SDK (in Python)

2. OCI Python SDK


3. OCI Console (UI)

Most data scientists prefer ADS SDK, as it automates the artifact generation and
introspection.

🧱 Model Management Operations

After saving to the Model Catalog, you can:


Operation Description

View / Edit Change metadata (name, description, tags, taxonomy).

Move models between compartments (e.g., from “Development” to


Move
“Production”).

Activate / Toggle model usability in deployments. Inactive models cannot be


Deactivate deployed but remain available.

Delete Permanently remove models (soft-deleted for 30 days).

Tagging Apply defined or free-form tags for organization.

🕒 Deleted models stay in the list for 30 days and can be filtered via the state filter.

📋 Metadata & Provenance Views in OCI Console

 Provenance View: Shows training source, notebook session, and Git details.

 Taxonomy View: Displays description, algorithm, framework, and custom


metadata.
 Schema View: Displays input/output schema definitions (read-only).

 Introspection Tests: Lists test results (Success, Failed, Not Tested).

✅ Always ensure all introspection tests pass before saving a model.

🧱 CLI / SDK Capabilities

Using CLI or Python SDK, you can:


 Create

 Update
 List

 Delete

…models within the Model Catalog programmatically.

🧱 Key Takeaways
✅ Serialization enables storing and reusing ML models
✅ OCI Model Catalog keeps artifacts immutable and versioned
✅ Supports both framework-specific and generic models
✅ Enables full lifecycle management — save → validate → deploy
✅ Integrates seamlessly with ADS SDK, OCI Console, and CLI

18. Model Deployment

🧱 OCI Data Science – Model Deployment

👨🏫 Instructor

Himanshu Raj – Senior Training Lead, AIML at Oracle

🔹 Overview

After training and evaluating models, the best candidates are stored in the Model
Catalog.
Model Deployment enables you to serve predictions using those models — either for
batch or real-time consumption.

⚙️ Model Deployment Flow

1. Client Application sends API calls to a deployed model endpoint.

2. The deployed model is hosted in OCI Data Science.


3. The service manages compute, environment, and scaling automatically.
Two types of prediction consumption:

 🕒 Batch consumption: Scheduled (hourly, daily, etc.)

 ⚡ Real-time consumption: Triggered instantly (e.g., fraud detection)

🧱 Model Deployment Architecture


Component Description

Load Balancer Distributes traffic across multiple model servers.

VM Instance Pool Hosts model server, conda env, and model artifact.

Model Artifact Contains the model file and prediction code ([Link]).

Conda Environment Includes all third-party dependencies (e.g., NumPy, XGBoost).

Logs Emit logs to OCI Logging for monitoring & debugging.

🛠️ Creating a Model Deployment

1⃣ From Console

Steps:
1. Enter deployment name

2. Select model from catalog

3. Choose compute shape and instances

4. Configure logging service

5. Set load balancer bandwidth

💡 Bandwidth Tip:
If payload = 1024 KB, requests = 120/sec
→ Bandwidth = 1024 × 120 × 8 / 1024 × 1.2 = 1152 Mbps

2⃣ From ADS SDK

[Link](deployment_properties)

or define properties directly using the .deploy() method.

3⃣ From OCI CLI

oci data-science model-deployment create --config-file [Link]

Optionally include log_config.json for access & prediction logs.


🚀 Invoking Model Deployments

 Send HTTP requests with feature vectors → get predictions in response.

 Can use:
o OCI CLI

o OCI Python SDK

o OCI Java SDK

 Payload limit: 10 MB

 Timeout: 60 seconds

 If latency critical → use streaming inference

 Must use Base64 encoding

🔧 Managing Model Deployments

From the OCI Console or via SDK/CLI, you can:

Operation Description

View/Edit Check OCID, compute, logs, etc.

Invoke Call the predict endpoint.

Update Change model, name, VM shape, or instances.

Deactivate / Reactivate Stop or restart deployments.

Delete Remove deployment (metadata preserved until deletion).

💡 When inactive, all compute billing stops but metadata is retained.

📊 Monitoring Model Deployments

🔸 Using OCI Logging

 Access Logs: Capture all HTTP requests.

 Predict Logs: Capture logs from [Link].


🔸 Using OCI Monitoring

 Built-in metrics:

o CPU utilization

o Memory utilization

o Network utilization

o Request count

o Latency

o Bandwidth

 You can:
o View metrics in Metrics Explorer

o Create Alarms for thresholds

✅ Summary

 Model Deployment allows real-time or batch inference via HTTP endpoints.

 Components: Load Balancer, VM Pool, Model Artifact, Conda Env, Logs.


 Create via Console, ADS SDK, or CLI.

 Monitor via Logging and Monitoring services.

 Deactivation halts billing, reactivation restores endpoint access.

19. Demo: Model Deployment

🧱 OCI Data Science – Demo: Model Deployment

👨🏫 Instructor

Himanshu Raj – Senior Training Lead, AIML at Oracle

🎯 Objective
Hands-on demo showing how to create, configure, and manage a model deployment
in the OCI Data Science project environment.

🧱 Steps to Create a Model Deployment

1⃣ Navigate to Model Deployments

 In your OCI Data Science Project, click on your project (e.g., test-ds).

 On the left sidebar, select Model Deployments.

 Click on Create Model Deployment.

2⃣ Basic Setup

 Compartment: Ensure you’re in the correct compartment (e.g., OCI Data


Science Compartment).
 Name: Enter a unique name (up to 255 characters).

o If not provided, OCI generates one automatically.

o Example: test-model-deploy
 Description: Optional. Example: This is to test model deployment.

3⃣ Select Model

 Choose an active model from the Model Catalog.

o Example: RF Classifier
 Click Select → Submit.

4⃣ Configure Compute

 Compute Shape: Choose the compute configuration for deployment.

o Example: 1 OCPU, 15 GB memory


 Number of Instances:

o Determines scalability.
o Example: 2 instances → handles more concurrent requests by distributing
load.

5⃣ Enable Logging (Optional)

 Click Select under Logging.

 Two types of logs:


o Access Logs: Capture request details.

o Predict Logs: Capture stdout and stderr outputs from prediction code.

 Select log names → click Next / Summary.

6⃣ Advanced Options – Load Balancing

 Define load balancer bandwidth (Mbps).

 Formula:

 Bandwidth = (Payload Size in KB × Requests/sec × 8 / 1024) × 1.2

 Example:

o Payload = 1024 KB
o Requests = 120/sec

o → Bandwidth = 1152 Mbps


 For demo: used 10 Mbps.

7⃣ Create Deployment

 Click Create.

 Wait for the deployment to initialize.


 Status changes to Active once ready.

🔍 Post-Deployment Details

📄 General Information
 Displays:

o Deployment name, OCID, description, compartment.


o Associated model, owner, and tags.

📊 Metrics Dashboard

 Shows:
o Success Rate

o Request Count

o CPU Utilization

o Memory Utilization

o Network Utilization

🧱 Logs & Work Requests

 Two log categories available:


o Predict Logs

o Access Logs

 Work Requests: Track creation status (e.g., Succeeded – 100% Complete).

🚀 Invoking the Model Deployment

Once active, you can invoke the model endpoint via:

Interface Command/Usage

HTTP Endpoint Use provided link for REST API call.

OCI CLI Use sample CLI command shown in console.

Python SDK Invoke through OCI Python SDK scripts.

Java SDK Use Java SDK to call endpoint.


🔧 Managing the Deployment

You can:
 Deactivate / Reactivate deployment (pauses or resumes instances).

 Delete deployment when no longer needed.

o Frees up resources and stops billing.

✅ Summary

 Demonstrated end-to-end model deployment process in OCI.

 Covered setup, configuration, compute selection, logging, and load balancing.


 Showed how to monitor metrics, view logs, and invoke predictions.

 Illustrated deactivation/reactivation for efficient resource management.

20. Demo: Model Deployment using Tensor Flow

🧱 OCI Data Science – Demo: Model Deployment using TensorFlow Model Class

👨🏫 Instructor

Himanshu Raj – Senior Training Lead, AIML at Oracle

🎯 Objective

Demonstrate end-to-end model deployment in Oracle Cloud Infrastructure (OCI)


Data Science using the Accelerated Data Science (ADS) library and TensorFlow
model class.

🧱 Overview

ADS provides framework-specific model classes (TensorFlow, PyTorch, Scikit-learn,


etc.) that help you register, prepare, verify, and deploy models into OCI Data Science
with minimal code.
⚙️ Step 1: Setup & Authentication

🧱 Imported Libraries

 ads → main interface to OCI Data Science

 logging → manage log output

 os → interact with OS paths


 pandas → tabular data handling

 tempfile → manage temporary directories

 tensorflow → build & train ML models

 tensorflow_datasets → load built-in datasets (e.g., Fashion-MNIST)

 warnings → suppress warnings

🔐 Authentication

Configured using Resource Principal authentication (recommended for OCI notebook


sessions).

🧱 Step 2: Dataset – Fashion-MNIST

 Training set: 60,000 images

 Test set: 10,000 images


 Each image: 28×28 grayscale, labeled with 10 classes (e.g., shirts, shoes).

 Visualized dataset using matplotlib.

🏗️ Step 3: Build and Train TensorFlow Model

Model Architecture

Layer Type Description

1 Flatten Converts 2D image into 1D vector

2 Dense(128, ReLU) Fully-connected layer


Layer Type Description

3 Dropout Prevents overfitting

4 Dense(10, Softmax) Output layer – 10 classes

Training Details
 Optimizer: Adam

 Loss: Sparse Categorical Cross-Entropy

 Metric: Accuracy

 Dataset scaled to [0, 1]

 Used first 10,000 samples to reduce compute time

 Example result: loss = 0.7899, accuracy = 0.7235

🧱 Step 4: Create Model Serialization Object

 TensorFlowModel() constructor wraps the trained model and creates an ADS


model object.

 This object provides helper methods to:


o Prepare artifacts

o Verify model

o Save to catalog

o Deploy

o Predict

from [Link].tensorflow_model import TensorFlowModel

tf_model = TensorFlowModel(estimator=model, artifact_dir="/tmp/artifacts")

🧱 Step 5: Prepare Model Artifacts

Files Generated by .prepare()


File Description

input_schema.json Defines input feature types and structure

model.h5 Serialized TensorFlow model

output_schema.json Defines output format

[Link] Runtime environment & conda setup

[Link] Contains load_model() and predict() functions

 Default model format: HDF5 (.h5)

 The .prepare() step also captures metadata like code provenance and model
parameters.

🧱 Step 6: Review Metadata

After preparation, the TensorFlow model object includes:


 runtime → environment name, conda pack, Python version
 model_provenance → training data and source code info

 _input / _output → schema details

 metadata_custom / metadata_taxonomy → key-value metadata for


classification, framework, and use case

🧱 Step 7: Verify the Model

 Use .verify() method to test [Link] without deploying.

 Ensures:

o load_model() loads correctly

o predict() works as expected

tf_model.verify()

✅ Verification successful → speeds up debugging and avoids failed deployments.


🗂️ Step 8: Save Model to Model Catalog

 Use .save() method to register model in OCI Model Catalog.

 Returns Model OCID.

 The model appears in console under Models → Active Status.

🚀 Step 9: Deploy Model

 Use .deploy() to create model deployment.

 You can specify:


o display_name

o description

o instance_type

o instance_count

o bandwidth

o logging groups

deployment = tf_model.deploy(display_name="demo-tf-model")

 Progress tracked using .summary_status()


 Once Active, the model becomes available as HTTPS endpoint.

🤖 Step 10: Invoke Predictions

Two modes:
1. Local – via model’s .predict() method before deployment

2. Deployed – via same .predict() method, which now sends requests to the live
endpoint

🧱 Step 11: Cleanup

Always remove resources after testing:

1. Delete deployment using .delete_deployment()


2. Delete model from catalog

3. Delete local artifact directory

This ensures no unnecessary billing and keeps workspace clean.

✅ Summary

Step Action Outcome

1 Import & Authenticate Setup OCI + ADS

2 Load Dataset Fashion-MNIST

3 Train Model TensorFlow Sequential

4 Create Model Object ADS TensorFlowModel

5 Prepare Artifacts 5 core files generated

6 Verify Local functional test

7 Save Register to Model Catalog

8 Deploy Create HTTPS endpoint

9 Predict Serve predictions

10 Cleanup Delete deployment + model

🧱 Key Takeaways

 ADS TensorFlowModel class abstracts all deployment complexity.

 Verify before deploy to save time and cost.

 Model Catalog acts as the central repository for reproducibility.

 Deployment cleanup is mandatory to avoid billing.

21. LLM Training & LangChain Integration


🧱 Lesson: Large Language Model (LLM) Training & LangChain Integration

Course: OCI Data Science


Topic: Training and integrating large language models using OCI Data Science Jobs
and ADS

🔹 Overview

 OCI Data Science Jobs provides fully managed infrastructure for training large
language models (LLMs) at scale.

 It supports both:
o Full-parameter fine-tuning

o Parameter-efficient fine-tuning

 Using the Accelerated Data Science (ADS) library, you can start training jobs
directly from GitHub repositories — without modifying the source code.

🔹 Steps for Fine-Tuning a Model

1. Access Pre-trained Model

o Obtain the model from Meta or Hugging Face.

2. Define the Training Job

o Use the ADS Python API to define the training configuration.

3. Create and Start Job Run

o Launch the job run via API.


o Stream the job run outputs in real-time.
4. Job Run Workflow

o Sets up the Conda environment and installs dependencies.

o Fetches source code from GitHub and checks out the specified commit.

o Runs the training script with defined arguments.

o Downloads model and dataset automatically.


o Saves outputs and checkpoints to OCI Object Storage when training
completes.

🔹 Infrastructure Handling

 No need to manually define:

o Number of nodes

o Number of GPUs
 ADS automatically configures compute resources based on:

o Replica count
o VM shape specified

🔹 Post-training Output

 Fine-tuning results (checkpoints) are saved to your OCI Object Storage


bucket.

 These outputs can be used for deployment or further evaluation.

🔹 Integration with LangChain

 OCI Generative AI Service supports:

o Text generation
o Summarization

o Embedding models
 These models can be integrated with LangChain through ADS.

Authentication

 By default, ADS uses the authentication method configured with:

 ads.set_auth()

 Optionally, you can specify authentication explicitly using the auth keyword (e.g.,
resource principal).
🧱 Key Takeaways

 OCI Data Science + ADS = streamlined, scalable, end-to-end LLM fine-tuning


and integration workflow.

 Fine-tune pre-trained models without modifying code.

 Automatic infrastructure management.


 Seamless integration with LangChain and OCI Generative AI for downstream
NLP tasks.

22. Demo: Deploy LangChain based RAG to OCI


Data Science

🤖 Lesson: Deploying LangChain-based RAG to OCI Data Science

Course: OCI Data Science


Topic: Deploying Retrieval-Augmented Generation (RAG) Applications

🔹 Overview

This demo demonstrates how to deploy a LangChain-based Retrieval-Augmented


Generation (RAG) application to OCI Data Science using the ADS library and OCI
Generative AI models.

The process covers:

 Building embeddings and retrievers

 Creating a LangChain retrieval QA pipeline

 Preparing and deploying the model as an OCI Data Science model

🔹 Steps in the Demo

1. Import Dependencies

 Import necessary classes from ADS and LangChain libraries.

2. Authenticate
 Use Resource Principal authentication for secure, OCI-native access.

 from [Link] import AuthType

 auth = AuthType.RESOURCE_PRINCIPAL

3. Create Embeddings and Model

 Use:
o GenerativeAIEmbeddings class to create vector embeddings

o GenerativeAI class to create the LLM model

 from [Link] import GenerativeAIEmbeddings, GenerativeAI

 embedding = GenerativeAIEmbeddings()

 llm = GenerativeAI()

4. Load and Process Documents

 Create a text loader to load documents.

 Split the document into manageable chunks for retrieval.

 from langchain.document_loaders import TextLoader

 loader = TextLoader("docs/ai_foundations.txt")

 documents = loader.load_and_split()

5. Create Vector Store

 Build a vector database using the previously generated embeddings and


documents.

 from [Link] import FAISS

 vector_store = FAISS.from_documents(documents, embedding)

6. Create Retriever and Chain

 Initialize a retriever from the vector store.


 Create a Retrieval QA Chain using:

o The retriever

o The LLM model

 from [Link] import RetrievalQA

 retriever = vector_store.as_retriever()

 chain = RetrievalQA.from_chain_type(llm=llm, retriever=retriever)

7. Prepare Model for Deployment

 Create a temporary directory for storing model artifacts.


 Use the ChainDeployment class to package the LangChain model.

 Prepare artifacts using:

 chain_deployment.prepare()

This generates files like:

o [Link]

o [Link]

o input_schema.json

o output_schema.json

8. Verify Model

 Test the prepared model locally before deployment.

 [Link](prompt="What is AI Foundations course?")

 ✅ The response is successfully generated → model verified.

9. Save and Deploy

 Save model to the OCI Model Catalog.

 Deploy the model using:

 [Link](display_name="LangChain_RAG_Model")
 Deployment typically takes 10–15 minutes.

 Once active, the model is accessible via HTTPS endpoint or SDK calls.

10. Test the Deployed Model

 Use the predict() method to query the deployed model.

 response = [Link]("Who are instructors for AI Foundation?")

 print(response)

 ✅ Output:

 Instructors for the Oracle Cloud Infrastructure AI Foundations course

 are Hemant Gahankari, Himansha Raj, and Nick Commisso.

🔹 Outcome

 Successfully deployed a LangChain RAG model to OCI Data Science.

 Verified real-time question-answering capability using retrieval-based


contextual data.

 Model integrates seamlessly with other LLM applications through OCI


endpoints.

🧱 Key Takeaways

 OCI Data Science + LangChain enables production-grade deployment of RAG


apps.
 Use ADS model classes for easy preparation, verification, saving, and
deployment.
 Embedding-based retrieval enhances accuracy and context relevance in
responses.
 End-to-end deployment requires minimal configuration thanks to ADS
automation.
23. Demo: OCI Data Science Operators

🧱 Lesson Summary — OCI Data Science Operators

Concept:
OCI Data Science Operators are low-code, prebuilt AI/ML tools that simplify tasks
like:

 ⏳ Forecasting — The Forecasting Operator leverages historical time series


data to generate accurate forecasts for future trends.(e.g., weather forecasts)

 ⚠️ Anomaly Detection — The Anomaly Detection Operator is a low-code tool


for integrating Anomaly Detection into any enterprise application.(Credit card
fraud)

 🔒 PII Detection — The PII Operator aims to detect and redact Personally
Identifiable Information (PII )in datasets by combining pattern match and
machine learning solution.
Features:

 Run via CLI or notebook session

 Executed inside Conda environments (preconfigured)

 Each operator comes with YAML templates for configuration


Process shown in demo:

1. Install operator Conda environment (e.g., Forecasting)


2. Initialize operator:

3. ads opctl init -t forecasting -o my-forecast


4. Configure [Link]

5. Activate environment:

6. conda activate forecast


7. Run operator:

8. ads opctl run -f my-forecast/[Link]


9. Check results:
Generated in /results → contains [Link], [Link], etc.
🧱 Lab / GitHub Resources

You can find all official OCI Data Science Operator labs and YAML examples on
GitHub here:

🔗 OCI Data Science Samples Repository:


👉 [Link]

Inside that repo, look for folders like:

 /operators/forecasting/

 /operators/anomaly-detection/

 /operators/pii/

Each directory contains:

 Example notebooks (.ipynb)

 YAML configuration files

 Quick start command examples (exactly like the demo)

24. Demo: OCI AI Quick Actions

🚀 Lesson Summary — Demo: OCI AI Quick Actions

🧱 What It Is

AI Quick Actions is a new low-code interface in OCI Data Science that lets you:

 Quickly deploy pre-trained Large Language Models (LLMs)


 Fine-tune them with your own dataset

 Evaluate model performance — all from the OCI Data Science Notebook UI

It’s designed for developers who want to use models like Mistral 7B, Meta LLaMA, or
Cohere models without deep infrastructure setup.

🧱 Demo Workflow
1. Open Notebook Session

o In OCI Data Science, launch your Notebook Session (JupyterLab


interface).
2. Click “AI Quick Actions”

o Found on the top toolbar of your notebook.


o Opens the AI Quick Actions window.

3. Tabs Overview

o Models Tab: Shows available out-of-the-box LLMs (e.g., Mistral 7B


Instruct).
o Deployments Tab: Manage your deployed models.

o Evaluations Tab: Assess models using custom datasets.

4. Deploy an LLM

o Click “Create Deployment”

o Choose:
 Model: e.g. Mistral 7B Instruct v0.02

 Compute Shape: [Link].8NR.1

 Log Group (optional)

o Then click Deploy

5. Use & Test

o Once active:
 You’ll get a REST endpoint

 You can test prompts directly in the notebook (e.g., “Tell us about
Las Vegas”)
 Adjust generation parameters (temperature, max tokens, etc.)

🧱 Related Labs & Resources

You can find OCI AI Quick Actions examples and documentation here:
🔗 OCI Data Science AI Quick Actions Labs / Samples
👉 [Link]
actions

Inside this directory, you’ll find:

 🧱 Example notebooks (e.g., deploy_llm_quick_actions.ipynb)

 ⚙️ Steps to configure policies for Quick Actions

 💡 Example of testing the model endpoint with Python

🧱 Official Documentation

For setup (policies, permissions, prerequisites):


📘 OCI Data Science — AI Quick Actions Documentation

5. MLOps Practices
1. MLOps Architecture

🧱 Lesson: MLOps Architecture

Instructor: Lyudmil Pelov — Product Manager, Oracle Cloud Infrastructure (OCI) Data
Science & AI Services
Course: OCI Data Science Professional
Topic: MLOps (Machine Learning Operations) on Oracle Cloud

🔍 What is MLOps?

MLOps (Machine Learning Operations) applies DevOps principles to machine


learning systems.
It standardizes, automates, and governs the entire ML lifecycle, from data preparation
to deployment and retraining.
Goal:
To make ML development and deployment more efficient, consistent, and scalable —
just like how DevOps transformed software delivery.

⚙️ Core Concepts

DevOps Practice MLOps Equivalent Description

Continuous Integration Model & Data Incorporates new datasets and model
(CI) Validation versions

Continuous Model Release to Automates safe rollout of trained


Deployment (CD) Production models

Continuous Training Retrains models automatically when


Unique to MLOps
(CT) new data arrives

🔄 Why Continuous Training Matters

Unlike traditional software, data constantly changes.


This causes data drift — when the statistical properties of input data evolve, reducing
model accuracy over time.
To prevent drift degradation, models must be:

 Continuously monitored,
 Retrained on fresh data,

 Validated before redeployment.

🧱 Maturity Levels in MLOps Automation

Level Description Example

1⃣ Manual Jupyter-based experimentation; data OCI Data Science


(Experimental) prep, & training done manually Notebook

2⃣ Automated ML Automated training/validation ADS pipelines or OCI


Pipeline triggered by new data DevOps triggers
Level Description Example

OCI DevOps + Model


3⃣ Full CI/CD Fully automated model retraining,
Catalog + Deployment
Pipeline testing, and redeployment
Pipelines

🏗️ MLOps Architecture in OCI

1. Data Ingestion

o Business or sensor data flows into OCI Object Storage or data lake.
2. Model Development

o Data scientists use OCI Data Science Notebook sessions for


exploration and training.
3. CI/CD Integration

o New data or notebook changes trigger OCI DevOps Build Pipelines.

4. Model Catalog

o The trained model is versioned and stored in the OCI Model Catalog.

5. Deployment Pipeline

o The OCI DevOps Deploy Pipeline takes the validated model, deploys it
to an endpoint for testing.

6. Testing & Approval


o The model is validated internally; if approved, it’s promoted to
production.

7. Monitoring

o OCI monitors model metrics and production data.


If performance drops, it triggers retraining — closing the continuous
learning loop.

🧱 End-to-End Flow Summary

Data → Notebook → CI/CD Pipeline → Model Catalog → Test Deployment → Approval


→ Production → Monitoring → Retrain
🧱 Key OCI Services Used

 OCI Data Science – Model development (Jupyter, ADS SDK)

 OCI DevOps – CI/CD pipelines for model automation

 OCI Model Catalog – Versioning and governance

 OCI Monitoring / Logging – Performance tracking

 OCI Object Storage – Data and model artifact storage

2. Data Science Jobs

🧱 Lesson: Data Science Jobs

Instructor: Lyudmil Pelov — Product Manager, Oracle Data Science & AI Services
Course: Oracle Cloud Infrastructure (OCI) Data Science Professional
Topic: Jobs Service — Automating and Scaling MLOps Tasks

⚙️ 1. What Are OCI Data Science Jobs?

The OCI Data Science Jobs service allows you to run repeatable, automated
machine learning and data tasks on fully managed infrastructure — only when
needed.
It is a key MLOps enabler because it:

 Automates parts of the ML lifecycle (data prep, model training, evaluation,


inference).
 Reduces cost by provisioning compute only for the duration of the job.

 Eliminates manual setup and maintenance of batch processing systems.

🧱 2. Key Concepts
Concept Description

Template defining what to run — includes the code (artifact), compute shape,
Job
environment, and configurations.

A specific execution instance of a Job. You can override parameters or


Job Run
environment variables for each run.

Each Job can have multiple runs, e.g., to test different model hyperparameters or
input data.

🧱 3. Components of a Job

1. Job Artifact – The code or script or instruction to execute (Python, Bash,


ZIP/TAR file).

o Immutable once uploaded.

o Required

o Defines the entry point for execution (JOB_RUN_ENTRYPOINT).


2. Compute Shape – Determines CPU/GPU and memory resources.

o Can be edited between runs.


3. Environment Variables & CLI Arguments –Parameters that can be customize
each job run.
4. Logging & Storage – Define log groups and block storage size.

5. VCN (Virtual Cloud Network) – Optional; enables access to internal or secure


resources.
6. Max Runtime – Up to 30 days per job run.

🚀 4. The Job Lifecycle

A job run is the actual processor that executes the instructions in the artifact and follows
the parameters set in a job and a job run.

Create Job → Configure Artifact & Compute → Start Job Run → Monitor & Log →
Complete or Cancel → Deprovision

 Create Job :
o Name
o Artifact
o Environment Variables
o Command Line Arguments
o Compute : CPU, GPU
o Logging
o Block Storage
o Max Run Time
o VCN
 Run Job :
o Name
o Environment Variables
o Command line Arguments
o Logging
o Max run time
 Moniter + Log :
o Compute Metrics
o Service Logging
o Custome Logging
 End :
o Finish/Cancel
o Deprovisioning
o Events

🧱 5. Supported Artifact Types

Artifact Type Description Use Case

Python / Bash
Single-file artifact Quick one-off tasks
Script

Full project, can include YAML runtime


ZIP / TAR Archive Complex ML pipelines
configuration

Defines environment, dependencies, and Reproducibility and


Runtime YAML
variables control
OCI provides Python preinstalled and supports Conda environments for dependency
isolation.

🔐 6. Integration and Access

 Jobs can securely access OCI resources (Object Storage, ADW, Databases).

 Jobs can also integrate with on-prem or third-party systems via VCN or OCI
Vault credentials.

 Supported interfaces include:


o OCI Console (UI)

o OCI CLI

o SDKs (Python, Java, Go, Ruby, JavaScript)

o Terraform

o CI/CD pipelines (e.g., Bitbucket, GitHub, Jenkins)

⚡ 7. Batch Inference Modes

Type Description Example Frequency

Processes full dataset


Regular Batch Daily model scoring Moderate
periodically

Processes smaller data Fraud detection every


Mini Batch High
slices more frequently few minutes

Distributed Splits massive datasets into Parallel model Long-running,


Batch parallel jobs training or analytics high-scale

Each approach balances speed, resource usage, and data volume.

📈 8. Scaling Resources

 You can scale up or down:

o Compute shapes (CPU/GPU cores, memory)

o Block storage size


 Scaling is available for Jobs and Notebook Sessions within OCI Data Science.

🧱 9. MLOps Use Cases for Jobs

✅ Data preprocessing pipelines


✅ Automated model training and evaluation
✅ Batch predictions (inference)
✅ Data validation or transformation
✅ Periodic retraining or scoring workflows

🧱 10. Summary

OCI Data Science Jobs provides:

 Fully managed, on-demand compute for ML workflows.

 A repeatable, secure, and cost-optimized way to automate tasks.

 Integration with OCI and third-party ecosystems.


 Flexibility for regular, mini, and distributed batch pipelines.

In essence:

“Jobs automate the heavy lifting of MLOps — from one-off scripts to full-scale pipelines
— while you pay only for what you use.”

3. Demo: Create Artifacts

🧱 Concept Recap: What’s a Job Artifact?

A job artifact in OCI Data Science is basically a Python script (or a zip file of a Python
project) that contains the code to execute inside an OCI Job.
It defines what happens when your job runs — e.g. data processing, model training,
batch inference, etc.

⚙️ Example: Simple Python Job Artifact


You can create a file named job_artifact.py:

from datatime import datetime


import os
import argparse

NAME = [Link]("NAME", "UNDIFINED")

parser = [Link]()
parser.add_argument("-g","--greeting", required=False, default="Hello")
args = parser.parse_args()

print(f'Job Run {[Link]().strftime("%Y-%m-%d %H:%M:%S")}')


print(f"{[Link]}, Your Environment Variable has value of : {NAME}")
print("Job Done.")

💻 Optional: Add OCI SDK for Resource Principal Authentication

If you want to use OCI SDK (for example, to access Object Storage), you can extend it
like this:

import argparse
import oci
import os

# Resource Principal
TENACY_OCID, dsc = None, None
COMPARTMENT_OCID = [Link]("PROJECT_COMPARTMENT_OCID","UNDEFINED")
RP = [Link]("OCI_RESOURCE_PRINCIPAL_VERSION","UNDEFINED")

if not RP == "UNDEFINED":
# LOCAL RUN
config = [Link].from_file("~/.oci/config","BIGDATA")
dsc = oci.data_science.DataScienceClient(config=config)
TENACY_OCID = config["tenancy"]
else:
# JOB RUN
singer = [Link].get_resource_principals_signer()
dsc = oci.data_science.DataScienceClient(config={}, signer=singer)
TENACY_OCID = singer.tenancy_id
# You 2 weeks ago * - adding base job wiht RP example

# Command Line Arguments


parser = [Link]()
parser.add_argument("-g", "--greeting", required=False, default="Hello")
args = parser.parse_args()

# print
print(
f'{[Link]} {[Link]("Name","Unknown")} in tenancy OCID
{TENACY_OCID}!'
)

# OCI SDK Clinet


if not COMPARTMENT_OCID or COMPARTMENT_OCID == "UNDEFINED":
shapes = dsc.list_job_shapes(compartment_id=COMPARTMENT_OCID)
print([Link][0])
else:
print("No PROJECT_COMPARTMENT_OCID set!")
print("Job Done")

🧱 Step-by-Step Lab (to simulate the demo)

1. Create the file locally

2. nano job_artifact.py

Paste the code above.


3. Test it locally

4. export NAME="Haroon"

5. python job_artifact.py -g "Welcome"

✅ Output:

Job started at: 2025-10-06 10:05:33

Welcome, Haroon!

Job finished at: 2025-10-06 10:05:34


6. Upload as a Job Artifact in OCI Console

o Go to OCI → Data Science → Jobs → Create Job

o Under Job Artifact, upload your job_artifact.py

o Set Environment Variable → NAME=Haroon

o Add Command-line argument → -g Welcome


o Choose a compute shape
o Click Run

4. Demo: Create and Manage Jobs

🧱 LAB: Create and Manage OCI Data Science Jobs

🧱 Objective

You’ll learn how to create a Data Science Job, upload your Python artifact, and
configure compute, logging, and networking.

🧱 Prerequisites

 Access to an Oracle Cloud tenancy

 A Data Science Project already created


 A Python artifact file (like job_artifact.py from the previous lab)

🧱 Step-by-Step Instructions

1⃣ Navigate to the Data Science Service

 Log in to the OCI Console

 Click the ☰ (hamburger menu) → Analytics & AI → Machine Learning →


Data Science

2⃣ Create a Project (if not created yet)

 Click “Create Project”

 Enter a name and description (e.g. DataScience_Jobs_Demo)


 Click Create

3⃣ Create a Job
 In your project dashboard, select Jobs in the left sidebar

 Click Create Job

Fill out the job details:

Setting Description

Name Optional — e.g. GreetingJob

Description e.g. “Simple Python job artifact demo”

Compartment Leave default or choose one

Artifact Upload your Python file job_artifact.py

Max upload size 100 MB via Console (use SDK for larger uploads)

4⃣ Configure Environment and Arguments

Under Environment Variables, add:

NAME = Haroon
Under Command-line Arguments, add:

-g Hey

This matches your artifact’s parameters (--greeting or -g and environment variable


NAME).

5⃣ Set Runtime and Compute

 Max runtime (minutes): 100

 Compute Shape: Choose one of:

o Fast Launch (pre-warmed): quicker startup

o Custom Configuration: allows GPU/Intel shapes (e.g., VM.GPU3.1)

6⃣ Enable Logging (Recommended)

 Under Logging, click Select


 Choose or create a Log Group

 Select “Automatically create log for every job run”

 Click Select

Logs will go to OCI Logging Service.

7⃣ Configure Storage

 Select Block Storage Size (e.g. 50 GB)

 This storage is automatically attached to your job run — adjust size if processing
large datasets.

8⃣ Configure Networking

 Choose Default Network if you don’t have custom VCN requirements.


(This provides basic internet and OCI service access.)
 Advanced users can select Custom Network with their own VCN/Subnet.

9⃣ Create the Job

 Review your settings


 Click Create

✅ Result:
Your job will be created — but not yet running.
This is a template describing the infrastructure and artifact.

🔟 Manage Your Job

After creation, you can:


 View General Info, Job Artifact, Logging, Infrastructure Config

 Edit job name, description, compute shape, or storage

 Download the artifact

 Move the job to another compartment


 Add tags

 Delete the job


To run the job, you’ll create a Job Run — that’s the next step in the lesson series.

🧱 Example Summary Configuration

Setting Value

Job Name GreetingJob

Artifact job_artifact.py

Env Var NAME=Haroon

Argument -g Hey

Max Runtime 100 mins

Compute Shape VM.Standard2.1 (Fast Launch)

Logging Enabled (auto-create)

Storage 50 GB

Network Default Network

5. Demo: Start and Manage a Job Run

🧱 LAB: Start and Manage a Job Run

🧱 Objective

Learn how to start, monitor, clone, and cancel OCI Data Science Job Runs.

🧱 Prerequisites

 You have already created a Data Science Project


 You have a Job created with a valid Python artifact (from the “Create and
Manage Jobs” lab)

🧱 Step-by-Step Instructions

1⃣ Go to Your Job

 Open OCI Console → Data Science Service

 Navigate to your Project → Jobs

 Select the job you created earlier (e.g. GreetingJob)

2⃣ Start a Job Run

 Click Start Job Run

You’ll see a configuration screen before the run starts.

3⃣ Configure Job Run Options

You can override parameters from the original job if you want:

Setting Example

Logging Keep default or choose another log group

Environment Variable NAME = Haroon Khan

Command-line Argument -g Hello again!

Max Runtime 60 minutes

After reviewing, click Start.

4⃣ Watch the Job Run Lifecycle

The job will progress through several states:


State Description

Accepted OCI acknowledges your job submission

Provisioning Infrastructure (VM + storage) is being created

Running Your Python artifact is executing

Succeeded / Failed Job completed successfully or with an error

Cancelled Job manually stopped

⏳ The provisioning and running phases are billed — metering stops automatically
when the run ends or is canceled.

5⃣ Monitor the Run

 You can monitor in the Jobs → Job Runs table

 View columns such as:


o Status

o Lifecycle Detail

o Created By

o Start Time

If a run fails, click it → Logs → check OCI Logging for error output (e.g. Python
exception, missing variable, etc.)

6⃣ Run Multiple Jobs in Parallel

You can start another job run while one is executing — for example, to test different
model parameters or hyperparameters.
 Click Start Job Run again

 Provide new environment variables or arguments


(e.g. different greeting or learning rate for ML model)

✅ Multiple job runs can execute simultaneously.


7⃣ Clone a Job Run

If you want to re-run a previous job with the same parameters:


 Select the previous Job Run

 Click Clone

 Change only what’s needed (like updating the greeting or timestamp)


 Click Clone Job Run

This saves time by keeping all prior environment variables and arguments intact.

8⃣ Cancel a Running Job

If you made a mistake or need to stop execution:


 While the job is Provisioning or Running, click Cancel

 Confirm cancellation when prompted

Once canceled:

 Infrastructure is destroyed
 Billing stops immediately

9⃣ Delete or Rename Jobs

You can also:


 Edit a job’s name or tags

 Delete unused jobs (to keep your workspace clean)

 Download artifacts for local modification or debugging

🧱 Example Lifecycle Summary

Action Description

Start Job Run Launches new job execution

Clone Job Run Duplicates settings from past runs


Action Description

Cancel Job Run Stops infrastructure + billing

View Logs Debug failed or completed runs

Parallel Runs Run multiple jobs simultaneously

🧱 Pro Tip

Use job runs to automate batch tasks or model retraining by scheduling them
through the OCI CLI, Python SDK, or DevOps CI/CD pipelines.

6. Demo: Scaling

Demo: Scaling – OCI Data Science

Presenter: Lyudmil Pelov, Product Manager – Oracle Data Science & AI Services

Overview:
This demo explains how to scale up or down jobs and notebook sessions in Oracle
Cloud Data Science Service to optimize CPU, memory, and storage usage.

🔹 Scaling Jobs

 If a job shows high CPU, memory, or storage utilization, you can edit the job
to change its compute shape or storage size.

 Steps:
1. Select the job → Click Edit.

2. Under Change Shape, select a new shape:

 Fast launch shape

 Standard shape (e.g., VM.Standard2.1, VM.Standard3.4, etc.)


 GPU shapes for model training tasks.
3. Optionally, increase block storage size.
4. Click Save Changes and start a new job run — it will use the updated
configuration immediately.

🔹 Scaling Notebooks

 You can monitor metrics such as CPU and memory utilization for your notebook
instance.
 If performance is low, scale up your notebook shape.

Steps to scale a notebook:

1. Deactivate the notebook first.

o Click Deactivate → Confirm → Notebook stops billing but keeps block


storage data safe.
2. Once inactive, click Activate again.

o During activation, choose a new compute shape (e.g., Intel, Flex, or


GPU) and optionally increase storage.
3. Click Activate to restart the notebook with new resources.

Note:
Deactivating a notebook stops billing but preserves data in block storage. When
reactivated, all previous files remain intact.

✅ Key Benefits

 Dynamically adjust resources for performance optimization.


 Cost-efficient: Pay only for active sessions or job runtimes.

 Flexible compute options for both CPU- and GPU-intensive tasks.

 Safe scaling: Data persistence ensured across shape changes.

7. Jobs Monitoring and Logging


Demo / Module: Jobs Monitoring and Logging – OCI Data Science

Presenter: Lyudmil Pelov, Product Manager – Oracle Data Science and AI Services

📘 Overview

This module focuses on monitoring and logging in OCI Data Science Jobs — the final
step in the job lifecycle before the infrastructure is deprovisioned.
It explains how to track job performance, metrics, logs, and events for effective
troubleshooting and optimization.

🔹 Monitoring

Purpose:
Monitoring helps check the health, capacity, and performance of cloud resources in
real time.
Components:

1. Metrics – Continuously emitted data points that measure:

o CPU and GPU utilization

o Memory usage

o Network bytes in/out

o Disk utilization
2. Alarms – Passive monitoring that triggers when metrics cross thresholds (e.g.,
CPU > 80%).
o Sends notifications via Slack, SMS, or Email through OCI Notifications
Service.

Use Cases:

 Identify resource bottlenecks.

 Scale up compute or storage when workloads increase.

 Debug performance issues.

 Track and maintain system health.


🔹 Logging

Purpose: Capture and record information about job execution and artifacts for
debugging and auditing.

Types of Logs:

1. Service Logs

o Automatically emitted by job runs to the OCI Logging Service.

o Capture both standard output (stdout) and standard error (stderr)


streams.
o Requires job run’s resource principal to have logging permissions.

o Recommended to enable for all jobs.


2. Custom Logs

o Defined by the user for specific contexts or outputs.

o User specifies where logs are stored.

o Multiple job runs can share the same log or use individual logs.
Automatic Logging:

 You can enable automatic log creation, letting the Data Science service create
logs within your log group automatically.
 Even if a job or job run is deleted, logs persist and must be managed manually.

🔹 Event Service

Purpose:
Detect and respond to changes in resources (like job or job run lifecycle events).

How It Works:

 Events represent Create, Read, Update, Delete (CRUD) operations.

 Users can create rules to monitor specific events and trigger actions.

 Actions include:

o Notifications (email, Slack, etc.)

o Oracle Functions
o Streaming

 Multiple actions can be tied to a single rule.


 OCI guarantees at least one delivery for each action.

✅ Key Takeaways

 Monitoring and logging provide visibility, diagnostics, and automation for Data
Science jobs.
 Metrics & Alarms → Detect and respond to performance issues.

 Service & Custom Logs → Capture outputs for debugging and auditing.
 Event Service → Automate actions based on resource state changes.

 Logs and events remain accessible even after job deletion for analysis and
record-keeping.

8. Data Science Pipeline

Lesson: Data Science Pipeline

Instructor: Hemant Gahankari, Senior Principal Training Lead – Oracle University

📘 Overview

This lesson introduces Data Science Pipelines in Oracle Cloud Infrastructure (OCI)
Data Science Service — a powerful feature that allows users to build and automate
end-to-end machine learning workflows composed of multiple tasks (called steps).

Pipelines help orchestrate complex data science processes such as data


preprocessing, model training, evaluation, and deployment in a structured,
reusable, and automated manner.

🔹 What Is a Data Science Pipeline?


 A Pipeline is a workflow made up of one or more steps.

 Each step performs a specific task — e.g.,


o Step 1: Data preprocessing

o Step 2: Model training (could include multiple models)

o Step 3: Model evaluation

o Step 4: Model deployment

 Steps can run in sequence (with dependencies) or in parallel to improve


efficiency.
 Pipelines allow integration of different environments and programming
languages within one workflow (e.g., Python for preprocessing, Java for model
training).

🔹 Pipeline Configuration

Each Pipeline and Pipeline Step has its own configuration:


Pipeline-Level Configuration

 Compute Shape – Defines the processing power and memory.

 Block Storage – Determines storage capacity.

 Environment Variables – Used to pass data or parameters (e.g., dataset paths).

 Logging Settings – Enable or define log destinations.


 Maximum Runtime – Defines time limits for execution.

 Default configurations apply to all steps unless overridden.


Step-Level Configuration

 Can override pipeline-level defaults.

 Each step can be implemented as either:


1. A Script (Python, Bash, or Java file — single or zipped).

2. An OCI Data Science Job (referenced using its OCID).

 Steps can access OCI resources (Object Storage, Database, etc.) if proper IAM
policies and VCN configurations are set.
🔹 Pipeline Lifecycle

1. Creating – Pipeline is being set up.

2. Active – Pipeline is ready for execution.

3. Pipeline Run – Each execution instance of a pipeline; multiple runs can be


created.
4. Deletion – Pipeline can be deleted once no longer needed.

🔹 Demonstration Scenario

In the demo setup:


Step 1 – Data Preprocessing

 Reads dataset from OCI Object Storage.

 Performs operations like:

o Dropping unnecessary columns

o Label encoding

o Scaling features
o Splitting into train/test datasets

 Saves processed data back to Object Storage.

 Updates Pipeline Variables with file locations for downstream steps.


Step 2 – Model Training

 Trains multiple models using algorithms such as:


o Linear Regression

o Random Forest

o XGBoost

 Stores all trained models in the OCI Model Catalog.

Step 3 – Model Evaluation and Deployment

 Retrieves all trained models from the Model Catalog.


 Evaluates and selects the best-performing model.

 Deploys the best model for inference.

✅ Key Takeaways

 A Pipeline automates and connects the full machine learning workflow.

 Steps can be independent or dependent, and run sequentially or in parallel.

 Configuration flexibility allows different environments and compute settings per


step.
 Integration with Object Storage, Model Catalog, and Jobs enables seamless
ML automation.
 Pipelines are essential for scalability, reproducibility, and efficiency in modern
data science projects.

9. Demo: Data Science Pipeline

☁️ Demo: Data Science Pipelines — OCI Data Science

🎯 Objective

To understand how to create and run a Data Science Pipeline in OCI to automate an
end-to-end machine learning workflow — from data preprocessing to model
deployment.

🧱 What Is a Data Science Pipeline?

A pipeline is an automated workflow that chains multiple steps of a machine learning


lifecycle — data processing, model training, evaluation, and deployment — into a single
executable sequence.

🗂️ Setup

 Use the Oracle Samples GitHub repository →


🔗 oci-data-science-ai-samples
 Navigate to:
pipelines/samples/employee-attrition/

 The folder contains multiple .zip files — each representing a step in the pipeline:
o [Link] → Data preprocessing

o [Link] → Train Linear Regression model

o [Link] → Train Random Forest model

o [Link] → Train XGBoost model

o evaluate_deploy.zip → Evaluate models & deploy best one

🏗️ Creating the Pipeline

1. Create a Data Science Project

 In OCI Console → Data Science → Create new Project

 Use the Pipeline section inside your project.

2. Create the Pipeline

 Click Create Pipeline

 Give it a name: e.g., Employee_Attrition_Pipeline

 Add description (optional)

 Configure:
o Environment variable:
data_location = oci://<your-bucket-name>@namespace/pipeline-temp-
bucket
o Compute Shape: VM.Standard2.2

o Block Volume Size: 50 GB

o Logging Configuration: Select an existing log group

⚙️ Defining Pipeline Steps


Step Description Depends On Artifact Entry Point

Step 1 Data Preprocessing — [Link] [Link]

Step
Train Linear Regression Step 1 [Link] [Link]
2a

Step
Train Random Forest Step 1 [Link] [Link]
2b

Step
Train XGBoost Step 1 [Link] [Link]
2c

Evaluate & Deploy Best Steps 2a, 2b,


Step 3 evaluate_deploy.zip evaluate_deploy.py
Model 2c

 All steps are “Build by Script” type.

 Steps 2a, 2b, and 2c run in parallel (to train multiple models simultaneously).

▶️ Running the Pipeline

 Click Start Pipeline Run

 Name the run (e.g., Run-11)

 Add environment variable:


CONDA_ENV_SLUG → select default service environment

 Check logs under configured log group

 Observe status transitions:

o Waiting → Accepted → In Progress → Succeeded

📊 Results

 AUC Scores (example):

o Linear Regression → 0.85 ✅ (Best model)

o Random Forest → 0.81

o XGBoost → 0.837
 The best-performing model (Linear Regression) is automatically deployed.

 You can verify deployment in:


o Models tab → shows the selected model and deployment status

🧱 Key Takeaways

 OCI Data Science Pipelines automate model training, evaluation, and


deployment.

 You can reuse step artifacts and integrate with OCI Logging & Object Storage.

 Supports parallel model training and end-to-end reproducibility.

 Reduces manual effort and ensures consistency across ML workflows.

10. Model Deployment: Autoscaling

🎯 Objective

To understand Autoscaling in OCI Data Science model deployments — a feature


that automatically adjusts compute resources to balance performance, availability,
and cost-efficiency.

⚙️ What Is Autoscaling?

Autoscaling allows a deployed model to automatically scale up or down the number


of compute instances based on real-time demand.

It helps maintain optimal performance while minimizing unnecessary cost —


especially useful for unpredictable workloads.

🚀 Why Autoscaling Matters

When deploying models, choosing the right compute shape and instance count can
be difficult.
Autoscaling solves this by dynamically managing resources based on usage thresholds.
🧱 Key Benefits of Autoscaling

# Benefit Description

Dynamic Resource Automatically increases or decreases compute instances


1⃣
Adjustment based on demand.

2⃣ Cost Efficiency You only pay for resources actually used — no idle cost.

Works with load balancers to reroute traffic to healthy


3⃣ Enhanced Availability
instances if one fails.

Supports user-defined NQL expressions to control when


4⃣ Customizable Triggers
scaling happens.

Load Balancer Load balancer bandwidth scales automatically to match


5⃣
Compatibility traffic.

Prevents too-frequent scaling by pausing actions for a set


6⃣ Cooldown Periods
duration after a scale event.

🧱 How Autoscaling Works

🔹 Metric-Based Autoscaling

 The only supported autoscaling type in model deployments.

 Triggered when a metric (like CPU or memory usage) meets/exceeds a defined


threshold.

 Metrics are collected by the Monitoring Service and aggregated over time.

Example:

If CPU utilization > 80% for 3 consecutive intervals → scale up instance count.

⏱️ Autoscaling Event Flow

1. Metrics (e.g., CPU%) are monitored.

2. If threshold is exceeded for several intervals → scaling event triggered.


3. A cooldown period begins → no further scaling during this time.

4. When cooldown ends → system reevaluates metrics and adjusts if needed.

🧱 Autoscaling Policy Types

Type Description

Predefined Choose from built-in metrics like CPU Utilization or Memory


Metric Utilization.

Use Monitoring Query Language (NQL) expressions for advanced


Custom Metric
control — combine multiple metrics, aggregations, and logical
(NQL)
conditions (AND/OR).

🧱 Setting Up Autoscaling

You can configure autoscaling:


 During model deployment creation

 Or for an existing deployment

Methods:

 OCI Console
 OCI CLI

 OCI Data Science API

🧱 Required IAM Policy

Add this to your tenancy:

Allow service datascience to read metrics in tenancy

where [Link] = 'oci_datascience_modeldeploy'


This allows autoscaling to access model deployment metrics.

🧱 Custom Metric Example


You can create an NQL expression using any available model deployment metrics, such
as:

 CPUUtilization
 MemoryUtilization

 ResponseTime

 RequestsPerSecond

 StatusCode, etc.
Example Query:

CpuUtilization[1m].mean() > 75 or MemoryUtilization[1m].mean() > 80

This query scales up if either CPU > 75% or memory > 80% in the last 1 minute.

🔄 Deployment States and Updates

Deployment
Allowed Updates
State

Can modify autoscaling policy independently (not combined with


Active
other changes).

All options (including autoscaling configuration) can be modified


Inactive
together.

📊 Monitoring Metrics

Metrics for model deployments are automatically emitted under namespace:

oci_datascience_modeldeploy
Common Metric Dimensions:

 resourceId

 statusCode

 statusFamily

 instanceId
 result
 networkType

No need to manually enable monitoring — it’s automatically active for every


deployment.

💡 Summary

Feature Description

Purpose Automatically scale model deployment resources.

Scaling Type Metric-based (CPU/Memory or custom).

Key Benefit Performance stability with cost efficiency.

Integrations Load balancer + OCI Monitoring.

Control NQL expressions for custom triggers.

11. Expert Tips: Pipelines

☁️ Expert Tips: Pipelines

(OCI Data Science Course — by Hemant Gahankari, Senior Principal Training


Lead, Oracle University)

🎯 Objective

To understand the Pipelines feature in OCI Data Science Service and how it
automates end-to-end machine learning workflows.

⚙️ What Are Pipelines?

Pipelines are a recently introduced feature in OCI Data Science that allow you to
automate complete ML workflows.

A pipeline consists of multiple steps that can be executed either:

 Sequentially (one after another), or


 In parallel (simultaneously).

🧱 Typical Steps in a Pipeline

Examples of workflow stages include:

1. Data Extraction — retrieving data from sources.

2. Data Validation — ensuring data quality and correctness.

3. Data Preparation / Preprocessing — cleaning and transforming data.

4. Model Training — applying ML algorithms.

5. Model Evaluation & Deployment — assessing and deploying the trained


model.

🚀 Key Purpose

 Automate repetitive machine learning tasks.

 Ensure consistency across model training and evaluation runs.

 Support complex workflows with dependencies and multiple environments.

🌟 Instructor’s Recommendation

Hemant emphasizes exploring OCI Pipelines through:

 Documentation

 Hands-on experimentation

This will help users fully leverage the feature to streamline and scale their
machine learning projects.

🧱 Summary Table

Feature Description

Service Oracle Cloud Infrastructure (OCI) Data Science


Feature Description

Feature Name Pipelines

Function Automate and manage multi-step ML workflows

Execution Mode Sequential or Parallel

Benefits Automation, efficiency, scalability, and reproducibility

Recommended By Hemant Gahankari (Oracle University)

6. Related OCI Services


1. Spark Applications, Data Flow, and Data
Science

🧱 Demo Notes: Spark Applications, Data Flow, and Data Science

👨🏫 Instructor

Jean-Rene Gauthier, Product Manager – OCI Data Science

🚀 1. Introduction to OCI Data Flow

Oracle Cloud Infrastructure (OCI) Data Flow

 A fully managed, serverless Apache Spark service for big data processing
and machine learning at scale.

 Enables you to run Spark applications (PySpark, SQL, Java, Scala) without
provisioning or managing infrastructure.
Key use cases:

 Data aggregation & transformation


 Feature engineering

 Data cleaning
 Model training (via MLlib or other libraries)
Why Spark?

 Popular engine for scalable, distributed data processing.


 Used for AI/ML workloads due to its high-speed parallel processing.

⚙️ 2. Data Flow Core Components

Component Description

Library Central repository of all Spark applications in Data Flow.

Reusable Spark app template (code + dependencies + parameters +


Application
runtime).

Execution instance of a Data Flow application — includes output, logs,


Run
stats, etc.

Logs Automatically stored in OCI Object Storage for debugging and auditing.

🔑 3. Key Capabilities

✅ Connect to Spark data sources and launch jobs in seconds.


✅ Manage all Spark apps from one interface.
✅ Create reusable Spark applications in any Spark language.
✅ Secure, isolated environment (no cluster sharing).
✅ Data encrypted in transit and at rest.
✅ Integrated with OCI IAM for secure access.
✅ Supports Object Storage and other connectors.

🔒 4. Security Model

 Each run executes on behalf of the user who launched it.

 Uses IAM policies for authorization.

 Jobs run in isolated pools (no shared clusters).

 Logs → Object Storage (securely stored).


 Data encrypted end-to-end.

🧱 5. Spark + Data Flow for AI/ML Workloads

Spark provides:
 ETL (Extract, Transform, Load) for large datasets

 Feature engineering

 Scalable model training

Spark includes MLlib, offering:

 Common ML algorithms (classification, regression, clustering)

 Training on distributed dataframes


 Ability to integrate with PySpark or Spark SQL

⚙️ 6. Spark Architecture Overview

Component Description

Driver Central coordinator that controls execution.

Cluster Manager Manages cluster resources and launches applications.

Executors Distributed worker nodes running tasks.

SparkContext Entry point for Spark functions.

When configuring resources:

 Choose driver/executor shape and number of executors.

 The more data or shorter processing time → more OCPUs needed.

o Example: 500 GB in 10 hrs → ~5 executor OCPUs.

🧱 7. Integration with OCI Data Science

 Data Flow integrates with OCI Data Science Notebooks (JupyterLab interface).

 Uses Accelerated Data Science (ADS) library.


 You can:

o Submit Spark jobs from notebooks

o Fetch logs

o Manage Data Flow runs

o Sync PySpark scripts

o Add custom Python libraries

📋 8. Prerequisites for Using Data Flow with Data Science

You need:
1. A Data Science Project + Notebook Session

2. An Object Storage bucket for logs & data

3. A PySpark / Spark SQL / Java application uploaded

4. Correct IAM policies for access permissions

💡 9. Best Practices for PySpark Development

1. Use sample data during development (not full dataset).

2. Break code into cells for iterative testing.

3. Convert Spark DF → Pandas DF for visualization.


4. Use Matplotlib or similar for analysis.

5. Remove unnecessary cells before converting notebook → script:

6. jupyter nbconvert --to script [Link]


7. Test locally:

8. spark-submit [Link]

9. Before uploading:

o Replace sample data paths with full dataset.

o Remove external library calls (like pandas/sklearn).

o Or include them in Data Flow environment manually.


📘 10. Summary

 Data Flow = Serverless Spark for large-scale batch & ML jobs.

 No infrastructure management – on-demand compute.

 Ideal for ETL, feature engineering, model training.

 Integrated with OCI Data Science notebooks and ADS SDK.

 Secure, scalable, and cost-effective Spark job execution.

2. Oracle Open Data

🌍 Lesson Notes: Oracle Open Data

👨🏫 Instructor

Wes Prichard, Senior Principal Product Manager – Oracle Data Science & AI Services

🧱 1. Overview

Oracle Open Data is a public data repository by Oracle for Research that provides
free access to curated datasets for researchers, scientists, and developers worldwide.

 Goal: To make large, high-quality datasets easily available for research and
AI/ML applications.

 Access:
🔓 No login or payment required.
🌐 Website: [Link]

🗂️ 2. What It Offers

Oracle Open Data includes datasets from multiple scientific and technical domains:
Domain Example Data Sources / Types

🛰️ Geospatial
Satellite imagery from GOES, MODIS, Landsat
Data

🧱 Life Sciences Chemical compounds, protein sequences, CT scans, genomic data

Datasets for training models — text corpora, images, audio,


🤖 AI & ML
annotated files

💡 3. Why It’s Valuable

Key Benefits:

✅ Trusted & Curated Data — All datasets come from reputable institutions like
NASA, Stanford, and DeepMind.
✅ Easy Access & Navigation — User-friendly platform for searching and
downloading.
✅ Regular Updates — Continuously refreshed with the latest research data.
✅ Ready to Use — Datasets include:

 Documentation

 Code samples

 Usage examples for reproducibility

🧱 4. How to Use It

1. Visit [Link]

2. Click “Explore Repository”

3. Browse datasets by category (Geospatial, Life Sciences, AI/ML, etc.)

4. Download data directly for analysis or model training.

📘 5. Summary

 Oracle Open Data = Free, curated, research-grade datasets from trusted


global sources.
 Supports AI, ML, and scientific research across multiple domains.

 Simple to use, no account needed.

 Accessible at: [Link]

3. OCI Data Labeling

🏷️ Lesson Notes: OCI Data Labeling

👨🏫 Instructor

Praveen Patil, Principal Product Manager – Data Science & AI Services, Oracle

🧱 1. What is Data Labeling?

Data labeling is the process of identifying and annotating properties (labels) for
image or text data.
It’s an essential step in preparing data for AI and machine learning (ML) training.

 Example: Labeling 100 images of tigers enables an AI system to identify tigers in


new, unlabeled images.
 Labeled data → Used to train supervised ML models.

🗂️ 2. Key Components

Term Description

A collection of data records (images, text files, or documents) and their


Data Set
associated labels.

Data
An individual data item (e.g., one image or one text file).
Record

Labels Assigned categories or properties (e.g., “tiger,” “signature”).

These datasets are interoperable across OCI Data Science and AI Services —
supporting supervised learning workflows.
⚙️ 3. OCI Data Labeling Service

A fully managed OCI service for quickly labeling raw data with minimal setup.
Steps:

1. Load your data

2. Label it using built-in UIs or APIs


3. Export labeled datasets for use in OCI Vision, Data Science, or other AI
services

Purpose: Enables seamless movement of labeled data across OCI for model training
and deployment.

👥 4. Who Uses It

Persona Role & Use

Build labeled datasets, use prebuilt UIs, export data for model
Data Scientists
training in OCI Data Science.

Developers & Label data to fine-tune AI models (e.g., image recognition) before
Engineers deployment.

🏭 5. Industry Applications

Industry Use Case

🏪 Retail & E-Commerce Analyze customer behavior, recommend products

🏥 Healthcare Label medical scans to detect anomalies

🏛️ Government Categorize documents, automate processes

🎬 Media & Entertainment Flag inappropriate content, analyze sentiment

⚙️ Manufacturing Detect defects or classify parts

🔁 6. Role in the AI/ML Lifecycle


 Data labeling occurs early in the ML pipeline — it’s foundational for quality
model training.
 Poor labeling → inaccurate models → wrong predictions.
 ML development is iterative — labeling may need updates to handle:

o Missing or underrepresented classes

o Model drift (changes in real-world data)


Key takeaway: High-quality labeled data = high-quality model.

💡 7. Current Capabilities

✅ Simple and fast labeling for images, text, and documents


✅ Interactive UI or API-based labeling
✅ Easy export to OCI Vision, Data Science, and other AI services
✅ Seamless integration for building and retraining models

📘 8. Summary

 Data labeling = foundational step for supervised learning.

 OCI Data Labeling Service simplifies annotation, management, and export.

 Used by data scientists and developers across multiple industries.

 Enables iterative, accurate ML model training with high-quality labeled data.

You might also like