Foundation of Data Science
Course Code: CT-202
(Module#1)
Dhawa Sang Dong, MSc Eng.
(Sr. Lecturer)
Kathmandu Engineering College
Kalimati, Kathmandu
December 2025
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 1 / 107
Chapter#1
Introduction to Data Science
✓ Class Outline
1 Introduction to Data Science
2 Terminologies in Data Science
3 Modern Data Ecosystem
4 Data Science Life Cycle
5 Trends, markets and applications of data science
6 Tools and Technologies in Data Science
7 Data Scientist and their Roles
8 Tools and Technologies in Data Science (Optional Content)
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 2 / 107
Course Evaluation
Theory (100)
I Internal weight (40/100)
- Assignments (in total 7) for each module [ 3 x 7 = 21 ]
- Class Activities and average of 2-Test [ 4 + (5 + 10) = 19 ]
II External weight (60/100)
- End Semester Exam by IOE, TU
Practical (50)
- Lab Attendance (0.5 per Lab) and Viva [ 7 + 1.5 x 12 = 25 ]
- Lab Report and Minor Data Science Project [ 10 + 15 = 25 ]
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 3 / 107
Course Overview
Marks Distribution
Reference Books
1 Principle of Data Science – Sinan Ozdemir, 2016 (Packt
Publishing)
2 Data Science from Scratch: First Principles with Python, Joel
Grus, 2017 (O’Reilly Media)
3 Introduction to Data Science – Laura Igual, 2017 (Springer)
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 4 / 107
Introduction to Data Science
What is Data? | Data Science
➠ Data is observation of process, events or phenomena
describing characteristics or properties expressed in the form
of facts or figures.
- Data is simply collection of information in either an organized
format or in unorganized format.
➠ Organized Data: this refers to data that is sorted into a
row/columns structure;
- every row represents a single observation and the columns
represents the characteristics of that observation.
➠ Unorganized Data: this is the type of data that is in the
free form, usually text or raw audio/signals that must be
parsed further to become organized.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 5 / 107
Introduction to Data Science
What is Data Science?
➠ Data Science is all about how we take data, use it to acquire
knowledge, and then use that knowledge to do the following:
✔ make decisions
✔ predict the future
✔ understand the past/present
✔ create new products/industries
➠ Fundamentally, data science includes Math and Statistics,
Computer Programming, and Domain Knowledge
Data Model is an organized, formal structure that defines the
relationships between data elements and represents real-world
entities or phenomena in a systematic way;
Usually, math is used to formalize relationship between
variables.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 6 / 107
Introduction to Data Science
What is Data Science?
✔ Data science is the field of exploring, manipulating, and
analyzing data, and use knowledge/insights to answer questions
or make recommendations or decision making.
✔ Data Science is the study of handling and extracting
meaningful insights from large data sets using modern tools
and algorithms.
✔ The meaningful insights drawn from the data help us in
decision-making.
➠ Gemini (prompt): Data Science is an interdisciplinary field that
uses scientific methods, processes, algorithms, and systems to
extract knowledge and instights from structured and
unstructured data.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 7 / 107
Introduction to Data Science
What is Data Science?
- Data alone holds limited value unless it is transformed into
actionable insights – essential for informed decision-making and
improving processes like design and manufacturing.
- Data Science Tools and Algorithms enable various data mining
tasks (✔descriptive, ✔predictive, ✔diagnostic, and
✔prescriptive analytics) providing insights within the data (past
and present) trends, forecast future outcomes, identify root
causes, and recommend potential actions.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 8 / 107
Introduction to Data Science
What is Data Science?
➠ Data science combines ✔math and statistics;
✔specialized programming; ✔advanced analytics,
✔artificial intelligence (AI) and ✔machine learning with
specific ✔domain (subject matter) expertise to uncover
actionable insights hidden in an organization’s data.
➠ These insights can be used to guide decision making and
strategic planning.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 9 / 107
Introduction ot Data Science
Importance of Data Science
- Data science is important because it transforms data into
knowledge, predictions, automation, and better decisions across
every field using data analytical tools.
❶ informed decision-making,
❷ enhancing customer experience,
❸ innovation and product development, and
❹ improving operational efficiency.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 10 / 107
Introduction ot Data Science
Importance of Data Science
1 Informed Decision-Making
- Because of data science, objective evidence (factual or
varifiable information) is replacing guesswork and intuition.
- predictive analytics can be used to forecast future trends,
market dynamics and customer behavior to predict outcomes.
- strategic decisions can be made on ①product development,
②marketing campaigns, ③resource allocation, and ④operational
changes based on varifiable insights to optimize strategies.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 11 / 107
Introduction ot Data Science
Importance of Data Science
2 Enhancing customer experience
- understanding customer is crucial for business success. Data
science helps companies to build customer-centric strategies.
- Analyzing preferences and behavior (past purchases, browsing
history), highly personalized product recommendations and
tailored marketing (?) can be delivered to individual
customer.
- Identifying the most relevant target audience for specific
products or services can make marketing efforts more effective -
targeted engagement.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 12 / 107
Introduction ot Data Scince
Importance of Data Science
3 Innovation and product Development
- Data driven insights are the foundation for innovation;
businesses use data science to identify market gaps uncovering
customer needs and opportunities.
- customer feed backs and service usage data can be used to
refine existing products or services improving products.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 13 / 107
Introduction ot Data Science
Importance of Data Science
4 Improving Operational Efficiency and Risk Mitigation
- By analyzing internal data, organization can improve their
operations and security;
- identifying bottlenecks, streamline workflows and complex
process optimization can lead increased efficiency and cost
reduction (process optimization).
- advanced algorithms can detect anomalies and patterns to
identify fraudulent activities, so risk minimization;
- even equipment or machinery fails can be predict, so predictive
maintenance can be scheduled proactively.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 14 / 107
Terminologies in Data Science
Terminologies in Data Science
Here, we will discuss some key words/terminologies/jargon related
to field of data science – a data analyst or data scientist may use
Algorithm: | Terminologies in Data Science
- An algorithm is a sequence of steps or guidelines designed to
accomplish a particular task (computational problems).
- They are especially valuable when dealing with big data or
machine learning.
- Data analysts often use algorithms to structure or examine
data, while data scientists employ them to make forecasts or
create models.
➠ GPT: An algorithm is a finite set of well-defined instructions
that takes some input, performs a sequence of steps and
produces an output.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 15 / 107
Terminologies in Data Science
Artificial Intelligence (AI): | Terminologies in Data Science
¬ “Artificial Intelligence is the science and engineering of making
machine intelligent.” – John McCarthy, Father of AI
(Darthmouth Conference 1956)
- Artificial intelligence (AI) uses algorithms and vast datasets
from computer science to enable machines to perform tasks
that typically require human intelligence, such as recognizing
patterns, making decisions, and solving complex problems.
- The intelligence is considered "artificial" because the computer
is programmed (implicitly) to carry out tasks typically
requiring or linked to human cognitive functions.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 16 / 107
Terminologies in Data Science
Big Data: | Terminologies in Data Science
¬ Big data is a vast set of information defined by the three V’s:
volume, velocity, and variety.
➠ Volume relates to the high amount of data – big data involves
handling large quantities;
➠ velocity refers to the speed at which data is generated and
gathered – data is collected rapidly, often streaming directly into
memory; and
➠ variety highlights the diversity of data types – big data
encompasses a wide range of structured, semi-structured, and
unstructured data, as well as various formats like numbers, text,
images, and audio.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 17 / 107
Terminologies in Data Science
Business intelligence (BI): | Terminologies in Data Science
➤ Business intelligence (BI) involves use of data analytics to
help organizations make informed/data-driven decisions.
➤ BI analysts examine business data such as revenue, sales, or
customer information and provide recommendations based on
their findings.
➤ By leveraging BI tools and techniques, businesses can identify
trends, uncover insights, and optimize operations.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 18 / 107
Terminologies in Data Science
Changelog: | Terminologies in Data Science
¬ A changelog is a record or log of all the changes made to a
project, software, or system over the time.
➠ It typically includes details about new features, bug fixes,
improvements, updates, or other modifications in each version
or release.
➠ Changelogs are useful for developers, users, and stakeholders
to track the evolution of a project and understand what has
been altered between different versions.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 19 / 107
Terminologies in Data Science
Classification: | Terminologies in Data Science
¬ Classification is a type of machine learning task that sorts data
into predefined categories.
➠ It can be applied, for instance, to develop email spam filters.
➠ Common algorithms used to build classification models
include logistic regression, decision trees, K-nearest neighbors
(KNN), and random forests.
Dashboard: | Terminologies in Data Science
¬ A dashboard is a tool used to display and monitor live data.
➠ dashboards are typically connected to databases and feature
visualizations that automatically update to reflect the most
current data in the database.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 20 / 107
Terminologies in Data Science
Data Analytics: | Terminologies in Data Science
✔ Data analytics involves gathering, transforming, and organizing
data to draw insights, make predictions, and support informed
decision-making.
➠ Data Analytic includes:
data engineering – developing data systems/architecture,
data science – using data to hypothesize and predict, and
data analysis – extracting meaningful information from data
➠ Professionals in this field include data analysts, data
scientists, and data engineers, all contributing to different
aspects of data analytics to draw insights from the data.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 21 / 107
Jargon of Data Science
Data Analytics: | Terminologies in Data Science
¬ There are four key types of data analytics, including:
✔ Descriptive analytics, ➠ what happened.
✔ Diagnostic analytics, ➠ why something happened.
✔ Predictive analytics, ➠ what will likely happen in the future.
✔ Prescriptive analytics, ➠ how to act
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 22 / 107
Terminologies in Data Science
Data Architecture: | Terminologies in Data Science
➠ Data architecture, or data design, is the strategic framework
for an organization’s data management system.
➠ It covers every stage of the data lifecycle, including data
collection, organization, usage, and disposal.
➠ Data architects are responsible for designing the plans that
guide how these systems are built and maintained.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 23 / 107
Terminologies in Data Science
Data Cleaning: | Terminologies in Data Science
➠ Data cleaning, also known as data cleansing or scrubbing, is
the process of preparing raw data for analysis.
➠ This involves ensuring the data is accurate, complete,
consistent, and free of bias.
➠ Clean data is essential before performing any analysis, as
unclean or flawed data can result in incorrect conclusions and
poor business decisions.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 24 / 107
Terminologies in Data Science
Data Engineering: | Terminologies in Data Science
➠ Data engineering involves creating systems that make data
accessible for analysis.
➠ Data engineers are responsible for building systems that
gather, manage, and transform raw data into usable insights.
➠ Their tasks often include developing algorithms to process data
into a more practical format, constructing database pipelines,
and designing new tools for data analysis.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 25 / 107
Terminologies in Data Science
Data Enrichment: | Terminologies in Data Science
➠ Data enrichment involves augmenting your existing dataset
with additional information.
➠ This process usually occurs during data transformation as you
prepare for analysis, particularly if you identify the need for
more data to effectively address your business questions.
Data Governance: | Terminologies in Data Science
➠ Data governance refers to the structured framework for how an
organization oversees its data management.
➠ It includes guidelines for data access and usage, as well as
rules related to accountability and compliance.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 26 / 107
Terminologies in Data Science
Data Lake: | Terminologies in Data Science
➠ A data lake is a storage repository designed to collect and
retain vast amounts of structured, semi-structured, and
unstructured raw data. (➠ data scientists are users)
➠ Data scientists utilize the data stored in data lakes for machine
learning or AI algorithms and models, or they may process the
data and transfer it to a data warehouse.
Data Mart: | Terminologies in Data Science
➠ A data mart is a smaller segment of a data warehouse that
contains all processed data relevant to a specific department.
➠ While a data warehouse might encompass information related
to finance, marketing, sales, and human resources, a data mart
focuses specifically on data pertinent to the finance team.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 27 / 107
Terminologies in Data Science
Data Mining: | Terminologies in Data Science
➠ Data mining involves thoroughly analyzing data to uncover
patterns and extract insights.
➠ It is a key component of data analytics, as the insights gained
during the mining process will guide your business decision or
recommendations.
Data Modeling: | Terminologies in Data Science
➠ Data modeling is the process of creating maps and
constructing data pipelines that link data sources for analysis.
➠ A data model serves as a tool to implement these pipelines
and organize data across various sources.
➠ Data modelers are systems analysts who collaborate with data
architects and database administrators to design databases and
data systems.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 28 / 107
Terminologies in Data Science
Data Visualization: | Terminologies in Data Science
➠ Data visualization is the process of presenting information and
data through charts, graphs, maps, and other visual aids.
➠ Effective data visualizations can enhance storytelling, make
data more accessible to a broader audience, reveal patterns
and relationships, and facilitate deeper exploration of the data.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 29 / 107
Terminologies in Data Science
Data Wrangling: | Terminologies in Data Science
➠ Data wrangling, also known as data munging or data
remediation which involves transforming raw data into a usable
data format for analysis, modeling, or visualization.
➠ The wrangling process consists of four stages:
✔ data discovery/understanding,
✔ data transformation,
✔ data validation, and
✔ data publishing.
➠ The data transformation stage can be further divided into
tasks such as data structuring, normalization or
denormalization, cleaning, and enrichment.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 30 / 107
Terminologies in Data Science
Data Warehouse: | Terminologies in Data Science
➠ A data warehouse is a centralized storage system that holds
processed and organized data from various sources.
➠ It may include a mix of current and historical data that has
been extracted, transformed, and loaded from both internal
and external databases.
More Jargon: | Data Science Jargon
✓ Data base, ✓ deep learning, ✓ machine learning,
✓ reinforcement learning, ✓ structured data, ✓ regression,
✓ structure query language (SQL), ✓ supervised learning,
✓ unsupervised learning, ✓ unstructured data
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 31 / 107
Terminologies in Data Science
Data Lake, Data Mart, Data Warehous:
Feature Data Lake Data Warehouse Data Mart
Store all raw data for future Store processed, structured Provide focused data for a
Purpose
analysis data for reporting & BI specific team/unit
Raw, unstructured, semi- Structured + some semi-
Data Type Highly structured
structured, structured structured
Specific departments (sales,
Users Data scientists, engineers Business analysts, BI teams
finance, HR)
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 32 / 107
Modern Data Ecosystem
What is Modern Data Ecosystem?
✍ A modern data ecosystem refers to the integrated set of
technologies, practices, and processes that an organizations
use to collect, store, process, analyze, and visualize data.
✍ The term data ecosystem refers to programming language,
packages, algorithms, cloud-computing services, general
infrastructure that the organization uses to collect, store,
analyze, and leverage data – Harvard business school.
✍ This ecosystem is designed to handle the complexities of
today’s data landscape, which includes vast amounts of
structured and unstructured data generated from various
sources, including IoT devices, social media, transactional
systems, and more.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 33 / 107
Modern Data Ecosystem
Some Key Elements of Modern Data Ecosystem:
1 Data Sources
2 Data Integration
3 Data Storage
4 Data Processing
5 Data Governance
6 Data Analytics Tools
7 Data Analysis Techniques
8 Data Visualization
9 Data Democratization
10 Advanced Data Analytics
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 34 / 107
Modern Data Ecosystem
Data Sources: | Elements of Modern Data Ecosystem
- these are the systems, applications, and devices that generates
or collect data for an organization.
- Data Sources can include ✔customer relationship management
(CRM), ✔web applications, ✔transactional databases, ✔social
media platform or Internet of Things (IOT) devices and more.
Data Integration: | Elements of Modern Data Ecosystem
- Data from different sources often needs to be integrated into a
unified format for analysis
- data integration tools and techniques are used to extract,
transform and load (ETL) data from diverse sources into a
centralized data repository or data lake.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 35 / 107
Modern Data Ecosystem
Data Storage: Elements of Modern Data Ecosystem
- A modern data ecosystem typically involves storing data in
scalable and flexible formats.
- this can include traditional relational database (like postgreSQL
or MySQL), cloud based data warehouses (like Amazon
Redshift or Google BigQuery) or distributed file systems like
Apache Hadoop (Disk-based).
Data Processing: Elements of Modern Data Ecosystem
- to analyze large volumes of data efficiently, distributed
computing frameworks like ✔Apache Spark (in-memory) or
✔Apache Hadoop MapReduce are commonly used.
- These frameworks enable parallel processing of data across
clusters of computers, allowing for high-performance data
processing and analytics.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 36 / 107
Modern Data Ecosystem
Data Governance: Elements of Modern Data Ecosystem
- data governance ensures the availability, integrity, privacy, and
security of data within an organization.
- it involves establishing policies, processes, and controls to
manage data effectively, comply with regulations, and maintain
data quality.
- data governance frameworks and tools help organizations
ensure the reliability of their data analytics processes.
Data Analytics Tools: Elements of Modern Data Ecosystem
- a wide range of tools and technologies exist for data analytics,
including programming languages (python or R), statistical
packages, business intelligence (BI) tools, data visualization
tools (tableau or power BI), and machine learning platform.
- these tools enable organizations to extract insights, discover
patterns and make data-driven decisions.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 37 / 107
Modern Data Ecosystem
Data Analysis Techniques: Modern Data Ecosystem Elements
- data analytics encompasses various techniques, including
✔ descriptive analytics (summarizing historical data),
✔ predictive analytics (forecasting future outcomes), and
✔ prescriptive analytics (providing recommendations).
- Organizations employ these techniques to gain insights, detect
anomalies, predict trends, and optimize business processes.
Data Visualization: Elements of Modern Data Ecosystem
- data visualization is crucial for effectively communicating
insights and findings from data analysis.
- modern data ecosystems leverage interactive and intuitive
visualization tools to present data in meaningful ways, enabling
stakeholders to understand and interpret the results easily.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 38 / 107
Modern Data Ecosystem
Data democratization: Elements of Modern Data Ecosystem
➠ data democratization aims to make data and analytics
accessible to a wider audience within an organization.
➠ this involves providing self-service ✔ analytics capabilities,
✔ intuitive dashboards, and ✔ tools that empower business
users and domain experts to explore and analyze data without
heavy reliance on IT or data science team.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 39 / 107
Modern Data Ecosystem
Advanced Analytics: Elements of Modern Data Ecosystem
➠ advanced analytics techniques such as machine learning,
artificial intelligence (AI) are increasingly being incorporated
into modern data ecosystems.
➠ these techniques enable organizations to leverage complex
algorithms and models to uncover patterns, make predictions,
automate recommendation & decision-making processes, and
gain competitive advantage.
Prompt (ChatGPT ?): Define Modern Data Ecosystem.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 40 / 107
Modern Data Ecosystem
What is Modern Data Ecosystem?
➠ ChatGPT (prompt): Define modern data ecosystem ?
✔ A modern data ecosystem is a collection of technologies,
processes, and people that work together to collect, store,
manage, process, analyze, and use data in a scalable, secure,
and efficient way—typically using cloud platforms,
automation, and advanced analytics (AI/ML).
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 41 / 107
Data Science Life Cycle
Data Science Life Cycle
Fig. 1 Data Science Life Cycle
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 42 / 107
Data Science Life Cycle
Discovery: | Data Science Life Cycle
➤ Understand the problem you’re trying to solve and the business
objectives – understanding the business problem
➠ In this phase, you will gather requirements and define key
metrics for success.
✔ Identify the problem.
✔ Understand the business context and goals.
✔ Define the data science objectives and questions.
✔ Specify success/performance criteria (metrics).
✔ Formulate a hypothesis to be tested with data.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 43 / 107
Data Science Life Cycle
Data Preparation: | Data Science Life Cycle
➤ Collect, clean, and organize relevant data needed to address
the problem (business problem).
✔ Data acquisition: from internal databases, external APIs,
web scraping, etc.
✔ Data cleaning: handle missing values, correct
inconsistencies, remove duplicates.
✔ Data transformation: standardization, normalization, and
encoding of variables.
✔ Data integration: from various sources.
✔ Exploratory Data Analysis(EDA): to understand data
distribution and relationships.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 44 / 107
Data Science Life Cycle
Model Plan: | Data Science Life Cycle
➠ Design a plan for model building by selecting algorithms and
creating a workflow for model development.
✔ Choose appropriate modeling techniques based on the data
and problem (e.g., regression, classification, clustering).
✔ Split the dataset into training, validation, and test sets.
✔ Decide on performance metrics to evaluate the model
performance: (e.g., accuracy, precision, recall, RMSE).
✔ Create a roadmap for feature selection, model iteration, and
testing.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 45 / 107
Data Science Life Cycle
Model Development: | Data Science Life Cycle
➠ Build and train the model on the prepared data.
✔ Develop models using the chosen algorithms.
✔ Train the model on the training data.
✔ Fine-tune hyperparameters to optimize performance.
✔ Use cross-validation to avoid overfitting.
✔ Feature engineering, if needed, to enhance model
performance.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 46 / 107
Data Science Life Cycle
Operationalize/Deployment: | Data Science Life Cycle
➠ Deploy the model into production and integrate it into the
business environment where it can provide predictions.
✔ Implement the model in a live system or application.
✔ Set up pipelines for data flow, allowing real-time predictions
or periodic batch processing.
✔ Ensure scalability and performance in a real-world setting.
✔ Monitor the model’s behavior and response time.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 47 / 107
Data Science Life Cycle
Communicate Results: | Data Science Life Cycle
➠ Present insights and results to stakeholders/audience to inform
decision-making or recommendation.
✔ Generate reports, dashboards, or visualizations that clearly
explain the model’s output and predictions.
✔ Provide actionable insights based on the model’s results.
✔ Communicate the business impact of the model’s findings.
✔ Suggest further iterations or improvements based on
feedback.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 48 / 107
Trends, markets and applications of data science
Trends, Application and job Market: | Data Science
✔ Data science is revolutionizing industries with trends and
applications focused on automation, AI integration, and
enhanced decision-making.
AI is the new electricity – Ng Andrew
Trends: | Data Science
➠ Artificial Intelligence and Machine Learning (AI/ML):
AI-powered models are advancing across fields like NLP
(Natural Language Processing), computer vision, and deep
learning, leading to automation in various sectors.
➠ Data Engineering and Real-time Analytics: As data grows,
real-time processing and data engineering are crucial for quick
insights, particularly in finance, e-commerce, and IoT.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 49 / 107
Trends, markets and applications of data science
Trends, Application and job Market: | Data Science
✔ Data science is revolutionizing industries with trends and
applications focused on automation, AI integration, and
enhanced decision-making.
Trends: | Data Science
➠ MLOps and AI Governance: Managing machine learning
operations (MLOps) for model deployment, scaling, and
governance is on the rise to streamline workflows and ensure
compliance (obeying required laws, regulations, and standards).
➠ Edge Computing: Processing data closer to the source is
enabling faster decision-making in IoT and autonomous
systems.
➠ Privacy-focused AI: With data privacy concerns, federated
learning and differential privacy are gaining attraction.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 50 / 107
Trends, markets and applications of data science
Trends, Application and job Market: | Data Science
✔ Data science is revolutionizing industries with trends and
applications focused on automation, AI integration, and
enhanced decision-making.
Trends: | Data Science
➠ Retail and E-commerce: Customer behavior analysis,
demand forecasting, and personalized recommendations are
core to the retail sector’s data strategy.
➠ Manufacturing and Supply Chain: Predictive maintenance,
inventory management, and quality control benefit from
real-time analytics.
➠ Telecommunications: Improving network management, churn
prediction, and customer segmentation are key applications.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 51 / 107
Trends, markets and applications of data science
Trends, Application and job Market: | Data Science
✔ Data science is revolutionizing industries with trends and
applications focused on automation, AI integration, and
enhanced decision-making.
Markets : | Data Science
➠ Healthcare: Predictive analytics, genomics, and personalized
medicine are transforming diagnostics and treatment plans.
➠ Finance and Banking: Risk assessment, fraud detection, and
algorithmic trading leverage data science for secure and
optimized financial services.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 52 / 107
Trends, markets and applications of data science
Trends, Application and job Market: | Data Science
✔ Data science is revolutionizing industries with trends and
applications focused on automation, AI integration, and
enhanced decision-making.
Applications: | Data Science
➠ Customer Segmentation and Personalization: Used in
marketing and e-commerce for targeted campaigns and
recommendations.
➠ Predictive Maintenance: Critical in manufacturing and
transportation to anticipate failures and reduce downtime.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 53 / 107
Trends, markets and applications of data science
Trends, Application and job Market: | Data Science
✔ Data science is revolutionizing industries with trends and
applications focused on automation, AI integration, and
enhanced decision-making.
Applications: | Data Science
➠ Fraud Detection and Risk Analysis: Vital in finance, where
data science models detect anomalies and reduce risk.
➠ Healthcare Analytics: From disease prediction to operational
optimization, data science is transforming healthcare delivery.
➠ Natural Language Processing: NLP applications are
essential in virtual assistants, chatbots, and sentiment analysis
across sectors.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 54 / 107
Trends, markets and applications of data science
Data Science Application and job Market: | Data Science
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 55 / 107
Tools and Technologies in Data Science
1. Programming Language
- Python: Most widely used; huge ecosystem (NumPy, pandas,
PyTorch, TensorFlow)
- R: Used heavily in statistics, biostatistics, epidemiology
- SQL: For database querying and relational data operations.
- Julia: Fast numerical computing and scientific computing.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 56 / 107
Tools and Technologies in Data Science
2. Libraries and Frameworks
- Pandas: tools for data structure and data analysis.
- NumPy: is library for numerical computing that supports array
and matrix operations.
- Scikit-learn: is machine learning library from where we can
import various machine learning algorithms and modules for
data mining and data analysis.
- TensorFLow: open-source framework for machine learning and
deep learning – neural networks.
- PyTorch : a deep learning framework similar to TensorFlow,
known for dynamic computation graph ans strong community.
- HuggingFace: ?
- Statsmodels: ?
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 57 / 107
Tools and Technologies in Data Science
3. Visualization Tools
- Matplotlib: popular python library for creating static, animated
and interactive visualization.
- Seabon: another python library built on top of Matplotlib, it
provides high level interface for drawing attractive and
informative statistical graphs.
- Tableau: A leading data visualization tool that allows users to
create interactive, shareable dashboards.
- Power BI: A business analytics tool from Microsoft for creating
interactive reports and visualizations.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 58 / 107
Tools and Technologies in Data Science
4. Big Data & Distributed Computing
- Apache Hadoop: An open-source framework for distributed
storage and processing of large datasets using a network of
computers (disk-based computing).
- Apache Spark: A fast, in-memory data processing engine with
support for a wide range of analytics tasks, including machine
learning and graph processing.
- Databricks: A unified data analytics and AI platform built on
Apache Spark that enables scalable data engineering, data
science, and machine learning workflows.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 59 / 107
Tools and Technologies in Data Science
5. Cloud Platforms
- AWS (Amazon Web Services): Offers scalable cloud storage,
computing power, and machine learning services.
- Azure: Microsoft’s cloud platform, offering a suite of tools for
building, deploying, and managing applications and services.
- Google Cloud: Provides tools for computing, data storage, and
machine learning, including services like BigQuery.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 60 / 107
Tools and Technologies in Data Science
6. Collaboration Tools
- Jupyter Notebook: An open-source web application for
creating and sharing documents that contain live code,
equations, visualizations, and narrative text.
- Google Colab: A cloud-based notebook environment that
allows you to write and execute Python code in a collaborative
manner.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 61 / 107
Data Scientist and their Roles
Data Scientist: | Data Science
✔ Data scientists are skilled professional who use a combination
of statistical, analytical, and machine learning techniques to
extract valuable insights from data.
- They play a key role in analyzing vast amounts of structured
and unstructured data to solve complex business problems,
drive decision-making, and forecast trends.
➠ Key Skills and Knowledge Areas:
✔ Statistics and Mathematics: Essential for creating predictive
models and understanding data patterns.
✔ Programming: Proficiency in languages like Python, R, and
SQL is common, as they enable data manipulation, analysis,
and model building.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 62 / 107
Data Scientist and their Roles
Data Scientist: | Data Science
➠ Key Skills and Knowledge Areas:
✔ Machine Learning and AI: Familiarity with machine
learning algorithms (e.g., regression, classification,
clustering) and deep learning for building complex models.
✔ Data Visualization: Ability to present findings through
visualizations (using tools like Tableau, Power BI, or Python
libraries) that are understandable to non-technical
stakeholders.
✔ Domain Knowledge: Understanding the industry they work
in (e.g., finance, healthcare, or retail) helps data scientists
apply relevant models and solutions.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 63 / 107
Data Scientist and their Roles
Role of Data Scientist: | Data Science
❶ Data Collection and Processing
- Role: Data scientists are responsible for identifying,
gathering, and cleaning data from various sources, ensuring
it’s reliable for analysis.
- Tasks: They preprocess and organize data using ETL
(Extract, Transform, Load) methods, which involve data
extraction, cleaning, transformation, and storage.
- Wrangling is the hands-on cleaning and preparation work
you do mainly during the data transformation step.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 64 / 107
Data Scientist and their Roles
Role of Data Scientist: | Data Science
❷ Exploratory Data Analysis (EDA)
- Role: Conducting EDA helps data scientists uncover
patterns, trends, and anomalies in the data.
- Tasks: They use statistical and visualization techniques to
summarize the main characteristics of datasets and draw
preliminary insights.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 65 / 107
Data Scientist and their Roles
Role of Data Scientist: | Data Science
❸ Feature Engineering
- Role: Data scientists create new features or modify existing
ones to enhance model accuracy and performance.
- Tasks: They select relevant features and apply techniques
such as scaling, encoding, and dimensionality reduction to
make data suitable for modeling.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 66 / 107
Data Scientist and their Roles
Role of Data Scientist: | Data Science
❹ Model Building and Evaluation
- Role: Data scientists build, train, and fine-tune machine
learning models to address specific business questions.
- Tasks: They select the right algorithms, train models,
validate performance using cross-validation and metrics like
accuracy, precision, and recall; and during model building,
they optimize the hyperparameters
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 67 / 107
Data Scientist and their Roles
Role of Data Scientist: | Data Science
❺ Machine Learning and Deep Learning
- Role: Advanced machine learning and deep learning skills
enable data scientists to tackle complex problems.
- Tasks: They implement algorithms for supervised,
unsupervised, and reinforcement learning, and may use deep
learning for tasks like image recognition or natural language
processing (NLP).
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 68 / 107
Data Scientist and their Roles
Role of Data Scientist: | Data Science
❻ Data Visualization and Reporting
- Role: Data scientists present data findings to stakeholders or
audiences in a clear, actionable format.
- Tasks: They create visualizations and dashboards using
tools like Tableau, Power BI, or Python libraries (e.g.,
Matplotlib, Seaborn) to make complex insights accessible to
non-technical audiences.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 69 / 107
Data Scientist and their Roles
Role of Data Scientist: | Data Science
❼ Experimentation and Testing
- Role: Data scientists design experiments to test hypotheses,
new products, or features.
- Tasks: They set up and analyze controlled experiments
(e.g., A/B testing) to measure the impact of changes and
optimize decisions based on statistically valid results.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 70 / 107
Data Scientist and their Roles
Role of Data Scientist: | Data Science
❽ Model Deployment and Monitoring
- Role: Once models are developed, data scientists often help
deploy them in production environments.
- Tasks: They work with data engineers and MLOps teams to
monitor model performance over time, ensuring continued
accuracy and detecting issues like data drift.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 71 / 107
Data Scientist and their Roles
Role of Data Scientist: | Data Science
❾ Collaboration and Communication
- Role: Data scientists work closely with cross-functional
teams to align data solutions with business goals.
- Tasks: They communicate insights and results to teams and
executives, translating technical findings into business
strategies and actionable insights.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 72 / 107
Data Scientist and their Roles
Role of Data Scientist: | Data Science
❿ Ethics and Data Privacy
- Role: Data scientists ensure their work adheres to ethical
guidelines and data privacy laws, especially when handling
personal data.
- Tasks: They address issues like data security, bias in models,
and regulatory compliance, particularly in fields like finance
and healthcare where data privacy is crucial.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 73 / 107
Module Assignment – As You Go
Module#1 Assignment is available at MS-Team.
Submission Deadline: 19th December 2025 (Before 3:00 PM)
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 74 / 107
Tools and Technologies in Data Science
NOTE: Following content is detailed study on Tools and
Technologies in Data Science; it is optional content if anyone
wants go more detail on the topic.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 75 / 107
Tools and Technologies in Data Science
Tools and Technologies: | Data Science
➠ To make use of data,
Raw data should be passed through data science tasks:
1 Data Management
2 Data Integration and Transformation
3 Data Visualization
4 Model Building
5 Model Deployment
6 Model Monitoring and Assessment
To perform above tasks explained, one needs following workflow
backbone:
✔ data asset management, ✔ code asset management,
✔ execution environments, and ✔ development environments
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 76 / 107
Tools and Technologies in Data Science
Tools and Technologies: | Data Science
➠ Data asset management
- It refers to the systematic process of
✔ organizing, ✔storing,
✔ securing, and ✔ maintaining
data as a valuable asset for an organization.
- It involves the practices and tools needed to ensure that data
is ✔ accessible, ✔ high-quality, ✔ reliable, and ✔ effectively
leveraged to support decision-making, analytics, and business
strategies, recommendations.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 77 / 107
Tools and Technologies in Data Science
Tools and Technologies: | Data Science
➠ Data Asset Management
- Here are some open source Data Asset Management tools:
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 78 / 107
Tools and Technologies in Data Science
Tools and Technologies: | Data Science
➠ Code asset management
- it involves ✔organizing, ✔tracking, and ✔maintaining the
code or scripts, and data models used in data projects.
- Effective code asset management ensures that the work is
✔ reproducible, ✔ version-controlled, and
✔ easy to collaborate on.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 79 / 107
Tools and Technologies in Data Science
Tools and Technologies: | Data Science
➠ Code Asset Management
- Here are some open source Code Asset Management tools:
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 80 / 107
Tools and Technologies in Data Science
Tools and Technologies: | Data Science
➠Execution environments
- these are setups where code, models, and data workflows
➠ run, enabling experimentation, development, testing,
evaluation and production.
- Each environment type offers specific advantages depending
on the project’s needs, such as processing power, scalability,
and reproducibility.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 81 / 107
Tools and Technologies in Data Science
Tools and Technologies: | Data Science
➠ Execution Environment
- Here are some open source Execution Environment:
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 82 / 107
Tools and Technologies in Data Science
Tools and Technologies: | Data Science
➠ Development environments
- these are platforms, tools, and setups used to
✔ write code,
✔ test code, and
✔ debug code for data projects.
- The right environment can
✔ streamline experimentation,
✔ improve productivity, and
✔ enhance collaboration.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 83 / 107
Tools and Technologies in Data Science
Tools and Technologies: | Data Science
➠ Development Environment
- Here are some open source Development Environment:
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 84 / 107
Tools and Technologies in Data Science
Tools and Technologies: | Data Science
❶ Data Management
- it is the process of collecting, persisting, and retrieving data
securely, efficiently, and cost-effectively.
- Data is collected from many sources, like Twitter, Flipkart,
Media, Sensors, and more.
- Store collected data in persistent storage so it is available
whenever you need it.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 85 / 107
Tools and Technologies in Data Science
Tools and Technologies: | Data Science
❶ Data Management
- Here are some open source data management tools:
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 86 / 107
Tools and Technologies in Data Science
Tools and Technologies: | Data Science
❷ Data Integration and Transformation
- It is the early process (on collected raw data) of Extracting,
Transforming, and Loading data. ➠ “ETL”.
- Some of this data is distributed in multiple repositories.
- For example, a database, a data cube, and flat files.
- Use the Extraction process to extract data from these
numerous repositories and save to a central repository like a
Data Warehouse.
- Data Warehouses are primarily used to collect and store
massive amounts of data for data analysis.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 87 / 107
Tools and Technologies in Data Science
Tools and Technologies: | Data Science
❷ Data Integration and Transformation
- Next, Data Transformation is the process of transforming the
values, structure, and format of data.
- After extraction, the next step is to transform the data;
- And once the data is transformed, it’s time to load the data.
- Transformed data is loaded back to the Data Warehouse.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 88 / 107
Tools and Technologies in Data Science
Tools and Technologies: | Data Science
❷ Data Integration and Transformation
- Here are some open source data Integration and
Transformation tools:
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 89 / 107
Tools and Technologies in Data Science
Tools and Technologies: | Data Science
❸ Data visualization
- It is the graphical representation of data and information.
- You can use visualization to represent data in the form of
charts, plots, maps, animations, etc.
- And data visualization conveys data more effectively for
decision-makers.
- It is a crucial step in the data science process.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 90 / 107
Tools and Technologies in Data Science
Tools and Technologies: | Data Science
❸ Data visualization
- Various forms of data visualizations include:
a bar chart ✔ a treemap ✔
➠ which compares the size of ➠ which displays hierarchy
each component, data,
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 91 / 107
Tools and Technologies in Data Science
Tools and Technologies: | Data Science
❸ Data visualization
- Various forms of data visualizations include:
a line chart ✔ a map chart ✔
➠ which plots a series of data ➠ which displays data by
points over time location.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 92 / 107
Tools and Technologies in Data Science
Tools and Technologies: | Data Science
❸ Data Visualization
- Here are some open source Data Visualization tools:
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 93 / 107
Tools and Technologies in Data Science
Tools and Technologies: | Data Science
❹ Model Building:
- This is where you train the data and analyze patterns with
machine learning algorithms.
- The system ‘learns’ how to provide predictions or decisions
by itself; you can then use this model to make predictions on
new, unseen data.
- Model building can be done using a service called IBM
Watson Machine Learning; it provides a full range of tools
and services for building models.
➠ Some other machine learning model building platform:
✔ Google Cloud AI Platform, ✔ Amazon SageMaker (AWS),
✔ Microsoft Azure Machine Learning, ✔ Databricks, BigML etc
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 94 / 107
Tools and Technologies in Data Science
Tools and Technologies: | Data Science
❺ Model Deployment:
- The process of integrating a developed model into a
production environment.
- In model deployment, a machine learning model is made
available to third-party applications via APIs.
- Business users can access and interact with the data through
these third-party applications.
- and, so this helps them make data-driven decisions.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 95 / 107
Tools and Technologies in Data Science
Tools and Technologies: | Data Science
❺ Model Deployment
- Here are some open source Model Deployment tools:
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 96 / 107
Tools and Technologies in Data Science
Tools and Technologies: | Data Science
❻ Model monitoring and assessment
- It runs continuous quality checks to ensure a model’s
accuracy, fairness, and robustness.
- Model monitoring uses tools like Fiddler to track the
performance of deployed models in a production
environment.
- Now, model assessment uses evaluation metrics like the
✔F1-score, ✔ true positive rate, or ✔ the sum of
squared error to understand a model’s performance.
- A well-known example is the IBM Watson Open scale, which
continuously monitors deployed machine learning and deep
learning models.
- It will improve the accuracy and quality of your predictions.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 97 / 107
Tools and Technologies in Data Science
Tools and Technologies: | Data Science
❻ Model monitoring and assessment
- Here are some open source Model monitoring and assessment
tools:
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 98 / 107
Tools and Technologies in Data Science
Cloud Based Tools and Technologies: | Data Science
➠ Fully Integrated Visual Tools and Platform
✔ Watson Studio and Watson OpenScale: It covers the
complete development life cycle for all data science,
machine learning, and artificial intelligence (AI) tasks.
✔ Microsoft Azure Machine Learning: It is also a full
cloud-hosted offering supporting the complete development
life cycle of all data science, machine learning, and AI tasks.
✔ H2O Driverless AI: Although it is a product you download
and install, there exists a one-click deployment for the
standard cloud service providers; This cloud provider does
not do operations and maintenance, as with Watson Studio,
Open Scale, and Azure Machine Learning,
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 99 / 107
Tools and Technologies in Data Science
Cloud Based Tools and Technologies: | Data Science
❶ Data Management
- software-as-a-service (SaaS) versions of existing open source
and commercial tools exist;
- The cloud provider operates the tool for you in the cloud;
- The cloud provider operates the product by backing up your
data and configuring and installing updates.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 100 / 107
Tools and Technologies in Data Science
Cloud Based Tools and Technologies: | Data Science
❶ Data Management
✔ Amazon Web Services DynamoDB is a NoSQL database
(database as a service). It allows storage and retrieving data
in a key-value or a document store format. The most
prominent document data structure is JSON.
✔ Cloudant is another database as a service offering; but in
the background, it is based on the open-source Apache
CouchDB.
✔ IBM DB2 service provided by IBM; it is an example of a
commercial database made available as a SaaS offering in
the cloud, taking away operational tasks from the user.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 101 / 107
Tools and Technologies in Data Science
Cloud Based Tools and Technologies: | Data Science
❷ Data Integration and Transformation
➠ Two commercial data integration tools widely used are:
✔ Informatica Cloud Data Integration, and
✔ IBM’s Data Refinery.
- Data Refinery is part of IBM Watson Studio; it allows
transforming large amounts of raw data into consumable,
quality information in a spreadsheet-like user interface.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 102 / 107
Tools and Technologies in Data Science
Cloud Based Tools and Technologies: | Data Science
❸ Data Visualization
✔ Datameer:
- A smaller company offering a cloud-based data visualization
tool is Datameer.
✔ IBM Cognos Analytics
- It is the service (Business intelligence suite) IBM offers as a
cloud solution for data visualization.
- IBM Data Refinery also offers data exploration and
visualization functionality in Watson Studio
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 103 / 107
Tools and Technologies in Data Science
Cloud Based Tools and Technologies: | Data Science
❹ Model Building
✔ IBM Watson Machine Learning
- Watson Machine Learning can train and build models using
various open-source libraries.
✔ Google has a similar service on their cloud called AI
Platform Training.
- Every cloud provider has a solution for this task.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 104 / 107
Tools and Technologies in Data Science
Cloud Based Tools and Technologies: | Data Science
❺ Model Deployment
✔ IBM Watson Machine Learning
- Watson Machine Learning deploys a model and makes it
available to consumers using a REST interface.
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 105 / 107
Tools and Technologies in Data Science
Cloud Based Tools and Technologies: | Data Science
❻ Model Monitoring and Assessment
✔ Amazon SageMaker Model Monitor
- Amazon SageMaker Model Monitor is an example of a cloud
tool to monitor deployed machine learning and deep learning
models continuously.
✔ Watson OpenScale.
- It is another tool for model monitoring
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 106 / 107
Module Assignment – As You Go
Module#1 Assignment is available at MS-Team.
Submission Deadline: 19th December 2025 (Before 3:00 PM)
Dhawa Sang Dong, MSc Eng. | Kathmandu Engineering College 107 / 107