0% found this document useful (0 votes)
6 views2 pages

Question Bank

The document outlines key concepts in data warehousing, data mining, regression, classification, clustering, and association rule mining. It covers definitions, differences between concepts, processes like ETL, and various algorithms used in these fields. Additionally, it discusses real-world applications, advantages, and limitations of each concept.

Uploaded by

bhadurisoham182
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views2 pages

Question Bank

The document outlines key concepts in data warehousing, data mining, regression, classification, clustering, and association rule mining. It covers definitions, differences between concepts, processes like ETL, and various algorithms used in these fields. Additionally, it discusses real-world applications, advantages, and limitations of each concept.

Uploaded by

bhadurisoham182
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Concepts of data warehousing:

i. What is a Data Warehouse?


ii. How is a Data Warehouse different from a traditional database?
iii. What are the key characteristics of a Data Warehouse (according to Bill Inmon)?
iv. What is OLAP and how is it different from OLTP?
v. What are the common use cases of Data Warehousing?
vi. What are the main components of a Data Warehouse architecture?
vii. What is the staging area in a Data Warehouse?
viii. Explain the ETL process.
ix. What is the role of metadata in Data Warehousing?
x. What is a data mart and how does it differ from a data warehouse?
xi. What are the advantages and disadvantages of data warehouses.
xii. What are top down and bottom up approaches in data warehousing?

Concepts of data mining:

i. What is data mining? Give some real world examples.


ii. List and explain any three major tasks of data mining.
iii. What is classification in data mining? Give an example.
iv. What is clustering? How is it different from classification?
v. What is association rule mining? Provide a simple example.
vi. Explain the concept of support and confidence in association rules.
vii. What is anomaly detection? Give a real-world application.

Concepts of regression:

i. What is regression in statistics and machine learning?


ii. What is the difference between regression and classification?
iii. What is the purpose of a regression model?
iv. What are the assumptions of linear regression?
v. What is the difference between simple linear regression and multiple linear
regression?
vi. What is the role of the intercept and slope in a linear regression model?
vii. How do you interpret the R-squared value in regression analysis?
viii. What is the difference between R-squared and adjusted R-squared?

Concepts of regression:

i. What is classification in machine learning?


ii. How is classification different from regression?
iii. What are some real-world examples of classification problems?
iv. What is a binary classification problem? Give an example.
v. What is a multi-class classification problem?
vi. How does a decision tree classifier work?
vii. What is the K-nearest neighbors (KNN) algorithm used for?
viii. Elaborate the steps in k nearest neighbours classifier.
ix. Explain naïve bayes classifier.
x. What are the assumptions in naïve bayes classifier?
xi. Why the naïve bayes classifier is called naïve?
Concepts of Clustering:

i. What is clustering in data mining or machine learning?


ii. How does clustering differ from classification?
iii. What are the main objectives of clustering?
iv. What are some real-world applications of clustering?
v. How does the K-means clustering algorithm work?
vi. What is the main limitation of the K-means algorithm?
vii. What is hierarchical clustering and how is it different from K-means?
viii. What is an elbow method in the context of clustering?
ix. Why is clustering considered an unsupervised learning method?
x. Briefly describe the steps of the K-Means algorithm.
xi. How does K-Means decide the assignment of data points to clusters?
xii. How do you determine the optimal number of clusters (K)?
xiii. What is the elbow method in K-Means clustering?
xiv. What are the limitations of K-Means clustering?
xv. What is hierarchical clustering?
xvi. How is hierarchical clustering different from K-means clustering?
xvii. What are the two main types of hierarchical clustering?
xviii. What is the difference between agglomerative and divisive clustering?
xix. What is a dendrogram? What does it represent in hierarchical clustering?
xx. What is the first step in agglomerative hierarchical clustering?
xxi. What is the role of distance metrics in hierarchical clustering?
xxii. What are linkage criteria in hierarchical clustering?
xxiii. Name and explain any three linkage methods used in hierarchical clustering.
xxiv. How do you decide the number of clusters using a dendrogram?
xxv. What are the advantages of hierarchical clustering over K-means?
xxvi. What are the limitations of hierarchical clustering?

Concepts of association rule mining:

i. What is association rule mining?


ii. What is the main goal of association rule mining?
iii. Give a real-world application of association rule mining.
iv. What is a frequent itemset?
v. What is support in association rule mining?
vi. What is confidence in association rule mining?

You might also like