0% found this document useful (0 votes)
33 views9 pages

CSE Data Warehousing & Mining Question Bank

The document is a question bank for the B.E. CSE program at Hindusthan College of Engineering and Technology, focusing on Data Warehousing and Mining for the 2025-2026 academic year. It includes various questions categorized into parts A, B, and C, covering topics such as OLAP, data mining techniques, frequent pattern mining, and the differences between OLAP and OLTP systems. The document serves as a comprehensive guide for students to prepare for examinations in these subjects.

Uploaded by

anakinalt9
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
33 views9 pages

CSE Data Warehousing & Mining Question Bank

The document is a question bank for the B.E. CSE program at Hindusthan College of Engineering and Technology, focusing on Data Warehousing and Mining for the 2025-2026 academic year. It includes various questions categorized into parts A, B, and C, covering topics such as OLAP, data mining techniques, frequent pattern mining, and the differences between OLAP and OLTP systems. The document serves as a comprehensive guide for students to prepare for examinations in these subjects.

Uploaded by

anakinalt9
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Hindusthan College of Engineering and

Technology
An Autonomous Institution, Approved by AICTE, New Delhi Affiliated to Anna University
Accredited by NBA (AERO, AUTO, CIVIL, CSE, ECE, EEE, IT, MECH, MCT, AGRI, FOOD, MBA &
MCA) Accredited by NAAC with ‘A++’ Grade| An ISO Certified Institution
Valley Campus, Pollachi Highway, Coimbatore 641032.

QUESTION BANK
BRANCH: B.E. CSE SEMESTER: V
Course code & 22CS5371/DATA WAREHOUSING AND MINING Academic year: 2025-2026
Name:
UNIT – I: DATA WAREHOUSING, BUSINESS ANALYSIS AND ON-LINE ANALYTICAL
PROCESSING (OLAP)
Basic Concepts Data Warehousing Components Building a Data Warehouse Database Architectures for Parallel
Processing Parallel DBMS Vendors Multidimensional Data Model Data Warehouse Schemas for Decision Support,
Concept Hierarchies -Characteristics of OLAP Systems
(Regulation- 2019)
PART A (Answer All Questions) (10 X 2 = 20) BTL COs Marks
1. Define a data warehouse. 1 CO1 2

2. List any two components of a data warehouse. 1 CO1 2

3. What is the purpose of a data warehouse? 2 CO1 2

4. Name two database architectures used for parallel processing. 1 CO1 2

5. Explain the term 'multidimensional data model'. 2 CO1 2

6. What is a star schema? 2 CO1 2

7. what is a fact table? 2 CO1 2

8. What is a dimension table? 2 CO1 2

9. Give an example of a concept hierarchy. 2 CO1 2

10. List two characteristics of OLAP systems. 1 CO1 2

11. Define OLAP. 1 CO1 2

12. What is the difference between OLAP and OLTP? 2 CO1 2

13. State any two benefits of using a data warehouse for business analysis. 2 CO1 2

14. Explain 'parallel processing' in the context of databases. 2 CO1 2

15. Identify two parallel DBMS vendors. 1 CO1 2

16. Analyze the advantages of using a snowflake schema over a star schema in a data 4 CO1 2
warehouse.
17. Evaluate the importance of metadata in maintaining a data warehouse. 5 CO1 2

18. Design a simple data mart for a retail business. What tables and data would you 3 CO1 2
include?
19. Compare and contrast the use of OLAP cubes with traditional relational databases 4 CO1 2
for business analysis.
20. Propose a strategy for implementing parallel processing in a large-scale data 6 CO1 2
warehouse.

1|Page
PART B
21. Explain the basic concepts of data warehousing. What are the key components 2 CO1 14
involved in building a data warehouse?
Describe the major components of a data warehousing system. How do these 2 CO1 14
22.
components interact to support data analysis and reporting?
23. Discuss the steps involved in building a data warehouse. What challenges might you 3 CO1 14
encounter during this process, and how can they be addressed?
24. Explain the different database architectures used for parallel processing. How do 4 CO1 14
these architectures enhance the performance of a data warehouse?
25. Compare and contrast different parallel DBMS vendors. What are the key features 4 CO1 14
and advantages offered by these vendors?
26. Describe the multidimensional data model in the context of data warehousing. What 2 CO1 14
are its key components and how does it support OLAP operations?
27. Explain the different data warehouse schemas used for decision support. Provide 4 CO1 14
examples and discuss their advantages and disadvantages.
28. Define concept hierarchies in data warehousing. How are they constructed and 2 CO1 14
utilized in data analysis?
29. What are the characteristics of OLAP systems? How do these characteristics 4 CO1 14
facilitate complex analytical and ad-hoc queries with rapid execution times?
30. Evaluate the performance metrics used in data warehouses. How can these metrics 5 CO1 14
be used to optimize data warehouse operations and ensure high performance?
PART – C
(Use Higher Order Thinking)
31. Compare and contrast the benefits and limitations of using a snowflake schema 4 CO1 10
versus a star schema in a data warehouse. How do each of these schemas affect
query performance and ease of maintenance?
32. Evaluate the impact of metadata management on the efficiency and accuracy of a 5 CO1 10
data warehouse. How does effective metadata management contribute to data quality
and usability in decision-making processes?
33. Design a comprehensive data mart for a retail business, including tables and data. 6 CO1 10
Explain how your design addresses common business needs such as sales tracking,
inventory management, and customer analysis.
34. Analyze the trade-offs between using OLAP cubes and traditional relational 4 CO1 10
databases for business analysis. Consider factors such as query performance, ease of
use, and scalability in your analysis.
35. Propose a detailed strategy for implementing parallel processing in a large-scale data 5 CO1 10
warehouse. Evaluate the potential challenges and benefits of your proposed strategy,
and suggest ways to overcome the challenges.

2|Page
UNIT – II: Typical OLAP Operations, OLAP and OLTP. DATA MINING-INTRODUCTION
Introduction to Data Mining Systems - Knowledge Discovery Process - Data Mining Techniques Issues
applications- Data Objects and attribute types, Statistical description of data, Data Preprocessing-Cleaning,
Integration, Reduction, Transformation and discretization, Data Visualization, Data similarity and dissimilarity
measures.
(Regulation- 2019)
PART A (Answer All Questions) (10 X 2 = 20) BTL COs Marks
1. What is an OLAP operation? 1 CO2 2

2. What does OLTP stand for? 1 CO2 2

3. Define data mining. 1 CO2 2

4. What is the purpose of data preprocessing? 1 CO2 2

5. Name two types of data objects in data mining. 1 CO2 2

6. What does data cleaning involve? 1 CO2 2

7. What is data reduction in the context of data mining? 1 CO2 2

8. List two data visualization techniques. 1 CO2 2

9. Describe what is meant by data transformation. 2 CO2 2

10. Explain the difference between OLAP and OLTP systems. 2 CO2 2

11. What is the knowledge discovery process? 2 CO2 2

12. How does statistical description of data aid in data mining? 2 CO2 2

13. What is data discretization? 2 CO2 2

14. Explain the concept of data similarity measures. 2 CO2 2

15. How is data integration performed in data mining? 2 CO2 2

16. What is the role of data visualization in data mining? 2 CO2 2

17. Describe the purpose of data reduction. 2 CO2 2

18. Explain what is meant by data attribute types. 2 CO2 2

19. Given a dataset, how would you apply data cleaning techniques to prepare it for 3 CO2 2
analysis?
20. Compare and contrast the roles of OLAP and OLTP systems in a business 4 CO2 2
environment.
PART B (Answer All Questions) ( 5 X 14 = 70)
21. Describe the key differences between OLAP and OLTP systems. How do these 2 CO2 14
differences influence their design and functionality?
Analyze the impact of data preprocessing techniques, such as cleaning and 4 CO2 14
22. integration, on the quality of data mining results. How do these techniques affect the
overall accuracy and reliability of the analysis?
23. Design a data mining process for a company looking to enhance customer 6 CO2 14
segmentation. Include steps such as data collection, preprocessing, mining, and
interpretation. Justify your choices for each step.
24. 5 CO2 14
Evaluate the effectiveness of OLAP cubes versus traditional relational databases in

3|Page
handling complex queries. Discuss the strengths and weaknesses of each approach in
the context of business intelligence.
25. Explain the concept of a multidimensional data model and its role in OLAP systems. 2 CO2 14
How does this model support complex analytical queries?
26. Compare the advantages and disadvantages of using a star schema versus a 4 CO2 14
snowflake schema in a data warehouse. How do these schemas affect query
performance and ease of maintenance?
27. Propose a strategy for integrating multiple data sources into a single data warehouse. 6 CO2 14
Discuss how you would handle challenges such as data consistency and redundancy.
28. Evaluate the role of metadata in the data mining process. How does effective 5 CO2 14
metadata management improve data quality and facilitate more accurate insights
29. Define the knowledge discovery process in data mining. What are the key stages, 2 CO2 14
and how do they contribute to deriving meaningful insights from data?
30. Analyze the impact of different data similarity and dissimilarity measures on the 4 CO2 14
results of clustering algorithms. How do these measures influence the formation and
interpretation of clusters?
PART – C (1 X 10 = 10)
(Use Higher Order Thinking)
31. Compare and contrast the use of OLAP cubes and traditional relational databases for 4 CO2 10
handling complex analytical queries. Consider factors such as query performance,
ease of use, and data aggregation capabilities. How do these differences impact
decision-making processes in a business?
32. Evaluate the effectiveness of various data preprocessing techniques (cleaning, 5 CO2 10
integration, reduction, transformation) in improving the quality of data mining
results. Discuss how each technique contributes to the accuracy and efficiency of
data mining.
33. Design a data mining framework for analyzing customer purchase behavior in an e- 6 CO2 10
commerce platform. Include stages such as data collection, preprocessing, mining,
and interpretation. Justify your approach and discuss how each stage contributes to
understanding customer behavior and improving business strategies.
34. Analyze the advantages and disadvantages of using a star schema versus a snowflake 4 CO2 10
schema in a data warehouse. How do these schemas affect data query performance,
ease of maintenance, and data redundancy? Provide examples to support your
analysis.
35. Evaluate the role of metadata in enhancing data mining processes. Discuss how well- 5 CO2 10
managed metadata contributes to data quality, interpretability, and the overall
effectiveness of data mining. Provide examples to illustrate its impact on the mining
process.

UNIT – III: DATA MINING - FREQUENT PATTERN ANALYSIS


Mining Frequent Patterns, Associations and Correlations Mining Methods- Pattern Evaluation Method Pattern
Mining in Multilevel, Multi-Dimensional Space Constraint Based Frequent Pattern Mining, Classification using
Frequent Patterns.
(Regulation- 2019)
PART A (Answer All Questions) (10 X 2 = 20) BTL COs Marks
1. What is frequent pattern mining? 1 CO3 2

2. Define association rule mining. 1 CO3 2

3. What does the term 'correlation' mean in the context of data mining? 1 CO3 2

4. Name one method used for mining frequent patterns. 1 CO3 2

4|Page
5. What is the purpose of pattern evaluation in data mining? 1 CO3 2

6. What does 'multilevel pattern mining' refer to? 1 CO3 2

7. Define 'constraint-based frequent pattern mining'. 1 CO3 2

8. What is the role of classification in frequent pattern mining? 1 CO3 2

9. What is a support count in the context of association rules? 2 CO3 2

10. Name a commonly used algorithm for mining frequent patterns. 2 CO3 2

11. Explain the concept of support in frequent pattern mining. 2 CO3 2

12. How does pattern evaluation help in selecting significant patterns? 2 CO3 2

13. Describe what is meant by 'multi-dimensional space' in pattern mining. 2 CO3 2

14. What is the difference between frequent itemsets and association rules? 2 CO3 2

15. How does constraint-based mining improve pattern discovery? 2 CO3 2

16. Explain the term 'confidence' in association rule mining. 2 CO3 2

17. What are the challenges of mining patterns in multilevel spaces? 2 CO3 2

18. How does classification using frequent patterns help in predictive analytics? 2 CO3 2

19. Given a dataset with transactional data, describe the steps you would follow to mine frequent 3 CO3 2
patterns using the Apriori algorithm.
20. Analyze the impact of different support thresholds on the results of frequent pattern mining. 4 CO3 2
How does changing the support threshold affect the number and relevance of patterns
discovered?
PART B (Answer All Questions) ( 5 X 14 = 70)
21. Explain the concept of frequent pattern mining. How do support and confidence 2 CO3 14
contribute to discovering meaningful patterns in transactional data?
22. Apply the Apriori algorithm to a given transactional dataset. Outline the steps 3 CO3 14
involved and describe how you would interpret the frequent itemsets and association
rules generated.
23. Analyze the challenges and limitations of mining frequent patterns in a multi- 4 CO3 14
dimensional space. How do dimensionality and data complexity impact the
effectiveness of pattern mining algorithms?
24. Evaluate the effectiveness of different pattern evaluation methods (e.g., support, 5 CO3 14
confidence, lift) in assessing the significance of association rules. How do these
methods impact the quality and usefulness of the mined patterns?
25. Design a constraint-based frequent pattern mining approach for a retail dataset that 6 CO3 14
includes constraints based on item categories and customer demographics. How
would you ensure that the constraints improve the relevance and interpretability of
the patterns?
26. Describe the process of mining frequent patterns in a multilevel space. How does 2 CO3 14
multilevel pattern mining enhance the ability to discover patterns at different levels
of abstraction?
27. Given a dataset with hierarchical item categories, apply a pattern mining algorithm 3 CO3 14
to discover frequent itemsets at multiple levels of the hierarchy. Explain how the
hierarchical structure affects the mining process and results.
28. Analyze the differences between frequent pattern mining and correlation mining. 4 CO3 14
How do these methods complement each other in uncovering relationships within the
data?
29. Evaluate the impact of various support thresholds on the quality and quantity of 5 CO3 14
frequent patterns discovered. Discuss how choosing different thresholds can
influence the usefulness of the patterns for decision-making.

5|Page
30. Propose a classification strategy that leverages frequent patterns to predict customer 6 CO3 14
behavior in an e-commerce environment. How would you integrate frequent patterns
into your classification model, and what benefits would this approach provide?
PART – C (1 X 10 = 10)
(Use Higher Order Thinking)
31. Analyze the trade-offs between using the Apriori algorithm and the FP-Growth 4 CO3 10
algorithm for mining frequent patterns in large datasets. Consider factors such as
computational efficiency, memory usage, and ease of implementation. How does the
choice of algorithm impact the mining process?
32. Evaluate the effectiveness of various pattern evaluation metrics (such as support, 5 CO3 10
confidence, and lift) in determining the usefulness of association rules in a retail
dataset. Discuss how each metric contributes to understanding the relationships
between items and how it can influence business decisions.
33. Design a comprehensive data mining process for discovering frequent patterns in a 6 CO3 10
multi-dimensional dataset with constraints. Include steps such as data preprocessing,
pattern mining, and result evaluation. Explain how your approach addresses
challenges related to dimensionality and constraints.
34. Analyze the challenges of mining frequent patterns in a multilevel hierarchy (e.g., 4 CO3 10
product categories and sub-categories). How do different levels of abstraction affect
the mining results, and what strategies can be used to effectively handle hierarchical
data?
35. Evaluate the impact of different constraint-based mining approaches (e.g., item 5 CO3 10
constraints, user constraints) on the quality and relevance of discovered patterns.
How do constraints improve pattern discovery, and what potential drawbacks should
be considered?

UNIT – IV: CLASSIFICATION AND CLUSTERING


Decision Tree Induction Bayesian Classification Rule Based Classification Classification by Back Propagation -
Support Vector Machines. Clustering Techniques Cluster analysis-Partitioning Methods. Hierarchical Methods
Density Based Methods - Grid Based Methods Evaluation of clustering.
(Regulation- 2020)
PART A (Answer All Questions) (10 X 2 = 20) BTL COs Marks
1. What is a Decision Tree? 1 CO4 2

2. Mention one advantage of Decision Tree Induction. 1 CO4 2

3. Define Information Gain in the context of Decision Trees. 1 CO4 2

4. How is a Decision Tree pruned? 2 CO4 2

5. What is Bayesian Classification? 1 CO4 2

6. Define Bayes' Theorem. 1 CO4 2

7. What is a Naive Bayes Classifier? 1 CO4 2

8. Mention one limitation of Naive Bayes Classifier. 2 CO4 2

9. What is Rule-Based Classification? 1 CO4 2

10. Give an example of a rule in Rule-Based Classification. 3 CO4 2

11. What is Back Propagation in Neural Networks? 1 CO4 2

6|Page
12. Why is the learning rate important in Back Propagation? 2 CO4 2

13. What is the main goal of Support Vector Machines? 1 CO4 2

14. What is a support vector in SVM? 1 CO4 2

15. What is Clustering in data analysis? 1 CO4 2

16. Differentiate between classification and clustering. 4 CO4 2

17. What is the K-means algorithm? 1 CO4 2

18. How is the initial selection of centroids done in K-means? 2 CO4 2

19. What is Hierarchical Clustering? 1 CO4 2

20. Mention one advantage of Hierarchical Clustering. 2 CO4 2

21. What is DBSCAN in clustering? 1 CO4 2

22. Mention one advantage of DBSCAN. 2 CO4 2

23. What is a Grid-Based Clustering method? 1 CO4 2

24. Give an example of a Grid-Based Clustering algorithm. 1 CO4 2

25. What is Silhouette Coefficient? 1 CO4 2

26. Mention one method to evaluate clustering performance. 2 CO4 2

PART B (Answer All Questions) ( 5 X 14 = 70)


27. Explain the process of constructing a decision tree using the ID3 algorithm. Include 2 CO4 14
how information gain is calculated and used to select the best attribute at each node.
Describe the Naive Bayes classification algorithm. Explain its assumptions, the 2 CO4 14
28.
training process, and how classification is performed on new data.
29. Develop a rule-based classification system for a given dataset. Outline the steps from 3 CO4 14
rule extraction to classification, and discuss how rules are evaluated for accuracy.
30. Explain how back propagation works in neural networks. Include the mathematical 2 CO4 14
formulation, the role of the learning rate, and techniques to avoid overfitting.
31. Analyze the concept of Support Vector Machines (SVM). Explain how SVM finds the 4 CO4 14
optimal hyperplane for classification, and discuss the role of kernels in SVM.
32. Compare and contrast partitioning methods and hierarchical methods for clustering. 4 CO4 14
Highlight their advantages, disadvantages, and appropriate use cases.
33. Explain the K-means clustering algorithm. Discuss its steps, how the initial centroids 2 CO4 14
affect the results, and methods to evaluate the quality of clusters.
34. Describe agglomerative hierarchical clustering. Explain the different linkage criteria 2 CO4 14
used and how they influence the resulting dendrogram.
35. Explain the DBSCAN clustering algorithm. Discuss its parameters, advantages over 2 CO4 14
K-means, and how it handles noise and outliers.
36. Evaluate different methods for assessing the quality of clustering results. Discuss 4 CO4 14
internal and external evaluation measures, providing examples of each.
PART – C (1 X 10 = 10)
(Use Higher Order Thinking)
37. Analyze the impact of pruning in decision tree induction. Evaluate the trade-offs 5 CO4 10
between a fully grown tree and a pruned tree in terms of accuracy and overfitting.
38. Design a spam email classifier using the Naive Bayes algorithm. Critically evaluate 6 CO4 10
its performance and limitations, suggesting improvements for handling non-
independent features.
39. Critique the use of Support Vector Machines for image classification tasks. Analyze 5 CO4 10
the role of kernel functions and discuss the advantages and disadvantages of different
kernels.

7|Page
40. Compare and contrast the effectiveness of K-means and DBSCAN for clustering 4 CO4 10
customer data in a retail business. Provide a detailed analysis of their performance,
handling of noise, and computational efficiency.
41. Develop a comprehensive framework for evaluating clustering algorithms in a real- 6 CO4 10
world scenario. Include various internal and external evaluation metrics, and critically
assess their effectiveness.

UNIT – V: MINING OBJECT, SPATIAL, MULTIMEDIA, TEXT AND WEB DATA


Multidimensional Analysis and Descriptive Mining of Complex Data Objects - Spatial Data Mining Multimedia
Data Mining-Text Mining - Mining the World Wide Web Applications and Trends in Data Mining.
(Regulation- 2020)
PART A (Answer All Questions) (10 X 2 = 20) BTL COs Marks
1. What is multidimensional analysis in data mining? 1 CO5 2

2. Define a data cube. 1 CO5 2

3. What are complex data objects? 1 CO5 2

4. Explain the concept of a data warehouse. 2 CO5 2

5. What is spatial data mining? 1 CO5 2

6. Mention one application of spatial data mining. 2 CO5 2

7. What is spatial autocorrelation? 1 CO5 2

8. Differentiate between spatial and non-spatial data. 4 CO5 2

9. What is multimedia data mining? 1 CO5 2

10. Give an example of a multimedia data mining application. 2 CO5 2

11. What is feature extraction in multimedia data mining? 1 CO5 2

12. How is similarity measured in multimedia data mining? 3 CO5 2

13. Define text mining. 1 CO5 2

14. What is natural language processing (NLP)? 1 CO5 2

15. Explain the term ‘sentiment analysis’ in text mining. 2 CO5 2

16. What is tokenization in text mining? 1 CO5 2

17. What is web mining? 1 CO5 2

18. Differentiate between web content mining and web structure mining. 4 CO5 2

19. What is a web crawler? 2 CO5 2

20. Explain the role of PageRank in web mining. 2 CO5 2

PART B (Answer All Questions) ( 5 X 14 = 70)


21. Explain the concept of multidimensional analysis in data mining. Discuss how data 2 CO5 14
cubes are used in this context and provide an example illustrating their application.
Describe the process of descriptive mining of complex data objects. Explain how it 2 CO5 14
22.
differs from traditional data mining techniques.
23. Analyze the challenges and techniques involved in spatial data mining. Provide 4 CO5 14
examples of spatial data mining applications in real-world scenarios.
24. Explain the role of spatial autocorrelation in spatial data mining. Discuss methods for 2 CO5 14
measuring spatial autocorrelation and their implications.

8|Page
25. Describe the process of multimedia data mining. Discuss the challenges and 2 CO5 14
techniques associated with mining multimedia data, providing examples for each type
of multimedia data.
26. Analyze the importance of feature extraction in multimedia data mining. Explain 4 CO5 14
various feature extraction techniques used for different types of multimedia data.
27. Explain the steps involved in text mining. Discuss various techniques used for text 2 CO5 14
preprocessing and their significance in the text mining process.
28. Analyze the role of natural language processing (NLP) in text mining. Discuss key 4 CO5 14
NLP techniques and their applications in extracting meaningful information from text
data.
29. Describe the different types of web mining. Compare and contrast web content 4 CO5 14
mining, web structure mining, and web usage mining, providing examples of each.
30. Explain the concept of PageRank and its role in web mining. Discuss how PageRank 2 CO5 14
is calculated and its impact on search engine optimization (SEO).
PART – C (1 X 10 = 10)
(Use Higher Order Thinking)
31. Evaluate the effectiveness of using data cubes for multidimensional analysis in large-scale 5 CO5 10
data warehouses. Discuss potential challenges and propose solutions to optimize their
performance.
32. Critically analyze the role of spatial data mining in urban planning. Discuss specific 5 CO5 10
techniques and their application to solve real-world urban planning issues, such as
traffic management or land use planning.
33. Design a framework for multimedia data mining in the context of social media 6 CO5 10
analysis. Include techniques for feature extraction, pattern recognition, and sentiment
analysis, and evaluate the potential impact of your framework.
34. Propose a text mining approach to analyze customer reviews for product 6 CO5 10
improvement. Discuss the techniques for text preprocessing, sentiment analysis, and
topic modeling, and critically evaluate the effectiveness of this approach.
35. Develop a comprehensive strategy for web mining to improve personalized 6 CO5 10
recommendations on an e-commerce platform. Analyze the integration of web content
mining, web structure mining, and web usage mining, and evaluate the overall
effectiveness.

Course coordinator HOD

9|Page

You might also like