0% found this document useful (0 votes)
6 views5 pages

Data Mining Concepts and Techniques

Uploaded by

susandiya12
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views5 pages

Data Mining Concepts and Techniques

Uploaded by

susandiya12
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

DSE : Data Mining & Knowledge Discovery (Sem- V)

CHAPTER-2
Data
(Test Questionnaire)
ONE MARK QUESTIONS

1. Briefly explain the concept of a data set in data mining with the help of an example.
2. Give three synonyms to refer a record in a data set.
3. Give three synonyms to refer an attribute in a data set.
4. State any two examples of nominal/ordinal/interval/ratio attributes.
5. State any two statistical operations that can be applied on nominal/ordinal/interval/ratio
attributes.
6. What do you mean by a countably infinite set? Give an example.
7. Give one example of a binary ordinal attribute.
8. Give one example of a discrete nominal attribute.
9. Give one example of discrete asymmetric attribute.
10. Give one example of continuous transaction data set.
11. State one drawback of record data.
12. Briefly differentiate between data cleaning and data preprocessing.
13. State any three data collection issues.
14. Differentiate between noise and an outlier.
15. Define measurement error along with an example.
16. Define collection error along with an example.
17. What do you mean by legitimate duplicates?
18. Briefly explain the term of deduplication.
19. Briefly explain the disadvantage of using the aggregation approach for pre-processing the
data before data mining.
20. Give any two examples of high dimensional datasets.
21. Differentiate between discretization & binarization in data mining.
22. Differentiate between supervised and unsupervised discretization.
23. Explain the term normalization. State the name of any two techniques that can be used
for normalization in data pre-processing.

TWO MARKS QUESTIONS

24. State the definition of an attribute in a data set. And, then briefly explain the same with
the help of an example.
25. State the definition of measurement scale of an attribute in a data set. And, then briefly
explain the same with the help of an example.
26. “Properties of an attribute need not be the same as the properties of the values used to
measure it.” Justify the statement with the help of an example.
27. Explain different types of attributes categorized based on the permissible transformations
for each type.
28. Explain following terms/ Give one example of each of the following attribute type:
a. Categorical/Qualitative attribute
b. Numeric/Quantitative attribute
c. Discrete attribute
d. Continuous attribute
e. Asymmetric attribute
f. Curse of dimensionality
g. Dimensionality reduction
h. Market-basket data
i. Document-term matrix
j. Temporal auto-correlation
k. Spatial auto-correlation
29. Briefly justify the reasons for the fact that-“Temperatures measured at Kelvin scale fall
into the category of ratio type attribute values, whereas temperatures measured at
Celsius (or Fahrenheit) scale are considered interval type attribute values”.
30. In the following picture, which type of data set is being represented? Explain in detail.

31. Differentiate between sequential data and sequence data with the help of examples.
32. Differentiate between spatial data and time-series data with the help of examples.
33. Let’s assume that you have been given a set of 20 chemical compounds, with a list of 10
substructures which is common to all those 20 compounds. Represent this graph-based
data in the form of record data.
34. In the following picture, which type of data set is being represented? Explain in detail.

35. Explain the term artifact with the help of an example.


36. State any two scenarios of inconsistent values.
37. Which type of data quality issue has been represented in the following diagram? Explain
it.

38. Briefly explain any two data quality issues from the application standpoint.
39. State the names of any four pre-processing techniques. Explain any one of them in detail.
40. Define the term Sampling. What do you mean by a representative sample in sampling?
41. Explain the concept of progressive sampling with the help of an example.
42. Differentiate between dimensionality reduction algorithms and feature subset selection
algorithms in data pre-processing.
43. Briefly explain any two techniques for dimensionality reduction in data pre-processing.
44. In context of feature subset selection, briefly explain redundant features and irrelevant
features in a typical data set with the help of examples.
45. Briefly explain the process of Feature subset selection with the help of a diagram.
46. Explain the term Feature Weighting.
47. A data set contains a categorical attribute named “Product Color” which have three
possible values, namely {Red, Blue, Green}. Perform one-hot encoding to convert this
categorical attribute into binary attribute.
48. Briefly explain any two techniques of normalization.

THREE MARKS QUESTIONS

49. Explain different types of attributes categorized based on the permissible statistical
operations for each type. Also give at least one example for each attribute type.
50. Explain the three general features of a typical data set.
51. Differentiate among record data, graph-based data and ordered data. Also provide one
example of each data type.
52. Explain the terms precision, bias and accuracy. Also explain the relationship amongst
them with the help of an example.
53. Define missing values for a typical data set in data mining. Also explain three methods to
deal with the missing values.
54. Explain the concept of Aggregation in detail with the help of an example.
55. Discuss any three advantages of the Aggregation approach.
56. Discuss three sampling approaches in detail.
57. Explain the Curse of dimensionality in data mining with the help of an example.
58. Explain the three strategies used for feature subset selection with examples.
59. Differentiate between equal-width discretization and equal-depth discretization with the
help of a common example.
60. Explain the term “Variable transformation”. Explain any two types of variable
transformations.

FIVE MARKS QUESTIONS

61. What is feature creation? How it is different from dimensionality reduction algorithms?
Explain any two techniques used for feature creation in data pre-processing with
examples.
62. Numericals based on similarity and dissimarity particularly Euclidean distance.
63. Numericals based on discretization and binarization.

You might also like