0% found this document useful (0 votes)
2 views2 pages

Data Science Important Questions

The document outlines assignment questions related to Data Science, covering topics such as the Data Science process, challenges of big data, machine learning techniques, and the Hadoop ecosystem. It includes questions on exploratory data analysis, the significance of presenting models, and the role of NoSQL databases. Additionally, it features tutorial questions that emphasize the importance of big data in Data Science and the use of Python tools for data analysis.

Uploaded by

syedtanisha1606
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views2 pages

Data Science Important Questions

The document outlines assignment questions related to Data Science, covering topics such as the Data Science process, challenges of big data, machine learning techniques, and the Hadoop ecosystem. It includes questions on exploratory data analysis, the significance of presenting models, and the role of NoSQL databases. Additionally, it features tutorial questions that emphasize the importance of big data in Data Science and the use of Python tools for data analysis.

Uploaded by

syedtanisha1606
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

VASIREDDY VENKATADRI INSTITUTE OF TECHNOLOGY

(AUTONOMOUS)
Accredited by NBA, Approved by AICTE, Permanently Affiliated to JNTUK
NAAC Accredited with ‘A’ Grade, ISO 9001:2015 Certified,
Nambur (V), Pedakakani (M), Guntur (Dt.), Andhra Pradesh – 522 508, [Link]
INTRODCTION TO DATA SCIENCE
Assignment Questions
Unit-1
1 Outline the steps involved in the Data Science process, from defining goals
and creating a project charter to presenting findings and building applications
on top of them. Provide a practical example for each step.
2 Identify the primary challenges in managing big data and describe how
Python-based solutions can mitigate these challenges.
3 Briefly explain the characteristics of Big data.
4 Discuss the different facets of data in Data Science. How do structured, semi-
structured, and unstructured data play a role in the data analysis process?

5 Explain the key challenges of big data processing and the role of Python tools
in overcoming them.
6 Briefly explain the steps in the data science process with neat diagram.
7 What is the significance of exploratory data analysis (EDA) in the Data Science
process? How would you use EDA to identify patterns and relationships in a
dataset? Discuss the methods and tools that can be used in this phase.

8 Briefly explain the significance of presenting and Automation of the model in


data science process.
9 Explain the role of NoSQL databases in big data ecosystems in detail.
Unit-2

1 Classify and explain the each type of machine learning techniques.


2 Write a sort note on the modeling process in machine learning.
3 Briefly explain general programming tips for dealing with large data sets.
4 Define machine learning (ML) and explain its role in data science (DS). List
and briefly describe the applications of machine learning.
5 Discuss the challenges involved in handling large datasets and optimizing
data processing workflows.
6 Compare supervised, unsupervised, and semi-supervised learning in machine
learning. Evaluate the advantages and limitations of semi-supervised learning
in handling real-world datasets where labeled data is scarce.

7 Explain the major approaches for addressing challenges associated with large-
scale data.
8 Define Machine Learning (ML) and explain its role in Data science (DS). List
and briefly describe the types of machine learning used in DS applications.

9 How is a machine learning model validated after training? Explain with


suitable examples.
10 Where machine learning is used in the data science process? Explain.
11 Discuss the python packages for working with data in memory.
Unit-3

1 Describe the core components of the Hadoop ecosystem and their


functions.
2 Explain how spark replaces map reduce for better performance.
3 Discuss the step-by-step process of MapReduce with a simple example,
and explain its limitations.
4 Explain the Hadoop framework and its components. How does Hadoop
enable distributed data storage and processing?
5 What are the core components of the Hadoop ecosystem? Explain
briefly.
6 Explain how Apache Spark improves the performance of data processing
frameworks with suitable examples.

****

Tutorial Questions
1 How can an understanding of the Big Data ecosystem—specifically distributed
file systems, machine learning frameworks, and NoSQL databases—enable
organizations to efficiently store, process, and analyze large-scale data for
better insights and decision-making?
2 Why is big data important in Data Science? Explain how the big data
ecosystem enables data science workflows and discuss the challenges involved.
3. Analyze the importance of data cleansing and transformation in the Data
Science process. Why is it necessary to integrate and transform data before
starting exploratory analysis and model building?
4 Evaluate the role of big data in Data Science. How does the big data ecosystem
support the data science process, and what challenges arise when working
with big data?
5 Highlight the Python tools and libraries (like scikit-learn) that would be useful
in data science.
6 Design and develop a simple movie recommender system using locality
sensitive Hashing (LSH) and Hamming Distance.. Explain the steps you would
follow, and provide code implementation for the system.
7. You are tasked with building a model to predict the likelihood of URL being
malicious or not. Outline the steps in the machine learning workflow.
8. Design and implement a machine learning model to predict malicious URLs.
Explain the steps you would take, the dataset you would use, and the features
you would extract. Provide a code example if applicable.
9 Why is Apache Spark employed for the preparation and transformation of
large-scale loan data, and how does Hive facilitate efficient storage and
querying of the transformed data? Justify your answer.
10 Explain the importance of clearly defining the research goal in a data analytics
project. How does retrieving data from external sources and uploading it to
HDFS support the achievement of this goal? Justify your answer.
*****

You might also like