0% found this document useful (0 votes)
2 views1 page

Short Notes Data Science

The document provides an overview of data science, including its processes and the big data ecosystem, highlighting tools like Hadoop and Spark. It covers machine learning types and techniques, emphasizing supervised and unsupervised learning, feature engineering, and model validation using Python's Scikit-learn. Additionally, it introduces NoSQL concepts and types, focusing on their characteristics and principles such as ACID and CAP.

Uploaded by

siddhardhar471
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views1 page

Short Notes Data Science

The document provides an overview of data science, including its processes and the big data ecosystem, highlighting tools like Hadoop and Spark. It covers machine learning types and techniques, emphasizing supervised and unsupervised learning, feature engineering, and model validation using Python's Scikit-learn. Additionally, it introduces NoSQL concepts and types, focusing on their characteristics and principles such as ACID and CAP.

Uploaded by

siddhardhar471
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

UNIT 1: INTRODUCTION TO DATA SCIENCE

- Data Science: Extracting knowledge from data


- Facets of Data: Volume, Velocity, Variety, Veracity, Value
- DS Process: Problem -> Collection -> Cleaning -> Integration -> EDA -> Feature Engg ->
Modeling -> Evaluation -> Deployment
- Big Data Ecosystem: Hadoop, HDFS, Spark, NoSQL
- EDA used to find patterns, trends, outliers
UNIT 2: MACHINE LEARNING IN DS
- Types of ML: Supervised, Unsupervised, Semi-supervised, Reinforcement
- Supervised: Regression, Classification
- Unsupervised: Clustering, PCA
- Feature Engineering: Selection, Extraction, Scaling
- Model Validation: Train-test split, Cross-validation
- Python Tool: Scikit-learn
- Handling Large Data: Sampling, Distributed processing
- Case Studies: Malicious URL Detection, Recommender System
UNIT 3: BIG DATA & NoSQL (HALF)
- Hadoop: Distributed Storage and Processing
- ACID: Atomicity, Consistency, Isolation, Durability
- CAP: Consistency, Availability, Partition Tolerance
- BASE: Basically Available, Soft State, Eventual Consistency
- NoSQL Types: Key-Value, Document, Column Family, Graph

You might also like