UNIT 1: INTRODUCTION TO DATA SCIENCE
- Data Science: Extracting knowledge from data
- Facets of Data: Volume, Velocity, Variety, Veracity, Value
- DS Process: Problem -> Collection -> Cleaning -> Integration -> EDA -> Feature Engg ->
Modeling -> Evaluation -> Deployment
- Big Data Ecosystem: Hadoop, HDFS, Spark, NoSQL
- EDA used to find patterns, trends, outliers
UNIT 2: MACHINE LEARNING IN DS
- Types of ML: Supervised, Unsupervised, Semi-supervised, Reinforcement
- Supervised: Regression, Classification
- Unsupervised: Clustering, PCA
- Feature Engineering: Selection, Extraction, Scaling
- Model Validation: Train-test split, Cross-validation
- Python Tool: Scikit-learn
- Handling Large Data: Sampling, Distributed processing
- Case Studies: Malicious URL Detection, Recommender System
UNIT 3: BIG DATA & NoSQL (HALF)
- Hadoop: Distributed Storage and Processing
- ACID: Atomicity, Consistency, Isolation, Durability
- CAP: Consistency, Availability, Partition Tolerance
- BASE: Basically Available, Soft State, Eventual Consistency
- NoSQL Types: Key-Value, Document, Column Family, Graph