Skip to content

Latest commit

 

History

20 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation


Python Jupyter scikit-learn TensorFlow Docker MLflow License



65+ Notebooks · 50+ Handwritten Theory PDFs · 5 Deployment-Ready Projects

A structured, hands-on journey through the entire data science stack — from linear regression to transformers.


📚 Modules   🛠️ Tech Stack   🚀 Get Started   🏗️ Projects   🤝 Contribute


 About This Repository

A comprehensive, self-contained curriculum covering the complete data science and machine learning pipeline — designed for learners who want depth, not just breadth.

📝 Handwritten Theory Notes Deep mathematical intuition and derivations across 50+ PDFs
💻 Practical Notebooks 65+ hands-on implementations with real-world datasets
🚀 Deployment Projects End-to-end ML pipelines deployed on AWS, Azure & Docker
🧪 MLOps Tooling Experiment tracking, model versioning & serving with MLflow, DVC & BentoML

The curriculum spans 28 structured modules organized in a progressive learning path — from foundational concepts through advanced deep learning and transformer architectures.


🗂️ Repository Structure

Complete-Data-Science-With-Machine-Learning-And-NLP/
│
├── 📁 02-Introduction/                          # Foundations: AI vs ML vs DL
├── 📁 03-Complete Linear Regression/            # Simple, Multiple, Polynomial
├── 📁 04-Ridge Lasso And Elasticnet/            # Regularization & Cross-Validation
├── 📁 05-Step By Step Project Implementation/   # ML Project Lifecycle
├── 📁 06-Logistic Regression/                   # Binary & Multiclass Classification
├── 📁 07-SVM/                                   # Support Vector Machines & Kernels
├── 📁 08-Naive Bayes/                           # Probabilistic Classification
├── 📁 09-K Nearest Neighbor/                    # Instance-Based Learning (KNN)
├── 📁 10-Decision-Tree/                         # Tree-Based Methods
├── 📁 11-Random Forest/                         # Bagging Ensemble Methods
├── 📁 12-Adaboost/                              # Adaptive Boosting
├── 📁 13-Gradient Boosting/                     # Gradient-Based Boosting
├── 📁 14-XgBoost/                               # Extreme Gradient Boosting
├── 📁 16-PCA/                                   # Principal Component Analysis
├── 📁 17-K-Mean/                                # K-Means Clustering
├── 📁 18-Hierarchical Cluster/                  # Agglomerative & Divisive
├── 📁 19-DBSCAN Clustering/                     # Density-Based Clustering
├── 📁 20-Silhouette Clustering/                 # Cluster Validation
├── 📁 21-Anomaly Detection ML/                  # Isolation Forest & LOF
├── 📁 22-Dockers/                               # Docker & Docker Compose
├── 📁 23-Git And Github/                        # Version Control
├── 📁 24-End To End ML Project With Deployment/ # AWS, Azure & Local Deployment
├── 📁 25-MLFlow Dagshub and BentoML/            # MLOps: Tracking & Serving
├── 📁 26-Complete NLP For Machine Learning/     # Text Processing & Embeddings
├── 📁 27-Deep Learning Bonus/                   # ANN, CNN, RNN, LSTM, GRU
├── 📁 28-Transformer/                           # Attention & Transformers
├── 📄 LICENSE
└── 📄 README.md

📚 Curriculum & Module Breakdown

Part I

📘 Module 02 · Introduction to Data Science & Machine Learning
Content Type Description
AI vs ML vs DL vs Data Science 📄 PDF Comparative study of fields and their intersections
Types of ML Techniques 📄 PDF Supervised, unsupervised, reinforcement learning overview
Instance-Based vs Model-Based Learning 📄 PDF Fundamental learning paradigm differences
Equation of a Straight Line & Hyperplane 📄 PDF Mathematical prerequisites for ML algorithms
📘 Module 03 · Complete Linear Regression

Theory (7 PDFs): Simple & Multiple Linear Regression, OLS Derivation, Polynomial Regression, Performance Metrics (MSE, RMSE, MAE, R²), Overfitting & Underfitting

Practicals (12 Notebooks):

Notebook Key Topics
Simple Linear Regression.ipynb Line fitting, gradient descent, cost function
Multiple Linear Regression.ipynb Multi-feature regression, economics dataset
Polynomial Regression Implementation.ipynb Non-linear curve fitting
Linear Regression ML Implementation.ipynb Complete sklearn pipeline
Model Training.ipynb Train/test split, model evaluation
Handling Imbalance Dataset.ipynb SMOTE, undersampling, oversampling
Ridge, Lasso Regression.ipynb Regularization comparison

Datasets: Height-Weight, Economic Index, Algerian Forest Fires

📘 Module 04 · Ridge, Lasso & ElasticNet Regression

Theory (2 PDFs): Ridge, Lasso & ElasticNet mathematical formulations; Types of Cross-Validation (K-Fold, Stratified, Leave-One-Out)

Notebook Key Topics
Ridge, Lasso Regression.ipynb L1/L2 regularization, hyperparameter tuning
Model Training.ipynb Cross-validation strategies, model selection
📘 Module 05 · ML Project Lifecycle — Step-by-Step Implementation
Notebook Key Topics
Simple Linear Regression.ipynb Foundation implementations
Multiple Linear Regression.ipynb Multi-variate analysis
EDA And FE Algerian Forest Fires.ipynb Exploratory data analysis, feature engineering
Model Training.ipynb Complete pipeline: preprocessing → training → evaluation

Includes: Flask web application (application.py), AWS Elastic Beanstalk configuration, trained model artifacts (.pkl files), and HTML templates for prediction interface.

📘 Module 06 · Logistic Regression

Theory (3 PDFs): Logistic Regression mathematics, Sigmoid function, Performance Metrics (Confusion Matrix, Precision, Recall, F1-Score, AUC-ROC), One-vs-Rest (OVR) Classification

Notebook Key Topics
Logistic Regression Implementation.ipynb Binary/multiclass classification, decision boundaries, evaluation metrics
📘 Module 07 · Support Vector Machines (SVM)

Theory (3 PDFs): SVC mathematical formulation, Support Vector Regression (SVR), SVM Kernels (Linear, RBF, Polynomial, Sigmoid)

Notebook Key Topics
Basic SVC Implementation.ipynb Linear SVM classification
SVM Kernels Implementation.ipynb Kernel trick, non-linear decision boundaries
Support Vector Regression Implementation.ipynb SVR for continuous prediction
📘 Module 08 · Naive Bayes Classifier

Theory (3 PDFs): Bayes' theorem derivation, Naive Bayes in-depth intuition, Variants (Gaussian, Multinomial, Bernoulli)

Notebook Key Topics
Naive Bayes Implementation.ipynb All three variants with real-world classification
📘 Module 09 · K-Nearest Neighbors (KNN)

Theory (2 PDFs): KNN classification & regression, Distance metrics (Euclidean, Manhattan, Minkowski), KD-Tree & Ball Tree optimization

Notebook Key Topics
KNNClassifier.ipynb Classification with optimal K selection
KNNRegressor.ipynb Regression using distance-weighted neighbors
📘 Module 10 · Decision Trees
Notebook Key Topics
Decision Tree Classifier Practical Implementation.ipynb Gini, entropy, information gain, tree visualization
Diabetes Prediction Using Decision Tree Regressor.ipynb Medical dataset regression with pruning

Part II

📗 Module 11 · Random Forest
Notebook Key Topics
Random Forest Classification Implementation.ipynb Bagging, feature importance, AUC-ROC visualization
Random Forest Regression Implementation.ipynb Ensemble regression, out-of-bag error

Datasets: Travel dataset for classification, Boston-style data for regression

📗 Module 12 · AdaBoost
Notebook Key Topics
Adaboost Classification Implementation.ipynb Adaptive boosting, sample weight updates
Adaboost Regression Implementation.ipynb Boosted regression trees
📗 Module 13 · Gradient Boosting
Notebook Key Topics
GradientBoost Classification Implementation.ipynb Gradient descent in function space
Gradientboost Regression Implementation.ipynb Residual learning, learning rate tuning
📗 Module 14 · XGBoost
Notebook Key Topics
XgboostBoost Classification Implementation.ipynb Regularized boosting, early stopping, feature importance
Xgboost Regression Implementation.ipynb XGBoost hyperparameter optimization

Part III

📙 Module 16 · Principal Component Analysis (PCA)
Notebook Key Topics
Principal Component Analysis (PCA) Implementation.ipynb Dimensionality reduction, explained variance ratio, scree plots
📙 Module 17 · K-Means Clustering
Notebook Key Topics
K Means Clustering.ipynb Elbow method, inertia, cluster visualization, centroid convergence
📙 Module 18 · Hierarchical Clustering
Notebook Key Topics
Hierarichal Clustering Implementation.ipynb Dendrograms, agglomerative clustering, linkage methods
📙 Module 19 · DBSCAN Clustering

Theory (1 PDF): DBSCAN algorithm — core points, border points, noise, epsilon & minPts

Notebook Key Topics
DBSCAN Implementation.ipynb Density-based clustering, handling arbitrary-shaped clusters
📙 Module 20 · Silhouette Analysis

Theory (1 PDF): Silhouette score interpretation, optimal cluster count validation

📙 Module 21 · Anomaly Detection

Theory (2 PDFs): Isolation Forest algorithm, Local Outlier Factor (LOF)

Notebook Key Topics
Isolation Anamoly Detection.ipynb Tree-based anomaly detection on healthcare data
DBSCAN Implementation.ipynb DBSCAN-based outlier detection

Datasets: Healthcare anomaly detection dataset


Part IV

🔧 Module 22 · Docker for Machine Learning

Theory (2 PDFs): Docker fundamentals, Dockerfile best practices, Docker Compose multi-container orchestration

Practicals: Hello World containerized application

🔧 Module 23 · Git & GitHub

Theory (1 PDF): Git workflow, branching strategies, collaboration best practices

🚀 Module 24 · End-to-End ML Projects with Deployment

Three complete, production-ready ML projects with full CI/CD pipelines:

Project Stack Deployment
Student Performance Predictor Flask, scikit-learn, CatBoost Local + Docker
Student Performance — Azure Flask, scikit-learn Azure Web App + GitHub Actions CI/CD
Student Performance — AWS Flask, scikit-learn AWS Elastic Beanstalk + AWS CodePipeline

Each project includes:

  • 📊 EDA & Model Training notebooks
  • 🏗️ Modular src/ package (data ingestion, transformation, model trainer, prediction pipeline)
  • 🌐 Flask web app with HTML templates
  • 📦 setup.py for package installation
  • 🐳 Dockerfile for containerization
  • ⚡ CI/CD workflows (.github/workflows/)
🔧 Module 25 · MLOps — MLflow, DagsHub, BentoML & DVC
Tool Content
MLflow Experiment tracking, model registry, metric logging (app.py with ElasticNet)
BentoML Model serving, API creation (service.py), model packaging (bentofile.yaml)
DVC Data version control, pipeline management

Part V

📕 Module 26 · Complete NLP for Machine Learning

Theory (1 PDF): Comprehensive NLP concepts (130+ pages)

Practicals (12 Notebooks):

# Notebook Key Topics
4 Tokenization Example Using NLTK.ipynb Word & sentence tokenization
5 Stemming And Its Types.ipynb Porter, Lancaster, Snowball stemmers
6 Lemmatization.ipynb WordNet lemmatizer, POS-aware lemmatization
7 Stopwords With NLTK.ipynb Stopword removal, custom stopword lists
8 Parts Of Speech Tagging.ipynb POS tagging with NLTK
9 Named Entity Recognition.ipynb NER with spaCy/NLTK
15 Bag Of Words Practical's.ipynb CountVectorizer, document-term matrix
16 TF-IDF Practical.ipynb Term frequency-inverse document frequency
26 Word2vec Practical Implementation.ipynb Word embeddings, Gensim Word2Vec
27 Spam Ham Classification (TF-IDF).ipynb Project: SMS spam detection pipeline
28-29 Spam Ham (Word2Vec, AvgWord2Vec).ipynb Project: Spam detection with embeddings
30-31 Kindle Review Sentiment Analysis.ipynb Project: Sentiment analysis on Kindle reviews

Part VI

📕 Module 27 · Deep Learning (Bonus)

Theory (11 PDFs):

Topic Content
Deep Learning Fundamentals Perceptron, multi-layer networks, forward propagation
Activation Functions Sigmoid, ReLU, Leaky ReLU, Tanh, Softmax, GELU
Loss Functions MSE, Cross-Entropy, Hinge Loss, Focal Loss
Optimizers SGD, Momentum, RMSProp, Adam, AdaGrad, AdaDelta
Weight Initialization Xavier, He, LeCun initialization techniques
CNNs Convolutional layers, pooling, architectures (VGG, ResNet)
RNNs Simple RNN, backpropagation through time, vanishing gradients
LSTM & GRU Long short-term memory, gated recurrent units
Bidirectional RNN Bi-directional sequence modeling

Practicals — RNN Project (IMDB Sentiment Analysis):

File Description
simplernn.ipynb Building RNN from scratch for IMDB classification
embedding.ipynb Word embeddings for sequence models
prediction.ipynb Inference pipeline
main.py Streamlit web application
simple_rnn_imdb.h5 Pre-trained model weights

Practicals — LSTM Project (Next Word Prediction):

File Description
experiemnts.ipynb LSTM model training on Shakespeare's Hamlet
app.py Streamlit app for next-word prediction
next_word_lstm.h5 Trained LSTM model
tokenizer.pickle Fitted tokenizer for inference
📕 Module 28 · Transformers
Topic Content
Encoder-Decoder & Seq2Seq Sequence-to-sequence architecture fundamentals
Attention Mechanism Self-attention, multi-head attention, scaled dot-product
Transformers Complete transformer architecture ("Attention Is All You Need")
Claude Architecture Modern LLM architecture analysis

🛠️ Technology Stack

Domain Technologies
Core ML & Data Science Python · NumPy · Pandas · Matplotlib · Seaborn · Plotly · scikit-learn · SciPy · StatsModels
Deep Learning TensorFlow · Keras · RNN · LSTM · GRU · CNN · Word Embeddings · Transformers
NLP Libraries NLTK · spaCy · Gensim (Word2Vec) · CountVectorizer · TF-IDF
Boosting XGBoost · CatBoost · AdaBoost · GradientBoosting
MLOps & Deployment MLflow · DagsHub · DVC · BentoML · Docker · Flask · Streamlit
Cloud & CI/CD AWS Elastic Beanstalk · AWS CodePipeline · Azure Web Apps · GitHub Actions
DevOps Docker · Docker Compose · Git · GitHub · Virtual Environments
Notebooks & IDEs Jupyter Notebook · VS Code · Google Colab Compatible

🚀 Getting Started

Prerequisites

Requirement Version
Python 3.8+
pip Latest
Jupyter Notebook or JupyterLab
Git 2.x+

Installation

# 1️⃣  Clone the repository
git clone https://github.com/Marmikgaj/Complete-Data-Science-With-Machine-Learning-And-NLP.git
cd Complete-Data-Science-With-Machine-Learning-And-NLP

# 2️⃣  Create a virtual environment
python -m venv venv
source venv/bin/activate        # Linux/Mac
# venv\Scripts\activate         # Windows

# 3️⃣  Install core dependencies
pip install numpy pandas matplotlib seaborn scikit-learn jupyter

# 4️⃣  Install additional libraries (as needed per module)
pip install xgboost catboost nltk spacy gensim
pip install tensorflow keras
pip install mlflow bentoml dvc
pip install flask streamlit

# 5️⃣  Launch Jupyter Notebook
jupyter notebook

🧭 Recommended Learning Path

 START
  │
  ▼
 Introduction (02) ──► Linear Regression (03) ──► Regularization (04)
                                                        │
  ┌─────────────────────────────────────────────────────┘
  ▼
 ML Lifecycle (05) ──► Logistic Regression (06) ──► SVM (07) ──► Naive Bayes (08)
                                                                       │
  ┌────────────────────────────────────────────────────────────────────┘
  ▼
 KNN (09) ──► Decision Trees (10) ──► Random Forest (11) ──► AdaBoost (12)
                                                                   │
  ┌────────────────────────────────────────────────────────────────┘
  ▼
 Gradient Boosting (13) ──► XGBoost (14) ──► PCA (16) ──► K-Means (17)
                                                               │
  ┌────────────────────────────────────────────────────────────┘
  ▼
 Hierarchical (18) ──► DBSCAN (19) ──► Silhouette (20) ──► Anomaly Detection (21)
                                                                    │
  ┌─────────────────────────────────────────────────────────────────┘
  ▼
 Docker (22) ──► Git (23) ──► End-to-End Projects (24) ──► MLOps (25)
                                                                │
  ┌─────────────────────────────────────────────────────────────┘
  ▼
 NLP (26) ──► Deep Learning (27) ──► Transformers (28)
                                           │
                                           ▼
                                        FINISH 🎓

🏗️ End-to-End Deployment Projects

1 · Student Performance — Local

cd "24-End To End ML Project With Deployment/mlproject-main"
pip install -r requirements.txt
python app.py
# → http://localhost:5000

2 · Student Performance — Azure

cd "24-End To End ML Project With Deployment/Student-Performance-Azure-deployment-main"
pip install -r requirements.txt
python app.py
# → CI/CD via GitHub Actions → Azure Web App

3 · Student Performance — AWS

cd "24-End To End ML Project With Deployment/AWS-CI-CD-Projects-main"
pip install -r requirements.txt
python app.py
# → CI/CD via AWS CodePipeline → Elastic Beanstalk

4 · RNN IMDB Sentiment App

cd "27-Deep Learning Bonus/02_RNN_Project"
pip install -r requirements.txt
streamlit run main.py

5 · LSTM Next-Word Prediction App

cd "27-Deep Learning Bonus/03_LSTM"
pip install -r requirements.txt
streamlit run app.py

📊 Repository Statistics

Metric Count
📓 Jupyter Notebooks 65+
📄 Theory PDFs (Handwritten Notes) 50+
📂 Structured Modules 26
🗃️ Datasets (CSV) 65+
🐍 Python Scripts 2,400+
🚀 Deployable Projects 5
☁️ Cloud Deployments (AWS/Azure) 3
🤖 Pre-trained Models (.h5, .pkl) 5+

🧭 Algorithm Coverage Map

╔═══════════════════════════════════════════════════════════════════════════╗
║                        SUPERVISED LEARNING                              ║
║  ┌────────────────┐  ┌────────────────┐  ┌──────────────────────────┐   ║
║  │   Regression   │  │ Classification │  │    Ensemble Methods      │   ║
║  ├────────────────┤  ├────────────────┤  ├──────────────────────────┤   ║
║  │ Linear         │  │ Logistic       │  │ Random Forest (Bagging)  │   ║
║  │ Multiple       │  │ SVM / SVC      │  │ AdaBoost                 │   ║
║  │ Polynomial     │  │ Naive Bayes    │  │ Gradient Boosting        │   ║
║  │ Ridge / Lasso  │  │ KNN            │  │ XGBoost                  │   ║
║  │ ElasticNet     │  │ Decision Tree  │  │ CatBoost                 │   ║
║  │ SVR            │  │                │  │                          │   ║
║  └────────────────┘  └────────────────┘  └──────────────────────────┘   ║
╚═══════════════════════════════════════════════════════════════════════════╝

╔═══════════════════════════════════════════════════════════════════════════╗
║                       UNSUPERVISED LEARNING                             ║
║  ┌────────────────────────┐  ┌──────────────────────────────────────┐   ║
║  │      Clustering        │  │    Dimensionality Reduction          │   ║
║  ├────────────────────────┤  ├──────────────────────────────────────┤   ║
║  │ K-Means                │  │ PCA (Principal Component Analysis)   │   ║
║  │ Hierarchical           │  │                                      │   ║
║  │ DBSCAN                 │  │                                      │   ║
║  │ Silhouette Analysis    │  │                                      │   ║
║  └────────────────────────┘  └──────────────────────────────────────┘   ║
║  ┌────────────────────────┐                                             ║
║  │   Anomaly Detection    │                                             ║
║  ├────────────────────────┤                                             ║
║  │ Isolation Forest       │                                             ║
║  │ Local Outlier Factor   │                                             ║
║  │ DBSCAN-based           │                                             ║
║  └────────────────────────┘                                             ║
╚═══════════════════════════════════════════════════════════════════════════╝

╔═══════════════════════════════════════════════════════════════════════════╗
║                          DEEP LEARNING                                  ║
║  ┌────────────────┐  ┌────────────────┐  ┌──────────────────────────┐   ║
║  │      ANN       │  │      CNN       │  │    Sequence Models       │   ║
║  ├────────────────┤  ├────────────────┤  ├──────────────────────────┤   ║
║  │ Perceptron     │  │ ConvNets       │  │ RNN                      │   ║
║  │ MLP            │  │ Pooling        │  │ LSTM                     │   ║
║  │ Backprop       │  │ VGG, ResNet    │  │ GRU                      │   ║
║  │ Optimizers     │  │                │  │ Bidirectional RNN        │   ║
║  │ Activations    │  │                │  │ Seq2Seq + Attention      │   ║
║  └────────────────┘  └────────────────┘  │ Transformers             │   ║
║                                           └──────────────────────────┘   ║
╚═══════════════════════════════════════════════════════════════════════════╝

╔═══════════════════════════════════════════════════════════════════════════╗
║                              NLP                                        ║
║  ┌────────────────────────┐  ┌──────────────────────────────────────┐   ║
║  │  Text Preprocessing    │  │     Feature Engineering              │   ║
║  ├────────────────────────┤  ├──────────────────────────────────────┤   ║
║  │ Tokenization           │  │ Bag of Words (BoW)                   │   ║
║  │ Stemming               │  │ TF-IDF                               │   ║
║  │ Lemmatization          │  │ Word2Vec                              │   ║
║  │ Stopword Removal       │  │ Average Word2Vec                     │   ║
║  │ POS Tagging            │  │                                      │   ║
║  │ Named Entity Recog.    │  │                                      │   ║
║  └────────────────────────┘  └──────────────────────────────────────┘   ║
╚═══════════════════════════════════════════════════════════════════════════╝

🤝 Contributing

Contributions are welcome! Here's how you can help improve this repository:

  1. Fork the repository
  2. Create a feature branch (git checkout -b feature/add-new-module)
  3. Commit your changes (git commit -m 'Add: New module on feature engineering')
  4. Push to the branch (git push origin feature/add-new-module)
  5. Open a Pull Request
💡 Contribution Ideas
  • 🐛 Fix typos in notebooks or PDFs
  • 📓 Add new practical notebooks for existing modules
  • 📊 Add more datasets and real-world case studies
  • 🧪 Add unit tests for deployment projects
  • 📝 Improve inline documentation and comments
  • 🆕 Add modules on Feature Engineering, Time Series, or Reinforcement Learning

📝 License

This project is licensed under the MIT License — see the LICENSE file for details.

MIT License · Copyright (c) 2025 Marmik Gajbhiye

🙏 Acknowledgments

  • Built as a comprehensive learning resource during the Complete Machine Learning, NLP & MLOps Bootcamp
  • Theory notes are handwritten for deeper mathematical understanding
  • All deployment projects are production-tested on AWS and Azure


If this repository helped you, consider giving it a ⭐



GitHub LinkedIn


About

My learning journey through Complete Machine Learning & NLP Bootcamp with MLOps & Deployment

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages