65+ Notebooks · 50+ Handwritten Theory PDFs · 5 Deployment-Ready Projects
A structured, hands-on journey through the entire data science stack — from linear regression to transformers.
📚 Modules 🛠️ Tech Stack 🚀 Get Started 🏗️ Projects 🤝 Contribute
A comprehensive, self-contained curriculum covering the complete data science and machine learning pipeline — designed for learners who want depth, not just breadth.
| 📝 | Handwritten Theory Notes | Deep mathematical intuition and derivations across 50+ PDFs |
| 💻 | Practical Notebooks | 65+ hands-on implementations with real-world datasets |
| 🚀 | Deployment Projects | End-to-end ML pipelines deployed on AWS, Azure & Docker |
| 🧪 | MLOps Tooling | Experiment tracking, model versioning & serving with MLflow, DVC & BentoML |
The curriculum spans 28 structured modules organized in a progressive learning path — from foundational concepts through advanced deep learning and transformer architectures.
Complete-Data-Science-With-Machine-Learning-And-NLP/
│
├── 📁 02-Introduction/ # Foundations: AI vs ML vs DL
├── 📁 03-Complete Linear Regression/ # Simple, Multiple, Polynomial
├── 📁 04-Ridge Lasso And Elasticnet/ # Regularization & Cross-Validation
├── 📁 05-Step By Step Project Implementation/ # ML Project Lifecycle
├── 📁 06-Logistic Regression/ # Binary & Multiclass Classification
├── 📁 07-SVM/ # Support Vector Machines & Kernels
├── 📁 08-Naive Bayes/ # Probabilistic Classification
├── 📁 09-K Nearest Neighbor/ # Instance-Based Learning (KNN)
├── 📁 10-Decision-Tree/ # Tree-Based Methods
├── 📁 11-Random Forest/ # Bagging Ensemble Methods
├── 📁 12-Adaboost/ # Adaptive Boosting
├── 📁 13-Gradient Boosting/ # Gradient-Based Boosting
├── 📁 14-XgBoost/ # Extreme Gradient Boosting
├── 📁 16-PCA/ # Principal Component Analysis
├── 📁 17-K-Mean/ # K-Means Clustering
├── 📁 18-Hierarchical Cluster/ # Agglomerative & Divisive
├── 📁 19-DBSCAN Clustering/ # Density-Based Clustering
├── 📁 20-Silhouette Clustering/ # Cluster Validation
├── 📁 21-Anomaly Detection ML/ # Isolation Forest & LOF
├── 📁 22-Dockers/ # Docker & Docker Compose
├── 📁 23-Git And Github/ # Version Control
├── 📁 24-End To End ML Project With Deployment/ # AWS, Azure & Local Deployment
├── 📁 25-MLFlow Dagshub and BentoML/ # MLOps: Tracking & Serving
├── 📁 26-Complete NLP For Machine Learning/ # Text Processing & Embeddings
├── 📁 27-Deep Learning Bonus/ # ANN, CNN, RNN, LSTM, GRU
├── 📁 28-Transformer/ # Attention & Transformers
├── 📄 LICENSE
└── 📄 README.md
📘 Module 02 · Introduction to Data Science & Machine Learning
| Content | Type | Description |
|---|---|---|
| AI vs ML vs DL vs Data Science | Comparative study of fields and their intersections | |
| Types of ML Techniques | Supervised, unsupervised, reinforcement learning overview | |
| Instance-Based vs Model-Based Learning | Fundamental learning paradigm differences | |
| Equation of a Straight Line & Hyperplane | Mathematical prerequisites for ML algorithms |
📘 Module 03 · Complete Linear Regression
Theory (7 PDFs): Simple & Multiple Linear Regression, OLS Derivation, Polynomial Regression, Performance Metrics (MSE, RMSE, MAE, R²), Overfitting & Underfitting
Practicals (12 Notebooks):
| Notebook | Key Topics |
|---|---|
Simple Linear Regression.ipynb |
Line fitting, gradient descent, cost function |
Multiple Linear Regression.ipynb |
Multi-feature regression, economics dataset |
Polynomial Regression Implementation.ipynb |
Non-linear curve fitting |
Linear Regression ML Implementation.ipynb |
Complete sklearn pipeline |
Model Training.ipynb |
Train/test split, model evaluation |
Handling Imbalance Dataset.ipynb |
SMOTE, undersampling, oversampling |
Ridge, Lasso Regression.ipynb |
Regularization comparison |
Datasets: Height-Weight, Economic Index, Algerian Forest Fires
📘 Module 04 · Ridge, Lasso & ElasticNet Regression
Theory (2 PDFs): Ridge, Lasso & ElasticNet mathematical formulations; Types of Cross-Validation (K-Fold, Stratified, Leave-One-Out)
| Notebook | Key Topics |
|---|---|
Ridge, Lasso Regression.ipynb |
L1/L2 regularization, hyperparameter tuning |
Model Training.ipynb |
Cross-validation strategies, model selection |
📘 Module 05 · ML Project Lifecycle — Step-by-Step Implementation
| Notebook | Key Topics |
|---|---|
Simple Linear Regression.ipynb |
Foundation implementations |
Multiple Linear Regression.ipynb |
Multi-variate analysis |
EDA And FE Algerian Forest Fires.ipynb |
Exploratory data analysis, feature engineering |
Model Training.ipynb |
Complete pipeline: preprocessing → training → evaluation |
Includes: Flask web application (application.py), AWS Elastic Beanstalk configuration, trained model artifacts (.pkl files), and HTML templates for prediction interface.
📘 Module 06 · Logistic Regression
Theory (3 PDFs): Logistic Regression mathematics, Sigmoid function, Performance Metrics (Confusion Matrix, Precision, Recall, F1-Score, AUC-ROC), One-vs-Rest (OVR) Classification
| Notebook | Key Topics |
|---|---|
Logistic Regression Implementation.ipynb |
Binary/multiclass classification, decision boundaries, evaluation metrics |
📘 Module 07 · Support Vector Machines (SVM)
Theory (3 PDFs): SVC mathematical formulation, Support Vector Regression (SVR), SVM Kernels (Linear, RBF, Polynomial, Sigmoid)
| Notebook | Key Topics |
|---|---|
Basic SVC Implementation.ipynb |
Linear SVM classification |
SVM Kernels Implementation.ipynb |
Kernel trick, non-linear decision boundaries |
Support Vector Regression Implementation.ipynb |
SVR for continuous prediction |
📘 Module 08 · Naive Bayes Classifier
Theory (3 PDFs): Bayes' theorem derivation, Naive Bayes in-depth intuition, Variants (Gaussian, Multinomial, Bernoulli)
| Notebook | Key Topics |
|---|---|
Naive Bayes Implementation.ipynb |
All three variants with real-world classification |
📘 Module 09 · K-Nearest Neighbors (KNN)
Theory (2 PDFs): KNN classification & regression, Distance metrics (Euclidean, Manhattan, Minkowski), KD-Tree & Ball Tree optimization
| Notebook | Key Topics |
|---|---|
KNNClassifier.ipynb |
Classification with optimal K selection |
KNNRegressor.ipynb |
Regression using distance-weighted neighbors |
📘 Module 10 · Decision Trees
| Notebook | Key Topics |
|---|---|
Decision Tree Classifier Practical Implementation.ipynb |
Gini, entropy, information gain, tree visualization |
Diabetes Prediction Using Decision Tree Regressor.ipynb |
Medical dataset regression with pruning |
📗 Module 11 · Random Forest
| Notebook | Key Topics |
|---|---|
Random Forest Classification Implementation.ipynb |
Bagging, feature importance, AUC-ROC visualization |
Random Forest Regression Implementation.ipynb |
Ensemble regression, out-of-bag error |
Datasets: Travel dataset for classification, Boston-style data for regression
📗 Module 12 · AdaBoost
| Notebook | Key Topics |
|---|---|
Adaboost Classification Implementation.ipynb |
Adaptive boosting, sample weight updates |
Adaboost Regression Implementation.ipynb |
Boosted regression trees |
📗 Module 13 · Gradient Boosting
| Notebook | Key Topics |
|---|---|
GradientBoost Classification Implementation.ipynb |
Gradient descent in function space |
Gradientboost Regression Implementation.ipynb |
Residual learning, learning rate tuning |
📗 Module 14 · XGBoost
| Notebook | Key Topics |
|---|---|
XgboostBoost Classification Implementation.ipynb |
Regularized boosting, early stopping, feature importance |
Xgboost Regression Implementation.ipynb |
XGBoost hyperparameter optimization |
📙 Module 16 · Principal Component Analysis (PCA)
| Notebook | Key Topics |
|---|---|
Principal Component Analysis (PCA) Implementation.ipynb |
Dimensionality reduction, explained variance ratio, scree plots |
📙 Module 17 · K-Means Clustering
| Notebook | Key Topics |
|---|---|
K Means Clustering.ipynb |
Elbow method, inertia, cluster visualization, centroid convergence |
📙 Module 18 · Hierarchical Clustering
| Notebook | Key Topics |
|---|---|
Hierarichal Clustering Implementation.ipynb |
Dendrograms, agglomerative clustering, linkage methods |
📙 Module 19 · DBSCAN Clustering
Theory (1 PDF): DBSCAN algorithm — core points, border points, noise, epsilon & minPts
| Notebook | Key Topics |
|---|---|
DBSCAN Implementation.ipynb |
Density-based clustering, handling arbitrary-shaped clusters |
📙 Module 20 · Silhouette Analysis
Theory (1 PDF): Silhouette score interpretation, optimal cluster count validation
📙 Module 21 · Anomaly Detection
Theory (2 PDFs): Isolation Forest algorithm, Local Outlier Factor (LOF)
| Notebook | Key Topics |
|---|---|
Isolation Anamoly Detection.ipynb |
Tree-based anomaly detection on healthcare data |
DBSCAN Implementation.ipynb |
DBSCAN-based outlier detection |
Datasets: Healthcare anomaly detection dataset
🔧 Module 22 · Docker for Machine Learning
Theory (2 PDFs): Docker fundamentals, Dockerfile best practices, Docker Compose multi-container orchestration
Practicals: Hello World containerized application
🔧 Module 23 · Git & GitHub
Theory (1 PDF): Git workflow, branching strategies, collaboration best practices
🚀 Module 24 · End-to-End ML Projects with Deployment
Three complete, production-ready ML projects with full CI/CD pipelines:
| Project | Stack | Deployment |
|---|---|---|
| Student Performance Predictor | Flask, scikit-learn, CatBoost | Local + Docker |
| Student Performance — Azure | Flask, scikit-learn | Azure Web App + GitHub Actions CI/CD |
| Student Performance — AWS | Flask, scikit-learn | AWS Elastic Beanstalk + AWS CodePipeline |
Each project includes:
- 📊 EDA & Model Training notebooks
- 🏗️ Modular
src/package (data ingestion, transformation, model trainer, prediction pipeline) - 🌐 Flask web app with HTML templates
- 📦
setup.pyfor package installation - 🐳 Dockerfile for containerization
- ⚡ CI/CD workflows (
.github/workflows/)
🔧 Module 25 · MLOps — MLflow, DagsHub, BentoML & DVC
| Tool | Content |
|---|---|
| MLflow | Experiment tracking, model registry, metric logging (app.py with ElasticNet) |
| BentoML | Model serving, API creation (service.py), model packaging (bentofile.yaml) |
| DVC | Data version control, pipeline management |
📕 Module 26 · Complete NLP for Machine Learning
Theory (1 PDF): Comprehensive NLP concepts (130+ pages)
Practicals (12 Notebooks):
| # | Notebook | Key Topics |
|---|---|---|
| 4 | Tokenization Example Using NLTK.ipynb |
Word & sentence tokenization |
| 5 | Stemming And Its Types.ipynb |
Porter, Lancaster, Snowball stemmers |
| 6 | Lemmatization.ipynb |
WordNet lemmatizer, POS-aware lemmatization |
| 7 | Stopwords With NLTK.ipynb |
Stopword removal, custom stopword lists |
| 8 | Parts Of Speech Tagging.ipynb |
POS tagging with NLTK |
| 9 | Named Entity Recognition.ipynb |
NER with spaCy/NLTK |
| 15 | Bag Of Words Practical's.ipynb |
CountVectorizer, document-term matrix |
| 16 | TF-IDF Practical.ipynb |
Term frequency-inverse document frequency |
| 26 | Word2vec Practical Implementation.ipynb |
Word embeddings, Gensim Word2Vec |
| 27 | Spam Ham Classification (TF-IDF).ipynb |
Project: SMS spam detection pipeline |
| 28-29 | Spam Ham (Word2Vec, AvgWord2Vec).ipynb |
Project: Spam detection with embeddings |
| 30-31 | Kindle Review Sentiment Analysis.ipynb |
Project: Sentiment analysis on Kindle reviews |
📕 Module 27 · Deep Learning (Bonus)
Theory (11 PDFs):
| Topic | Content |
|---|---|
| Deep Learning Fundamentals | Perceptron, multi-layer networks, forward propagation |
| Activation Functions | Sigmoid, ReLU, Leaky ReLU, Tanh, Softmax, GELU |
| Loss Functions | MSE, Cross-Entropy, Hinge Loss, Focal Loss |
| Optimizers | SGD, Momentum, RMSProp, Adam, AdaGrad, AdaDelta |
| Weight Initialization | Xavier, He, LeCun initialization techniques |
| CNNs | Convolutional layers, pooling, architectures (VGG, ResNet) |
| RNNs | Simple RNN, backpropagation through time, vanishing gradients |
| LSTM & GRU | Long short-term memory, gated recurrent units |
| Bidirectional RNN | Bi-directional sequence modeling |
Practicals — RNN Project (IMDB Sentiment Analysis):
| File | Description |
|---|---|
simplernn.ipynb |
Building RNN from scratch for IMDB classification |
embedding.ipynb |
Word embeddings for sequence models |
prediction.ipynb |
Inference pipeline |
main.py |
Streamlit web application |
simple_rnn_imdb.h5 |
Pre-trained model weights |
Practicals — LSTM Project (Next Word Prediction):
| File | Description |
|---|---|
experiemnts.ipynb |
LSTM model training on Shakespeare's Hamlet |
app.py |
Streamlit app for next-word prediction |
next_word_lstm.h5 |
Trained LSTM model |
tokenizer.pickle |
Fitted tokenizer for inference |
📕 Module 28 · Transformers
| Topic | Content |
|---|---|
| Encoder-Decoder & Seq2Seq | Sequence-to-sequence architecture fundamentals |
| Attention Mechanism | Self-attention, multi-head attention, scaled dot-product |
| Transformers | Complete transformer architecture ("Attention Is All You Need") |
| Claude Architecture | Modern LLM architecture analysis |
| Domain | Technologies |
|---|---|
| Core ML & Data Science | Python · NumPy · Pandas · Matplotlib · Seaborn · Plotly · scikit-learn · SciPy · StatsModels |
| Deep Learning | TensorFlow · Keras · RNN · LSTM · GRU · CNN · Word Embeddings · Transformers |
| NLP Libraries | NLTK · spaCy · Gensim (Word2Vec) · CountVectorizer · TF-IDF |
| Boosting | XGBoost · CatBoost · AdaBoost · GradientBoosting |
| MLOps & Deployment | MLflow · DagsHub · DVC · BentoML · Docker · Flask · Streamlit |
| Cloud & CI/CD | AWS Elastic Beanstalk · AWS CodePipeline · Azure Web Apps · GitHub Actions |
| DevOps | Docker · Docker Compose · Git · GitHub · Virtual Environments |
| Notebooks & IDEs | Jupyter Notebook · VS Code · Google Colab Compatible |
| Requirement | Version |
|---|---|
| Python | 3.8+ |
| pip | Latest |
| Jupyter | Notebook or JupyterLab |
| Git | 2.x+ |
# 1️⃣ Clone the repository
git clone https://github.com/Marmikgaj/Complete-Data-Science-With-Machine-Learning-And-NLP.git
cd Complete-Data-Science-With-Machine-Learning-And-NLP
# 2️⃣ Create a virtual environment
python -m venv venv
source venv/bin/activate # Linux/Mac
# venv\Scripts\activate # Windows
# 3️⃣ Install core dependencies
pip install numpy pandas matplotlib seaborn scikit-learn jupyter
# 4️⃣ Install additional libraries (as needed per module)
pip install xgboost catboost nltk spacy gensim
pip install tensorflow keras
pip install mlflow bentoml dvc
pip install flask streamlit
# 5️⃣ Launch Jupyter Notebook
jupyter notebook START
│
▼
Introduction (02) ──► Linear Regression (03) ──► Regularization (04)
│
┌─────────────────────────────────────────────────────┘
▼
ML Lifecycle (05) ──► Logistic Regression (06) ──► SVM (07) ──► Naive Bayes (08)
│
┌────────────────────────────────────────────────────────────────────┘
▼
KNN (09) ──► Decision Trees (10) ──► Random Forest (11) ──► AdaBoost (12)
│
┌────────────────────────────────────────────────────────────────┘
▼
Gradient Boosting (13) ──► XGBoost (14) ──► PCA (16) ──► K-Means (17)
│
┌────────────────────────────────────────────────────────────┘
▼
Hierarchical (18) ──► DBSCAN (19) ──► Silhouette (20) ──► Anomaly Detection (21)
│
┌─────────────────────────────────────────────────────────────────┘
▼
Docker (22) ──► Git (23) ──► End-to-End Projects (24) ──► MLOps (25)
│
┌─────────────────────────────────────────────────────────────┘
▼
NLP (26) ──► Deep Learning (27) ──► Transformers (28)
│
▼
FINISH 🎓
cd "24-End To End ML Project With Deployment/mlproject-main"
pip install -r requirements.txt
python app.py
# → http://localhost:5000 |
cd "24-End To End ML Project With Deployment/Student-Performance-Azure-deployment-main"
pip install -r requirements.txt
python app.py
# → CI/CD via GitHub Actions → Azure Web App |
cd "24-End To End ML Project With Deployment/AWS-CI-CD-Projects-main"
pip install -r requirements.txt
python app.py
# → CI/CD via AWS CodePipeline → Elastic Beanstalk |
cd "27-Deep Learning Bonus/02_RNN_Project"
pip install -r requirements.txt
streamlit run main.py |
cd "27-Deep Learning Bonus/03_LSTM"
pip install -r requirements.txt
streamlit run app.py |
|
| Metric | Count | |
|---|---|---|
| 📓 | Jupyter Notebooks | 65+ |
| 📄 | Theory PDFs (Handwritten Notes) | 50+ |
| 📂 | Structured Modules | 26 |
| 🗃️ | Datasets (CSV) | 65+ |
| 🐍 | Python Scripts | 2,400+ |
| 🚀 | Deployable Projects | 5 |
| ☁️ | Cloud Deployments (AWS/Azure) | 3 |
| 🤖 | Pre-trained Models (.h5, .pkl) | 5+ |
╔═══════════════════════════════════════════════════════════════════════════╗
║ SUPERVISED LEARNING ║
║ ┌────────────────┐ ┌────────────────┐ ┌──────────────────────────┐ ║
║ │ Regression │ │ Classification │ │ Ensemble Methods │ ║
║ ├────────────────┤ ├────────────────┤ ├──────────────────────────┤ ║
║ │ Linear │ │ Logistic │ │ Random Forest (Bagging) │ ║
║ │ Multiple │ │ SVM / SVC │ │ AdaBoost │ ║
║ │ Polynomial │ │ Naive Bayes │ │ Gradient Boosting │ ║
║ │ Ridge / Lasso │ │ KNN │ │ XGBoost │ ║
║ │ ElasticNet │ │ Decision Tree │ │ CatBoost │ ║
║ │ SVR │ │ │ │ │ ║
║ └────────────────┘ └────────────────┘ └──────────────────────────┘ ║
╚═══════════════════════════════════════════════════════════════════════════╝
╔═══════════════════════════════════════════════════════════════════════════╗
║ UNSUPERVISED LEARNING ║
║ ┌────────────────────────┐ ┌──────────────────────────────────────┐ ║
║ │ Clustering │ │ Dimensionality Reduction │ ║
║ ├────────────────────────┤ ├──────────────────────────────────────┤ ║
║ │ K-Means │ │ PCA (Principal Component Analysis) │ ║
║ │ Hierarchical │ │ │ ║
║ │ DBSCAN │ │ │ ║
║ │ Silhouette Analysis │ │ │ ║
║ └────────────────────────┘ └──────────────────────────────────────┘ ║
║ ┌────────────────────────┐ ║
║ │ Anomaly Detection │ ║
║ ├────────────────────────┤ ║
║ │ Isolation Forest │ ║
║ │ Local Outlier Factor │ ║
║ │ DBSCAN-based │ ║
║ └────────────────────────┘ ║
╚═══════════════════════════════════════════════════════════════════════════╝
╔═══════════════════════════════════════════════════════════════════════════╗
║ DEEP LEARNING ║
║ ┌────────────────┐ ┌────────────────┐ ┌──────────────────────────┐ ║
║ │ ANN │ │ CNN │ │ Sequence Models │ ║
║ ├────────────────┤ ├────────────────┤ ├──────────────────────────┤ ║
║ │ Perceptron │ │ ConvNets │ │ RNN │ ║
║ │ MLP │ │ Pooling │ │ LSTM │ ║
║ │ Backprop │ │ VGG, ResNet │ │ GRU │ ║
║ │ Optimizers │ │ │ │ Bidirectional RNN │ ║
║ │ Activations │ │ │ │ Seq2Seq + Attention │ ║
║ └────────────────┘ └────────────────┘ │ Transformers │ ║
║ └──────────────────────────┘ ║
╚═══════════════════════════════════════════════════════════════════════════╝
╔═══════════════════════════════════════════════════════════════════════════╗
║ NLP ║
║ ┌────────────────────────┐ ┌──────────────────────────────────────┐ ║
║ │ Text Preprocessing │ │ Feature Engineering │ ║
║ ├────────────────────────┤ ├──────────────────────────────────────┤ ║
║ │ Tokenization │ │ Bag of Words (BoW) │ ║
║ │ Stemming │ │ TF-IDF │ ║
║ │ Lemmatization │ │ Word2Vec │ ║
║ │ Stopword Removal │ │ Average Word2Vec │ ║
║ │ POS Tagging │ │ │ ║
║ │ Named Entity Recog. │ │ │ ║
║ └────────────────────────┘ └──────────────────────────────────────┘ ║
╚═══════════════════════════════════════════════════════════════════════════╝
Contributions are welcome! Here's how you can help improve this repository:
- Fork the repository
- Create a feature branch (
git checkout -b feature/add-new-module) - Commit your changes (
git commit -m 'Add: New module on feature engineering') - Push to the branch (
git push origin feature/add-new-module) - Open a Pull Request
💡 Contribution Ideas
- 🐛 Fix typos in notebooks or PDFs
- 📓 Add new practical notebooks for existing modules
- 📊 Add more datasets and real-world case studies
- 🧪 Add unit tests for deployment projects
- 📝 Improve inline documentation and comments
- 🆕 Add modules on Feature Engineering, Time Series, or Reinforcement Learning
This project is licensed under the MIT License — see the LICENSE file for details.
MIT License · Copyright (c) 2025 Marmik Gajbhiye
- Built as a comprehensive learning resource during the Complete Machine Learning, NLP & MLOps Bootcamp
- Theory notes are handwritten for deeper mathematical understanding
- All deployment projects are production-tested on AWS and Azure