Skip to content

Latest commit

Β 

History

18 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

🌸 Data Classification Using AI

Iris Flower Classification with Machine Learning

A clean, reproducible internship project demonstrating an end-to-end supervised machine learning classification pipeline using the classic Iris dataset, feature scaling, K-Nearest Neighbors, model selection, and evaluation.


Python Scikit--learn Tests Status


πŸ“Œ About the Project

Data Classification Using AI is an internship project focused on understanding and implementing the fundamentals of supervised machine learning through a practical classification problem.

The project uses the well-known Iris flower dataset to classify flowers into three species based on four numerical measurements:

  • 🌿 Sepal length
  • 🌿 Sepal width
  • 🌸 Petal length
  • 🌸 Petal width

Rather than treating the task as a black-box prediction problem, the project demonstrates the complete machine learning workflow: data validation β†’ preprocessing β†’ train/test splitting β†’ feature scaling β†’ model selection β†’ training β†’ prediction β†’ evaluation.

The implementation intentionally stays focused on the assignment requirements and avoids unnecessary frameworks, databases, APIs, or cloud services.


🎯 Project Objective

The main objective is to build a reliable classification model that can learn from labeled Iris flower examples and predict the species of previously unseen samples.

The project demonstrates several important machine learning concepts:

Data β†’ Preprocessing β†’ Training β†’ Prediction β†’ Evaluation

It also demonstrates why evaluation should go beyond a single accuracy number by using a confusion matrix, precision, recall, and F1 score.


🧠 Machine Learning Approach

The project uses K-Nearest Neighbors (KNN) as the classification algorithm.

KNN predicts the class of a new observation by examining the labels of its nearest training examples. Because KNN is distance-based, feature scaling is an important part of the pipeline.

Model pipeline

                    IRIS DATASET
                         β”‚
                         β–Ό
                Data Validation
                         β”‚
                         β–Ό
             Stratified 80/20 Split
                  β”Œβ”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”
                  β”‚             β”‚
               TRAIN          TEST
                  β”‚             β”‚
                  β–Ό             β”‚
            StandardScaler      β”‚
                  β”‚             β”‚
                  β–Ό             β”‚
             K Selection        β”‚
                  β”‚             β”‚
                  β–Ό             β”‚
               KNN Model        β”‚
                  β”‚             β”‚
                  β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜
                         β–Ό
                      Predict
                         β”‚
              β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
              β–Ό                     β–Ό
       Confusion Matrix         Macro F1

Why StandardScaler?

The four Iris measurements are expressed on different numerical ranges. Standardization puts the features on a comparable scale so that distance calculations used by KNN are not dominated by a feature simply because of its numerical magnitude.

The scaler is placed inside the scikit-learn Pipeline, meaning it is fitted only using the training data. This prevents test-data leakage.


πŸ“Š Dataset

The Iris benchmark contains:

Property Value
Total samples 150
Classes 3
Features 4
Samples per class 50
Training samples 120
Testing samples 30

Target classes

setosa       50 samples
versicolor   50 samples
virginica    50 samples

βš™οΈ Model Selection

The assignment demonstrates K = 5 for KNN, and the implementation supports that exact demonstration.

In addition, the project includes a small model-selection step that evaluates K = 1 through K = 15 using:

  • 5-fold cross-validation
  • Macro F1 scoring
  • Smallest-K tie breaking

The resulting selected value is K = 5, matching the assignment demonstration.

To reproduce the exact fixed-K demonstration:

python -m src.main --no-tune-k --k 5

πŸ“ˆ Verified Results

The complete pipeline has been executed successfully.

Metric Result
Accuracy 93.33%
Macro Precision 94.44%
Macro Recall 93.33%
Macro F1 93.27%
Test samples 30
Correct predictions 28 / 30
Selected K 5

Confusion Matrix

                 Predicted
              Setosa  Versicolor  Virginica

Setosa           10       0          0
Versicolor        0      10          0
Virginica         0       2          8

The classifier correctly recognizes all 10 Setosa and 10 Versicolor test samples. Two Virginica samples are classified as Versicolor, giving an overall accuracy of 93.33%.


πŸ§ͺ Testing

The repository includes an automated test suite covering the important parts of the pipeline.

pytest -q

Verified result:

5 passed

Tests cover:

  • Dataset schema and dimensions
  • Class distribution
  • Stratified 80/20 split
  • Model construction
  • StandardScaler + KNN pipeline
  • K selection
  • End-to-end classification performance

πŸš€ Getting Started

1. Clone the repository

git clone https://github.com/HadeedJalani/Data-classification.git
cd Data-classification

2. Create a virtual environment

Windows PowerShell

python -m venv .venv
.\.venv\Scripts\Activate.ps1

macOS / Linux

python3 -m venv .venv
source .venv/bin/activate

3. Install dependencies

python -m pip install --upgrade pip
pip install -r requirements.txt

4. Run the tests

pytest -q

5. Run the complete pipeline

python -m src.main

6. Run the exact K=5 assignment demonstration

python -m src.main --no-tune-k --k 5

πŸ“ Project Structure

Data-classification/
β”‚
β”œβ”€β”€ πŸ“‚ data/
β”‚   └── iris.csv                    # Iris dataset
β”‚
β”œβ”€β”€ πŸ“‚ src/
β”‚   β”œβ”€β”€ __init__.py
β”‚   β”œβ”€β”€ config.py                   # Project configuration
β”‚   β”œβ”€β”€ data.py                     # Loading, validation and splitting
β”‚   β”œβ”€β”€ evaluate.py                 # Metrics and reports
β”‚   β”œβ”€β”€ main.py                     # Main executable pipeline
β”‚   β”œβ”€β”€ model.py                    # Scaling, KNN and K selection
β”‚   └── visualize.py                # Evaluation visualizations
β”‚
β”œβ”€β”€ πŸ“‚ tests/
β”‚   └── test_pipeline.py            # Automated tests
β”‚
β”œβ”€β”€ πŸ“‚ docs/
β”‚   β”œβ”€β”€ demo.py
β”‚   β”œβ”€β”€ PRESENTATION_OUTLINE.md
β”‚   β”œβ”€β”€ SUBMISSION_CHECKLIST.md
β”‚   └── TECHNICAL_REPORT.md
β”‚
β”œβ”€β”€ πŸ“‚ artifacts/                   # Generated locally at runtime
β”œβ”€β”€ πŸ“‚ reports/                     # Generated locally at runtime
β”‚
β”œβ”€β”€ .gitignore
β”œβ”€β”€ requirements.txt
└── README.md

πŸ› οΈ Technology Stack

Technology Purpose
Python Core programming language
Pandas Dataset loading and manipulation
NumPy Numerical operations
Scikit-learn Scaling, KNN, model selection and evaluation
Matplotlib Confusion matrix and K-selection visualizations
Joblib Model persistence
Pytest Automated testing

πŸ” Key Engineering Decisions

Reproducibility

A fixed random_state=42 is used for the train/test split so the experiment can be reproduced consistently.

Stratification

The 80/20 split uses stratification to preserve the class distribution between training and testing data.

Leakage prevention

StandardScaler is part of the scikit-learn Pipeline, so preprocessing is learned from training data rather than from the complete dataset.

Practical scope

The project follows the supplied internship assignment closely. Computer vision and CNNs shown in the assignment as future directions are not unnecessarily added to this implementation.


πŸ“¦ Generated Artifacts

After running the pipeline, the following outputs are generated locally:

artifacts/
β”œβ”€β”€ confusion_matrix.png
β”œβ”€β”€ iris_knn_model.joblib
β”œβ”€β”€ k_selection.png
└── metrics.json

reports/
└── classification_report.txt

These files are intentionally ignored by Git because they are reproducible outputs rather than source code.


πŸŽ“ Internship Project Context

This repository represents the implementation of Project 2: Data Classification Using AI from an Artificial Intelligence internship assignment.

The project focuses on demonstrating practical understanding of:

  • Supervised learning
  • Classification
  • Data preprocessing
  • Feature scaling
  • Train/test methodology
  • K-Nearest Neighbors
  • Hyperparameter selection
  • Confusion matrices
  • Precision, recall and F1 score
  • Reproducible ML engineering

The goal is not simply to produce a prediction, but to demonstrate the reasoning and engineering workflow behind a complete machine learning solution.


πŸ“œ License

This project is intended primarily for educational and internship demonstration purposes.


πŸ‘¨β€πŸ’» Author

Hadeed Jalani

Artificial Intelligence / Machine Learning Project

⭐ If you found this project useful, consider giving the repository a star.

About

Iris flower classification using Scikit-learn KNN, StandardScaler, model selection, and comprehensive ML evaluation.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages