Skip to content

Use a portable dataset path and document the training workflow - #1

Merged
AtulJ505 merged 1 commit into
mainfrom
improve/reliability-and-project-guide
Sep 9, 2026
Merged

AtulJ505 merged 1 commit into
mainfrom
improve/reliability-and-project-guide

Conversation

@AtulJ505

@AtulJ505 AtulJ505 commented Sep 9, 2026

Copy link
Copy Markdown
Owner

Data ingestion uses a Windows-specific backslash path, preventing a clean Unix checkout from locating the included dataset. The README contains only a heading.

Use os.path.join for the dataset path and document environment setup, training, artifacts, the Flask interface and evaluation limitations.

Validation: Python syntax checks passed and the dataset exists at the documented path. The full multi-model training search was not rerun.

Copilot AI lite review requested due to automatic review settings September 9, 2026 08:13

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The updated README documents running the Flask app, but application.py currently uses a string port (port='8000') which is likely to fail at runtime, so the documented workflow is not actually runnable as written.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Pull request overview

This PR makes dataset ingestion cross-platform by replacing a Windows-specific path with a portable join, and significantly expands the README to document local environment setup, training, artifacts, and the Flask prediction UI.

Changes:

  • Use os.path.join(...) for the dataset CSV path in data ingestion.
  • Replace the placeholder README with end-to-end setup, training, and usage documentation.
  • Add notes on reproducibility, limitations, and safety considerations for pickle artifacts.
File summaries
File Description
src/components/data_ingestion.py Updates the dataset path to be OS-portable.
README.md Adds detailed setup/training/run instructions and project limitations.
Review details
  • Files reviewed: 2/2 changed files
  • Comments generated: 2
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread README.md
Comment on lines +30 to +34
python application.py
```

Training evaluates multiple model families and hyperparameters, so it can take time. The web application listens on port 8000 and expects the trained `artifacts/model.pkl` and `artifacts/preprocessor.pkl` files. Open `http://127.0.0.1:8000/predictdata` for the form.

logging.info("Entered the data ingestion method")
try:
df = pd.read_csv("notebook\data\stud.csv")
df = pd.read_csv(os.path.join("notebook", "data", "stud.csv"))
@AtulJ505
AtulJ505 merged commit 3ac4b00 into main Sep 9, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants