Repository navigation
Use a portable dataset path and document the training workflow - #1
Merged
Merged
Conversation
There was a problem hiding this comment.
🟡 Changes recommended
The updated README documents running the Flask app, but application.py currently uses a string port (port='8000') which is likely to fail at runtime, so the documented workflow is not actually runnable as written.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Pull request overview
This PR makes dataset ingestion cross-platform by replacing a Windows-specific path with a portable join, and significantly expands the README to document local environment setup, training, artifacts, and the Flask prediction UI.
Changes:
- Use
os.path.join(...)for the dataset CSV path in data ingestion. - Replace the placeholder README with end-to-end setup, training, and usage documentation.
- Add notes on reproducibility, limitations, and safety considerations for pickle artifacts.
File summaries
| File | Description |
|---|---|
src/components/data_ingestion.py |
Updates the dataset path to be OS-portable. |
README.md |
Adds detailed setup/training/run instructions and project limitations. |
Review details
- Files reviewed: 2/2 changed files
- Comments generated: 2
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Comment on lines
+30
to
+34
| python application.py | ||
| ``` | ||
|
|
||
| Training evaluates multiple model families and hyperparameters, so it can take time. The web application listens on port 8000 and expects the trained `artifacts/model.pkl` and `artifacts/preprocessor.pkl` files. Open `http://127.0.0.1:8000/predictdata` for the form. | ||
|
|
| logging.info("Entered the data ingestion method") | ||
| try: | ||
| df = pd.read_csv("notebook\data\stud.csv") | ||
| df = pd.read_csv(os.path.join("notebook", "data", "stud.csv")) |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Data ingestion uses a Windows-specific backslash path, preventing a clean Unix checkout from locating the included dataset. The README contains only a heading.
Use os.path.join for the dataset path and document environment setup, training, artifacts, the Flask interface and evaluation limitations.
Validation: Python syntax checks passed and the dataset exists at the documented path. The full multi-model training search was not rerun.