Skip to content
View irfanalidv's full-sized avatar

Organizations

@brainsfeed @re-sources-io

Block or report irfanalidv

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
irfanalidv/README.md
Irfan Ali — AI Engineer. I build tools that turn messy information into data you can use.

Website LinkedIn PyPI ORCID Email


AI Engineer, based in Bengaluru. I build tools that turn messy information into data you can actually use — reading web pages and documents, pulling out what matters, and checking the result is right before anyone relies on it.

Master's in Data Science & AI, IISER Tirupati, 2026.

Looking for AI engineering work — Bengaluru or remote. Available immediately. irfan.ali@datacortex.in

Experience

Role Org When
AI Engineer DataCortex · India Dec 2025 – present
Head of AI Kuration AI · Hong Kong 2024–2025
Senior Manager – Data & AI, R&D Luminous Power Technologies · India 2023–2024
Data Analytics & Automation Associate Lynk · India 2022–2023
Head of Data & Analytics brainsfeed · Hong Kong 2018–2022

Things I've built

Full writeups at datacortex.in. Everything below has a public repo you can check.

Company-intelligence tool — Reads a company's website and works out what products they sell, returning organised data rather than plain text. Five steps: search, read the pages, pull out candidate products, merge duplicates, then decide which are real. Results that aren't confident get flagged for a person to check instead of being trusted automatically.

I also built the test set: a group of companies where I'd worked out the correct answer by hand, kept fixed, so I could tell whether a change made the tool better or worse. The writeup is mostly about where it got things wrong — including a change I was confident about that made accuracy worse and got reverted.

Case study · Test set and scoring script

FastAPI Python Pydantic

Voice check-in system — Handles check-in phone calls and writes up what was said. Safety checks run on the raw transcript before any model is involved, and if one provider fails it falls back to another, then to fixed rules. Context from earlier calls comes from searching those transcripts rather than stuffing everything into one long prompt.

Case study

Next.js Groq pgvector

Trade operations app — Stock, billing, ledgers, credit limits and field visits for FMCG distributors, built for a phone screen. In daily use by a distributor in Nepal. No AI in the critical path here — getting the numbers right is about careful data handling and permissions, and the writeup is honest about where I'd do that differently.

Case study

Next.js TypeScript Supabase

Open source

Python libraries published on PyPI.

Library What it does
RAGNav · src Search that combines keyword matching with meaning-based matching, so a question can be found either way. On 500 test questions, the right passage was in the top three 95.6% of the time — script and results in the repo
ragfallback · src Stops a search-and-answer system failing quietly. Rewrites weak queries, scores how confident the results are, and falls back when they're poor
AgentEnsemble · src Running several AI agents together — routing work between them, planning, and tracking what it costs
nepal-gov-agent · src Answers questions about Nepal government policy and legal documents, in Nepali and English, citing the sentence each answer came from
AgentCare · src Voice AI for healthcare — call intake, pulling out the details that matter, chasing missing information, booking appointments
scrapeflow-py · src Reading data off websites with Playwright — handles sessions, rate limits and pages that try to block you
AskPandas · src Ask questions about a CSV in plain English, using a model running on your own machine. No API keys, nothing leaves your computer
lingo-nlp-toolkit · src Small text-processing utilities that work with both older pipelines and newer models
PyroChain · src Using AI agents to work out which features matter in a dataset
toxic-comment-classifier · src Detects toxic comments, with a score per category
socialmediaextractor · src Pulls social media profile links out of websites
trustpilot-scraper · src Collects Trustpilot reviews into a structured file

All packages: pypi.org/user/irfanalidv

Papers

  • Cross-validation framework for mental-health AI on MentalChat16K — BERT and neural networks · IJAINN, Dec 2025 · DOI
  • Neural-symbolic topic evolution on Yelp reviews — multi-aspect temporal topic modelling · IJAINN, Oct 2025 · DOI

ORCID: 0000-0003-0022-3047

Contact

irfan.ali@datacortex.in · LinkedIn · datacortex.in

Pinned Loading

  1. ragfallback ragfallback Public

    ragfallback is a Python library that prevents silent RAG failures — chunk quality, retrieval fallback, adaptive querying, and answer evaluation in one package.

    Python 2

  2. AgentEnsemble AgentEnsemble Public

    AgentEnsemble is a Production-ready multi-agent orchestration for Python. ReAct, Swarm, Pipeline, Debate, Router, Planner, WorkflowGraph. Observability, cost tracking, human-in-loop. Structured out…

    Python 1

  3. AskPandas AskPandas Public

    AI-powered data engineering and analytics assistant for querying CSV data using natural language—locally and intelligently

    Python 1

  4. scrapeflow-py scrapeflow-py Public

    Production-ready web scraping engine on Playwright. LLM extraction, hybrid selectors, session persistence, rate limiting, anti-detection.

    Python 1

  5. AgentCare AgentCare Public

    Python framework for voice-AI workflows: healthcare front-desk booking, care coordination, follow-up, and workplace burnout check-ins.

    Python 1

  6. RAGNav RAGNav Public

    RAGNav is a Hybrid RAG retrieval — BM25 + embeddings + document graph. Runs fully offline. SQuAD R@3: 0.956, zero API calls.

    Python 1