AI Engineer, based in Bengaluru. I build tools that turn messy information into data you can actually use — reading web pages and documents, pulling out what matters, and checking the result is right before anyone relies on it.
Master's in Data Science & AI, IISER Tirupati, 2026.
Looking for AI engineering work — Bengaluru or remote. Available immediately. irfan.ali@datacortex.in
| Role | Org | When |
|---|---|---|
| AI Engineer | DataCortex · India | Dec 2025 – present |
| Head of AI | Kuration AI · Hong Kong | 2024–2025 |
| Senior Manager – Data & AI, R&D | Luminous Power Technologies · India | 2023–2024 |
| Data Analytics & Automation Associate | Lynk · India | 2022–2023 |
| Head of Data & Analytics | brainsfeed · Hong Kong | 2018–2022 |
Full writeups at datacortex.in. Everything below has a public repo you can check.
Company-intelligence tool — Reads a company's website and works out what products they sell, returning organised data rather than plain text. Five steps: search, read the pages, pull out candidate products, merge duplicates, then decide which are real. Results that aren't confident get flagged for a person to check instead of being trusted automatically.
I also built the test set: a group of companies where I'd worked out the correct answer by hand, kept fixed, so I could tell whether a change made the tool better or worse. The writeup is mostly about where it got things wrong — including a change I was confident about that made accuracy worse and got reverted.
Case study · Test set and scoring script
FastAPI Python Pydantic
Voice check-in system — Handles check-in phone calls and writes up what was said. Safety checks run on the raw transcript before any model is involved, and if one provider fails it falls back to another, then to fixed rules. Context from earlier calls comes from searching those transcripts rather than stuffing everything into one long prompt.
Next.js Groq pgvector
Trade operations app — Stock, billing, ledgers, credit limits and field visits for FMCG distributors, built for a phone screen. In daily use by a distributor in Nepal. No AI in the critical path here — getting the numbers right is about careful data handling and permissions, and the writeup is honest about where I'd do that differently.
Next.js TypeScript Supabase
Python libraries published on PyPI.
| Library | What it does |
|---|---|
| RAGNav · src | Search that combines keyword matching with meaning-based matching, so a question can be found either way. On 500 test questions, the right passage was in the top three 95.6% of the time — script and results in the repo |
| ragfallback · src | Stops a search-and-answer system failing quietly. Rewrites weak queries, scores how confident the results are, and falls back when they're poor |
| AgentEnsemble · src | Running several AI agents together — routing work between them, planning, and tracking what it costs |
| nepal-gov-agent · src | Answers questions about Nepal government policy and legal documents, in Nepali and English, citing the sentence each answer came from |
| AgentCare · src | Voice AI for healthcare — call intake, pulling out the details that matter, chasing missing information, booking appointments |
| scrapeflow-py · src | Reading data off websites with Playwright — handles sessions, rate limits and pages that try to block you |
| AskPandas · src | Ask questions about a CSV in plain English, using a model running on your own machine. No API keys, nothing leaves your computer |
| lingo-nlp-toolkit · src | Small text-processing utilities that work with both older pipelines and newer models |
| PyroChain · src | Using AI agents to work out which features matter in a dataset |
| toxic-comment-classifier · src | Detects toxic comments, with a score per category |
| socialmediaextractor · src | Pulls social media profile links out of websites |
| trustpilot-scraper · src | Collects Trustpilot reviews into a structured file |
All packages: pypi.org/user/irfanalidv
- Cross-validation framework for mental-health AI on MentalChat16K — BERT and neural networks · IJAINN, Dec 2025 · DOI
- Neural-symbolic topic evolution on Yelp reviews — multi-aspect temporal topic modelling · IJAINN, Oct 2025 · DOI
ORCID: 0000-0003-0022-3047





