Unstructured complaints in → clustered, scored, ranked root-cause categories out. No manual tagging. No keyword lists maintained by hand.
InsightWell takes a raw stream of customer complaint text (currently the Twitter US Airline Sentiment dataset as a public proxy for a support inbox) and turns it into an operational dashboard that answers three questions a support/ops team actually cares about:
- What are people complaining about, grouped how a human would group them — not raw keyword buckets.
- How bad is each category, right now — sentiment mix, volume share, and whether it's rising or fading.
- Where do I look first — a single ranked list, weighted by volume, negativity, trend, and label confidence.
The pipeline runs offline in Python and emits one static JSON artifact (pipeline/output/insights.json). The frontend is a fully static Next.js dashboard that reads that artifact — no backend, no database, no live inference in the browser.
┌─────────────────┐ ┌──────────────────────┐ ┌───────────────────┐
│ Tweets.csv │ ───▶ │ run_pipeline.py │ ───▶ │ insights.json │
│ (raw complaints)│ │ clean → embed → │ │ (single artifact) │
│ │ │ cluster → sentiment │ │ │
│ │ │ → aggregate │ │ │
└─────────────────┘ └──────────────────────┘ └─────────┬──────────┘
│
▼
┌───────────────────┐
│ Next.js dashboard │
│ (static, client) │
└───────────────────┘
pipeline/run_pipeline.py — a single, linear, six-step script. No orchestrator, no DAG framework; the whole thing runs top to bottom in one process.
| Step | What happens |
|---|---|
| 1. Download | Pulls the raw dataset if not already cached locally. |
| 2. Clean | Strips URLs, @mentions, HTML entities; collapses whitespace. |
| 3. Topic model | Embeds every complaint with all-MiniLM-L6-v2, clusters with BERTopic, reduces to ~15 topics, then maps each topic's top c-TF-IDF keywords to a human-readable label ("Flight Delays", "Lost Baggage", "Rude Staff", …) via ordered keyword-matching rules. |
| 4. Sentiment | Scores every complaint with cardiffnlp/twitter-roberta-base-sentiment-latest (positive / neutral / negative), batched, on GPU if available. Accuracy is validated against the dataset's own ground-truth labels and reported in the output. |
| 5. Aggregate | Per category: volume, volume share, sentiment mix, a 14-bin time trend, weekday distribution, and a composite radar score (volume / negativity / trend / confidence). |
| 6. Write | Emits pipeline/output/insights.json — the single contract between pipeline and frontend. |
Design choices worth knowing:
- The BERTopic outlier bucket (
topic_id == -1) is kept as its own"Other / Uncategorized"category rather than force-merged viareduce_outliers()— so no real category's volume/severity gets inflated by forced reassignment. confidence_scoreis not model-derived — it's the mean of the dataset's own human-labeledairline_sentiment_confidencefor each topic's tweets. Real signal, not a synthetic proxy.trend_scoreis the literal share of a topic's volume in the second half of the time range (>50 = rising, <50 = fading) — computed from real timestamps, not a fitted slope.
cd pipeline
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python run_pipeline.pyweb/ — Next.js 16 (App Router) + React 19 + Tailwind v4, reading the static insights.json at build/runtime with zero backend calls.
|
Components
|
Stack
|
cd web
npm install
npm run dev # http://localhost:3000InsightWell/
├── pipeline/
│ ├── run_pipeline.py single-file ETL: clean → topic model → sentiment → aggregate
│ ├── requirements.txt
│ ├── data/Tweets.csv raw dataset (downloaded on first run)
│ └── output/insights.json ← the one contract with the frontend
│
└── web/
├── app/ Next.js App Router entry (layout, page, globals)
├── components/ dashboard visual components (charts, tiles, tables)
├── lib/
│ ├── insights.ts loads/types the pipeline JSON
│ ├── adapters.ts insights.json → component-shaped view models
│ └── types.ts
└── public/ favicon / app icon
Everything the frontend renders comes from one JSON shape:
Swap pipeline/data/Tweets.csv for any complaint dataset with a text column, timestamps, and (optionally) a ground-truth sentiment label — the pipeline and dashboard don't hardcode airline-specific logic beyond the topic label rules.
Built with BERTopic, a RoBERTa sentiment model, and Next.js — no manual labeling in the loop.
{ "generated_at": "2026-07-10T19:11:20Z", "methodology": { "dataset": "Twitter US Airline Sentiment (public proxy dataset)", "total_complaints_analyzed": 14640, "topic_model": "BERTopic (all-MiniLM-L6-v2 embeddings)", "sentiment_model": "cardiffnlp/twitter-roberta-base-sentiment-latest", "sentiment_accuracy_vs_ground_truth": 0.7719 }, "overview": { "total_volume": 14640, "overall_negative_pct": 51.95, "...": "..." }, "timeseries": [ { "date": "Feb 16", "positive": 0, "neutral": 3, "negative": 1 } ], "heatmap": { "...": "category × weekday matrix" }, "categories": [ { "id": "customer-service", "label": "Customer Service", "keywords": ["flight", "bag", "service", "..."], "volume": 6821, "volume_pct": 46.59, "sentiment": { "positive": 17.55, "neutral": 26.52, "negative": 55.93 }, "radar": { "volume_score": 100.0, "negative_score": 55.93, "trend_score": 62.48, "confidence_score": 90.87 }, "trend": [249, 415, 387, "..."], "sample_complaints": [{ "text": "...", "created_at": "...", "sentiment": "negative" }] } ] }