Data Engineer / Applied AI · Villahermosa, Mexico (open to Remote / Hybrid, relocate MX + abroad)
M.Sc. Applied Artificial Intelligence · Tecnológico de Monterrey (in progress)
English C2 (Cambridge CAE) · Spanish native
I build and operate production data pipelines and ML systems end to end: ingest, transform, serve, monitor. Most of my recent work sits at the intersection of AWS data platforms, SQL/Spark-class processing, and applied ML / retrieval when the product needs it.
| Area | What I actually ship |
|---|---|
| Data engineering | AWS Glue, Step Functions, Airflow/Prefect, Athena, Redshift Spectrum, Docker, GitLab CI/CD |
| Processing | Python, SQL, PySpark / Spark, Databricks (working use), dbt where modeling fits |
| Serving / APIs | FastAPI, REST, gRPC |
| Applied ML / GenAI | Classical ML in production, RAG (hybrid retrieval + citations), LLM tooling with human review gates |
Looking at roles in Data Engineering / Senior DE / Data Platform, and AI-adjacent seats where pipelines, quality, and reliability matter as much as the model.
- Ran AWS pricing pipelines (Glue + Step Functions + Airflow) for millions of vehicle valuations.
- Cut deploy time ~60% with GitLab CI/CD for feature versioning and model/pipeline releases.
- Tuned ETL paths on Athena / Redshift Spectrum / SageMaker for cost and iteration speed.
- Inventory loss-mitigation work that recovered $104M MXN from negative-margin stock and blocked ~$1M MXN/month in further losses.
- Automated SQL / Python / Tableau reporting (~80% less manual reporting time).
- Production RAG over technical engineering docs (ChromaDB + BM25 hybrid, Spanish semantic search).
- Ops classifier on ~140,000 daily movement records → 96% accuracy across 44 concepts (hours → minutes).
- Prefect ETL: three pipelines from 68 → 10 minutes (~70%) via parallel execution and fuzzy ID matching.
- Water-breakthrough research framework (tree models, sequence models, Transformers, PINNs) on multi-year production history; event timing within ~2 days on held-out wells.
Python · SQL · PySpark/Spark · FastAPI
AWS: Glue · Athena · Redshift Spectrum · Step Functions · SageMaker · Lambda
Orchestration: Airflow · Prefect · GitLab CI/CD
Databricks · dbt · Docker · Kafka (concepts + listed use)
RAG: LangChain / LlamaIndex patterns · Chroma · embeddings · eval/review habits
Postgres · MySQL · SQL Server · MongoDB
Power BI · Tableau
- M.Sc. Applied Artificial Intelligence · Tecnológico de Monterrey · 2025–2027 (current)
Recent coursework includes Big Data, NLP, MLOps (TC5061), software design (TC5062). - Certificate, Data Analytics and Visualization · Tecnológico de Monterrey · 2022
- B.S. Petroleum Engineering · Universidad Olmeca · 2018
Coursework and public MSc materials: github.com/carloshgalvan95/MNA
| Repo | Why it exists |
|---|---|
| MNA | Applied AI master's work |
| personal_finance_gui | Small TypeScript finance tracker |
| Square_Meter_Value_Real_Estate | Feature analysis for price / m² |
| World_Weather_Analysis | API-driven weather / travel analysis |
Proprietary employer systems stay private; the bullets above are the parts I can describe publicly.
- LinkedIn: linkedin.com/in/carloshgalvan
- Email:
carloshgalvan95@gmail.com - Phone:
+52 993 323 4114