Skip to content
View carloshgalvan95's full-sized avatar

Block or report carloshgalvan95

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
carloshgalvan95/README.md

Carlos Galvan

Data Engineer / Applied AI · Villahermosa, Mexico (open to Remote / Hybrid, relocate MX + abroad)
M.Sc. Applied Artificial Intelligence · Tecnológico de Monterrey (in progress)
English C2 (Cambridge CAE) · Spanish native

I build and operate production data pipelines and ML systems end to end: ingest, transform, serve, monitor. Most of my recent work sits at the intersection of AWS data platforms, SQL/Spark-class processing, and applied ML / retrieval when the product needs it.

LinkedIn · Email · MSc repo


Focus

Area What I actually ship
Data engineering AWS Glue, Step Functions, Airflow/Prefect, Athena, Redshift Spectrum, Docker, GitLab CI/CD
Processing Python, SQL, PySpark / Spark, Databricks (working use), dbt where modeling fits
Serving / APIs FastAPI, REST, gRPC
Applied ML / GenAI Classical ML in production, RAG (hybrid retrieval + citations), LLM tooling with human review gates

Looking at roles in Data Engineering / Senior DE / Data Platform, and AI-adjacent seats where pipelines, quality, and reliability matter as much as the model.


Selected results

Kavak · Pricing Data Engineer (2022–2025)

  • Ran AWS pricing pipelines (Glue + Step Functions + Airflow) for millions of vehicle valuations.
  • Cut deploy time ~60% with GitLab CI/CD for feature versioning and model/pipeline releases.
  • Tuned ETL paths on Athena / Redshift Spectrum / SageMaker for cost and iteration speed.

Kavak · Senior Pricing Analyst (2021–2022)

  • Inventory loss-mitigation work that recovered $104M MXN from negative-margin stock and blocked ~$1M MXN/month in further losses.
  • Automated SQL / Python / Tableau reporting (~80% less manual reporting time).

PEMEX · Data Scientist, Reservoir Engineering (2025–present)

  • Production RAG over technical engineering docs (ChromaDB + BM25 hybrid, Spanish semantic search).
  • Ops classifier on ~140,000 daily movement records → 96% accuracy across 44 concepts (hours → minutes).
  • Prefect ETL: three pipelines from 68 → 10 minutes (~70%) via parallel execution and fuzzy ID matching.
  • Water-breakthrough research framework (tree models, sequence models, Transformers, PINNs) on multi-year production history; event timing within ~2 days on held-out wells.

Stack (short list)

Python · SQL · PySpark/Spark · FastAPI
AWS: Glue · Athena · Redshift Spectrum · Step Functions · SageMaker · Lambda
Orchestration: Airflow · Prefect · GitLab CI/CD
Databricks · dbt · Docker · Kafka (concepts + listed use)
RAG: LangChain / LlamaIndex patterns · Chroma · embeddings · eval/review habits
Postgres · MySQL · SQL Server · MongoDB
Power BI · Tableau

Education

  • M.Sc. Applied Artificial Intelligence · Tecnológico de Monterrey · 2025–2027 (current)
    Recent coursework includes Big Data, NLP, MLOps (TC5061), software design (TC5062).
  • Certificate, Data Analytics and Visualization · Tecnológico de Monterrey · 2022
  • B.S. Petroleum Engineering · Universidad Olmeca · 2018

Coursework and public MSc materials: github.com/carloshgalvan95/MNA


Public repos

Repo Why it exists
MNA Applied AI master's work
personal_finance_gui Small TypeScript finance tracker
Square_Meter_Value_Real_Estate Feature analysis for price / m²
World_Weather_Analysis API-driven weather / travel analysis

Proprietary employer systems stay private; the bullets above are the parts I can describe publicly.


Contact

Pinned Loading

  1. PyBer_Analysis PyBer_Analysis Public

    Analysis to address the disparities found between the average fares for every type of city using Pandas Python library, Jupiter Notebooks and Matplotlib for data visualization.

    Jupyter Notebook

  2. School_District_Analysis School_District_Analysis Public

    School district analysis based on overall grade performance to determine how much budget impacts the grades of students using Pandas Python library and Jupiter Notebooks.

    Jupyter Notebook

  3. World_Weather_Analysis World_Weather_Analysis Public

    World weather analysis to determine correlations between latitude and weather conditions to use as parameters on trip recommendations given the desired weather conditions using Python Pandas, SciPy…

    Jupyter Notebook 1

  4. Craigslist_Used_Cars_Analysis Craigslist_Used_Cars_Analysis Public

    Craigslist used cars database analysis to determine correlations between price depreciation and manufacturer, model and age of used cars

    Jupyter Notebook

  5. Square_Meter_Value_Real_Estate Square_Meter_Value_Real_Estate Public

    Most important factor that influence the price per square meter of a property.

    Jupyter Notebook

  6. personal_finance_gui personal_finance_gui Public

    A personal finance application with GUI for tracking income, expenses, and financial goals

    TypeScript