Skip to content
View rbhughes's full-sized avatar

Block or report rbhughes

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
rbhughes/README.md

Hi, I'm Bryan πŸ‘‹

I'm a Senior AI Engineer in Chicago. I've spent 20+ years wrangling petroleum and geoscience data β€” the messy, vendor-locked kind β€” and I now build AI systems on top of it that answer questions from real data without making things up. My portfolio lives at purr.io.

πŸ› οΈ Recent contract work

I built the data pipeline, AI backend and CLI test framework for a commercial oil & gas analytics product: a fact-gating harness turns English queries or spatial AOIs into insight that carries only the numbers the deterministic stats support. Dataset-centric agents run ad-hoc SQL, generate and edit charts, and monitor their own quality. The harness detects any ephemeral drops in LLM quality and retries. The full Oracle-to-Parquet pipeline opens well, production, spatial, legal, financial and forecasting data to AI-assisted investigation.

Proprietary work β€” the client is unnamed and there is no code to show. The same ideas, in public and with code you can read, are below.

πŸ”¬ In the open at purr.io β€” findings, not features

  • walker.purr.io (agentic_dog_walker) β€” watch a small LLM plan dog-walking routes live: every tool call, validation bounce, and audit veto streamed as it happens. A hand-rolled agent loop with a deterministic referee and auditor, an MCP facade, and a model picker whose entries must pass an automated qualification gauntlet. Served from a retired laptop in a closet for about $1/month.
  • spacing.purr.io (well-spacing-playbook) β€” how close is too close for horizontal wells? The naive regression has the wrong sign; holding geology fixed flips it. A calibrated P10/P50/P90 interference model on public Alberta data, all 105,724 laterals mapped, and an honest account of where public data ends.
  • methane.purr.io (methane-outliers) β€” who flares and vents the most, relative to their own production? Alberta and Texas side by side on identical rolling windows, refreshed on schedule by a zero-cost pipeline (GitHub Actions + Cloudflare R2).
  • comed.purr.io (power_puddle) β€” in 2025 this project called PJM's data-center load forecast a puddle of wishful thinking. Volume 2 grades that call a year later, with receipts.
  • kingfisher.purr.io (kingfisher_wells) β€” three commercial data vendors, one Oklahoma county, and well locations that can't all be right. A medallion pipeline preserved as an interactive data-quality museum.
  • clay-ai-hyperscale β€” a failed attempt to fine-tune the Clay foundation model to spot data centers from satellite imagery. The post-mortem is the point; the hand-labeled dataset is published for anyone who wants to beat the baseline.

πŸ€– How I work with AI

Every commit here is co-authored with Claude. I'd rather tell you how than let you guess.

The same rule governs the code and the process: language at the boundaries, deterministic code in the middle. Anything with an exact answer β€” route order, safety thresholds, schema validation, which well an identifier refers to β€” is plain code with tests; the model translates intent and narrates results. That's the stated architecture of agentic_dog_walker, it's why geo-mini-rag replaced embedding search with exact lookup once I measured that embeddings cannot resolve identifiers, and on contract work it is a grounding layer that checks every generated number against the rows behind it before anyone sees it.

How I direct the work is written down per project, in the CLAUDE.md files committed next to the code β€” including one that overrides autonomous mode outright: explain before writing, one step at a time, and leave the parts that carry the learning for me to type.

None of it ships on vibes. The fact-gated harness described above is one example; agentic_dog_walker does the same to models, where none reaches the picker until it clears an automated qualification gauntlet.

AI makes me faster. The harness is what makes me willing to ship.

πŸ›’οΈ Earlier: energy data tooling

pg_ppdm (the PPDM 3.9 well-data model, Oracle β†’ PostgreSQL) Β· purr_petra_cli Β· purr_geographix / purr_petra β€” extraction tooling for vendor-locked geoscience project databases.

πŸ“« Reach me

Pinned Loading

  1. power_puddle power_puddle Public

    A Data Engineering workflow to assess ComEd (Northern Illinois) readiness for hyperscaler growth using Prefect, dbt and Grafana

    Python

  2. agentic_dog_walker agentic_dog_walker Public

    A small LLM plans dog walks live at walker.purr.io β€” hand-rolled agent loop, deterministic referee + auditor, structured finish line, and a model-qualification harness. Becoming an honest way to me…

    Python

  3. geo-mini-rag geo-mini-rag Public

    A small RAG pipeline for cheap models, and what it takes to read oil and gas file formats into one. Embeddings cannot locate identifiers; the fix is exact lookup, measured.

    Python

  4. well-spacing-playbook well-spacing-playbook Public

    How close is too close for horizontal wells? A measured answer, a calibrated P10/P50/P90 interference model over 105,724 Alberta laterals, and an honest account of where public data runs out

    Python

  5. methane-outliers methane-outliers Public

    Peer-expectation outlier scores for vented and flared gas β€” Alberta facilities and Texas leases, same lens, two regulatory regimes β€” on a zero-cost GitHub Actions + DuckDB + Cloudflare R2 pipeline

    Python

  6. clay-ai-hyperscale clay-ai-hyperscale Public

    A (failed) attempt to train the Clay AI Foundation model to detect data centers

    Python