From adccfdcc9e08c2cb0c986f94771ed33d74156122 Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 14 Jun 2026 11:27:04 +0000 Subject: [PATCH 1/3] Add research and plan for aligning with dbt, FHIR, and other tools Maps Modelable's domain/model/projection/lineage concepts onto dbt (model contracts, versions, exposures, meta) and FHIR (StructureDefinition, profiles, extensions, canonical URLs), proposes phased export/import emitters, and surveys other alignment candidates (OpenLineage, Iceberg/Delta, Snowplow/Segment, OMOP CDM, GraphQL federation). https://claude.ai/code/session_011UdKKjqh8GXnoKTZyYTbbZ --- ROADMAP.md | 4 + docs/README.md | 1 + docs/dbt-fhir-tool-alignment.md | 282 ++++++++++++++++++++++++++++++++ 3 files changed, 287 insertions(+) create mode 100644 docs/dbt-fhir-tool-alignment.md diff --git a/ROADMAP.md b/ROADMAP.md index 5785409d..ac1bfa2b 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -16,6 +16,10 @@ item is not committed until it has an issue and an accepted design. - External artifact registry integration. - Catalog and governance-system integration. - Open Data Contract Standard interchange. +- dbt schema/source export and import (see + [docs/dbt-fhir-tool-alignment.md](docs/dbt-fhir-tool-alignment.md)). +- FHIR R4 profile export and import for a small base-resource set (see + [docs/dbt-fhir-tool-alignment.md](docs/dbt-fhir-tool-alignment.md)). - Additional artifact formats driven by concrete consumers. ## Later diff --git a/docs/README.md b/docs/README.md index 14a76b7c..faaea7ed 100644 --- a/docs/README.md +++ b/docs/README.md @@ -31,6 +31,7 @@ example. The documents here describe the language and the larger system model. - [Platform usage scenarios](platform-usage-scenarios-spec.md) - [Technology evaluation](technology-evaluation.md) - [Data-model language research](data-model-languages.md) +- [dbt, FHIR, and other tool alignment research](dbt-fhir-tool-alignment.md) Specifications include future phases where clearly labelled. Current release scope is summarized in the root README and [ROADMAP.md](../ROADMAP.md). diff --git a/docs/dbt-fhir-tool-alignment.md b/docs/dbt-fhir-tool-alignment.md new file mode 100644 index 00000000..5635ec58 --- /dev/null +++ b/docs/dbt-fhir-tool-alignment.md @@ -0,0 +1,282 @@ +# Research and Plan: Aligning with dbt, FHIR, and Other External Tools + +> **Status:** Research and planning. No emitter, importer, or CLI work is +> committed by this document. It extends +> [external-tools-data-modelling.md](external-tools-data-modelling.md) and +> [migration-guide.md](migration-guide.md) with concept mappings and phased +> proposals for dbt, FHIR, and other ecosystems. Concrete work requires an +> issue and an accepted design per [ROADMAP.md](../ROADMAP.md). + +## 1. Purpose + +Teams adopting Modelable rarely start from a blank slate. Analytics teams +already run dbt; healthcare and life-sciences teams already exchange FHIR +resources; most organizations already have a catalog, lineage, or table-format +standard in place. This document maps Modelable's core concepts (domain, +model, model version, projection, lineage, classification) onto the concepts +of dbt and FHIR, proposes alignment work in phases consistent with the +existing roadmap, and surveys other tools worth evaluating. + +Alignment here means two things, in priority order: + +1. **Conceptual alignment** — using compatible terminology and version/derivation + semantics so teams can map their existing dbt/FHIR artifacts onto `.mdl` + without re-learning a new model of the world. +2. **Artifact alignment** — optional emitters/importers that generate or + consume dbt and FHIR artifacts from the normalized Modelable graph. + +Per [modelable-system-spec.md](modelable-system-spec.md) §2.6 +(framework-first integration), Modelable should wrap and interoperate with +these tools, not replace or execute them. + +## 2. dbt (data build tool) + +### 2.1 Why dbt matters + +dbt is the de facto standard for SQL-based transformation and documentation +in the analytics/warehouse layer. Domains that own canonical models in +Modelable frequently also own one or more dbt models that materialize those +models (or projections of them) into a warehouse. dbt's newer governance +features — model contracts, model versions, groups, access modifiers, and +exposures — cover much of the same ground as Modelable's model versions, +domain ownership, and projections, but scoped to the warehouse. + +### 2.2 Concept mapping + +| dbt concept | Modelable concept | Notes | +|---|---|---| +| `sources:` (raw, unmanaged tables) | External upstream not yet modeled, or a `binding` source | dbt sources describe data Modelable does not own; a Modelable domain may still declare a `binding` pointing at the same table once it has a canonical model | +| `models:` with `contract: {enforced: true}` | Published Model Version | Both are schema-enforced and intended to be stable contracts for downstream consumers | +| `versions:` / `latest_version` (dbt model versioning) | `Model @ N (additive\|breaking)` | Both express "this contract changed; old and new versions can coexist"; dbt expresses the diff per version, Modelable redeclares the full version | +| `columns:` with `name`, `data_type`, `constraints` | Field declarations with types and `@key`/constraints | Direct field-level mapping; dbt `data_type` is warehouse-specific and needs a type-mapping table similar to the existing JSON Schema mapping | +| `meta:` arbitrary key/value | `@classification(...)`, `@owner(...)`, custom annotations | dbt has no fixed governance vocabulary; Modelable's annotations are more structured | +| `group:` + `access: private\|protected\|public` | Domain ownership + projection visibility | dbt groups approximate domain ownership; `access` approximates "is this a published projection or an internal model" | +| `exposures:` (dashboards, applications, ML models that consume a model) | Projections / declared consumers | Both describe "who consumes this contract and how" | +| generic/singular `tests` | `constraints` on a Model Version | Both express validation rules attached to a schema | +| `semantic_models:` / `metrics:` (MetricFlow / dbt Semantic Layer) | No current equivalent | See §2.4 — potential future "metric" or aggregation-projection construct | +| `manifest.json` / `catalog.json` (dbt docs) | Normalized model graph / `modelable lineage` output | Both are machine-readable graphs of models, columns, and dependencies | + +### 2.3 Alignment plan + +**Phase A — dbt schema/source export (extends Phase 1/4 of +[external-tools-data-modelling.md](external-tools-data-modelling.md)):** + +Add a `dbt-yaml` (working name) compile target that generates dbt +`schema.yml` fragments for a model or projection: + +```bash +modelable compile customer.Customer@2 --target dbt-yaml --out ./dist/dbt +``` + +Generated output should map: + +```text +Modelable Model/Projection -> dbt model/source entry +field name + type -> column name + data_type (per-adapter type map) +@key -> constraints: [{type: primary_key}] +@pii / @classification -> meta: {modelable_classification: ...} +owner -> meta: {modelable_owner: ...} +model version -> versions: [{v: N, ...}] or a versioned model name +lineage (projection) -> meta: {modelable_lineage: [...]} +``` + +This lets a Modelable canonical model be dropped into an existing dbt project +as a documented, contract-enforced source or model stub without hand-writing +YAML. + +**Phase B — dbt import (extends +[migration-guide.md](migration-guide.md) §3 source-format table):** + +`modelable generate --from --output +models/.mdl` to bootstrap `.mdl` models from an existing dbt project, +following the same "review the generated output" workflow as other +LLM-assisted imports. + +**Phase C — exposure/lineage stitching:** + +Treat dbt `exposures` as external consumers in the lineage graph, so +`modelable lineage` can show "this field flows into dbt exposure X" even when +the exposure itself lives outside `.mdl`. This feeds +[distributed-lineage-spec.md](distributed-lineage-spec.md) rather than the +local compiler. + +**Phase D — semantic layer (deferred, see §2.4).** + +### 2.4 Open questions + +- dbt model versioning expresses only the *diff* between versions + (`versions: [{v: 1, columns: [...]}]` reusing a shared base); Modelable + redeclares each version in full. The emitter should generate full + per-version `columns:` blocks rather than attempt diff-based output — + simpler and avoids inferring dbt's diff format from Modelable's diff engine. +- dbt `data_type` is adapter-specific (Snowflake vs. BigQuery vs. Postgres + types differ). A `dbt-yaml` emitter needs either a single target-adapter + type map (configurable) or to omit `data_type` and rely on `contract: + {enforced: false}` for the initial pass. +- MetricFlow `semantic_models`/`metrics` have no Modelable equivalent today. + Do not add a "metric" model kind speculatively; revisit only if a concrete + consumer needs aggregation-as-contract beyond the existing `group by` + projection aggregation (idl-design-spec.md §3.4). + +## 3. FHIR (Fast Healthcare Interoperability Resources) + +### 3.1 Why FHIR matters + +FHIR (HL7) is the dominant interoperability standard for healthcare data +exchange. Any domain dealing with clinical or health-plan data will need its +canonical models to interoperate with FHIR Resources, Profiles, and +Extensions. FHIR's profiling mechanism — constraining or extending a base +Resource via a `StructureDefinition` with `derivation: constraint` — is +conceptually close to a Modelable projection that derives from a source model +with field-level lineage. + +### 3.2 Concept mapping + +| FHIR concept | Modelable concept | Notes | +|---|---|---| +| Resource (e.g., `Patient`, `Observation`, `Encounter`) | Canonical `entity`/`event` model | Base resources are externally owned canonical contracts, analogous to a model owned by an "HL7" domain | +| `StructureDefinition` (`derivation: specialization`) | Model Version schema | Defines fields (elements), cardinality, types, and terminology bindings | +| Profile (`StructureDefinition`, `derivation: constraint`, `baseDefinition: `) | Projection deriving from a source model | A profile narrows cardinality/types and adds extensions on top of a base resource — same shape as `projection X from domain.Model@N as m { ... }` | +| Extension (`StructureDefinition`, `type: Extension`, stable `url`) | Additive optional field (`(additive)` version, `?` field) | FHIR extensions are versionless, stable-URL additive fields — close to Modelable's additive-change discipline | +| Canonical URL + `version` (business version) | `domain.Model@version` | Both are globally unique, versioned identifiers for a contract | +| `ValueSet` / `CodeSystem` + `binding.strength` | `enum(...)` (or `ref` for shared vocabularies) | Controlled vocabularies; binding strength (`required`/`extensible`/`preferred`/`example`) has no direct Modelable equivalent today | +| `Reference(ResourceType)` | `ref` | Cross-resource/cross-domain reference | +| `meta.security` / `meta.tag` | `@classification(...)`, `@pii` | Security labels and classification tags | +| Implementation Guide (bundle of profiles + narrative + examples) | Workspace + generated Markdown docs | An IG is a versioned, documented bundle of profiles — analogous to a Modelable workspace's generated docs output | +| `CapabilityStatement` | Adapter binding capabilities | Declares what operations/resources a server supports | + +### 3.3 Alignment plan + +**Phase A — FHIR profile export (export-only, R4 first):** + +Add a `fhir-profile` (working name) compile target that, given a Modelable +projection whose lineage traces to a declared FHIR base resource, generates a +FHIR R4 `StructureDefinition` with `derivation: constraint`: + +```bash +modelable compile clinical.PatientSummary@1 --target fhir-profile --out ./dist/fhir +``` + +Mapping: + +```text +Modelable Projection -> StructureDefinition (derivation: constraint) +projection name + version -> StructureDefinition.url + .version +baseDefinition -> declared via projection source annotation +field <- source.field -> ElementDefinition with sliced/renamed path +enum(...) -> ElementDefinition.binding (valueSet, strength: required) +ref -> ElementDefinition type Reference(ResourceType) +@pii / @classification -> meta.security coding +``` + +Start with a small, explicitly supported set of base resources (e.g. +`Patient`, `Observation`, `Encounter`) rather than attempting full coverage of +the FHIR resource catalog. + +**Phase B — terminology and reference mapping:** + +- Map Modelable `enum(...)` to a FHIR `binding` (`valueSet` + `strength`). + Modelable does not need its own ValueSet/CodeSystem resources initially — + reference external canonical URLs. +- Map `ref` to FHIR `Reference(ResourceType)` when the target + model corresponds to a known FHIR resource. + +**Phase C — FHIR import (extends +[migration-guide.md](migration-guide.md) §3):** + +`modelable generate --from --output +models/.mdl` to draft a starting `.mdl` model/projection from an +existing profile, with the same human-review workflow as other imports. + +**Phase D (Later, roadmap) — Implementation Guide packaging:** + +Generate an IG-shaped documentation bundle (profiles + narrative + examples) +from a Modelable workspace, reusing the Markdown emitter. + +### 3.4 Open questions / caveats + +- FHIR's type system includes deeply nested `BackboneElement`s and recursive + `Extension` structures, which are richer than Modelable's flat + field/value-object model. Deep nesting should map to nested Modelable + `value` models; very deep or recursive FHIR structures may not be fully + representable and should fail with a clear `EMIT003`/`EMIT002`-style + diagnostic (per [emitter-spec.md](emitter-spec.md) §10) rather than partial + output. +- FHIR has multiple concurrently active versions (R4, R4B, R5, and R6 in + ballot). An emitter must target one FHIR version explicitly. **Recommend + R4** as the first target — it remains the most widely deployed version in + production health systems. +- This is export/import of static profile artifacts only. FHIR servers (e.g. + HAPI FHIR), `CapabilityStatement`-driven runtime conformance, and FHIR + Subscriptions are runtime concerns and stay out of scope, consistent with + the Phase 5 boundary in + [external-tools-data-modelling.md](external-tools-data-modelling.md) and + [technology-evaluation.md](technology-evaluation.md). + +## 4. Other tools to evaluate for alignment + +| Tool / standard | What it is | Why relevant to Modelable | Suggested alignment | Phase | +|---|---|---|---|---| +| **OpenLineage** | Open standard for lineage events (job/run/dataset/column facets); adopted by Airflow, Spark, dbt, OpenMetadata, and major cloud catalogs | Modelable's internal lineage graph could be exported as OpenLineage `ColumnLineageDatasetFacet` events, letting catalogs that already consume OpenLineage ingest Modelable lineage without a bespoke integration | Add an OpenLineage export alongside the planned OpenMetadata export (Phase 3) | 3 | +| **Open Data Contract Standard (ODCS) / Data Contract CLI** | Vendor-neutral data contract interchange format | Already on the roadmap (Phase 4); reaffirm — dbt model contracts and FHIR profiles both have partial overlap with ODCS fields (owner, classification, quality) | No change — keep as Phase 4 | 4 | +| **Apache Iceberg / Delta Lake (table formats)** | Open table formats with schema evolution (add/rename/widen columns with stable field IDs) | Schema evolution semantics (stable column IDs, additive-only safe changes) closely mirror Modelable's additive/breaking model and the field-ID concern already flagged for Protobuf | Potential `--target iceberg-schema` emitter reusing the same field-ID stability mechanism proposed for Protobuf | 5 | +| **Snowplow / Segment tracking plans** | Versioned event-schema governance for product analytics | "Event model + classification + versioning" maps closely to Modelable's `event` kind and `@classification` | Potential compile target for analytics/event-tracking teams | 5 | +| **OMOP CDM** | Common Data Model for healthcare observational research (alternative to FHIR for analytics use cases) | Worth a follow-up evaluation if FHIR's operational profile model proves too heavyweight for analytics-only healthcare domains | Evaluate only if a concrete healthcare-analytics consumer emerges; do not build speculatively | Later | +| **GraphQL SDL / federation (Apollo subgraphs)** | Schema definition language with subgraph ownership and composition | Subgraph ownership and composed schema concepts parallel Modelable domain ownership and cross-domain projections | Potential future compile target alongside OpenAPI (Phase 5) | 5 | + +Tools already evaluated and not repeated here: JSON Schema, Avro, Protobuf, +OpenAPI, AsyncAPI, Apicurio, OpenMetadata, LinkML — see +[external-tools-data-modelling.md](external-tools-data-modelling.md) and +[data-model-languages.md](data-model-languages.md). + +## 5. Recommended sequencing + +This slots into the existing phased plan from +[external-tools-data-modelling.md](external-tools-data-modelling.md) and +[ROADMAP.md](../ROADMAP.md): + +| Phase | Existing focus | New additions from this document | +|---|---|---| +| 1 — Local modelling compiler | JSON Schema, Markdown, TypeScript | none | +| 2 — Artifact registry | Apicurio | none | +| 3 — Catalog/governance sync | OpenMetadata | + OpenLineage export | +| 4 — Contract interchange | ODCS, Data Contract CLI | + dbt `schema.yml`/source export and import | +| 4b (new) — Domain-specific interchange | — | FHIR R4 profile export (small resource set) and import | +| 5 — Event/API/runtime targets | Avro, Protobuf, OpenAPI, AsyncAPI, runtime stack | + Iceberg/Delta schema target, analytics tracking-plan target, GraphQL SDL | + +## 6. Non-goals + +- Executing dbt, running a FHIR server, or collecting OpenLineage runtime + events — these are runtime/execution concerns, consistent with + [modelable-system-spec.md](modelable-system-spec.md) §2.6 and the "Do Not + Incorporate Yet" list in + [external-tools-data-modelling.md](external-tools-data-modelling.md). +- Redesigning the core `.mdl` type system or projection model to match dbt's + or FHIR's type systems. All mapping happens in emitters/importers, not in + the IDL or normalized graph. +- Adding a "metric"/semantic-layer model kind speculatively to mirror + MetricFlow. Revisit only against a concrete requirement. + +## 7. Open decisions + +- Whether dbt and FHIR emitters/importers are first-party (in `cli/`) or + third-party plugins, pending the plugin-registry decision already open in + [emitter-spec.md](emitter-spec.md) §11. +- Which FHIR base resources are in scope for Phase 4b (proposed starting set: + `Patient`, `Observation`, `Encounter`). +- Which warehouse dialect's `data_type` vocabulary the dbt emitter targets by + default (or whether it omits `data_type` until `contract.enforced` is + requested). + +## 8. Dependencies + +- [external-tools-data-modelling.md](external-tools-data-modelling.md) — + phased external-tool roadmap this document extends +- [migration-guide.md](migration-guide.md) — import/source-format table this + document extends +- [emitter-spec.md](emitter-spec.md) — emitter interface, diagnostics, and + open decisions referenced for new targets +- [distributed-lineage-spec.md](distributed-lineage-spec.md) — cross-tool + lineage stitching for dbt exposures and OpenLineage +- [ownership-permissions-spec.md](ownership-permissions-spec.md) — + classification/ownership mapping for dbt `meta` and FHIR `meta.security` From 9fa23a54d9113048bb864e2d1adeae5b155396bb Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 14 Jun 2026 12:02:30 +0000 Subject: [PATCH 2/3] Add modelable attach: track external dbt/FHIR drift as new model versions Imports a dbt schema.yml model or FHIR R4 StructureDefinition, diffs its fields against an existing published model version with the existing compatibility-diff engine, and appends a new additive/breaking version plus an .attachments.json provenance record when the source has drifted. --- cli/pyproject.toml | 2 + cli/src/modelable/commands/llm.py | 85 +++++++- cli/src/modelable/llm/engine.py | 185 ++++++++++++++++++ cli/src/modelable/llm/importers.py | 165 +++++++++++++++- cli/src/modelable/llm/provenance.py | 30 +++ cli/tests/test_llm_features.py | 290 ++++++++++++++++++++++++++++ cli/uv.lock | 28 +++ docs/cli-spec.md | 30 +++ docs/dbt-fhir-tool-alignment.md | 37 +++- 9 files changed, 842 insertions(+), 10 deletions(-) diff --git a/cli/pyproject.toml b/cli/pyproject.toml index 0319d1bd..377a77da 100644 --- a/cli/pyproject.toml +++ b/cli/pyproject.toml @@ -43,6 +43,7 @@ dependencies = [ "referencing>=0.35", "pygls>=2.1.1,<3", "psycopg[binary]>=3.2", + "pyyaml>=6.0.3", ] [project.urls] @@ -111,6 +112,7 @@ module = [ "pygls.*", "lsprotocol.*", "psycopg.*", + "yaml.*", ] ignore_missing_imports = true diff --git a/cli/src/modelable/commands/llm.py b/cli/src/modelable/commands/llm.py index 7c38078c..70f3588e 100644 --- a/cli/src/modelable/commands/llm.py +++ b/cli/src/modelable/commands/llm.py @@ -1,5 +1,6 @@ from __future__ import annotations +from dataclasses import asdict from pathlib import Path import click @@ -12,11 +13,13 @@ from modelable.llm.context import build_workspace_summary from modelable.llm.engine import ( answer_model_question_cli, + attach_external_version, describe_path_or_ref, explain_validation, generate_entity_from_prompt, import_definition, recommend_cli, + render_attach_audit_summary, render_update_audit_summary, render_write_audit_summary, suggest_projection, @@ -24,7 +27,12 @@ update_definition, validate_generated_text, ) -from modelable.llm.provenance import build_write_provenance, write_provenance_sidecar +from modelable.llm.provenance import ( + AttachmentRecord, + build_write_provenance, + write_attachment_record, + write_provenance_sidecar, +) from modelable.llm.providers import build_provider console = Console() @@ -36,6 +44,7 @@ def register_llm_commands(cli_group: click.Group) -> None: cli_group.add_command(import_model) cli_group.add_command(diff) cli_group.add_command(update) + cli_group.add_command(attach) cli_group.add_command(transform) cli_group.add_command(suggest_projection_cmd) cli_group.add_command(ask) @@ -240,6 +249,80 @@ def update( console.print(render_update_audit_summary(result)) +@click.command() +@click.argument("ref") +@click.option( + "--source", + "source", + type=click.Path(exists=True, path_type=Path), + required=True, + help="External dbt schema.yml or FHIR StructureDefinition JSON file.", +) +@click.option("--source-format", "source_format", type=click.Choice(["dbt", "fhir"]), required=True) +@click.option( + "--source-name", + "source_name", + default=None, + help="dbt model name or FHIR resource name to match, if the source defines multiple.", +) +@click.option("--path", "path", type=click.Path(exists=True, path_type=Path), required=True) +@click.option("--output", "output", type=click.Path(path_type=Path), default=None) +@click.option("--preview", is_flag=True, help="Show the new version diff without writing changes.") +def attach( + ref: str, + source: Path, + source_format: str, + source_name: str | None, + path: Path, + output: Path | None, + preview: bool, +) -> None: + """Attach a model version to an external dbt or FHIR source and record drift as a new version.""" + try: + result = attach_external_version( + path, ref, source, source_format, source_name=source_name, output=output, write=not preview + ) + except ValueError as exc: + raise click.ClickException(str(exc)) from exc + + for warning in result.warnings: + console.print(f"[yellow]WARN[/yellow] {warning}") + + if not result.attached: + console.print( + f"[green]OK[/green] {ref} already matches {result.source_descriptor} " + f"({result.source_format}); no new version created" + ) + return + + if preview: + from modelable.llm.chat import _render_update_preview + + console.print(_render_update_preview(result)) + return + + write_attachment_record( + result.path, + AttachmentRecord( + ref=result.ref, + source_format=result.source_format, + source_name=result.source_name, + source_path=result.source_descriptor, + source_hash=result.source_hash, + from_version=result.from_version, + to_version=result.to_version, + change_kind=result.change_kind, + changes=[asdict(change) for change in result.changes], + ), + ) + console.print(result.content.rstrip()) + console.print( + f"[green]OK[/green] attached {ref} to {result.source_descriptor} " + f"({result.source_format}); new version {result.to_version} ({result.change_kind})" + ) + console.print(render_attach_audit_summary(result)) + + @click.command() @click.argument("ref") @click.option("--path", "path", type=click.Path(exists=True, path_type=Path), required=True) diff --git a/cli/src/modelable/llm/engine.py b/cli/src/modelable/llm/engine.py index 033aa47e..60818f75 100644 --- a/cli/src/modelable/llm/engine.py +++ b/cli/src/modelable/llm/engine.py @@ -1,10 +1,12 @@ from __future__ import annotations +import hashlib import re from dataclasses import dataclass from os import environ from pathlib import Path +from modelable.compat.diff import FieldChange, compare_model_versions from modelable.compiler.workspace import load_workspace from modelable.diagnostics.model import render_diagnostic from modelable.emitters.csharp import emit_csharp @@ -37,6 +39,7 @@ from modelable.llm.validation_help import explain_validation_errors from modelable.parser.ir import ( AnnKey, + ChangeKind, DirectMapping, FieldDef, MdlFile, @@ -82,6 +85,25 @@ class UpdatePlanResult: diagnostics_repaired: int +@dataclass(frozen=True) +class AttachResult: + path: Path + source_path: Path + ref: str + original_content: str + content: str + warnings: list[str] + attached: bool + from_version: int + to_version: int | None + change_kind: str | None + changes: list[FieldChange] + source_format: str + source_name: str + source_descriptor: str + source_hash: str + + def describe_path_or_ref(path: Path | None = None, ref: str | None = None) -> str: if ref and path is not None: workspace = load_workspace(path) @@ -365,6 +387,169 @@ def update_definition( ) +_BREAKING_ATTACH_CHANGE_KINDS = {"removed_field", "type_changed", "enum_changed", "identity_changed"} + + +def attach_external_version( + path: Path, + ref: str, + source: Path | str, + source_format: str, + *, + source_name: str | None = None, + output: Path | None = None, + write: bool = True, +) -> AttachResult: + """Attach a model version to an external dbt or FHIR source. + + If the external source's fields differ from the referenced model version, append a + new `.mdl` version block with a computed `additive`/`breaking` change kind. + """ + workspace = load_workspace(path) + model_ref = parse_model_ref(ref) + source_path = _find_source_path_for_ref(workspace, model_ref.domain, model_ref.name) + if source_path is None: + raise ValueError(f"Could not find source file for {ref}") + + mdl_text = source_path.read_text(encoding="utf-8") + mdl = parse_text_to_ir(mdl_text) + domain = next((item for item in mdl.domains if item.name == model_ref.domain), None) + if domain is None: + raise ValueError(f"Unknown domain: {model_ref.domain}") + versions = domain.models.get(model_ref.name) + if not versions: + raise ValueError(f"Unknown model: {ref}") + current = next((item for item in versions if item.version == model_ref.version), None) + if current is None: + raise ValueError(f"Unknown model version: {ref}") + + if isinstance(source, Path): + source_text = source.read_text(encoding="utf-8") + source_descriptor = str(source) + else: + source_text = source + source_descriptor = "inline" + imported = import_from_text(source_text, source_format, domain_name=model_ref.domain, source_name=source_name) + source_hash = hashlib.sha256(source_text.encode("utf-8")).hexdigest() + + new_fields = _build_attached_fields(current.fields, imported.model_version.fields) + candidate_version = ModelVersion( + model_kind=current.model_kind, + version=current.version, + change_kind=current.change_kind, + fields=new_fields, + ) + changes = compare_model_versions(current, candidate_version) + + if not changes: + return AttachResult( + path=output or source_path, + source_path=source_path, + ref=ref, + original_content=mdl_text, + content=mdl_text, + warnings=imported.warnings, + attached=False, + from_version=current.version, + to_version=None, + change_kind=None, + changes=[], + source_format=source_format, + source_name=imported.source_name, + source_descriptor=source_descriptor, + source_hash=source_hash, + ) + + change_kind = _classify_attach_change_kind(changes) + next_version_number = max(item.version for item in versions) + 1 + new_version = ModelVersion( + model_kind=current.model_kind, + version=next_version_number, + change_kind=ChangeKind(change_kind), + fields=new_fields, + ) + versions.append(new_version) + + new_text = render_mdl(mdl) + _, errors = validate_generated_text(new_text) + if errors: + raise ValueError("Attached definition failed validation: " + "; ".join(errors)) + + out_path = output or source_path + if write: + out_path.write_text(new_text, encoding="utf-8") + + return AttachResult( + path=out_path, + source_path=source_path, + ref=ref, + original_content=mdl_text, + content=new_text, + warnings=imported.warnings, + attached=True, + from_version=current.version, + to_version=next_version_number, + change_kind=change_kind, + changes=changes, + source_format=source_format, + source_name=imported.source_name, + source_descriptor=source_descriptor, + source_hash=source_hash, + ) + + +def _build_attached_fields(old_fields: list[FieldDef], candidate_fields: list[FieldDef]) -> list[FieldDef]: + """Combine the current field set with imported fields, preserving existing annotations.""" + candidate_by_name = {field.name: field for field in candidate_fields} + old_names = {field.name for field in old_fields} + new_fields: list[FieldDef] = [] + for old_field in old_fields: + candidate = candidate_by_name.get(old_field.name) + if candidate is None: + continue + new_fields.append( + FieldDef( + name=old_field.name, + type=candidate.type, + optional=candidate.optional, + default=old_field.default, + annotations=list(old_field.annotations), + ) + ) + for candidate in candidate_fields: + if candidate.name not in old_names: + new_fields.append( + FieldDef( + name=candidate.name, + type=candidate.type, + optional=candidate.optional, + default=candidate.default, + annotations=list(candidate.annotations), + ) + ) + return new_fields + + +def _classify_attach_change_kind(changes: list[FieldChange]) -> str: + for change in changes: + if change.kind in _BREAKING_ATTACH_CHANGE_KINDS: + return "breaking" + if change.kind == "nullability_changed" and change.from_optional and not change.to_optional: + return "breaking" + return "additive" + + +def render_attach_audit_summary(result: AttachResult) -> str: + return render_write_audit_summary( + provider="local", + model="modelable-local", + validation_status="passed", + files_written=str(result.path), + inputs=f"ref={result.ref} source={result.source_descriptor} format={result.source_format}", + diagnostics_repaired=0, + ) + + def _summarize_update_target(workspace, ref: str) -> str: model_ref = parse_model_ref(ref) domain = next((item for item in workspace.mdl.domains if item.name == model_ref.domain), None) diff --git a/cli/src/modelable/llm/importers.py b/cli/src/modelable/llm/importers.py index 9b151c9f..8ed2ff58 100644 --- a/cli/src/modelable/llm/importers.py +++ b/cli/src/modelable/llm/importers.py @@ -4,17 +4,24 @@ import re from dataclasses import dataclass, field from pathlib import Path +from typing import Any + +import yaml from modelable.llm.render import render_model_version from modelable.parser.ir import ( + AnnClassification, AnnKey, + AnnOwner, AnnPii, + Annotation, ArrayType, ChangeKind, DecimalType, DomainDef, EnumType, FieldDef, + FieldType, MdlFile, ModelKind, ModelVersion, @@ -41,7 +48,9 @@ def to_workspace(self) -> MdlFile: return MdlFile(domains=[DomainDef(name=self.domain_name, models={self.model_name: [self.model_version]})]) -def import_from_text(source_text: str, source_format: str, *, domain_name: str | None = None) -> ImportedModel: +def import_from_text( + source_text: str, source_format: str, *, domain_name: str | None = None, source_name: str | None = None +) -> ImportedModel: source_format = source_format.lower() if source_format == "json-schema": return _import_json_schema(source_text, domain_name=domain_name) @@ -53,11 +62,19 @@ def import_from_text(source_text: str, source_format: str, *, domain_name: str | return _import_protobuf(source_text, domain_name=domain_name) if source_format in {"sql", "ddl"}: return _import_sql(source_text, domain_name=domain_name) + if source_format == "dbt": + return _import_dbt(source_text, domain_name=domain_name, source_name=source_name) + if source_format == "fhir": + return _import_fhir(source_text, domain_name=domain_name, source_name=source_name) raise ValueError(f"Unsupported source format: {source_format}") -def import_from_path(path: str | Path, source_format: str, *, domain_name: str | None = None) -> ImportedModel: - return import_from_text(Path(path).read_text(encoding="utf-8"), source_format, domain_name=domain_name) +def import_from_path( + path: str | Path, source_format: str, *, domain_name: str | None = None, source_name: str | None = None +) -> ImportedModel: + return import_from_text( + Path(path).read_text(encoding="utf-8"), source_format, domain_name=domain_name, source_name=source_name + ) def _import_json_schema(source_text: str, *, domain_name: str | None) -> ImportedModel: @@ -167,6 +184,148 @@ def _import_sql(source_text: str, *, domain_name: str | None) -> ImportedModel: return ImportedModel("sql", table_name, domain, _sanitize_ident(_basename_name(table_name)), version, warnings) +def _import_dbt(source_text: str, *, domain_name: str | None, source_name: str | None = None) -> ImportedModel: + doc = yaml.safe_load(source_text) or {} + models = doc.get("models") or [] + if not models: + raise ValueError("dbt schema document does not declare any models") + if source_name is not None: + model = next((item for item in models if item.get("name") == source_name), None) + if model is None: + raise ValueError(f"dbt model '{source_name}' not found in source") + else: + model = models[0] + name = model.get("name") or "DbtModel" + domain = domain_name or _guess_domain_name(name) + fields: list[FieldDef] = [] + warnings: list[str] = [] + for column in model.get("columns") or []: + fields.append(_field_from_dbt_column(column, warnings)) + version = ModelVersion(model_kind=ModelKind.entity, version=1, change_kind=ChangeKind.additive, fields=fields) + return ImportedModel("dbt", name, domain, _sanitize_ident(name), version, warnings) + + +def _field_from_dbt_column(column: dict[str, Any], warnings: list[str]) -> FieldDef: + name = column["name"] + data_type = column.get("data_type") + field_type: FieldType + if data_type: + field_type = _sql_type_to_field_type(data_type) + else: + warnings.append(f"Column '{name}' has no data_type; defaulting to string") + field_type = PrimitiveType(kind="string") + + constraint_types = { + constraint.get("type") for constraint in column.get("constraints") or [] if isinstance(constraint, dict) + } + annotations: list[Annotation] = [] + if "primary_key" in constraint_types: + annotations.append(AnnKey()) + optional = "not_null" not in constraint_types and "primary_key" not in constraint_types + + meta = column.get("meta") or {} + if meta.get("modelable_pii"): + annotations.append(AnnPii()) + classification = meta.get("modelable_classification") + if classification: + annotations.append(AnnClassification(level=str(classification))) + owner = meta.get("modelable_owner") + if owner: + annotations.append(AnnOwner(team=str(owner))) + + return FieldDef(name=name, type=field_type, optional=optional, annotations=annotations) + + +_FHIR_PRIMITIVE_TYPES = { + "string": "string", + "code": "string", + "id": "string", + "markdown": "string", + "uri": "string", + "url": "string", + "canonical": "string", + "oid": "string", + "boolean": "bool", + "integer": "int", + "integer64": "int", + "positiveInt": "int", + "unsignedInt": "int", + "decimal": "float", + "dateTime": "timestamp", + "instant": "timestamp", + "date": "date", + "time": "time", + "base64Binary": "binary", +} + + +def _import_fhir(source_text: str, *, domain_name: str | None, source_name: str | None = None) -> ImportedModel: + doc = json.loads(source_text) + if doc.get("resourceType") != "StructureDefinition": + raise ValueError("FHIR source must be a StructureDefinition resource") + resource_type = doc.get("type") or doc.get("name") or "FhirResource" + name = doc.get("name") or resource_type + domain = domain_name or _guess_domain_name(resource_type) + + elements = (doc.get("snapshot") or {}).get("element") or (doc.get("differential") or {}).get("element") or [] + fields: list[FieldDef] = [] + warnings: list[str] = [] + for element in elements: + path = element.get("path", "") + segments = path.split(".") + if len(segments) != 2: + continue + field_name = segments[1] + if field_name.endswith("[x]"): + field_name = field_name[: -len("[x]")] + warnings.append(f"Choice-type element '{path}' flattened to '{field_name}'") + fields.append(_field_from_fhir_element(field_name, element, warnings)) + + version = ModelVersion(model_kind=ModelKind.entity, version=1, change_kind=ChangeKind.additive, fields=fields) + return ImportedModel("fhir", name, domain, _sanitize_ident(resource_type), version, warnings) + + +def _field_from_fhir_element(field_name: str, element: dict[str, Any], warnings: list[str]) -> FieldDef: + path = element.get("path", field_name) + types = element.get("type") or [] + field_type: FieldType + if not types: + warnings.append(f"Element '{path}' has no declared type; defaulting to string") + field_type = PrimitiveType(kind="string") + else: + if len(types) > 1: + codes = ", ".join(str(item.get("code", "?")) for item in types) + warnings.append(f"Element '{path}' has multiple types ({codes}); using the first") + field_type = _fhir_type_to_field_type(types[0], path, warnings) + + binding = element.get("binding") or {} + if binding.get("strength") == "required" and binding.get("valueSet"): + warnings.append( + f"Element '{path}' has a required binding to {binding['valueSet']}; " + "represent as enum(...) manually if a fixed value set is known" + ) + + optional = element.get("min", 0) == 0 + annotations: list[Annotation] = [] + if field_name == "id": + annotations.append(AnnKey()) + return FieldDef(name=field_name, type=field_type, optional=optional, annotations=annotations) + + +def _fhir_type_to_field_type(type_entry: dict[str, Any], path: str, warnings: list[str]) -> FieldType: + code = type_entry.get("code", "") + if code in _FHIR_PRIMITIVE_TYPES: + return PrimitiveType(kind=_FHIR_PRIMITIVE_TYPES[code]) + if code == "Reference": + targets = type_entry.get("targetProfile") or [] + if targets: + return RefType(target=str(targets[0]).rsplit("/", 1)[-1]) + warnings.append(f"Element '{path}' is an untyped Reference; falling back to named type") + return NamedType(name="Reference") + warnings.append(f"Element '{path}' has unsupported FHIR type '{code}'; falling back to named type") + return NamedType(name=code or "Unknown") + + def _fields_from_json_schema(schema: dict) -> tuple[list[FieldDef], list[str]]: warnings: list[str] = [] properties = schema.get("properties", {}) diff --git a/cli/src/modelable/llm/provenance.py b/cli/src/modelable/llm/provenance.py index 03adbbb3..020e51a5 100644 --- a/cli/src/modelable/llm/provenance.py +++ b/cli/src/modelable/llm/provenance.py @@ -52,3 +52,33 @@ def write_provenance_sidecar(artifact_path: Path, provenance: WriteProvenance) - sidecar_path = provenance_sidecar_path(artifact_path) sidecar_path.write_text(render_write_provenance(provenance), encoding="utf-8") return sidecar_path + + +@dataclass(frozen=True) +class AttachmentRecord: + ref: str + source_format: str + source_name: str + source_path: str + source_hash: str + from_version: int + to_version: int | None + change_kind: str | None + changes: list[dict[str, object]] + + def as_dict(self) -> dict[str, object]: + return asdict(self) + + +def attachment_sidecar_path(artifact_path: Path) -> Path: + return artifact_path.with_name(f"{artifact_path.name}.attachments.json") + + +def write_attachment_record(artifact_path: Path, record: AttachmentRecord) -> Path: + sidecar_path = attachment_sidecar_path(artifact_path) + records: list[dict[str, object]] = [] + if sidecar_path.exists(): + records = json.loads(sidecar_path.read_text(encoding="utf-8")) + records.append(record.as_dict()) + sidecar_path.write_text(json.dumps(records, indent=2, ensure_ascii=False) + "\n", encoding="utf-8") + return sidecar_path diff --git a/cli/tests/test_llm_features.py b/cli/tests/test_llm_features.py index bd5f7b6f..18e9b874 100644 --- a/cli/tests/test_llm_features.py +++ b/cli/tests/test_llm_features.py @@ -177,6 +177,97 @@ def test_sql_importer_marks_primary_key(): assert "name?: string" in text +def test_dbt_importer_maps_columns_and_meta(): + imported = import_from_text( + """ +version: 2 +models: + - name: Customer + columns: + - name: customerId + data_type: text + constraints: + - type: primary_key + - name: email + data_type: text + constraints: + - type: not_null + meta: + modelable_pii: true + modelable_classification: restricted + - name: loyaltyTier + data_type: text +""", + "dbt", + domain_name="customer", + ) + text = imported.to_mdl() + assert "domain customer" in text + assert "entity Customer @ 1 (additive)" in text + assert "@key customerId: string" in text + assert '@pii @classification("restricted") email: string' in text + assert "loyaltyTier?: string" in text + + +def test_dbt_importer_selects_named_model(): + source = """ +version: 2 +models: + - name: Customer + columns: + - name: customerId + data_type: text + - name: Order + columns: + - name: orderId + data_type: text +""" + imported = import_from_text(source, "dbt", domain_name="orders", source_name="Order") + text = imported.to_mdl() + assert "entity Order @ 1 (additive)" in text + assert "orderId?: string" in text + + +def test_fhir_importer_maps_elements_to_fields(): + source = json.dumps( + { + "resourceType": "StructureDefinition", + "name": "Patient", + "type": "Patient", + "snapshot": { + "element": [ + {"path": "Patient", "min": 0, "max": "*"}, + {"path": "Patient.id", "min": 0, "max": "1", "type": [{"code": "id"}]}, + {"path": "Patient.active", "min": 0, "max": "1", "type": [{"code": "boolean"}]}, + {"path": "Patient.birthDate", "min": 0, "max": "1", "type": [{"code": "date"}]}, + { + "path": "Patient.managingOrganization", + "min": 0, + "max": "1", + "type": [ + { + "code": "Reference", + "targetProfile": ["http://hl7.org/fhir/StructureDefinition/Organization"], + } + ], + }, + {"path": "Patient.name", "min": 0, "max": "*", "type": [{"code": "HumanName"}]}, + ] + }, + } + ) + imported = import_from_text(source, "fhir", domain_name="clinical") + text = imported.to_mdl() + assert "domain clinical" in text + assert "entity Patient @ 1 (additive)" in text + assert "@key id?: string" in text + assert "active?: bool" in text + assert "birthDate?: date" in text + assert "managingOrganization?: ref" in text + assert "name?: HumanName" in text + assert any("HumanName" in warning for warning in imported.warnings) + + def test_cli_generate_describe_ask_and_recommend(tmp_path): mdl = tmp_path / "workspace.mdl" mdl.write_text( @@ -487,6 +578,205 @@ def test_cli_update_projection_field(tmp_path): assert "status <- c.name" in updated +def _attachments_path(path: Path) -> Path: + return path.with_name(f"{path.name}.attachments.json") + + +def test_cli_attach_dbt_creates_breaking_version(tmp_path): + mdl = tmp_path / "workspace.mdl" + mdl.write_text( + """ +domain customer { + owner: "customer-team" + entity Customer @ 1 (additive) { + @key customerId: uuid + @pii email: string + name: string + } +} +""", + encoding="utf-8", + ) + + schema = tmp_path / "customer_schema.yml" + schema.write_text( + """ +version: 2 +models: + - name: Customer + columns: + - name: customerId + data_type: text + constraints: + - type: primary_key + - name: email + data_type: text + meta: + modelable_pii: true + constraints: + - type: not_null + - name: name + data_type: text + constraints: + - type: not_null + - name: loyaltyTier + data_type: text +""", + encoding="utf-8", + ) + + runner = CliRunner() + result = runner.invoke( + cli, + [ + "attach", + "customer.Customer@1", + "--source", + str(schema), + "--source-format", + "dbt", + "--source-name", + "Customer", + "--path", + str(tmp_path), + ], + ) + assert result.exit_code == 0, result.output + updated = mdl.read_text(encoding="utf-8") + assert "entity Customer @ 1 (additive)" in updated + assert "entity Customer @ 2 (breaking)" in updated + assert "@key customerId: string" in updated + assert "loyaltyTier?: string" in updated + assert "new version 2 (breaking)" in result.output + + attachments = json.loads(_attachments_path(mdl).read_text(encoding="utf-8")) + assert len(attachments) == 1 + record = attachments[0] + assert record["ref"] == "customer.Customer@1" + assert record["source_format"] == "dbt" + assert record["source_name"] == "Customer" + assert record["from_version"] == 1 + assert record["to_version"] == 2 + assert record["change_kind"] == "breaking" + change_kinds = {change["kind"] for change in record["changes"]} + assert "type_changed" in change_kinds + assert "added_field" in change_kinds + + +def test_cli_attach_no_diff_skips_new_version(tmp_path): + mdl = tmp_path / "workspace.mdl" + mdl.write_text( + """ +domain customer { + owner: "customer-team" + entity Customer @ 1 (additive) { + @key customerId: string + name: string + } +} +""", + encoding="utf-8", + ) + + schema = tmp_path / "customer_schema.yml" + schema.write_text( + """ +version: 2 +models: + - name: Customer + columns: + - name: customerId + data_type: text + constraints: + - type: primary_key + - name: name + data_type: text + constraints: + - type: not_null +""", + encoding="utf-8", + ) + + runner = CliRunner() + result = runner.invoke( + cli, + [ + "attach", + "customer.Customer@1", + "--source", + str(schema), + "--source-format", + "dbt", + "--source-name", + "Customer", + "--path", + str(tmp_path), + ], + ) + assert result.exit_code == 0, result.output + assert "no new version created" in result.output + assert mdl.read_text(encoding="utf-8").count("entity Customer @") == 1 + assert not _attachments_path(mdl).exists() + + +def test_cli_attach_preview_does_not_write(tmp_path): + mdl = tmp_path / "workspace.mdl" + original = """ +domain customer { + owner: "customer-team" + entity Customer @ 1 (additive) { + @key customerId: string + name: string + } +} +""" + mdl.write_text(original, encoding="utf-8") + + schema = tmp_path / "customer_schema.yml" + schema.write_text( + """ +version: 2 +models: + - name: Customer + columns: + - name: customerId + data_type: text + constraints: + - type: primary_key + - name: name + data_type: text + constraints: + - type: not_null + - name: phone + data_type: text +""", + encoding="utf-8", + ) + + runner = CliRunner() + result = runner.invoke( + cli, + [ + "attach", + "customer.Customer@1", + "--source", + str(schema), + "--source-format", + "dbt", + "--source-name", + "Customer", + "--path", + str(tmp_path), + "--preview", + ], + ) + assert result.exit_code == 0, result.output + assert "@@" in result.output + assert "entity Customer @ 2 (additive)" in result.output + assert mdl.read_text(encoding="utf-8") == original + assert not _attachments_path(mdl).exists() + + _TRANSFORM_MDL = """ domain customer { owner: "test-team" diff --git a/cli/uv.lock b/cli/uv.lock index cb3d2cc6..73e195d5 100644 --- a/cli/uv.lock +++ b/cli/uv.lock @@ -275,6 +275,7 @@ dependencies = [ { name = "psycopg", extra = ["binary"] }, { name = "pydantic" }, { name = "pygls" }, + { name = "pyyaml" }, { name = "referencing" }, { name = "rich" }, ] @@ -310,6 +311,7 @@ requires-dist = [ { name = "pytest", marker = "extra == 'dev'", specifier = ">=9.0.3" }, { name = "pytest-cov", marker = "extra == 'dev'", specifier = ">=4.0" }, { name = "pytest-lsp", marker = "extra == 'dev'", specifier = ">=1.0.0" }, + { name = "pyyaml", specifier = ">=6.0.3" }, { name = "referencing", specifier = ">=0.35" }, { name = "rich", specifier = ">=13.0" }, { name = "ruff", marker = "extra == 'dev'", specifier = ">=0.15.17" }, @@ -570,6 +572,32 @@ wheels = [ { url = "https://files.pythonhosted.org/packages/dd/0f/b8564e8bbec03a6efe76fe450f0d328e8df96c603340055fef7e4428208b/pytest_lsp-1.0.0-py3-none-any.whl", hash = "sha256:36d002eda8d9bcd3ff9a5b33c382dd29dfd054ad9e475ad31d3aac53fdbf516b", size = 26424, upload-time = "2025-10-25T12:10:09.495Z" }, ] +[[package]] +name = "pyyaml" +version = "6.0.3" +source = { registry = "https://pypi.org/simple" } +sdist = { url = "https://files.pythonhosted.org/packages/05/8e/961c0007c59b8dd7729d542c61a4d537767a59645b82a0b521206e1e25c2/pyyaml-6.0.3.tar.gz", hash = "sha256:d76623373421df22fb4cf8817020cbb7ef15c725b9d5e45f17e189bfc384190f", size = 130960, upload-time = "2025-09-25T21:33:16.546Z" } +wheels = [ + { url = "https://files.pythonhosted.org/packages/9d/8c/f4bd7f6465179953d3ac9bc44ac1a8a3e6122cf8ada906b4f96c60172d43/pyyaml-6.0.3-cp314-cp314-macosx_10_13_x86_64.whl", hash = "sha256:8d1fab6bb153a416f9aeb4b8763bc0f22a5586065f86f7664fc23339fc1c1fac", size = 181814, upload-time = "2025-09-25T21:32:35.712Z" }, + { url = "https://files.pythonhosted.org/packages/bd/9c/4d95bb87eb2063d20db7b60faa3840c1b18025517ae857371c4dd55a6b3a/pyyaml-6.0.3-cp314-cp314-macosx_11_0_arm64.whl", hash = "sha256:34d5fcd24b8445fadc33f9cf348c1047101756fd760b4dacb5c3e99755703310", size = 173809, upload-time = "2025-09-25T21:32:36.789Z" }, + { url = "https://files.pythonhosted.org/packages/92/b5/47e807c2623074914e29dabd16cbbdd4bf5e9b2db9f8090fa64411fc5382/pyyaml-6.0.3-cp314-cp314-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:501a031947e3a9025ed4405a168e6ef5ae3126c59f90ce0cd6f2bfc477be31b7", size = 766454, upload-time = "2025-09-25T21:32:37.966Z" }, + { url = "https://files.pythonhosted.org/packages/02/9e/e5e9b168be58564121efb3de6859c452fccde0ab093d8438905899a3a483/pyyaml-6.0.3-cp314-cp314-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:b3bc83488de33889877a0f2543ade9f70c67d66d9ebb4ac959502e12de895788", size = 836355, upload-time = "2025-09-25T21:32:39.178Z" }, + { url = "https://files.pythonhosted.org/packages/88/f9/16491d7ed2a919954993e48aa941b200f38040928474c9e85ea9e64222c3/pyyaml-6.0.3-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:c458b6d084f9b935061bc36216e8a69a7e293a2f1e68bf956dcd9e6cbcd143f5", size = 794175, upload-time = "2025-09-25T21:32:40.865Z" }, + { url = "https://files.pythonhosted.org/packages/dd/3f/5989debef34dc6397317802b527dbbafb2b4760878a53d4166579111411e/pyyaml-6.0.3-cp314-cp314-musllinux_1_2_aarch64.whl", hash = "sha256:7c6610def4f163542a622a73fb39f534f8c101d690126992300bf3207eab9764", size = 755228, upload-time = "2025-09-25T21:32:42.084Z" }, + { url = "https://files.pythonhosted.org/packages/d7/ce/af88a49043cd2e265be63d083fc75b27b6ed062f5f9fd6cdc223ad62f03e/pyyaml-6.0.3-cp314-cp314-musllinux_1_2_x86_64.whl", hash = "sha256:5190d403f121660ce8d1d2c1bb2ef1bd05b5f68533fc5c2ea899bd15f4399b35", size = 789194, upload-time = "2025-09-25T21:32:43.362Z" }, + { url = "https://files.pythonhosted.org/packages/23/20/bb6982b26a40bb43951265ba29d4c246ef0ff59c9fdcdf0ed04e0687de4d/pyyaml-6.0.3-cp314-cp314-win_amd64.whl", hash = "sha256:4a2e8cebe2ff6ab7d1050ecd59c25d4c8bd7e6f400f5f82b96557ac0abafd0ac", size = 156429, upload-time = "2025-09-25T21:32:57.844Z" }, + { url = "https://files.pythonhosted.org/packages/f4/f4/a4541072bb9422c8a883ab55255f918fa378ecf083f5b85e87fc2b4eda1b/pyyaml-6.0.3-cp314-cp314-win_arm64.whl", hash = "sha256:93dda82c9c22deb0a405ea4dc5f2d0cda384168e466364dec6255b293923b2f3", size = 143912, upload-time = "2025-09-25T21:32:59.247Z" }, + { url = "https://files.pythonhosted.org/packages/7c/f9/07dd09ae774e4616edf6cda684ee78f97777bdd15847253637a6f052a62f/pyyaml-6.0.3-cp314-cp314t-macosx_10_13_x86_64.whl", hash = "sha256:02893d100e99e03eda1c8fd5c441d8c60103fd175728e23e431db1b589cf5ab3", size = 189108, upload-time = "2025-09-25T21:32:44.377Z" }, + { url = "https://files.pythonhosted.org/packages/4e/78/8d08c9fb7ce09ad8c38ad533c1191cf27f7ae1effe5bb9400a46d9437fcf/pyyaml-6.0.3-cp314-cp314t-macosx_11_0_arm64.whl", hash = "sha256:c1ff362665ae507275af2853520967820d9124984e0f7466736aea23d8611fba", size = 183641, upload-time = "2025-09-25T21:32:45.407Z" }, + { url = "https://files.pythonhosted.org/packages/7b/5b/3babb19104a46945cf816d047db2788bcaf8c94527a805610b0289a01c6b/pyyaml-6.0.3-cp314-cp314t-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:6adc77889b628398debc7b65c073bcb99c4a0237b248cacaf3fe8a557563ef6c", size = 831901, upload-time = "2025-09-25T21:32:48.83Z" }, + { url = "https://files.pythonhosted.org/packages/8b/cc/dff0684d8dc44da4d22a13f35f073d558c268780ce3c6ba1b87055bb0b87/pyyaml-6.0.3-cp314-cp314t-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:a80cb027f6b349846a3bf6d73b5e95e782175e52f22108cfa17876aaeff93702", size = 861132, upload-time = "2025-09-25T21:32:50.149Z" }, + { url = "https://files.pythonhosted.org/packages/b1/5e/f77dc6b9036943e285ba76b49e118d9ea929885becb0a29ba8a7c75e29fe/pyyaml-6.0.3-cp314-cp314t-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:00c4bdeba853cc34e7dd471f16b4114f4162dc03e6b7afcc2128711f0eca823c", size = 839261, upload-time = "2025-09-25T21:32:51.808Z" }, + { url = "https://files.pythonhosted.org/packages/ce/88/a9db1376aa2a228197c58b37302f284b5617f56a5d959fd1763fb1675ce6/pyyaml-6.0.3-cp314-cp314t-musllinux_1_2_aarch64.whl", hash = "sha256:66e1674c3ef6f541c35191caae2d429b967b99e02040f5ba928632d9a7f0f065", size = 805272, upload-time = "2025-09-25T21:32:52.941Z" }, + { url = "https://files.pythonhosted.org/packages/da/92/1446574745d74df0c92e6aa4a7b0b3130706a4142b2d1a5869f2eaa423c6/pyyaml-6.0.3-cp314-cp314t-musllinux_1_2_x86_64.whl", hash = "sha256:16249ee61e95f858e83976573de0f5b2893b3677ba71c9dd36b9cf8be9ac6d65", size = 829923, upload-time = "2025-09-25T21:32:54.537Z" }, + { url = "https://files.pythonhosted.org/packages/f0/7a/1c7270340330e575b92f397352af856a8c06f230aa3e76f86b39d01b416a/pyyaml-6.0.3-cp314-cp314t-win_amd64.whl", hash = "sha256:4ad1906908f2f5ae4e5a8ddfce73c320c2a1429ec52eafd27138b7f1cbe341c9", size = 174062, upload-time = "2025-09-25T21:32:55.767Z" }, + { url = "https://files.pythonhosted.org/packages/f1/12/de94a39c2ef588c7e6455cfbe7343d3b2dc9d6b6b2f40c4c6565744c873d/pyyaml-6.0.3-cp314-cp314t-win_arm64.whl", hash = "sha256:ebc55a14a21cb14062aa4162f906cd962b28e2e9ea38f9b4391244cd8de4ae0b", size = 149341, upload-time = "2025-09-25T21:32:56.828Z" }, +] + [[package]] name = "referencing" version = "0.37.0" diff --git a/docs/cli-spec.md b/docs/cli-spec.md index a86fe0ac..965778b1 100644 --- a/docs/cli-spec.md +++ b/docs/cli-spec.md @@ -761,3 +761,33 @@ modelable lineage export --format ndjson --output Exports the full lineage graph from `registry.db` as NDJSON for external catalog ingestion. **Defined in:** `distributed-lineage-spec.md` §6. + +### 10.8 `attach` — Attach a model version to an external dbt or FHIR source + +```text +modelable attach --source --source-format [--source-name NAME] --path PATH [--output FILE] [--preview] +``` + +Imports fields from an external dbt `schema.yml` model or a FHIR R4 `StructureDefinition` +and compares them against the referenced model version using the same field-by-field +comparison as `diff`. `--source-name` selects a specific dbt model when the source file +declares more than one; it is ignored for FHIR sources, which describe a single resource. + +- If the imported fields match the referenced version, no changes are made. +- Otherwise, a new model version is appended to the `.mdl` source with fields derived + from the external source (existing field annotations such as `@key`, `@pii`, and + `@classification` are preserved by field name) and a `change_kind` of `additive` or + `breaking` computed from the same rules as `diff` (removed fields, type changes, enum + changes, identity changes, and optional-to-required narrowing are breaking; everything + else is additive). + +By default the command writes the source file for the referenced definition; `--output` +directs the result to an alternate path. `--preview` shows the rendered diff without +writing changes. + +When the command writes a file, it appends a record to a `.attachments.json` sidecar +next to the `.mdl` file describing the source format, matched source name, source +content hash, the version transition, the computed change kind, and the field-level +changes, and it prints the standard audit summary. + +**Defined in:** `dbt-fhir-tool-alignment.md` §2.3, §3.3. diff --git a/docs/dbt-fhir-tool-alignment.md b/docs/dbt-fhir-tool-alignment.md index 5635ec58..f80e2d54 100644 --- a/docs/dbt-fhir-tool-alignment.md +++ b/docs/dbt-fhir-tool-alignment.md @@ -1,11 +1,16 @@ # Research and Plan: Aligning with dbt, FHIR, and Other External Tools -> **Status:** Research and planning. No emitter, importer, or CLI work is -> committed by this document. It extends -> [external-tools-data-modelling.md](external-tools-data-modelling.md) and -> [migration-guide.md](migration-guide.md) with concept mappings and phased -> proposals for dbt, FHIR, and other ecosystems. Concrete work requires an -> issue and an accepted design per [ROADMAP.md](../ROADMAP.md). +> **Status:** Research and planning. Most of this document's phased proposals +> (emitters, catalog/lineage integration, additional artifact targets) are not +> committed and require an issue and an accepted design per +> [ROADMAP.md](../ROADMAP.md). One slice has shipped: the `modelable attach` +> command (see [cli-spec.md](cli-spec.md) §10.8) imports a dbt `schema.yml` +> model or a FHIR `StructureDefinition` (§2.3 Phase B / §3.3 Phase C below) and +> records field-level drift against an existing Modelable model version as a +> new `additive`/`breaking` version plus an attachment record. This document +> extends [external-tools-data-modelling.md](external-tools-data-modelling.md) +> and [migration-guide.md](migration-guide.md) with concept mappings and +> phased proposals for dbt, FHIR, and other ecosystems. ## 1. Purpose @@ -92,6 +97,15 @@ models/.mdl` to bootstrap `.mdl` models from an existing dbt project, following the same "review the generated output" workflow as other LLM-assisted imports. +**Implemented (partial):** `modelable attach --source + --source-format dbt [--source-name NAME]` imports a dbt model's +`columns:` (with `data_type`, `constraints`, and `modelable_*` `meta` keys) and +compares them to an existing published model version, appending a new version +with a computed `additive`/`breaking` change kind when they differ. See +[cli-spec.md](cli-spec.md) §10.8. `modelable generate --from ` +bootstrapping for brand-new models and `manifest.json` input remain +unimplemented. + **Phase C — exposure/lineage stitching:** Treat dbt `exposures` as external consumers in the lineage graph, so @@ -188,6 +202,17 @@ the FHIR resource catalog. models/.mdl` to draft a starting `.mdl` model/projection from an existing profile, with the same human-review workflow as other imports. +**Implemented (partial):** `modelable attach --source + --source-format fhir` imports the direct child +elements of a FHIR R4 `StructureDefinition` (primitive types, `Reference` +targets, and cardinality) and compares them to an existing published model +version, appending a new version with a computed `additive`/`breaking` change +kind when they differ. Elements with complex FHIR types (e.g. +`BackboneElement`, `HumanName`, `CodeableConcept`) fall back to a named type +with a warning, per §3.4. See [cli-spec.md](cli-spec.md) §10.8. `modelable +generate --from ` bootstrapping for brand-new models +remains unimplemented. + **Phase D (Later, roadmap) — Implementation Guide packaging:** Generate an IG-shaped documentation bundle (profiles + narrative + examples) From 63bd2a5da10b46fa86ea37973cf4bb3c24de986e Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 14 Jun 2026 12:05:17 +0000 Subject: [PATCH 3/3] Fix ruff import-sort error in importers.py --- cli/src/modelable/llm/importers.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/cli/src/modelable/llm/importers.py b/cli/src/modelable/llm/importers.py index 8ed2ff58..b42f825d 100644 --- a/cli/src/modelable/llm/importers.py +++ b/cli/src/modelable/llm/importers.py @@ -12,9 +12,9 @@ from modelable.parser.ir import ( AnnClassification, AnnKey, + Annotation, AnnOwner, AnnPii, - Annotation, ArrayType, ChangeKind, DecimalType,