Skip to content

feat: refresh Foundry model catalog - #77

Merged
graphite-app[bot] merged 1 commit into
mainfrom
feat/model-refresh-0924
Sep 25, 2026
Merged

graphite-app[bot] merged 1 commit into
mainfrom
feat/model-refresh-0924

Conversation

@anandpant

@anandpant anandpant commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Bring the catalog in line with foundry-cli models list --json as of September 24, 2026.

  • Add GPT-6 Astra (GA), GPT-6 Luna, GPT-6 Sol, Codex Auto Review, text-embedding-ada-002, Claude Opus 5.5, Grok 4.7, and DeepSeek V4.1 Flash. Promote gemini-3.8-flash from Experimental to GA. Remove claude-opus-4.1, now Sunset with a 2027-01-07 retirement, and add it to the removed-alias list.
  • Four listed candidates were probed and left out. GPT-5.2 Pro and GPT-5.4 Pro return DisabledForUser on Responses with no Chat Completions route; o1 (Azure OpenAI) and Schematic 7B (Palantir Hub) return LanguageModelNotAvailable because their backends serve neither OPEN_AI_RESPONSES nor GPT_CHAT_COMPLETION. They are recorded under docs Findings rather than dropped silently.
  • DeepSeek V4.1 Flash is cataloged as third-party with the verified openai-chat transport. Native @ai-sdk/deepseek exposes reasoning text but breaks tool continuation at the proxy, so no optional peer was added; the comparison lands with the results docs.
  • inputTypes follow each family's siblings and were confirmed by probes: the GPT-6 trio matches gpt-5.6-*, Codex Auto Review is Responses-only because the Chat Completions route rejects it, and Grok 4.7 matches its siblings.
  • ModelPerformance.cost becomes optional alongside speed, and ModelClass gains SPECIALIZED_EMBEDDING, so the catalog can mirror what the listing publishes instead of inventing values. Grok 4.7 now omits cost, and the embeddings carry their real SPECIALIZED_EMBEDDING class.
  • Re-diffed all 69 entries against the listing and corrected 28 stale fields on pre-existing models: Claude Opus 4.5-4.8 cost, Opus 4.7/4.8, Opus 5 and Sonnet 5 cutoffs, GPT-5.5 cutoff, GPT-5.3 Codex class, Gemini 3.5 Flash cost/class/URL, Gemini 3.5 Flash-Lite display name, embedding costs and URLs, and realtime cutoffs and URLs. The diff is now zero.
  • The CLI catalog fixture now carries displayName, trainingCutoffDate and performance as well, and the enrollment test checks them, so future metadata drift fails instead of passing silently. The harness embedding probe list derives from the catalog so new embeddings are surveyed automatically. Opus 5.5 joins the adaptive-thinking set; the proxy rejects budget-based thinking for it with HTTP 400.

Final catalog: 69 entries — 63 language aliases, three embeddings, three realtime.

Validation

pnpm run verify passed, covering release policy, lint, tests, typecheck, build, skill validation, and the package audit. Each candidate was probed against its raw RID before cataloging, and the eight additions then passed a focused live harness run. The new drift assertions were confirmed load-bearing by perturbing a display name and watching the enrollment test fail. No package version or changelog changes.

Copy link
Copy Markdown
Contributor Author

How to use the Graphite Merge Queue

Add either label to this PR to merge it via the merge queue:

  • merge - adds this PR to the back of the merge queue
  • fast - for urgent changes, fast-track this PR to the front of the merge queue

You must have a Graphite account in order to use the merge queue. Sign up using this link.

An organization admin has enabled the Graphite Merge Queue in this repository.

Please do not merge from GitHub as this will restart CI on PRs being processed by the merge queue.

This stack of pull requests is managed by Graphite. Learn more about stacking.

@anandpant
anandpant marked this pull request as ready for review September 25, 2026 05:21
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Repo admins can enable using credits for code reviews in their settings.

@anandpant
anandpant force-pushed the feat/model-refresh-0924 branch from f475b56 to d69e0b6 Compare September 25, 2026 05:44
@graphite-app

graphite-app Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Merge activity

## Summary

Bring the catalog in line with `foundry-cli models list --json` as of September 24, 2026.

- Add GPT-6 Astra (GA), GPT-6 Luna, GPT-6 Sol, Codex Auto Review, `text-embedding-ada-002`, Claude Opus 5.5, Grok 4.7, and DeepSeek V4.1 Flash. Promote `gemini-3.8-flash` from Experimental to GA. Remove `claude-opus-4.1`, now Sunset with a 2027-01-07 retirement, and add it to the removed-alias list.
- Four listed candidates were probed and left out. GPT-5.2 Pro and GPT-5.4 Pro return `DisabledForUser` on Responses with no Chat Completions route; o1 (Azure OpenAI) and Schematic 7B (Palantir Hub) return `LanguageModelNotAvailable` because their backends serve neither `OPEN_AI_RESPONSES` nor `GPT_CHAT_COMPLETION`. They are recorded under docs Findings rather than dropped silently.
- DeepSeek V4.1 Flash is cataloged as third-party with the verified `openai-chat` transport. Native `@ai-sdk/deepseek` exposes reasoning text but breaks tool continuation at the proxy, so no optional peer was added; the comparison lands with the results docs.
- `inputTypes` follow each family's siblings and were confirmed by probes: the GPT-6 trio matches `gpt-5.6-*`, Codex Auto Review is Responses-only because the Chat Completions route rejects it, and Grok 4.7 matches its siblings.
- `ModelPerformance.cost` becomes optional alongside `speed`, and `ModelClass` gains `SPECIALIZED_EMBEDDING`, so the catalog can mirror what the listing publishes instead of inventing values. Grok 4.7 now omits `cost`, and the embeddings carry their real `SPECIALIZED_EMBEDDING` class.
- Re-diffed all 69 entries against the listing and corrected 28 stale fields on pre-existing models: Claude Opus 4.5-4.8 cost, Opus 4.7/4.8, Opus 5 and Sonnet 5 cutoffs, GPT-5.5 cutoff, GPT-5.3 Codex class, Gemini 3.5 Flash cost/class/URL, Gemini 3.5 Flash-Lite display name, embedding costs and URLs, and realtime cutoffs and URLs. The diff is now zero.
- The CLI catalog fixture now carries `displayName`, `trainingCutoffDate` and `performance` as well, and the enrollment test checks them, so future metadata drift fails instead of passing silently. The harness embedding probe list derives from the catalog so new embeddings are surveyed automatically. Opus 5.5 joins the adaptive-thinking set; the proxy rejects budget-based thinking for it with HTTP 400.

Final catalog: 69 entries — 63 language aliases, three embeddings, three realtime.

## Validation

`pnpm run verify` passed, covering release policy, lint, tests, typecheck, build, skill validation, and the package audit. Each candidate was probed against its raw RID before cataloging, and the eight additions then passed a focused live harness run. The new drift assertions were confirmed load-bearing by perturbing a display name and watching the enrollment test fail. No package version or changelog changes.
@graphite-app
graphite-app Bot force-pushed the feat/model-refresh-0924 branch from d69e0b6 to afaca1b Compare September 25, 2026 05:49
@graphite-app
graphite-app Bot merged commit afaca1b into main Sep 25, 2026
9 checks passed
@graphite-app
graphite-app Bot deleted the feat/model-refresh-0924 branch September 25, 2026 05:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant