feat: refresh Foundry model catalog - #77
Merged
Merged
Conversation
Contributor
Author
How to use the Graphite Merge QueueAdd either label to this PR to merge it via the merge queue:
You must have a Graphite account in order to use the merge queue. Sign up using this link. An organization admin has enabled the Graphite Merge Queue in this repository. Please do not merge from GitHub as this will restart CI on PRs being processed by the merge queue. This stack of pull requests is managed by Graphite. Learn more about stacking. |
This was referenced Sep 25, 2026
anandpant
marked this pull request as ready for review
September 25, 2026 05:21
|
Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits. |
anandpant
force-pushed
the
feat/model-refresh-0924
branch
from
September 25, 2026 05:44
f475b56 to
d69e0b6
Compare
Merge activity
|
## Summary Bring the catalog in line with `foundry-cli models list --json` as of September 24, 2026. - Add GPT-6 Astra (GA), GPT-6 Luna, GPT-6 Sol, Codex Auto Review, `text-embedding-ada-002`, Claude Opus 5.5, Grok 4.7, and DeepSeek V4.1 Flash. Promote `gemini-3.8-flash` from Experimental to GA. Remove `claude-opus-4.1`, now Sunset with a 2027-01-07 retirement, and add it to the removed-alias list. - Four listed candidates were probed and left out. GPT-5.2 Pro and GPT-5.4 Pro return `DisabledForUser` on Responses with no Chat Completions route; o1 (Azure OpenAI) and Schematic 7B (Palantir Hub) return `LanguageModelNotAvailable` because their backends serve neither `OPEN_AI_RESPONSES` nor `GPT_CHAT_COMPLETION`. They are recorded under docs Findings rather than dropped silently. - DeepSeek V4.1 Flash is cataloged as third-party with the verified `openai-chat` transport. Native `@ai-sdk/deepseek` exposes reasoning text but breaks tool continuation at the proxy, so no optional peer was added; the comparison lands with the results docs. - `inputTypes` follow each family's siblings and were confirmed by probes: the GPT-6 trio matches `gpt-5.6-*`, Codex Auto Review is Responses-only because the Chat Completions route rejects it, and Grok 4.7 matches its siblings. - `ModelPerformance.cost` becomes optional alongside `speed`, and `ModelClass` gains `SPECIALIZED_EMBEDDING`, so the catalog can mirror what the listing publishes instead of inventing values. Grok 4.7 now omits `cost`, and the embeddings carry their real `SPECIALIZED_EMBEDDING` class. - Re-diffed all 69 entries against the listing and corrected 28 stale fields on pre-existing models: Claude Opus 4.5-4.8 cost, Opus 4.7/4.8, Opus 5 and Sonnet 5 cutoffs, GPT-5.5 cutoff, GPT-5.3 Codex class, Gemini 3.5 Flash cost/class/URL, Gemini 3.5 Flash-Lite display name, embedding costs and URLs, and realtime cutoffs and URLs. The diff is now zero. - The CLI catalog fixture now carries `displayName`, `trainingCutoffDate` and `performance` as well, and the enrollment test checks them, so future metadata drift fails instead of passing silently. The harness embedding probe list derives from the catalog so new embeddings are surveyed automatically. Opus 5.5 joins the adaptive-thinking set; the proxy rejects budget-based thinking for it with HTTP 400. Final catalog: 69 entries — 63 language aliases, three embeddings, three realtime. ## Validation `pnpm run verify` passed, covering release policy, lint, tests, typecheck, build, skill validation, and the package audit. Each candidate was probed against its raw RID before cataloging, and the eight additions then passed a focused live harness run. The new drift assertions were confirmed load-bearing by perturbing a display name and watching the enrollment test fail. No package version or changelog changes.
graphite-app
Bot
force-pushed
the
feat/model-refresh-0924
branch
from
September 25, 2026 05:49
d69e0b6 to
afaca1b
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Summary
Bring the catalog in line with
foundry-cli models list --jsonas of September 24, 2026.text-embedding-ada-002, Claude Opus 5.5, Grok 4.7, and DeepSeek V4.1 Flash. Promotegemini-3.8-flashfrom Experimental to GA. Removeclaude-opus-4.1, now Sunset with a 2027-01-07 retirement, and add it to the removed-alias list.DisabledForUseron Responses with no Chat Completions route; o1 (Azure OpenAI) and Schematic 7B (Palantir Hub) returnLanguageModelNotAvailablebecause their backends serve neitherOPEN_AI_RESPONSESnorGPT_CHAT_COMPLETION. They are recorded under docs Findings rather than dropped silently.openai-chattransport. Native@ai-sdk/deepseekexposes reasoning text but breaks tool continuation at the proxy, so no optional peer was added; the comparison lands with the results docs.inputTypesfollow each family's siblings and were confirmed by probes: the GPT-6 trio matchesgpt-5.6-*, Codex Auto Review is Responses-only because the Chat Completions route rejects it, and Grok 4.7 matches its siblings.ModelPerformance.costbecomes optional alongsidespeed, andModelClassgainsSPECIALIZED_EMBEDDING, so the catalog can mirror what the listing publishes instead of inventing values. Grok 4.7 now omitscost, and the embeddings carry their realSPECIALIZED_EMBEDDINGclass.displayName,trainingCutoffDateandperformanceas well, and the enrollment test checks them, so future metadata drift fails instead of passing silently. The harness embedding probe list derives from the catalog so new embeddings are surveyed automatically. Opus 5.5 joins the adaptive-thinking set; the proxy rejects budget-based thinking for it with HTTP 400.Final catalog: 69 entries — 63 language aliases, three embeddings, three realtime.
Validation
pnpm run verifypassed, covering release policy, lint, tests, typecheck, build, skill validation, and the package audit. Each candidate was probed against its raw RID before cataloging, and the eight additions then passed a focused live harness run. The new drift assertions were confirmed load-bearing by perturbing a display name and watching the enrollment test fail. No package version or changelog changes.