Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -267,6 +267,7 @@ jobs:
packages/integrations/deepagents-sdk/dist/**
packages/integrations/fx-sdk/dist/**
packages/integrations/cursor-sdk/dist/**
packages/integrations/grok-build-sdk/dist/**
packages/evals/dist/**
retention-days: 1

Expand Down
1 change: 1 addition & 0 deletions packages/docs/docs.json
Original file line number Diff line number Diff line change
Expand Up @@ -51,6 +51,7 @@
"v4/integrations/overview",
"v4/integrations/claude-code",
"v4/integrations/codex",
"v4/integrations/grok-build",
"v4/integrations/eve",
"v4/integrations/deep-agents",
"v4/integrations/crewai",
Expand Down
1 change: 1 addition & 0 deletions packages/docs/images/integrations/grok-build.svg
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
85 changes: 85 additions & 0 deletions packages/docs/v4/integrations/grok-build.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,85 @@
---
Comment thread
antonvishal marked this conversation as resolved.
title: "Grok Build"
description: "Give Grok Build persistent Stagehand browser tools through its CLI and project MCP configuration."
---

Grok Build can load the Stagehand facade as a project MCP server. The facade owns one persistent browser and exposes `run`, `snapshot`, and `screenshot` for the full CLI run. The integration uses Grok's headless streaming output directly.

<Note>
Stagehand ships this experimental integration from the repository rather than publishing it as a standalone adapter.
</Note>

## Prerequisites

- Node.js 24 or newer
- pnpm 11.10.0
- The Grok Build CLI and either `XAI_API_KEY` or an existing `grok login`
- Google Chrome for local browser mode

## Quickstart

<Steps>
<Step title="Clone and build Stagehand">
```bash
git clone https://github.com/browserbase/stagehand.git
cd stagehand
pnpm install --frozen-lockfile
pnpm exec turbo run build --filter @browserbasehq/stagehand-integrations
```
</Step>
<Step title="Install and authenticate Grok Build">
```bash
npm install --global @xai-official/grok
grok login
# or: export XAI_API_KEY="your-xai-api-key"
```
</Step>
<Step title="Configure the MCP server">
Copy `packages/integrations/grok-build/.grok/config.toml` into the project where Grok will run. Replace the absolute facade path and browser credentials.
</Step>
<Step title="Run a browser task">
```bash
cd packages/integrations/grok-build
grok -p \
--output-format streaming-json \
--always-approve \
--tools search_tool,use_tool \
--disallowed-tools Agent \
--no-plan \
--no-subagents \
--disable-web-search \
"Open https://example.com, snapshot it, and report the heading."
```
</Step>
</Steps>

## Eval harness

The registered `grok_build` harness follows the same CLI pattern as the Cursor harness. It creates an isolated temporary home and workspace, copies cached Grok authentication only when no API key is present, writes the selected Stagehand MCP mount to project configuration, runs `grok -p --output-format streaming-json`, and maps native tool and usage events into the shared verifier trajectory.

```bash
evals run b:webvoyager \
--harness grok_build \
--tool stagehand_facade \
-l 1 -t 1 -e browserbase
```

| Variable | Purpose |
| --- | --- |
| `XAI_API_KEY` | Grok credential. The Stagehand MCP server receives only its own configured environment. |
| `GROK_HOME` | Source for cached `auth.json` when `XAI_API_KEY` is unset. |
| `EVAL_GROK_BUILD_PATH` | Optional path to the `grok` binary for eval runs. |
| `EVAL_GROK_BUILD_MAX_TURNS` | Grok turn limit. Defaults to 50. |
| `EVAL_GROK_BUILD_SANDBOX` | Optional Grok sandbox profile. |
| `STAGEHAND_BROWSER` | `local` or `browserbase`. |
| `BROWSERBASE_API_KEY` | Required for Browserbase. |

The harness supports `stagehand_facade`, `playwright_mcp`, and `chrome_devtools_mcp`. It restricts Grok to its MCP discovery and invocation tools, disables subagents, plan mode, and web search, and asks Grok not to edit repository files. The temporary Grok home disables compatibility MCP imports, memory, and background leader reuse so user configuration cannot add unrelated tools.

<Warning>
`run` executes model-authored JavaScript in the browser. Use Browserbase for untrusted tasks and review the [integration security boundary](/v4/integrations/overview#security-boundary).
</Warning>

<Card title="Grok Build integration source" icon="github" href="https://github.com/browserbase/stagehand/tree/main/packages/integrations/grok-build">
Open the CLI configuration and eval harness instructions.
</Card>
9 changes: 6 additions & 3 deletions packages/docs/v4/integrations/overview.mdx
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
---
title: "Integrations"
sidebarTitle: "Overview"
description: "Connect Claude Code, Codex, CrewAI, Deep Agents, Eve, Mastra, fx, Pi, or the Vercel AI SDK to a persistent Stagehand browser."
description: "Connect Claude Code, Codex, Grok Build, CrewAI, Deep Agents, Eve, Mastra, fx, Pi, or the Vercel AI SDK to a persistent Stagehand browser."
---

Each integration gives your agent one persistent browser and three tools: `run`, `snapshot`, and `screenshot`. Your agent decides how to navigate and interact while Stagehand manages the browser session.
Expand All @@ -19,6 +19,9 @@ Stagehand ships these experimental integrations from the monorepo and does not p
<Card title="Codex" icon="/images/integrations/codex.svg" href="/v4/integrations/codex">
Give a Codex agent persistent Stagehand browser tools over MCP/stdio.
</Card>
<Card title="Grok Build" icon="/images/integrations/grok-build.svg" href="/v4/integrations/grok-build">
Give a Grok Build agent persistent Stagehand browser tools.
</Card>
<Card title="Eve by Vercel" icon="/images/integrations/eve.svg" href="/v4/integrations/eve">
Give Vercel's framework for building durable agents native Stagehand tools.
</Card>
Expand Down Expand Up @@ -78,7 +81,7 @@ Every integration exposes the same browser capabilities.

The tools share a browser for the lifetime of the integration's client session. A navigation performed by `run` is visible to the next `snapshot`, and authentication and page state remain available across calls.

The Claude Code, Codex, CrewAI, Mastra, fx, Vercel AI SDK, and local Deep Agents examples keep one MCP client session open so the stdio server and browser stay alive. Eve, Pi, and Managed Deep Agents bind equivalent tools in-process.
The Claude Code, Codex, Grok Build, CrewAI, Mastra, fx, Vercel AI SDK, and local Deep Agents examples keep one MCP client session open so the stdio server and browser stay alive. Eve, Pi, and Managed Deep Agents bind equivalent tools in-process.

The private [`core/` workspace package](https://github.com/browserbase/stagehand/tree/main/packages/integrations/core) owns the shared TypeScript tool contract, browser runtime, native bindings, and stdio MCP server. The TypeScript integrations and CrewAI use this package.

Expand All @@ -88,7 +91,7 @@ Do not create a new MCP process for every tool call. Doing so starts a new brows

## Run the integrations from source

Claude Code, Codex, CrewAI, Mastra, fx, and the Vercel AI SDK use the shared TypeScript Stagehand facade MCP server. Eve and Pi use native in-process bindings from the same package. These integrations require Node.js 24 or newer and [pnpm](https://pnpm.io/installation) 11.10.0. CrewAI also requires Python 3.11–3.13 and [uv](https://docs.astral.sh/uv/).
Claude Code, Codex, Grok Build, CrewAI, Mastra, fx, and the Vercel AI SDK use the shared TypeScript Stagehand facade MCP server. Eve and Pi use native in-process bindings from the same package. These integrations require Node.js 24 or newer and [pnpm](https://pnpm.io/installation) 11.10.0. CrewAI also requires Python 3.11–3.13 and [uv](https://docs.astral.sh/uv/).

<Steps>
<Step title="Clone and install Stagehand">
Expand Down
11 changes: 11 additions & 0 deletions packages/evals/framework/benchHarness.ts
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,8 @@ import { runFxAgent } from "./fxRunner.js";
import { FX_TOOL_SURFACES, prepareFxToolAdapter } from "./fxToolAdapter.js";
import { runCursorAgent } from "./cursorRunner.js";
import { CURSOR_TOOL_SURFACES, prepareCursorToolAdapter } from "./cursorToolAdapter.js";
import { runGrokBuildAgent } from "./grokBuildRunner.js";
import { GROK_BUILD_TOOL_SURFACES, prepareGrokBuildToolAdapter } from "./grokBuildToolAdapter.js";
import {
buildExternalHarnessTaskPlan,
type ExternalHarnessTaskPlan,
Expand Down Expand Up @@ -346,6 +348,14 @@ export const cursorHarness = defineExternalHarness({
runAgent: runCursorAgent,
});

export const grokBuildHarness = defineExternalHarness({
harness: "grok_build",
supportedToolSurfaces: GROK_BUILD_TOOL_SURFACES,
defaultModels: ["grok-build/auto" as AvailableModel],
prepareToolAdapter: prepareGrokBuildToolAdapter,
runAgent: runGrokBuildAgent,
});

const harnessRegistry = new Map<Harness, BenchHarness>([
["stagehand", stagehandHarness],
["claude_code", claudeCodeHarness],
Expand All @@ -356,6 +366,7 @@ const harnessRegistry = new Map<Harness, BenchHarness>([
["deepagents", deepagentsHarness],
["fx", fxHarness],
["cursor", cursorHarness],
["grok_build", grokBuildHarness],
]);

export function registerBenchHarness(harness: BenchHarness): () => void {
Expand Down
187 changes: 187 additions & 0 deletions packages/evals/framework/grokBuildRunner.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,187 @@
import {
buildGrokBuildTranscript,
extractGrokBuildToolCall,
runGrokBuildSession,
stringifyError,
type GrokBuildProcessRunner,
type GrokBuildTokenUsage,
} from "@browserbasehq/stagehand-integrations-grok-build-sdk";
import type { AvailableModel } from "stagehand-v3";
import type { EvalLogger } from "../logger.js";
import type { ExternalHarnessTaskPlan } from "./externalHarnessPlan.js";
import type { PreparedGrokBuildToolAdapter } from "./grokBuildToolAdapter.js";
import { grokBuildAdapter } from "./harnesses/grokBuildAdapter.js";
import {
buildExternalHarnessPrompt,
metricValue,
parseEvalResult,
runExternalHarnessTask,
type ExternalHarnessToolAdapterLike,
type MetricValue,
type ParsedEvalResult,
} from "./harnesses/externalRunner.js";
import type { TaskResult } from "./types.js";
import type { ExternalHarnessVerifierConfig } from "./verifierAdapter.js";

export type { GrokBuildProcessRunner } from "@browserbasehq/stagehand-integrations-grok-build-sdk";

export interface GrokBuildRunnerInput {
plan: ExternalHarnessTaskPlan;
model: AvailableModel;
logger: EvalLogger;
toolAdapter?: PreparedGrokBuildToolAdapter;
signal?: AbortSignal;
runProcess?: GrokBuildProcessRunner;
verifier?: ExternalHarnessVerifierConfig;
}

export interface ParsedGrokBuildResult extends ParsedEvalResult {}

const MCP_ONLY_LINE =
"Your only browser access is the MCP server configured in this workspace. Never launch a browser yourself or run shell commands to browse.";

function composeGrokBuildToolInstructions(toolInstructions?: string): string {
return [
toolInstructions ?? "Use the available browser tools to complete the task.",
MCP_ONLY_LINE,
"Do not edit repository files.",
].join("\n");
}

export function buildGrokBuildPrompt(
plan: ExternalHarnessTaskPlan,
toolInstructions?: string,
): string {
return buildExternalHarnessPrompt({
plan,
toolInstructions: composeGrokBuildToolInstructions(toolInstructions),
resultContract: "marker",
});
}

export function parseGrokBuildResult(raw: string): ParsedGrokBuildResult {
return parseEvalResult(raw);
}

export async function runGrokBuildAgent({
plan,
model,
logger,
toolAdapter,
signal,
runProcess,
verifier,
}: GrokBuildRunnerInput): Promise<TaskResult> {
const adapterLike: ExternalHarnessToolAdapterLike = {
promptInstructions: composeGrokBuildToolInstructions(toolAdapter?.promptInstructions),
captureEvidence: toolAdapter?.captureEvidence,
drainStepObservations: toolAdapter?.drainStepObservations,
observedToolMatcher: toolAdapter?.observedToolMatcher,
};
return runExternalHarnessTask({
harness: "grok_build",
plan,
logger,
toolAdapter: adapterLike,
verifier,
resultContract: "marker",
fallbackErrorMessage: "Grok Build did not report success",
parseResult: parseGrokBuildResult,
runSession: async (prompt) => {
const sessionResult = await runGrokBuildSession({
prompt,
model,
logger,
signal,
runProcess,
session: {
...(toolAdapter?.cwd && { cwd: toolAdapter.cwd }),
...(toolAdapter?.env && { env: toolAdapter.env }),
...(process.env.EVAL_GROK_BUILD_PATH && {
binaryPath: process.env.EVAL_GROK_BUILD_PATH,
}),
maxTurns: readGrokBuildMaxTurns(),
...(process.env.EVAL_GROK_BUILD_SANDBOX && {
sandbox: process.env.EVAL_GROK_BUILD_SANDBOX,
}),
},
onToolResult: toolAdapter?.onToolResult
? (name) => toolAdapter.onToolResult!(name)
: undefined,
});
const usage = sessionResult.tokenUsage;
return {
raw: sessionResult,
resultText: sessionResult.resultText,
transcriptText: buildGrokBuildTranscript(sessionResult.events),
iterationError: sessionResult.iterationError,
status: sessionResult.status,
stopReason:
sessionResult.stopReason ||
(sessionResult.status === "sdk_error"
? stringifyError(sessionResult.iterationError) || undefined
: undefined),
usage: {
inputTokens: usage.inputTokens,
outputTokens: usage.outputTokens,
totalTokens: usage.totalTokens,
...(usage.reported && {
cachedInputTokens: usage.cachedInputTokens,
cacheCreationInputTokens: usage.cacheCreationInputTokens,
reasoningOutputTokens: usage.reasoningOutputTokens,
}),
},
costUsd: sessionResult.costUsd,
metrics: buildGrokBuildMetrics(usage, sessionResult.endEvent, sessionResult.events),
};
},
toTrajectory: (
{ raw, parsed, finalObservation, stepObservations, observedToolName, status },
taskSpec,
) =>
grokBuildAdapter.fromHarnessResult(
{
events: raw.events,
...(finalObservation && { finalObservation }),
...(stepObservations?.length && { stepObservations }),
...(observedToolName && { observedToolName }),
finalAnswer: parsed.finalAnswer ?? raw.resultText,
status,
usage: {
input_tokens: raw.tokenUsage.inputTokens,
output_tokens: raw.tokenUsage.outputTokens,
cached_input_tokens: raw.tokenUsage.cachedInputTokens,
reasoning_tokens: raw.tokenUsage.reasoningOutputTokens,
},
},
taskSpec,
),
});
}

export function readGrokBuildMaxTurns(): number {
for (const key of ["EVAL_GROK_BUILD_MAX_TURNS", "AGENT_EVAL_MAX_STEPS"]) {
const parsed = Number.parseInt(process.env[key] ?? "", 10);
if (Number.isFinite(parsed) && parsed > 0) return parsed;
}
return 50;
}

function buildGrokBuildMetrics(
usage: GrokBuildTokenUsage,
endEvent: Record<string, unknown> | undefined,
events: Array<Record<string, unknown>>,
): Record<string, MetricValue> {
const toolSteps = events.filter(
(event) => extractGrokBuildToolCall(event)?.subtype === "completed",
).length;
return {
grok_build_input_tokens: metricValue(usage.inputTokens),
grok_build_output_tokens: metricValue(usage.outputTokens),
grok_build_total_tokens: metricValue(usage.totalTokens),
grok_build_cached_input_tokens: metricValue(usage.cachedInputTokens),
grok_build_reasoning_tokens: metricValue(usage.reasoningOutputTokens),
grok_build_num_turns: metricValue(endEvent?.num_turns),
grok_build_tool_steps: metricValue(toolSteps),
};
}
Loading
Loading