Skip to content

[bot] Google GenAI Batches API is not instrumented (mis-tagged as a generic LLM span) #165

Description

@braintrust-bot

Summary

The genai_1_18_0 module instruments Google's google-genai Java SDK by subclassing its package-private ApiClient (com.google.genai.BraintrustApiClient) and swapping it into every service object on the Client, including client.batches and client.async.batches (BraintrustInstrumentation.wrapClient, lines 36–52). This means a call to client.batches.create(...)/.get(...)/.list(...)/.cancel(...)/.createEmbeddings(...) does produce a span — but BraintrustApiClient.tagSpan() has no awareness of the Batches API's request/response shape, so the span is mis-tagged: it's marked span_attributes.type = "llm" as if it were a real generation call, but the request/response field extraction (built for generateContent-shaped bodies) finds essentially nothing meaningful to populate.

What is missing

In braintrust-sdk/instrumentation/genai_1_18_0/src/main/java/com/google/genai/BraintrustApiClient.java, tagSpan() (lines 50–161):

  • Request metadata extraction (lines 65–76) looks for top-level model, systemInstruction, tools, toolConfig, safetySettings, cachedContent. A batchGenerateContent request body wraps everything under a batch object ({"batch": {"displayName": ..., "inputConfig": {...}}}), and a createEmbeddings batch job body is shaped around EmbeddingsBatchJobSource/CreateEmbeddingsBatchJobConfig — none of the expected top-level fields exist, so metadata stays essentially empty.
  • input_json construction (lines 97–113) only populates model/contents/config (from generationConfig) — a batch-create body has none of these at top level (the model is only present in the URL path, e.g. {model}:batchGenerateContent, and getModel(genAIEndpoint) (used as a fallback at line 102) is the only path by which model would end up in input_json at all); the batch's actual per-item contents/generation requests (whether inline or file/GCS-referenced) are never captured.
  • Response handling (lines 116–150) dumps the whole response body as output_json (line 125) — for a batch create/get call, that whole body is just BatchJob resource metadata (name, state, createTime, etc.), not generative output — and reads usageMetadata for metrics (line 128), which a BatchJob response never has. There is no instrumentation at all of retrieving a completed batch's actual per-request results (where real generation outputs and usageMetadata become available), so even a fully successful batch job produces a span with type: "llm" but no meaningful input, output, or token metrics.
  • No test or example anywhere in the repo exercises client.batches in any form (confirmed via grep for batch under braintrust-sdk/instrumentation/genai_1_18_0/ — the only match is BraintrustInstrumentation.java's own field-swap wiring, not a test).

Braintrust docs status: not_found

Checked https://www.braintrust.dev/docs/integrations/ai-providers/gemini in full: no mention of batchGenerateContent, Batches, or batch/async generation jobs anywhere, in any language. The only "batch"-adjacent mention is batchEmbedContents, and only in the unrelated context of AI proxy/gateway passthrough support for embeddings ("The gateway also supports Gemini's native embedContent and batchEmbedContents endpoints") — not span/tracing coverage, and not the Batches job API this issue is about.

Note: this repo's own gap-audit history already treats "Batch API not instrumented, mis-tagged as a generic LLM span" as a valid, in-scope finding — see the already-filed and still-open #155 for Anthropic's Message Batches API, which this issue mirrors for the Google GenAI provider (a distinct upstream API/SDK/module, not a duplicate).

Upstream sources

Local repo files inspected

  • braintrust-sdk/instrumentation/genai_1_18_0/src/main/java/com/google/genai/BraintrustApiClient.java — lines 50–161 (tagSpan)
  • braintrust-sdk/instrumentation/genai_1_18_0/src/main/java/com/google/genai/BraintrustInstrumentation.java — lines 22–57 (wrapClient), confirming client.batches/client.async.batches are explicitly wired to the same instrumented ApiClient (lines 37, 47) and therefore do produce (mis-tagged) spans rather than bypassing instrumentation entirely
  • Repo-wide grep for batch/Batch/batchGenerateContent/BatchJob under braintrust-sdk/instrumentation/genai_1_18_0/ — only match is the field-swap wiring above; no test or example exercises the Batches API

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions