Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 3 additions & 1 deletion skills/bigquery-ai-ml/SKILL.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,9 @@
---
name: bigquery-ai-ml
license: Apache-2.0
metadata:
version: v1
version: v2
publisher: google
description: >-
Leverages BigQuery's built-in machine learning and GenAI capabilities
for advanced data analytics. Use when you need to write SQL queries
Expand Down
4 changes: 3 additions & 1 deletion skills/bigquery-bigframes/SKILL.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,9 @@
---
name: bigquery-bigframes
license: Apache-2.0
metadata:
version: v3
version: v4
publisher: google
description: >-
Generates Python code using BigQuery DataFrames (BigFrames). Use by default for any Python data task involving BigQuery, including data processing, analysis, and machine learning. Don't use for SQL-first workflows or the google-cloud-bigquery client library — use bigquery-basics.

Expand Down
4 changes: 3 additions & 1 deletion skills/bigquery-sql/SKILL.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,9 @@
---
name: bigquery-sql
license: Apache-2.0
metadata:
version: v1
version: v2
publisher: google
description: >-
Provides BigQuery SQL query optimization techniques, execution best practices, and performance tuning rules for high-efficiency querying. Use when optimizing BigQuery SQL queries, reducing query costs, or designing performant SQL transformations.
---
Expand Down
161 changes: 78 additions & 83 deletions skills/dak-setup/scripts/dak-setup.js

Large diffs are not rendered by default.

31 changes: 17 additions & 14 deletions skills/gcp-spark/SKILL.md
Original file line number Diff line number Diff line change
@@ -1,22 +1,22 @@
---
name: gcp-spark
description: |
Develops, optimizes and executes Spark code on Managed Spark on Google Cloud (Dataproc Clusters and Serverless).
Reads and writes data using BigLake Iceberg catalogs, BigQuery and Spanner.
Debugs execution failures.
Develops, optimizes and runs PySpark/Spark code on Managed Spark (Dataproc
clusters and Serverless) on Google Cloud.
Use when:
- Writing Spark ETL pipelines on Google Cloud Platform.
- Optimizing PySpark or Spark SQL code for performance, memory, or OOM risks.
- Preparing Spark workloads for production submission.
- Training or running inference with Machine Learning models with spark on Google Cloud Platform.
- Authoring or running Spark/PySpark notebooks (any kernel), incl. data
analysis, reports and visualizations.
- Writing Spark ETL pipelines or preparing workloads for production.
- Training or running inference with ML models on Spark.
- Optimizing PySpark or Spark SQL for performance, memory, or OOM risks.
- Managing Spark clusters, jobs, batches, and interactive sessions.
Don't use when:
- Writing generic Python scripts that don't use Spark.
- Writing generic Python that doesn't use Spark.
- Performing simple SQL queries that can be done directly in BigQuery.
- Troubleshooting failed Spark workloads or analyzing logs (use @skill:gcp-spark-troubleshooting).
- Troubleshooting failed Spark workloads (use @skill:gcp-spark-troubleshooting).
license: Apache-2.0
metadata:
version: v20
version: v22
publisher: google
---

Expand Down Expand Up @@ -132,10 +132,13 @@ metadata:
.location(<REGION>)
.getOrCreate()
```
* **Production Logging**: In all production PySpark jobs and scripts,
you MUST use the standard Python `logging` module instead of `print()`
statements for job lifecycle, progress, and record counts. Configure it
with timestamps:
* **Production Logging**: In standalone
production PySpark batch jobs and scripts (`.py`), you MUST use the
standard Python `logging` module instead of `print()` statements for job
lifecycle, progress, and record counts. Do **NOT** use the `logging`
module in notebooks (`.ipynb`); use `print()`, `.show()`, and standard
cell outputs instead. Configure `logging` for `.py` scripts with
timestamps:

```python
import logging
Expand Down
2 changes: 1 addition & 1 deletion skills/gcs-security-assessment/references/phases/output.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ Present your assessment in a scannable, action-oriented format.

The target user is a Storage Admin with limited security expertise who needs to
quickly understand: what's wrong, how bad is it, and how to fix it. They may
action remediations via the Google Cloud Console or gcloud CLI — both paths
action remediations via the Cloud Console (Pantheon) or gcloud CLI — both paths
should be clear.

## Output Structure
Expand Down
22 changes: 17 additions & 5 deletions skills/google-cloud-auth-verification/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ description: >-
Use whenever interacting with GCP resources, running Spark/PySpark pipelines, BigQuery queries, GCS paths (gs://), or creating/running notebooks.
license: Apache-2.0
metadata:
version: v2
version: v3
publisher: google
---

Expand All @@ -17,10 +17,22 @@ metadata:
> [!IMPORTANT] **Pre-Flight Execution Priority Order**: Before generating code,
> implementation plans, or executing tasks for any GCP or Notebook workload:
>
> 1. **Verify Shell, Script & Notebook Credentials**: If shell-based commands,
> local Python scripts, or notebook kernels (`gs://...`, BigQuery, Dataproc)
> are required, verify credentials via bundled probe (`gcloud auth list &&
> gcloud config list`) or Application Default Credentials (ADC).
> 1. **Skip Pre-Flight Shell Probes When Live Notebook Execution or MCP Tools**
> **Are Available**:
> - When a live notebook cell execution tool (`execute_cell` /
> `notebook_execute_cell` / `notebook__execute_cell`) is available, or
> when the task is handled directly by a connected MCP tool
> (`bigquery__execute_sql`, `bigquery__get_table_info`,
> `bigquery__list_table_ids`, etc.), **do NOT run a pre-flight
> `gcloud auth list && gcloud config list` shell probe**. Instead,
> verify authentication reactively only if a cell execution or MCP tool
> call returns an authentication or credentials error (`401`, `403`,
> `DefaultCredentialsError`, `Reauthentication is needed`).
> - When `gcloud` or `bq` shell commands are required, or when generating
> standalone scripts or notebooks without a live cell execution tool
> (`execute_cell`), verify credentials upfront via the single bundled
> probe (`gcloud auth list && gcloud config list`) or Application
> Default Credentials (ADC).
> 2. **Distinguish Authentication vs. IAM Permissions**:
> - If `gcloud auth list` returns `No credentialed accounts`, **HARD
> STOP** immediately and instruct the user to run `gcloud auth login`
Expand Down
151 changes: 132 additions & 19 deletions skills/google-cloud-storage-basics/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,19 +3,19 @@ name: google-cloud-storage-basics
description: >-
Stores, retrieves, and manages data as objects in Cloud Storage (Google
Cloud Storage, or GCS) buckets. Use when you need to interact with Cloud
Storage — create or configure buckets, upload, download, stream, or transfer
data, organize objects with folders, generate signed URLs, control access
(IAM, ACLs, public access prevention), set storage classes and tiering
(Standard, Nearline, Coldline, Archive), manage cost and lifecycle, protect
data (versioning, encryption/CMEK, retention and Bucket Lock, object holds,
soft delete), host static websites, trigger Pub/Sub notifications on object
changes, mount buckets as a file system (gcsfuse), or optimize storage
performance at any scale. Covers the gcloud storage / gsutil CLI, JSON and
XML APIs, client libraries, Terraform, and Cloud Storage MCP servers. Don't
use for block storage (Persistent Disk), data warehousing/analytics
(BigQuery), or databases (Cloud SQL, Spanner, Bigtable, Firestore).
Storage — set up a Storage MCP server (remote or local Toolbox), create or
configure buckets, upload, download, stream, or transfer data, organize
objects with folders, generate signed URLs, control access (IAM, ACLs,
public access prevention), set storage classes (Standard, Nearline,
Coldline, Archive), manage lifecycle and cost, protect data (versioning,
CMEK, retention, Bucket Lock, holds, soft delete), host static websites,
trigger Pub/Sub notifications, mount buckets (gcsfuse), or optimize
performance. Covers gcloud storage / gsutil, JSON/XML APIs, client
libraries, Terraform, and Cloud Storage MCP servers. Don't use for
non-Storage MCP servers, block storage (Persistent Disk), BigQuery, or
databases (Cloud SQL, Spanner, Bigtable, Firestore).
metadata:
version: v2
version: v4
publisher: google
tags:
- gcs
Expand All @@ -38,6 +38,104 @@ balance cost, performance, and availability, and integrates with IAM for
fine-grained access control. GCS serves workloads ranging from website content
and backups to data lakes and high-throughput AI/ML training.

## Tool Execution Priority (MCP Toolbox First)

Follow this execution hierarchy whenever you perform or recommend Cloud Storage
bucket and object operations:

### 1. Priority 1 — Use `cloud-storage` MCP Tools When Connected

Before running `gcloud storage` or `curl` in a shell, check your available tools
for the `cloud-storage` MCP server (either the local
[MCP Toolbox](references/mcp-usage.md) or the Google-hosted remote Cloud Storage
MCP server). **Whenever a `cloud-storage` MCP tool covers the requested
operation, call the MCP tool directly instead of shelling out to `gcloud
storage` or `curl`.**

**Why MCP first:** MCP tools are built for agents. Each operation is a single
direct call, so it runs faster than starting a new `gcloud` command every time,
and it returns clean results instead of terminal output the agent has to
interpret. It also doesn't need the `gcloud` CLI installed and signed in. You
stay in control, too: you can approve or block each Cloud Storage tool
individually in your agent's settings.

Operation | Tool | Server
:--------------------- | :------------------------------------ | :-----
List buckets | `cloud-storage:list_buckets` | Both
Bucket metadata | `cloud-storage:get_bucket_metadata` | Local
Bucket IAM policy | `cloud-storage:get_bucket_iam_policy` | Local
Create a basic bucket | `cloud-storage:create_bucket` | Both
Delete an empty bucket | `cloud-storage:delete_bucket` | Both
List objects | `cloud-storage:list_objects` | Both
Object metadata | `cloud-storage:get_object_metadata` | Both
Read object content | `cloud-storage:read_object` | Both
Download to local file | `cloud-storage:download_object` | Local
Upload a local file | `cloud-storage:upload_object` | Local
Write text to object | `cloud-storage:write_object` | Local
Write text to object | `cloud-storage:write_text` | Remote
Copy an object | `cloud-storage:copy_object` | Local
Move / rename object | `cloud-storage:move_object` | Local
Delete an object | `cloud-storage:delete_object` | Both

**Server:** *Local* is the local MCP Toolbox (14 tools), *Remote* is the
Google-hosted remote server (8 tools), and *Both* means either server.

- `read_object` is capped at 8 MiB. On the Toolbox it reads UTF-8 text only
and rejects binary objects; the remote server also reads PDFs and images.
For binary or larger objects, use `download_object`.
- `upload_object` handles binary files and files larger than 8 MiB.
- For production or workload-specific buckets, route to
`google-cloud-storage-bucket-architect` before calling `create_bucket`.

> [!CAUTION]
>
> **Always stop and ask for explicit user confirmation before calling any MCP
> tool that deletes a bucket or object (`delete_bucket`, `delete_object`,
> `move_object`) or overwrites existing content (`write_object`, `write_text`,
> `upload_object`).**

### 2. Priority 2 — Fall Back to `gcloud storage` / JSON API

Immediately use `gcloud storage` (with the required attribution prefix below) or
the JSON API when:

1. **No `cloud-storage` MCP server is connected** in your environment, **or**
2. **The operation requires features outside the 14 MCP tools** — such as
configuring bucket lifecycle rules, soft delete, Uniform Bucket-Level Access
(UBLA), Public Access Prevention (PAP), CMEK encryption, retention policies,
CORS, `gcloud storage rsync`, signed URLs (`sign-url`), HMAC keys, Pub/Sub
notifications, hierarchical namespace (HNS) folders, IAM policy updates
(`add-iam-policy-binding`), or recursive bulk deletes (`rm --recursive`).

### 3. Gentle Recommendation When MCP Is Not Connected

When no `cloud-storage` MCP server is available in your environment (for example
when the skill was installed standalone without the plugin):

- **Never disrupt or delay the user's workflow.** Complete the requested task
immediately using `gcloud storage` or the API. Never pause or ask the user
to install the MCP server before answering.
- **Include a brief, one-line tip at the very end of your response ONLY when
all three conditions hold:**
1. No `cloud-storage` MCP tools are connected in the environment,
2. The task involved an interactive bucket or object operation that the
local MCP Toolbox natively handles (one of the 14 local Toolbox tools in
the table above), **and**
3. You have **not** already mentioned the MCP server earlier in the
conversation (**mention at most once per conversation**).
- **Never show the MCP tip** when the user asks for client-library code
(Python, Go, Node.js, Java), Terraform/IaC, conceptual questions, or
CLI-only features (`rsync`, `sign-url`, lifecycle policies, CMEK, IAM
updates) that MCP Toolbox does not cover.
- **Use this quiet, non-intrusive footnote format:**

> *Tip: To let me run these operations as structured tools without shell
> commands, install the
> [Google Cloud Storage plugin](https://github.com/gemini-cli-extensions/google-cloud-storage)
> (it asks for your project ID), or
> [add the Cloud Storage MCP server](https://github.com/gemini-cli-extensions/google-cloud-storage/blob/main/skills/google-cloud-storage-basics/references/mcp-usage.md)
> to your agent's MCP settings.*

## Attribution

Tag every Cloud Storage command you run or provide to the user while using this
Expand Down Expand Up @@ -118,10 +216,11 @@ npx skills add gemini-cli-extensions/google-cloud-storage --skill <skill-name>

## Quick Start

If a Cloud Storage MCP server is connected, prefer its structured tools (such as
`create_bucket`, `list_objects`, `read_object`, and `upload_object`) over the
CLI and API commands below — see [MCP Usage](references/mcp-usage.md). Fall back
to `gcloud storage` and the JSON API when no MCP server is available.
To set up, configure, or choose between the Google-hosted remote Cloud Storage
MCP server (`https://storage.googleapis.com/storage/mcp`) and the local MCP
Toolbox (`cloud-storage`), read [MCP Usage](references/mcp-usage.md). When a
`cloud-storage` MCP server is connected, prefer its tools (see
[Tool Execution Priority](#tool-execution-priority-mcp-toolbox-first)).

1. **Enable the Cloud Storage API:**

Expand All @@ -140,7 +239,10 @@ to `gcloud storage` and the JSON API when no MCP server is available.
For a production or workload-specific bucket, route to
`google-cloud-storage-bucket-architect` before creating a bucket (see
[Routing to Specialized GCS Skills](#routing-to-specialized-gcs-skills)).
The commands below create a basic default bucket.
The options below create a basic default bucket.

Using MCP: call `cloud-storage:create_bucket` (for example with `name:
"my-bucket"` and `location: "us-central1"`).

Using the gcloud CLI:

Expand All @@ -161,6 +263,11 @@ to `gcloud storage` and the JSON API when no MCP server is available.

3. **Upload an Object:**

Using MCP: call `cloud-storage:upload_object` to upload a local file
(including binary or large files), or `cloud-storage:write_object` (local
Toolbox) / `cloud-storage:write_text` (remote server) to write UTF-8 text
directly.

Using the gcloud CLI:

```bash
Expand All @@ -178,7 +285,11 @@ to `gcloud storage` and the JSON API when no MCP server is available.
"https://storage.googleapis.com/upload/storage/v1/b/my-bucket/o?uploadType=media&name=my-file.txt"
```

4. **Download an Object:**
4. **Download or Read an Object:**

Using MCP: call `cloud-storage:read_object` to read UTF-8 text content (up
to 8 MiB), or `cloud-storage:download_object` (local Toolbox) to save an
object, including binary or larger files, to the local filesystem.

Using the gcloud CLI:

Expand Down Expand Up @@ -224,7 +335,9 @@ to `gcloud storage` and the JSON API when no MCP server is available.
- [Data Management](references/data-management.md): IAM roles, authentication
(including signed URLs and HMAC), access control, routing for 403 error
troubleshooting, network security, automated security assessment, data
protection, and pricing and cost optimization (lifecycle rules, Autoclass).
protection, pricing and cost optimization (lifecycle rules, Autoclass), and
Cloud Audit Logs (enabling Data Access logs, exempting principals, and
estimating log volume and ingestion cost).

- [Storage Intelligence](references/storage-intelligence.md): The subscription
for managing storage at scale — Storage Insights datasets (BigQuery metadata
Expand Down
Loading
Loading