AI Gateway Google Vertex AI

Google Vertex AI

Use Gemini, Claude and open models from your own Google Cloud project, with every call going through Forgebench's budgets, metering and tracing.

You bring the Google Cloud project; Google bills it directly. Forgebench never serves Vertex AI on its own account.

How it works

You give Forgebench one way to authenticate to your project. On each request, Forgebench attaches it and calls Vertex AI for you. Your code only ever holds a Forgebench API key.

  • Vertex models are named vertex-<model>, for example vertex-gemini-2.5-flash.
  • Forgebench calls the global location. A model that isn't served at global returns 404.
  • Every call comes from one egress IP, 13.207.123.116, which you can allowlist.

Choose how to connect

MethodModelsWhat you give Forgebench
Service account keyAllA service-account JSON key
Express API keyGemini, and open models such as Llama. Not Claude.An API key, plus a project ID for open models
Workload Identity FederationAllA configuration file that contains no secret
  • Service account key is the usual choice: it covers every model and takes a few minutes.
  • Express API key is the quickest way to try Gemini, but Google doesn't accept API keys for Claude.
  • Workload Identity Federation avoids long-lived keys entirely. Use it if your organization disallows them. It needs an Enterprise workspace, because we create an AWS role dedicated to your workspace.

All three are added in the same place: API Keys → Credentials → Add credential (or Admin → Secrets), then Vertex AI. You need the admin role.

Service account key

  1. Enable Vertex AI and your models

    In your project, enable the Vertex AI API. Gemini works straight away. For partner models such as Claude, open each one in Model Garden and click Enable.

  2. Create a service account

    For example forgebench-vertex@<your-project>.iam.gserviceaccount.com. Grant it Vertex AI User (roles/aiplatform.user) and nothing else.

  3. Create a JSON key

    On the service account's Keys tab, add a JSON key and download it. If your organization blocks key creation (iam.disableServiceAccountKeyCreation), request an exception for this project or use Workload Identity Federation.

  4. Add it to Forgebench

    Choose Vertex AI, set Authentication to Service account JSON (all models), paste the whole file and choose Save provider key. The key is encrypted at rest and never shown again.

Authentication starts on Express API key. Switch it before you paste.
Authentication starts on Express API key. Switch it before you paste.

Express API key

  1. Create a key

    In the Google Cloud console, create an API key for Vertex AI in your project.

  2. Add it to Forgebench

    Choose Vertex AI, leave Authentication on Express API key, and paste the key. To use open models such as Llama, DeepSeek or Qwen, also enter your Project ID. Gemini doesn't need it.

Workload Identity Federation

Your project trusts one AWS IAM role, and Forgebench runs your workspace's gateway as that role. On each call the gateway proves its identity to Google and gets back a token that expires within the hour. There is no key to store, leak or rotate.

The role belongs to your workspace alone and has no AWS permissions. Only your workspace's gateway can assume it.

  1. Ask us for your role

    Contact Forgebench support. We'll create the role and send you its ARN, our AWS account ID, and a setup guide with your values filled in.

  2. Enable Vertex AI and your models

    The same as for a service account: enable the Vertex AI API, and enable partner models in Model Garden.

  3. Create a pool and an AWS provider

    Create a workload identity pool, then an AWS provider in it for our account ID. Restrict the provider to your role with this attribute condition: attribute.aws_role == "arn:aws:sts::<account id>:assumed-role/<role name>"

  4. Grant Vertex AI User

    Grant roles/aiplatform.user on the project to principalSet://iam.googleapis.com/projects/<project number>/locations/global/workloadIdentityPools/<pool>/attribute.aws_role/arn:aws:sts::<account id>:assumed-role/<role name>.

  5. Create the configuration file

    Run gcloud iam workload-identity-pools create-cred-config with --aws for your provider. The file names your pool and provider and contains no secret.

  6. Add it to Forgebench

    Choose Vertex AI, set Authentication to Workload Identity Federation (all models, no key), paste the file, enter your project ID and choose Save provider key.

To cut off access, remove the Vertex AI User grant in Google Cloud. Calls fail within about two minutes. Disabling the provider only blocks new token exchanges; tokens already issued keep working for up to an hour.

Send a test request

Saving a credential only checks its format. Google accepts or rejects it on the first request, so send one straight away:

curl -X POST https://api.forgebench.ai/v1/chat/completions \
-H "Authorization: Bearer sk_..." \
-H "Content-Type: application/json" \
-d '{
  "model": "vertex-gemini-2.5-flash",
  "messages": [{"role": "user", "content": "ping"}]
}'

A 200 with a completion means you're connected. If not, see Troubleshooting.

The Gateway catalog lists every Vertex model. Until you add a credential they show as Needs key.

Vertex models in the Gateway catalog before a credential is added.
Vertex models in the Gateway catalog before a credential is added.

Restrict access with VPC Service Controls

Optional. A VPC Service Controls perimeter makes your project refuse Vertex AI calls unless they come from Forgebench's egress IP and from the identity you gave Forgebench. You'll need an organization-level Google Cloud admin.

SettingValue
Egress IP13.207.123.116/32
ProtocolHTTPS (TCP 443)

We'll tell your workspace admin before this address changes.

  1. Create an access level

    In Access Context Manager, create an access level for the IP subnetwork 13.207.123.116/32.

  2. Create a perimeter in dry-run mode

    Put your project in a perimeter that restricts aiplatform.googleapis.com. Start it in dry-run mode, so it only logs what it would block.

  3. Add an ingress rule

    Allow Vertex AI calls from that access level by your Forgebench identity only: the service account for a key, or the principalSet://… member from the Workload Identity Federation steps.

  4. Test, then enforce

    Send a test request, check the audit logs show no violation for it, then enforce the perimeter. A new perimeter can take up to 30 minutes to start blocking.

The IP is shared by every Forgebench customer, so it shows a call came from Forgebench, not which workspace sent it. The ingress rule's identity is what limits access to your workspace. That's why the rule needs both.

Troubleshooting

ErrorCauseFix
404 publisher model "was not found or your project does not have access to it"The model isn't enabled in your project, or isn't served at globalEnable it in Model Garden, or choose a model served at global
Claude fails with an express API keyGoogle doesn't accept API keys for ClaudeUse a service account key or Workload Identity Federation
403 VPC_SERVICE_CONTROLSThe access level doesn't include our IP, or the ingress rule names a different identityCheck both. The error's vpcServiceControlsUniqueIdentifier finds the matching audit log entry.
Calls from other IPs still succeedThe perimeter isn't enforced yetWait up to 30 minutes after enforcing, then test again
403 from sts.googleapis.com (Workload Identity Federation)The provider's account ID or attribute condition doesn't match your roleCopy the role exactly as we sent it, including sts and assumed-role
403 "Permission 'aiplatform.endpoints.predict' denied" (Workload Identity Federation)Vertex AI User isn't granted to the principalSet://… member, or hasn't propagatedGrant it, then allow a few minutes
Saving says the key belongs to a different Forgebench organizationSomeone else has added the same key, so it may have leakedRotate it in Google Cloud and add the new key

Next