Guide
This guide walks through setting up Google Vertex AI language models for use with Onyx.Authentication Methods
Onyx supports two authentication methods for Google Vertex AI:- Service Account JSON — upload a downloaded service account key file. This works with Onyx deployed inside or outside Google Cloud and is the default flow covered in the guide below.
- Workload Identity (GKE) — authenticate using your GKE cluster’s ambient credentials, with no key file to generate or store. See Workload Identity (GKE) below.
Service Account JSON
Create a Service Account for Onyx


Create a New Key for the Service Account
Navigate to Language Models
Configure Google Vertex AI
Choose Visible Models
Designate Provider Access
Workload Identity (GKE)
For Onyx deployed on GKE, you can authenticate to Vertex AI using GKE Workload Identity Federation. With this method, Onyx picks up the pod’s ambient GCP credentials automatically (via Google’s Application Default Credentials) and no key file is ever created or stored.Prerequisites
- Onyx is already deployed on GKE.
- Billing is enabled in the project where you will use Vertex AI.
- For the CLI examples, authenticate the Google Cloud CLI
and configure
kubectlfor your Onyx cluster. The deployment administrator needs permission to enable APIs and grant IAM roles; cluster or Kubernetes service account permissions are only needed when changing those settings.
Check Workload Identity is enabled
PROJECT_ID.svc.id.goog using your GKE project’s ID.

If Workload Identity or the metadata server is disabled
If Workload Identity or the metadata server is disabled
NODE_POOL_NAME and repeat for every node pool that runs Onyx.Give the Onyx pod identity Vertex AI access
onyx with your namespace:roles/aiplatform.user, formerly Vertex AI User)
in the Vertex AI project to that account’s Workload Identity
principal.
This grants access directly to the pod identity, without creating a Google Cloud service account or JSON key.If the Kubernetes account is already linked to a Google Cloud service account,
grant the role to that Google Cloud account instead and keep the existing link.
Repeat the grant for each distinct identity used by the API server and workers.Grant access directly with gcloud
Grant access directly with gcloud
GKE_PROJECT_ID identifies the cluster’s project;
VERTEX_PROJECT_ID identifies the project where Vertex AI is enabled. They can be the same.
The Helm chart uses the default Kubernetes service account when serviceAccount.create is false and
serviceAccount.name is empty; use the account shown on your running pods.Alternative: use a linked Google Cloud service account
Alternative: use a linked Google Cloud service account
KSA_NAME and NAMESPACE. If an email is returned, reuse that account rather than replacing its link.
Otherwise, open the Service Accounts page
and create an account named onyx-vertex-ai, without creating a JSON key.Set the values below to your deployment. For an existing account,
set GSA_EMAIL to the returned email and GSA_PROJECT_ID to the project that owns that account.KSA_NAME="onyx" in the binding. If you change the account assigned to pods,
apply the Helm values and roll out the API server and workers.
Standard cluster pods must run on node pools with the GKE Metadata Server enabled.
See Google’s service account linking
guide.Navigate to Language Models
Configure Google Vertex AI

global;
this is independent of your GKE cluster’s location.
Choose Visible Models
Designate Provider Access
Verify the connection
If the connection fails
If the connection fails
iam.gke.io/gcp-service-account annotation and the roles/iam.workloadIdentityUser binding.
A permissions error can also mean that roles/aiplatform.user was granted in a different project from the GCP
Project ID in Onyx, or that the Vertex AI API is disabled there.
Allow a few minutes for new IAM grants to propagate before retrying.FAQ
Does Onyx need to run on Google Cloud or GKE to use Vertex AI?
No. Vertex AI is accessed through Google Cloud APIs, and Onyx can call those APIs from another cloud or an on-premises deployment using Service Account JSON, with the required permissions and network access. The GKE steps above apply to credentials supplied automatically to Onyx pods by the GKE metadata server. Remotely accessing a GKE cluster withkubectl does not give an external Onyx deployment those pod credentials.
Onyx’s Workload Identity (GKE) option uses Application Default Credentials under the hood.
Google also supports Workload Identity Federation for workloads outside Google
Cloud.
That requires a separate setup that makes credentials available to the Onyx API server and workers;
it is outside the scope of this GKE walkthrough.
Can I use a Vertex AI API key instead of a service account?
No. Today, Onyx’s Vertex AI provider supports only Service Account JSON or Workload Identity (GKE) as authentication methods — a standalone Vertex AI API key (Vertex Express Mode) is not accepted. If you just want to authenticate with a plain API key, you can instead add a Custom LLM provider using thegemini provider name,
which routes through the Google AI Studio Gemini Developer API using an API key.
See the Custom Inference Provider guide.