AI Gateway: Vertex gemini-* returns 500 unknown backend
Symptom
POST /v1/chat/completionswithmodelmatchinggemini-*returns 500 (body may be genericserver_error).- Envoy access log
response_code_detailsincludesext_proc_errorand text likeunknown backend: …/zelkor-platform-backend-vertex/route/…. - Other models on the same gateway (for example Ollama Cloud) still return 200.
AIServiceBackendfor Vertex may show Accepted even while requests fail.
Cause
Envoy AI Gateway builds the dataplane backend list from BackendSecurityPolicy auth. For Vertex with a service-account file, the controller must mint a token into Secret ai-eg-bsp-<backend-security-policy-name> (key gcpAccessToken). If rotation fails, the controller skips the Vertex backend in filter config while the route still references it — ext_proc then reports unknown backend.
Common rotation failures:
- Secret referenced by
existingSecretlacks keyservice_account.json(wrong key name). - Invalid, revoked, or deleted GCP service account key JSON.
- GCP OAuth error such as
invalid_grant/account not foundon theBackendSecurityPolicystatus.
Confirm
Run against your platform namespace (example: zelkor):
bash
kubectl -n zelkor get secret <vertex-existingSecret> -o go-template='{{range $k,$v := .data}}{{$k}}{{"\n"}}{{end}}'Expect service_account.json.
bash
kubectl -n zelkor get secret ai-eg-bsp-<release>-vertex-gcp -o go-template='{{range $k,$v := .data}}{{$k}}{{"\n"}}{{end}}'Expect gcpAccessToken. If the Secret is missing, rotation never succeeded.
bash
kubectl -n zelkor get backendsecuritypolicy <release>-vertex-gcp -o yamlCheck status.conditions for ReconciliationFailed and the rotation error message.
Search AI Gateway controller logs for vertex-gcp, service_account, or Skipping this backend.
Fix
- Create a valid GCP service account key JSON for the project in
workspace.models.providers.vertex.project. - Recreate the Secret with the required key name (no pod restart required):
bash
kubectl -n zelkor create secret generic zelkor-platform-vertex-sa \
--from-file=service_account.json=./sa.json \
--dry-run=client -o yaml | kubectl apply -f -- Wait until
BackendSecurityPolicyis Accepted andai-eg-bsp-*-vertex-gcpcontainsgcpAccessToken. - Retry
POST /v1/chat/completionswithmodel: gemini-2.5-flash(or your configured id).
Alternatives: set workspace.models.providers.vertex.credentialsJson in Helm instead of existingSecret, or leave both empty for ADC / Workload Identity on GKE.