gRPC Load Balancing for AI and LLM Inference on Kubernetes
Oct 2026
TL;DR
Kubernetes L4 load balancing pins a long-lived gRPC connection to one GPU pod. Linkerd balances each gRPC request across available inference pods without application code changes.
- The root cause lies in HTTP/2 multiplexing and Kubernetes connection-level routing.
- Linkerd proxies route individual requests, and the article provides working YAML for a multi-tenant inference platform.
- Per-route latency and success-rate metrics support SLOs for inference endpoints and streaming responses.
- Retry limits bound cascading GPU load, while mutual TLS secures pod-to-pod inference traffic through workload identity.
- A comparison table explains when to use Linkerd, Google Cloud’s GKE Inference Gateway, Istio, or Cilium.
Linkerd fits multi-tenant inference, Kubernetes GPU autoscaling, and token-generation streaming that requires balanced east-west gRPC traffic.
Why Kubernetes load balancing pins gRPC inference traffic to one GPU pod
Kubernetes Services usually balance traffic at Layer 4, which means kube-proxy assigns each TCP connection to a backend pod. The assignment happens when the client opens the connection. Kubernetes does not reconsider the selected pod for every request sent through that connection.
gRPC keeps those connections open because it runs over HTTP/2. A client can send many concurrent requests as separate HTTP/2 streams on one TCP connection. Every stream follows the endpoint that kube-proxy selected when the connection began, so one GPU pod may receive hundreds of inference requests while sibling pods receive none. A larger replica count cannot redistribute traffic already traveling over established connections.
GPU inference makes connection pinning especially expensive. An overloaded gRPC model server, such as NVIDIA Triton or a KServe gRPC predictor, can build a request queue while another GPU pod remains idle. HTTP-first servers such as vLLM only hit this when you run their gRPC endpoint rather than the default HTTP API. Clients then experience higher tail latency even though the deployment appears to have spare capacity. A CPU-backed CRUD service suffers the same imbalance, but an idle GPU wastes more expensive and often supply-constrained capacity.
Connection pinning also weakens Horizontal Pod Autoscaler behavior. An HPA may detect high GPU utilization or queue depth and create another replica. Existing clients keep sending requests through their established connections, so the new pod receives traffic only when a client opens a fresh connection and Kubernetes selects it. The deployment can therefore report a successful scale-out while the original pod stays saturated and the new GPU remains underused.
Token streaming extends the imbalance over longer periods. An LLM generation request may hold an HTTP/2 stream open for several seconds while the server emits tokens. Several long-running generations can consume one pod’s effective capacity, and later requests on the same pinned connection still reach that pod instead of an available replica.
Effective gRPC load balancing must inspect individual HTTP/2 requests at Layer 7. A Layer 7 proxy can maintain its own upstream connections and choose a ready inference pod for each new request. Connection-level balancing cannot make that decision after Kubernetes has assigned the original TCP connection.
Linkerd's per-request gRPC load balancing for inference pods
Linkerd distributes individual gRPC calls across ready inference pods, even when the client keeps one HTTP/2 connection open. Each meshed workload gets a Linkerd proxy that intercepts outbound traffic. The proxy accepts the application's persistent connection, discovers the Service's pod endpoints, and selects an endpoint for each new RPC.
Request-level balancing removes the connection pinning that leaves one GPU pod busy while others remain underused. When Kubernetes adds a ready inference pod, Linkerd can send new RPCs to it without waiting for clients to reconnect. Existing streaming RPCs remain on their original pods, while later streams can use the added capacity.
Linkerd has provided this behavior as a generally available feature on every meshed pod since 2018. You do not need to modify gRPC clients, add a language-specific resolver, or maintain balancing logic in each inference service. You inject the proxy through Kubernetes configuration and restart the workload so the meshed pods can handle traffic.
The same mechanism works across applications written in different languages because Linkerd handles balancing in the infrastructure layer. A Python inference router and a Go control service receive the same behavior without separate client libraries or balancing policies.
Architecture alone does not establish the proxy's performance cost. Buoyant's existing gRPC benchmarking post provides the relevant performance testing, so you can evaluate measured behavior separately from the request-routing mechanism.
Configuring a multi-tenant inference platform on Linkerd
A logical inference Service gives the router one destination while Linkerd handles tenant routing and model canaries. The example assumes each model Deployment exposes a container port named grpc and carries the labels referenced by its Service, plus a serving: gpu label so the Server below selects it.
apiVersion: v1
kind: Namespace
metadata:
name: inference
annotations:
linkerd.io/inject: enabled
---
apiVersion: v1
kind: ServiceAccount
metadata:
name: inference-router
namespace: inference
---
apiVersion: v1
kind: Service
metadata:
name: llm-inference
namespace: inference
spec:
ports:
- name: grpc
port: 8080
targetPort: grpc
---
apiVersion: v1
kind: Service
metadata:
name: tenant-a-v1
namespace: inference
spec:
selector:
tenant: tenant-a
model-version: v1
ports:
- name: grpc
port: 8080
targetPort: grpc
---
apiVersion: v1
kind: Service
metadata:
name: tenant-a-v2
namespace: inference
spec:
selector:
tenant: tenant-a
model-version: v2
ports:
- name: grpc
port: 8080
targetPort: grpc
---
apiVersion: v1
kind: Service
metadata:
name: tenant-b-v1
namespace: inference
spec:
selector:
tenant: tenant-b
model-version: v1
ports:
- name: grpc
port: 8080
targetPort: grpc
---
apiVersion: gateway.networking.k8s.io/v1
kind: GRPCRoute
metadata:
name: tenant-model-routing
namespace: inference
spec:
parentRefs:
- group: core
kind: Service
name: llm-inference
port: 8080
rules:
- matches:
- method:
type: Exact
service: inference.v1.Inference
headers:
- name: x-tenant
value: tenant-a
backendRefs:
- name: tenant-a-v1
port: 8080
weight: 90
- name: tenant-a-v2
port: 8080
weight: 10
- matches:
- method:
type: Exact
service: inference.v1.Inference
headers:
- name: x-tenant
value: tenant-b
backendRefs:
- name: tenant-b-v1
port: 8080
weight: 100
---
apiVersion: policy.linkerd.io/v1beta3
kind: Server
metadata:
name: gpu-models
namespace: inference
spec:
podSelector:
matchLabels:
serving: gpu
port: grpc
proxyProtocol: gRPC
---
apiVersion: policy.linkerd.io/v1alpha1
kind: MeshTLSAuthentication
metadata:
name: router-identity
namespace: inference
spec:
identityRefs:
- kind: ServiceAccount
name: inference-router
---
apiVersion: policy.linkerd.io/v1alpha1
kind: AuthorizationPolicy
metadata:
name: router-to-gpu-models
namespace: inference
spec:
targetRef:
group: policy.linkerd.io
kind: Server
name: gpu-models
requiredAuthenticationRefs:
- group: policy.linkerd.io
kind: MeshTLSAuthentication
name: router-identityThe router sends gRPC requests to llm-inference.inference.svc.cluster.local:8080 and sets x-tenant after authenticating the caller. Linkerd routes tenant A through a 90/10 model canary while keeping tenant B on its current version. Each GRPCRoute rule matches the gRPC service and then the x-tenant header, since a rule must carry an RPC match. inference.v1.Inference is a placeholder; use the fully-qualified name of your own gRPC service. The AuthorizationPolicy permits meshed requests only from the router's ServiceAccount, so clients cannot bypass tenant checks by calling model Services directly.
Best for multi-tenant inference platforms that use GPU autoscaling. When the HPA adds a pod behind one of these Services, Linkerd discovers the endpoint and can send it the next eligible gRPC request without waiting for clients to open new connections.
Measuring inference SLOs with per-route metrics
Linkerd exposes success rate and latency for each declared inference route without requiring application instrumentation. You can separate /Generate, /Embed, and model-specific gRPC methods instead of averaging every request into one service-level metric. Prometheus can retain those measurements, while Grafana or Buoyant Cloud can alert when a route consumes its error budget too quickly.
Route-level metrics let you define SLOs around the behavior users experience. For example, an interactive generation route might require a high success rate and a bounded latency percentile. A batch embedding route can use a different threshold because each request may process more input. Separate SLOs prevent slower batch work from hiding an interactive inference regression.
Streaming responses require a different interpretation of latency. A token stream can remain open for minutes while operating normally, so total duration alone cannot identify a problem. Track time to first token and token throughput through model instrumentation, then combine those signals with Linkerd’s route-level success and request metrics. Linkerd shows whether failures or slow responses concentrate on a specific gRPC method, model version, or tenant route.
A stalled stream may remain in flight and therefore may not appear as a completed failure immediately. You can enforce a maximum stream duration or idle timeout so abandoned streams close with an observable status. Alerts should compare completed success rate with stream age or application-level token activity. That combination distinguishes healthy long generations from connections that stopped producing tokens.
Best for token-generation streaming. Per-route metrics work well when several models or tenants share the same inference service. Each route gets its own latency and success-rate view, while model metrics supply token-specific signals such as first-token delay and generation rate.
Retries and mTLS for securing inference APIs
Retries need a bound because every repeated inference request consumes scarce GPU capacity. A client that retries each timeout independently can send duplicate generation work to an already saturated backend. Those retries increase queue depth, raise latency, and trigger more timeouts across other callers.
Linkerd bounds retries per route. retry.linkerd.io/limit on a GRPCRoute caps the attempts, and retry.linkerd.io/grpc names the gRPC statuses that are safe to retry. The limit stops proxies from multiplying load when a backend starts failing, and you enable retries only on the routes whose calls are safe to repeat. For streaming inference, retries should occur only before the backend begins returning tokens because a replacement stream cannot continue the original response.
Linkerd automatically establishes mTLS between meshed pods and derives workload identity from Kubernetes ServiceAccounts. The proxies authenticate both sides of every connection and encrypt inference traffic without requiring TLS code or certificates inside model servers. Pod IP changes do not affect identity because authorization follows the workload rather than its network address.
Multi-tenant isolation requires authorization policy in addition to encryption. You can assign separate ServiceAccounts to tenant routers or model backends, then allow each identity to call only its permitted inference routes. A compromised workload cannot reach another tenant’s model merely because both workloads share a cluster network. Linkerd enforces those decisions at the destination pod, where the proxy verifies the authenticated caller before forwarding the request.
Linkerd vs. GKE Inference Gateway vs. Istio for gRPC inference
GKE Inference Gateway and service meshes solve different routing stages. North-south traffic enters the cluster, while east-west traffic moves between internal services and inference pods.
GKE Inference Gateway fits a front door that serves several models. Its model-aware routing can choose a backend using inference-specific signals that a general service mesh does not inspect. However, an internal router calling model servers over gRPC still needs east-west request balancing. GKE Inference Gateway does not replace that layer, and it remains limited to GKE.
Istio sidecar mode provides mature per-request gRPC balancing and similar reliability controls through Envoy. The buying decision usually turns on operating complexity rather than missing gRPC features. Istio Ambient changes the comparison because ztunnel operates at L4. You must deploy waypoint proxies wherever inference traffic needs L7 routing, balancing, or observability.
Cilium suits clusters already built around its networking stack, but beta-stage gRPC balancing makes it a less settled choice for production inference paths.
Best for model-aware ingress. Use GKE Inference Gateway when GKE needs a north-south router that understands models and inference workload signals. Pair it with Linkerd when requests continue through internal gRPC services.
Best for east-west inference traffic. Use Linkerd when GPU pods need mature per-request balancing, Kubernetes portability, automatic mTLS, and route-level metrics without application changes.
Getting started with Linkerd on your inference platform
Install Linkerd, mesh the inference namespace, and apply the routing pattern shown earlier. Validate request distribution under representative inference traffic before enabling GPU autoscaling in production.
Next, compare your results with Buoyant’s gRPC benchmarking post and add the manifests to your platform repository. Linkerd fixes connection pinning in the infrastructure layer, so application owners do not need to rewrite inference clients or model servers.
FAQs
Does Linkerd require changes to inference application code?
No. Linkerd injects a proxy beside each selected Kubernetes workload and performs per-request gRPC load balancing in the infrastructure layer. You configure namespace or workload injection and routing policy, but the inference client and server code can remain unchanged.
Does Linkerd work with vLLM, NVIDIA Triton, and KServe?
Yes, when the selected serving path uses supported HTTP or gRPC traffic and the relevant workloads join the mesh. NVIDIA Triton exposes gRPC endpoints directly. vLLM and KServe deployments vary by protocol and architecture, so verify which service receives the request and mesh every internal hop that needs per-request balancing.
Can Linkerd run alongside GKE Inference Gateway?
Yes. GKE Inference Gateway can handle model-aware routing for traffic entering a GKE cluster, while Linkerd balances gRPC requests between internal services and inference pods. The two layers can complement each other when a front-door router calls meshed model backends.
Does Linkerd support EKS, AKS, and on-premises Kubernetes?
Yes. Linkerd runs on Kubernetes rather than depending on GKE-specific infrastructure. You can use the same gRPC load-balancing approach on EKS, AKS, supported on-premises clusters, and hybrid deployments, subject to Linkerd’s Kubernetes and networking compatibility requirements.
Will a newly scaled GPU pod receive traffic immediately?
The Linkerd proxy can send new requests to a ready endpoint without waiting for the client to establish another connection. Linkerd cannot move an inference request that is already streaming, but later requests multiplexed through the existing client connection can reach the new pod.
