gRPC load balancing guide for Linkerd on Kubernetes
Sep 2026
Why gRPC Load Balancing Breaks on Kubernetes
Kubernetes balances TCP connections, while gRPC sends many RPCs over one long-lived HTTP/2 connection. When a client connects to a Service, kube-proxy rules select a pod and network connection tracking preserves that choice. Every subsequent RPC on the same connection reaches the selected pod. Kubernetes does not make another backend decision for each RPC because its default load balancing cannot see HTTP/2 request boundaries. The Kubernetes project describes this behavior as connection pinning.
HTTP/2 multiplexing can keep that connection busy for hours. A gRPC client may send concurrent unary calls and streaming calls through the same socket, but kube-proxy sees one TCP flow. The selected pod receives all traffic from that client, while other replicas may receive none. A large number of clients can distribute connections more evenly by chance, but a few high-volume clients can still create severe imbalance.
HTTP/1.1 often avoids such persistent imbalance because clients tend to create more connections or recycle them more frequently. Each new TCP connection gives kube-proxy another opportunity to select a pod. HTTP/1.1 keep-alive can still pin several requests, but connection turnover often makes connection-level balancing approximate request-level balancing. HTTP/2 deliberately reduces that turnover.
The Horizontal Pod Autoscaler cannot redistribute established gRPC connections. Suppose one pinned pod reaches a CPU threshold and HPA adds two replicas. Kubernetes adds the new pod addresses to the Service, but existing clients remain connected to the original pod. The new replicas receive traffic only when clients establish new connections. HPA may therefore add available capacity while the overloaded pod continues handling the same RPC load.
What L7 Load Balancing Means for gRPC
Layer 4 load balancing assigns each TCP connection to a backend pod. Since a gRPC client can send many RPCs through one long-lived HTTP/2 connection, Kubernetes selects a pod when the connection opens and keeps every RPC on that pod. Kubernetes documents this connection-pinning behavior.
Layer 7 load balancing assigns each gRPC call separately. An L7 proxy understands HTTP/2 framing and can identify the stream that represents an individual RPC. The proxy selects a backend for that stream, even when the client continues using the same TCP connection. It can therefore distribute consecutive RPCs across different pods without forcing the client to reconnect.
HTTP/2 awareness prevents the proxy from treating a multiplexed connection as one indivisible flow. For unary gRPC, each request and response forms one RPC stream. For server-streaming or bidirectional calls, the proxy chooses a backend when the stream begins. Every message within that stream stays with the selected backend because moving an active stream would break its state and ordering.
How Linkerd Fixes gRPC Load Balancing
Linkerd moves gRPC load balancing into a proxy that runs beside every meshed application container. Pod-level network rules redirect outbound traffic through this sidecar. The proxy understands HTTP/2, so it can identify the stream boundaries that represent individual gRPC calls instead of treating the connection as one indivisible flow.
For each new RPC, the Linkerd proxy discovers the available pods behind the Kubernetes Service and selects a backend. Multiple RPCs carried over one client connection can therefore reach different pods. Linkerd preserves HTTP/2 connection reuse while avoiding the pod pinning caused by kube-proxy and iptables.
Linkerd uses an exponentially weighted moving average algorithm to choose among eligible backends. The algorithm incorporates recent response latency and current load, which lets the proxy reduce traffic to a slow or busy pod. Recent measurements receive more weight than older ones, so routing decisions respond to changing pod conditions without requiring a central load balancer.
Proxy injection provides this behavior without changes to application code, protocol definitions, or gRPC client configuration. You can continue using normal Kubernetes Service names. Linkerd intercepts the connection, discovers service endpoints, and balances RPCs before forwarding them. Existing clients do not need a custom resolver or client-side balancing library.
Linkerd has applied this proxy-based model to Kubernetes gRPC traffic since 2018. An official Kubernetes article published that year described L7 proxying as the practical answer to HTTP/2 connection pinning and presented Linkerd as an implementation. That production lineage matters because per-RPC routing sits directly in the request path for every meshed service.
gRPC Autoscaling: How Linkerd Makes HPA Work
Linkerd lets new replicas receive gRPC calls without waiting for clients to replace their existing HTTP/2 connections. Kubernetes normally assigns a connection to one backend when the connection opens. Every RPC on that connection remains pinned, so an HPA scale-out can add pods without reducing load on the original pod. The Kubernetes project documents this connection-level limitation.
A Linkerd proxy routes individual RPCs instead of relying on the backend selected for the client connection. After a new pod becomes Ready and Kubernetes publishes its endpoint, the client-side proxy can add that pod to its available destinations. Subsequent unary RPCs can then reach the new replica over a proxy-managed connection. The application keeps its existing gRPC channel, and no client reconnection or code change is required.
HPA still decides when to add replicas based on its configured metrics. Linkerd changes what happens after that decision. Added capacity can start receiving new RPCs as soon as endpoint updates reach the proxies, which gives overloaded replicas a chance to shed work. Readiness probes remain important because they prevent proxies from routing calls to a pod before the service can handle them.
Long-lived gRPC streams require a separate expectation. Linkerd can place newly created streams on new replicas, but it cannot move an active stream between pods without interrupting that stream. Workloads dominated by persistent streams may need connection rotation, stream-aware capacity planning, or application behavior that periodically establishes new streams.
gRPC Observability Out of the Box
Linkerd observes gRPC calls at the proxy, so your application does not need tracing libraries or custom metrics code. The proxy reads HTTP/2 request metadata and gRPC trailers, including the grpc-status value used to distinguish successful calls from failures. It records request rate, success rate, and latency distributions for traffic between meshed workloads.
Kubernetes network telemetry usually stops at the TCP connection. You can see connection counts and transferred bytes, but one HTTP/2 connection may carry thousands of RPCs with different methods, durations, and status codes. TCP metrics cannot tell you that /payments.PaymentService/Authorize is slow while another method remains healthy.
Linkerd can group metrics by gRPC method when route metadata identifies those methods. You can define that metadata through Linkerd routing policy or generate it from protobuf definitions without changing the service implementation. Prometheus stores the resulting latency histograms, which Grafana can display as percentile distributions rather than a single average that hides slow calls.
The Linkerd Viz extension exposes the same telemetry through dashboards and command-line tools. Grafana provides historical views for workloads and routes. linkerd viz stat summarizes request rate, success rate, and latency percentiles, while linkerd viz top shows live requests with fields such as source, destination, path, and status. Since a gRPC path contains the service and method name, linkerd viz top can reveal which RPC is failing while the traffic is active.
Reliability Primitives: Retries, Timeouts, and Circuit Breaking
Linkerd keeps reliability policy outside application code by reading Kubernetes resources that travel with the service deployment. A ServiceProfile can identify gRPC methods, mark eligible methods as retryable, assign timeouts, and define a retry budget. Newer Linkerd releases also support parts of this policy through Gateway API resources such as HTTPRoute, although available fields depend on the installed Linkerd and Gateway API versions.
Retry budgets limit retries as a percentage of original requests over a rolling period. With a retry ratio of 0.2, for example, 100 original RPCs create capacity for roughly 20 retries, subject to the configured minimum retry allowance. A fixed retry count could let every failed RPC produce several more calls during an outage. Percentage-based budgets cap that traffic amplification as the failure rate rises. You should mark only safe, idempotent gRPC methods as retryable unless the application can deduplicate repeated writes.
Timeouts bound how long the proxy waits for an RPC before returning an error to the client. You can apply them per route in a ServiceProfile or through supported Gateway API timeout fields. Long-lived server streams and bidirectional streams need separate treatment because a short request timeout can terminate healthy sessions.
Linkerd implements circuit-breaking behavior through failure accrual. The client-side proxy observes endpoint failures, temporarily removes an unhealthy pod from consideration, and tests it again after a backoff period. Linkerd versions expose failure-accrual controls through Service metadata annotations or supported policy resources rather than application code. Check the API reference for your installed version before committing those settings to deployment manifests.
Canary Deployments and Progressive Delivery with gRPC
Linkerd can divide gRPC traffic between stable and canary services at the RPC level. The proxy reads each new gRPC request carried over HTTP/2 and applies the configured traffic weight before selecting a destination. Applications keep calling the same Kubernetes Service, so proto files and client-side balancing code remain unchanged.
Per-request routing makes low canary weights meaningful. With a 5 percent canary weight, Linkerd can evaluate every new RPC independently rather than assigning 5 percent of long-lived client connections to the canary. Observed traffic approaches the configured weight as request volume grows. Connection-level routing can produce a much less even distribution because one busy connection may carry far more RPCs than another.
Argo Rollouts and Flagger can automate weight changes through traffic-splitting resources supported by Linkerd. A delivery controller starts with a small canary share, evaluates configured health metrics, and then increases the share or rolls it back. Linkerd proxies enforce each updated weight without requiring clients to reconnect or discover a separate canary endpoint.
Streaming RPCs require a narrower interpretation of the weight. Linkerd chooses a destination when each stream begins, but it does not redistribute individual messages within an active stream. A small number of long-lived streams may therefore produce traffic ratios that differ from the configured weight. Canary analysis should use enough newly started RPCs and should evaluate gRPC success rate and latency rather than relying only on connection counts.
gRPC Streaming and Long-Lived Connections
Linkerd balances a streaming RPC when the stream begins. A gRPC stream counts as one RPC, even when it carries thousands of messages over several hours. Linkerd selects a destination pod for that RPC and keeps the stream attached to the same pod until the stream closes or its connection fails.
Server streaming and bidirectional streaming follow the same placement rule. For server streaming, one backend sends multiple messages after the client opens the RPC. For bidirectional streaming, both peers exchange messages on the established stream. Linkerd does not distribute individual messages across pods because those messages share ordered state within one RPC.
Linkerd can place separate streams on different pods even when the client reuses one HTTP/2 connection. New streams can therefore reach pods added during an HPA scale-out. Existing streams remain on their original pods, so one high-volume stream can still create uneven load that request-level balancing cannot divide.
Rolling deployments require suitable shutdown and reconnection behavior. Removing a pod from endpoint discovery prevents new streams from selecting it, but removal does not migrate active streams. If Kubernetes terminates the pod before a stream finishes, the client may need to reconnect and rebuild any application state. You should set an adequate termination grace period and test client reconnection for workloads with long-running streams.
Linkerd vs. Istio Ambient vs. Cilium for gRPC
Linkerd
Linkerd balances unary calls and new streams through the proxy attached to each meshed pod. Because each source proxy selects among current destination pods, HPA scale-out does not depend on clients opening new HTTP/2 connections. Long-lived streams stay with their selected destination until they end.
Best for Choose Linkerd when you need mature gRPC balancing with per-method metrics and direct HPA compatibility.
Istio Ambient
Istio Ambient's ztunnel creates an L4 bottleneck before traffic reaches an L7 waypoint. Issue 56864, filed in July 2025 and still open in the supplied research, describes connections pinning traffic to one waypoint replica. New waypoint replicas can therefore receive little traffic after HPA scale-out. Ambient deployments commonly add a waypoint for each destination namespace that needs L7 handling. Istio sidecar mode avoids this specific problem because each sidecar balances individual RPCs at L7.
Best for Choose Istio sidecar mode when you already operate Istio and require predictable gRPC balancing. Test Ambient waypoint scaling under realistic connection lifetimes before adopting it.
Cilium
Cilium's gRPC load-balancing capability is described as beta or experimental in the supplied brief. The available sources do not verify per-RPC behavior, HPA response, streaming semantics, method-level metrics, or production readiness. Engineers should confirm those details against the Cilium version they plan to deploy and test long-lived connections during scale-out.
Best for Consider Cilium when you already use its networking stack and can validate the current gRPC implementation in your own workload.
Customer Evidence: gRPC at Scale with Linkerd
The available customer evidence confirms three published Linkerd case studies, but the supplied source contains only their index titles. The Buoyant case study index does not provide the underlying metrics, architecture details, quotes, or methodology needed to evaluate broader technical claims.
ZeroFlucs
ZeroFlucs used Linkerd to secure gRPC traffic on Kubernetes with mTLS identities. Its listing specifically names Go, gRPC, Kubernetes, Linkerd, and mTLS. The index does not confirm the number of services or describe how ZeroFlucs handled load balancing, autoscaling, or streaming.
Entain
Buoyant lists a case study reporting 10x throughput while reducing costs. The supplied index material associates the listing with Entain only indirectly, and it does not provide the baseline, test conditions, or workload profile behind the 10x figure. Engineers should treat the number as a reported case study outcome rather than a reproducible benchmark for gRPC deployments.
Plaid
Plaid’s case study reports pain-free deployments at global scale. The title supports Linkerd’s use in a large deployment environment, but the index does not identify Plaid’s protocol mix or confirm a gRPC-heavy architecture. It also provides no deployment frequency, service count, or traffic volume.
Together, the listings show Linkerd serving several production needs, including gRPC mTLS, higher throughput, lower costs, and large-scale deployments. The available evidence does not support more specific claims about fleet size, configuration, or the share of traffic carried over gRPC.
Getting Started: Meshing a gRPC Service with Linkerd
1. Install the Linkerd CLI
Install the CLI package that matches your operating system, then confirm that your shell can find it.
linkerd version --client
linkerd check --pre2. Install the Linkerd control plane
linkerd install --crds | kubectl apply -f -
linkerd install | kubectl apply -f -
linkerd checkRun these commands against the intended Kubernetes context. For production clusters, review the generated configuration before applying it.
3. Enable proxy injection
Annotate the namespace that contains both the gRPC client and server workloads.
kubectl annotate namespace grpc-demo \
linkerd.io/inject=enabledLinkerd injects proxies only when Kubernetes creates pods. Restart existing deployments to replace unmeshed pods.
kubectl rollout restart deployment -n grpc-demo
kubectl rollout status deployment -n grpc-demoAlternatively, inject a specific workload manifest before applying it.
kubectl get deployment grpc-server -n grpc-demo -o yaml \
| linkerd inject - \
| kubectl apply -f -4. Confirm that every pod has a proxy
kubectl get pods -n grpc-demo \
-o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.spec.containers[*].name}{"\n"}{end}'Each meshed pod should include a linkerd-proxy container. Run linkerd check again if pods fail to become ready.
5. Install Viz and inspect traffic
linkerd viz install | kubectl apply -f -
linkerd viz check
linkerd viz stat pods -n grpc-demoSend repeated unary RPCs through one reused client channel while the gRPC server runs multiple replicas. The per-pod request rates should show traffic reaching more than one server pod.
For a live request view, run the following command while the client generates traffic.
linkerd viz top deployment/grpc-server -n grpc-demoConclusion
Kubernetes L4 load balancing cannot distribute gRPC calls evenly because it assigns each long-lived HTTP/2 connection to one pod. Every RPC on that connection remains pinned, which can create hot pods and leave new HPA replicas idle. Kubernetes recommends an L7-aware proxy because the proxy can route individual RPCs across available pods.
Linkerd provides that L7 routing in the pod proxy without requiring changes to gRPC application code. Start by meshing one gRPC client and service. Then scale the service and use Linkerd Viz to confirm that new replicas receive requests and that per-method latency, request rate, and success rate remain visible.
FAQs
Does Linkerd support gRPC-Web and gRPC-Gateway?
Linkerd can proxy gRPC-Web and gRPC-Gateway traffic. gRPC-Web usually reaches Linkerd as HTTP traffic before a component translates it to gRPC. Linkerd can classify the translated backend calls as gRPC, while requests exposed by gRPC-Gateway appear as regular HTTP requests on the client-facing side.
Does enabling Linkerd require changes to my gRPC service definitions or proto files?
Linkerd requires no changes to service definitions, proto files, or generated clients. Its proxy intercepts network traffic outside the application. You inject the proxy into the relevant Kubernetes workloads and configure policy through Kubernetes resources.
What happens to in-flight gRPC streams during a rolling deployment?
Linkerd does not move an active stream to another pod. Kubernetes termination starts connection draining, which gives an existing stream time to finish before the pod exits. Your application must handle graceful shutdown, and the pod termination grace period must accommodate expected stream duration. Clients should reconnect when a stream outlives that period.
Can I use Linkerd gRPC load balancing alongside an existing Ingress controller?
Linkerd can operate behind an existing Ingress controller. The controller accepts external traffic, and Linkerd balances RPCs among meshed backend pods. Meshing the Ingress controller can add Linkerd metrics and encrypted traffic on the internal hop, but compatibility depends on how the controller manages HTTP/2 and TLS.
Does Istio sidecar mode fix gRPC load balancing the same way Linkerd does?
Istio sidecar mode can perform per-RPC gRPC load balancing because Envoy understands HTTP/2 at Layer 7. Istio Ambient without a waypoint provides only Layer 4 balancing, while its waypoint architecture can leave connections pinned to one waypoint replica, as described in Istio issue 56864.
