Skip to main content

Get Service Mesh Certified with Buoyant.

Enroll now!
close

Case Studies

How Censys Cuts GCP Cross-Zone Costs by $283K a Year with Buoyant Enterprise for Linkerd

The enterprise architect's guide to the service mesh

Download ebook

Linkerd Production Readiness Pre-Launch Checklist

Download Checklist

Lean team running mission-critical infrastructure

Censys is a cybersecurity company whose engineering organization runs primarily on Google Cloud Platform (GCP), with a significant and growing presence on its own colocation hardware. At this time around 50 engineers support Censys' current infrastructure and networking architecture.  Censys' SRE team has stayed lean for most of its history and currently is made of around 8 people. This lean team has built their internal model around self-service by building tooling and paved roads that let developers own as much of their services as possible, with minimal ongoing operational overhead on the platform team to maintain.

Their architecture is designed with hub-and-spoke topology with discrete networking for each major environment and runs ~470 microservices in production. At the center of that footprint is the data cluster that carries mission-critical traffic and generates significantly more traffic than any other cluster Censys runs. 

‍

The need for a service mesh and evaluating Linkerd

As Censys designed their current infrastructure to replace an older networking setup, they uncovered a costly pattern: 

  1. Redundant ingress overhead: Services were defining ingress resources even for purely internal, service-to-service traffic, and handling their own TLS termination. Combined with load balancers sitting in front of that ingress, this generated a large volume of unnecessary cross-zone hops. Censys was, in effect, getting billed twice for the same internal traffic.
  2. Expensive multi-cluster egress: : The problem was compounded by how Censys' clusters relate to one another. Their data cluster acts as a source of truth for multiple other clusters, so production reaches into data, and data sends heavy traffic back out to production and beyond. Censys needed a solution that could cut cross-zone egress costs consistently not just within a single cluster, but across clusters. 

‍

To address the technical needs behind this pattern, Censys decided to add a service mesh to their infrastructure; however, their evaluation of different service meshes wasn't a straight line. The team had briefly run Kuma in its legacy environment years earlier, and had a small, experimental Linkerd footprint in a few spots roughly three years before rollout. "More just for fun than anything," as Rob Northover, their Staff Site Reliability Engineer put it.

When the team seriously evaluated options for their newer architecture revamp, the team  looked at the enterprise distribution of Linkerd, Istio and at GCP's native service mesh offering. They found that Istio could not be made to work reliably across clusters the way Censys needed, and the learning curve and operational overhead were steep relative to the payoff. GCP's native mesh would also not work for the team and during evaluation,  it broke one of Censys' clusters so badly that the cluster had to be torn down and rebuilt which left the team frustrated and doubting its capabilities.

Once they learned about High Availability Zone Load Balancing (HAZL), available within Buoyant Enterprise for Linkerd (BEL), they immediately realized its potential to address this ever-growing cost problem. Additionally, because they had been using Linkerd previously, they knew deployment would be fast  and they wouldn’t need a large amount of overhead in order to maintain. 

“When we took a more in-depth look at BEL and saw how the HAZL capability worked, we knew it was what we need.” - Matt Farmer, Site Reliability Engineer, Censys

Deploying Linkerd

BEL's first real deployment landed on the data cluster while it was still a non-production environment because its egress bill was already climbing fast enough to justify meshing it early. That rollout happened somewhat organically, as the team tuned HAZL’s thresholds and deployment practices in real time. 

Rob drove the BEL rollout largely on his own from an infrastructure standpoint. Since most Censys services were relying on ingress resources and doing their own TLS termination even for purely internal traffic, the SRE team needed some additional coordination with the developers. Moving to the mesh meant reconfiguring services to talk to mesh endpoints instead of ingress, and turning off manual TLS termination in favor of Linkerd's built-in mTLS. Most services already had the flags to support this, though naming conventions weren't always consistent, and a handful needed minor changes.

$283,000 cost savings and other impactful results

The impact of BEL has been huge for Censys’ business. Using the HAZL feature, Censys now keeps roughly 85–90% of data cluster traffic within the same zone, versus its previous  baseline of about 33% On a conservative estimate, that represents an annual savings of $283,000 in cross-zone egress costs on GCP, without adding any operational complexity to their infrastructure. , . 

“The savings speak for themselves. I don't know that there was another solution out there — at least none that I found — that could have achieved that for us." - Robert Northover, Staff Site Reliability Engineer, Censys

While the savings by themselves have validated their decision to adopt BEL, it also unlocked other technical needs both current and to support their roadmap. 

Multi-cluster expansion

BEL gave Censys a way to handle certificate management and endpoint exposure across clusters that would otherwise have required tedious, suboptimal workarounds. It also opened the door to additional capabilities, like federated services and seamless cross-cluster traffic splitting, that Censys expects to be  valuable as it considers expanding to additional regions.

Improved observability

BELs advanced multicluster traffic telemetry, powered by Buoyant Cloud, has been valuable during incidents, and lower-traffic environments get rich, granular metrics. Today, it is giving Censys visibility into the data cluster through its dashboard where local tooling can't keep up. The team also noted that some of BEL's built-in observability has been useful for surfacing pre-existing issues elsewhere in the stack. For example, a case where all of Censys' telemetry was funneling through a single ingress pod despite multiple gateways being available.

Cleaner gRPC load balancing

Censys' services are almost entirely gRPC-based. Before BEL, teams had built their own load-shedding and traffic-distribution logic to work around GCP's load balancers. BEL's equal-weight strategy handles that traffic distribution automatically, letting many of those workarounds be retired.

A foundation for the future

Looking forward to upcoming projects, Censys isn't yet leaning heavily on Linkerd for compliance, but expects to soon. With service-to-service mTLS already available out of the box, the team is positioned to meet upcoming security and compliance requests, including authentication between services and tighter sandboxing of sensitive components, such as core account-management services. 

In the meantime, the Censys team will continue to mesh additional services and continue to see additional cost savings with HAZL. To learn more about how you can save on your GCP or AWS bill, contact us to learn more. 

Interested in what Buoyant Enterprise for Linkerd can do for your team? Contact us to learn more.

‍