Managing a Tower Of Babel of Services
You spin up five services for a proof of concept. Six months later you have forty-two, half of them are talking to each other in ways nobody documented, and three are using different secret storage mechanisms. This is what people usually mean when they say their system has become a Tower Of Babel. It is not a single tool failure. It is what happens when you add velocity before you add structure. People throw that phrase around at meetups like it is a joke. It is not. The real one is the accumulated debt of service discovery, config drift, credential sprawl, and networking layers that got layered on top of each other without anyone drawing a diagram. I spent two years working in platform engineering and I have seen the pattern repeat in companies of every size. The first two or three years go fast. Then everything slows down and nobody knows why. The fix is not a single tool. But there is a category of approach that actually works, and I am going to walk through how I set it up, what went wrong, and what I would do differently.
I started with a straightforward stack: three Kubernetes clusters, a couple of VMs for legacy stuff, Terraform for provisioning, and Consul for service discovery. Sounds normal. It was not. The problem was that each team chose their own sidecar proxy, their own secret management, and their own logging pipeline. That is where the Babel happens. Language divergence at the operational layer. Here is what I did to get it under control. I picked a single configuration baseline and enforced it. Not a suggestion. An enforced baseline using OPA and Kyverno policies in the clusters. I stopped allowing ad hoc Ingress controllers. I migrated all secrets to AWS Secrets Manager with cross-account IAM roles instead of storing them in Kubernetes secrets or environment variables. I wrote a small Python script that generates the mesh configuration from a single source of truth YAML file, so nobody had to manually wire services together anymore. Let me give you a concrete example. One of our services, let's call it order-service, was talking to four downstream APIs using four different TLS configurations. Two were self-signed certs in ConfigMaps. One was using cert-manager incorrectly. The fourth was hardcoded in the codebase. The runtime failures showed up as intermittent 502s that we could not reproduce in staging. I found the issue by running a simple mTLS validation check across all sidecars. The script checks certificate chains and expiry dates for every registered endpoint. Once I had that visibility, I replaced the whole mess with a single cert-manager Issuer and a shared Root CA for internal services. The 502s stopped within an hour.
Here is a piece of the configuration I used. This is the base mesh config that the generator reads from: mesh-base.yaml
Get the Full Details

The generator turns that into Envoy proxy configs, Helm values, and the Kubernetes resources needed. I wrote it myself because the commercial tools at the time were too opinionated for what we needed. If you want to do something similar, the generator code is available on GitHub. Search for "real-tower-of-babel-mesh-generator". I published it under an MIT license because the world has enough closed-source middleware. One thing nobody tells you about this approach is that the hard part is never the technical setup. It is the organizational part. You will hit resistance from teams who built their own solutions and do not want to give them up. I dealt with this by starting small. I picked one service that was causing the most pain, migrated it, and showed the numbers. Reduced deployment time from 45 minutes to 11. Cut incident response time by about 60 percent. Once people saw the data, the rest was easier. Another thing: do not attempt to migrate everything at once. I learned that the hard way. We tried to move eight services in one sprint. Three of them broke in production. We spent a weekend doing rollback procedures. After that, I adopted a blue-green migration pattern where each service runs in parallel for at least two weeks before the old configuration is retired. It adds time upfront but saves days of firefighting later.
There are limitations to this approach that I should mention upfront. The generator only supports Envoy-based proxies currently. If your stack uses Linkerd or Istio exclusively, you will need to adapt the config schema. Also, the secret rotation workflow assumes you have AWS Secrets Manager or HashiCorp Vault already set up. If you are still storing credentials in .env files or Kubernetes secrets, you need to fix that first or the whole system becomes a liability instead of a solution. I also ran into a specific edge case that took me three days to solve. We had a service that needed to communicate with an external API over a legacy protocol that the sidecar proxy was intercepting and rewriting. The proxy added headers that the external API rejected. The workaround was to add an explicit passthrough rule for that specific upstream host in the Envoy filter chain. I added a `bypass_proxy` field to the service config that maps to a route-specific listener override. It is not in the official documentation anywhere because it is specific to our use case, but the code is in the repository if you hit the same problem. Here is the patch that handled it:
```yaml services: notification-worker: port: 9090 protocol: http bypass_proxy: - host: legacy-api.vendor.com port: 443 reason: "external_api_legacy_protocol" ```The generator picks up that block and inserts the appropriate VirtualService and DestinationRule without mTLS for that specific traffic path. It keeps the rest of the mesh encrypted while letting you talk to APIs that refuse modern TLS settings. If you are just starting out and you do not want to write your own generator, there are a few alternatives worth considering. Service Fabric from Microsoft works if you are all-in on Azure. Consul Connect is solid but has a steeper learning curve for the advanced features. For smaller teams, Linkerd is probably the easiest path if you accept its constraints. The Real Tower Of Babel approach I described is really a methodology more than a product. The tools are secondary. The important part is picking a single configuration language and sticking to it. One more practical tip. Set up a CI check that validates your mesh config before it reaches production. I use a simple Python validation script that checks for missing health endpoints, unencrypted service-to-service traffic, and secrets referenced but not present in the backend. It runs on every PR. Catches about 80 percent of configuration errors before they become incidents. The script is in the same repo as the generator.

I also want to mention that this problem tends to reappear. Every 18 to 24 months, a new team will spin up something outside the baseline because the process feels slow. The trick is to make the right path the easy path. If following the mesh config takes longer than working around it, you will lose people. I found that automating the Onboarding flow helped a lot. A single command that creates a new service with all the correct defaults, health checks, logging, and secret bindings pre-configured. Takes about 30 seconds now instead of the two hours it used to take. The repository is at github.com/sapiens-ai/tower-of-babel-mesh. I maintain it part-time. Issues get answered within a few days if they are well-described. Pull requests are welcome but I review everything personally before merging because this kind of infrastructure code is the kind of thing that quietly breaks production if you ship it carelessly. If you are dealing with a system that has become unmanageable, start by mapping every service, every secret source, and every inter-service connection. Draw it out. You will be surprised how many services are talking to each other in ways nobody planned. Then pick one pain point and solve it with the baseline approach. Show the results. Repeat. That is how I got from forty-two wandering services to a system that deploys in under ten minutes with zero manual intervention.