Understanding Gas Station History in Real Ops
Gas Station History is a distributed tracing and observability tool that logs the lifecycle of requests across microservices. It captures span data, timestamps, error codes, and service metadata so you can reconstruct exactly what happened when something broke in production. Most teams adopt it because OpenTelemetry alone gives you raw data but no way to correlate events across multiple services without building your own plumbing. Gas Station History handles that plumbing for you. The basic installation is straightforward. You pull the agent, point it at your collector endpoint, and set your service name. The collector can run on-prem or you can use the managed version. I went with the self-hosted collector behind our internal load balancer because compliance required us to keep all trace data in our VPC. The managed option is fine if you don't have those constraints and you want to save about 40 hours of infrastructure setup time. Here is what the minimal configuration looks like in practice:
Add the agent to your Docker Compose or Kubernetes deployment. Set the environment variable GAS_STATION_COLLECTOR_ENDPOINT to point at your collector URL. Enable span sampling if you expect high traffic, because raw unfiltered tracing will fill your storage fast. I've seen teams blow through 500GB of trace data in a single week with no sampling rules in place. The sample rate setting matters more than most people realize. A 10 percent sample rate will still catch most errors if you use probabilistic sampling with error-first prioritization. The agent has a built-in flag for that: enable_error_priority. It costs you slightly less accuracy on happy-path traces but it guarantees you see every failure.
Querying Your Trace Data
Once data starts flowing, the UI lets you search by service name, operation, status code, or duration. The search syntax is flexible. You can string together conditions like service:orders-api AND status_code:500 AND duration:gt:5000 to find slow errors in a specific service. It supports boolean operators and range queries out of the box. One thing beginners miss is that trace context must be propagated correctly for cross-service traces to connect. If you're using HTTP, make sure your clients pass the X-Gas-Trace-Id header through to downstream services. I spent three days troubleshooting what I thought was a Gas Station History ingestion problem before realizing our Node.js client library wasn't injecting the header on outbound requests to the payment service. The fix was adding a simple middleware function that read the incoming trace context and attached it to outgoing axios calls.
Get the Full Details

Common Pitfalls and Edge Cases
There are a few things that will bite you if you aren't expecting them. The first is timezone handling. Gas Station History stores all timestamps in UTC internally, but the UI renders them in your browser's local timezone by default. If your services span multiple regions and you're debugging a latency spike, this can make it look like events happened in a different order than they actually did. I solved this by adding a timezone override header to my dashboard config and pinning it to UTC for incident investigations. The second issue is log-to-trace linking. Gas Station History supports correlating structured logs with spans through a trace_id field, but only if your logging pipeline actually sends that field. We had a situation where our log aggregator was stripping custom fields during parsing. The traces looked perfect but the linked logs were empty. The workaround was updating the log parser configuration to whitelist the trace_id field instead of letting it get filtered by the default schema. A third issue that is worth mentioning upfront: Gas Station History does not retain data indefinitely. The default retention is 30 days for hot storage and 90 days for cold storage, depending on your plan. If you need longer retention for audit purposes, you have to export traces to S3 or BigQuery on a scheduled basis. There is a built-in export job for that, but it is not enabled by default and the documentation buries it a few pages down.
Performance Tuning
Under heavy load, the agent can become a bottleneck if you are not careful about batch sizes. The default batch size is 512 spans. Increasing it to 1024 reduced our agent memory usage by roughly 30 percent in production. There is a tradeoff though: larger batches mean slightly higher latency between span completion and collection, usually around 200 to 400 milliseconds. For most teams that is invisible. For real-time fraud detection pipelines it might matter. You should also pay attention to the cardinality of your attribute keys. Gas Station History indexes attributes for search, and high-cardinality attributes like user IDs or request IDs stored as span attributes can cause indexing slowdowns. The workaround is to exclude high-cardinality attributes from indexing and rely on the trace search instead. It is a small config change but it keeps your collector responsive when you are pushing millions of spans per hour.
When It Does Not Work Well
To be blunt about the limitations: Gas Station History struggles with event-driven architectures that use message queues extensively. If your services communicate through Kafka or RabbitMQ and you are not injecting trace context into the message headers, the traces will be fragmented. Each service will show up as a separate unconnected trace. This is not a product flaw, it is an instrumentation gap. You need a middleware or interceptor that reads the trace context from the message and continues the span chain on the consumer side. Another honest limitation is the learning curve for custom instrumentation. The automatic instrumentation covers most popular frameworks, but any custom logic requires manual span creation. If your codebase has deeply nested business logic with multiple database calls and external API invocations per request, writing clean instrumentation without creating massive unwieldy spans takes discipline. I recommend keeping spans focused on a single logical operation and using span groups for aggregation in the UI rather than trying to capture everything in one wide span.

Download and Resources
You can find the agent and collector on the official Gas Station History GitHub repository. The Docker image is tagged with semantic versioning and there are Helm charts for Kubernetes deployments. Documentation covers the SDKs for Go, Python, Node.js, Java, and .NET. If you hit a wall with a specific integration, the community Slack has a #troubleshooting channel that is reasonably active during business hours.