
Shipping OpenTelemetry From ECS Fargate to Dash0: The Sidecar That Also Routes Logs
Published on Sep 27, 2026
Table of Contents
- The Shape We Landed On
- Sidecar, Not Agentless, Not a Central Collector
- Contrib, Not ADOT
- The Collector Is Its Own Log Router
- Redaction Is a Processor, Not a Promise
- Zero-Code Instrumentation, and Two Things That Bite
- What It Costs
- Scope Tokens by Dataset, Because Permissions Won't
- The Alert We Deleted
- What We'd Change
- FAQ
- The Shape We Landed On
- Sidecar, Not Agentless, Not a Central Collector
- Contrib, Not ADOT
- The Collector Is Its Own Log Router
- Redaction Is a Processor, Not a Promise
- Zero-Code Instrumentation, and Two Things That Bite
- What It Costs
- Scope Tokens by Dataset, Because Permissions Won't
- The Alert We Deleted
- What We'd Change
- FAQ
A Node.js API moving off a PaaS onto ECS Fargate needs somewhere for its telemetry to go. Choosing Dash0 as the backend was the easy part. The decisions that actually took work were all downstream of it: where the collector runs, what it is allowed to see, and what it strips before anything leaves the AWS account.
The task now runs two containers. The application, instrumented without a line of code change, and an OpenTelemetry Collector that is simultaneously the OTLP endpoint, the log router, the redaction point and the CloudWatch writer. That last combination is the part worth writing down, because each of the obvious alternatives costs something real.
One constraint forced every decision below: the application writes personal data into its own log lines, and none of it may leave the AWS account. That rules out redacting at the observability vendor, and it rules out any path where a raw log line reaches the internet first.
The Shape We Landed On
Traces and metrics go over OTLP to localhost. Logs go over the FireLens socket. Both meet in the same processor pipeline, and only then does anything fan out.
ECS task (private subnet, eu-central-1)
api ── traces + metrics ── OTLP/HTTP 127.0.0.1:4318 ──┐
└─ stdout/stderr ──── FireLens socket ───────────┤
ECS task metadata ── awsecscontainermetrics ─────────┤
▼
otel collector (contrib, essential, FireLens log router)
memory_limiter → filter → transform/redact → resourcedetection → batch
├─► Dash0 OTLP (TLS · Bearer ingest token · dataset header)
└─► CloudWatch Logs (redacted copy)Egress runs task → NAT gateway → internet → backend, over TLS, Frankfurt to Ireland. Both endpoints sit inside the EU, which is what the data processing agreement required. PrivateLink to the backend exists and we did not use it: an interface endpoint is a paid, always-on cost, and the environment emits a few megabytes a day.
Sidecar, Not Agentless, Not a Central Collector
There are four ways to get telemetry out of a Fargate task. Three of them fail the constraint above or cost more than they return.
The deciding factor is that logs and traces converge on one pipeline. A regex that strips an email address from a log body and an OTTL statement that deletes a query string from url.full live in the same config file, get reviewed in the same pull request, and are proven by the same probe. Split them across two processes and they drift.
Contrib, Not ADOT
AWS ships its own collector distribution, ADOT. On an AWS-native stack it is the natural default: mirrored in public ECR, supported, and already familiar to anyone who has used the ECS integration. We did not use it, for one blunt reason.
ADOT does not register the transform processor, which is where the OTTL redaction statements live, and it has no fluentforward receiver, which is how the task's stdout arrives. Both are load-bearing here, so the choice made itself: otelcol-contrib.
Check this against your own required components before assuming a distribution has them. The failure mode is not a helpful error at build time — it is the collector refusing to start on an unknown config key, in a container you cannot easily shell into.
Contrib has no public ECR mirror, so the image pulls from GHCR through the NAT gateway, and it is pinned by index digest rather than by tag. A moving tag on the component that redacts personal data is not a risk worth taking for the convenience.
The Collector Is Its Own Log Router
Traces and metrics are simple: the SDK exports to 127.0.0.1:4318 and the collector receives them. Logs are the awkward signal, because the application writes to stdout and ECS hands stdout to a log driver, not to your pipeline.
The documented answer is FireLens with a Fluent Bit sidecar and two outputs. That means a third container, a Fluent Bit config that has to live somewhere — an S3 bucket or a custom image — and redaction rules written twice in two different syntaxes.
Instead, point FireLens at the collector container itself. ECS does not verify that the log router is really Fluent Bit; it just mounts a unix socket into every other container and sets their log driver to write to it.
{
"name": "otel-collector",
"essential": true,
"user": "0",
"firelensConfiguration": { "type": "fluentbit" },
"command": ["--config=env:OTELCOL_CONFIG"]
}
# and in the collector config
receivers:
fluentforward:
endpoint: unix:///var/run/fluent.sock
otlp:
protocols: { http: { endpoint: 127.0.0.1:4318 } }Three details make it work. firelensConfiguration.type must be fluentbit even though it is not; the fluentforward receiver speaks the same wire protocol; and the container needs user: "0" to bind the socket. This is an established pattern rather than a discovery — several vendor sidecars ship it — but it is not in the AWS documentation, so most teams reach for Fluent Bit by default.
The collector is marked essential. That is the honest trade-off of this design: the CloudWatch copy now depends on the collector staying alive, where before it was the platform's job. Marking it essential converts the failure from silent log loss into a task replacement, which is loud and visible.
We kept the CloudWatch copy deliberately. It is written by the awscloudwatchlogs exporter with raw_log: true, after redaction, so both destinations hold byte-identical scrubbed lines. One is the working surface, the other is the AWS-native audit trail nobody has to hold a vendor login to read.
Redaction Is a Processor, Not a Promise
"We'll be careful about what we log" is not a control. A transform processor with OTTL statements is, because it runs on every record whether or not the developer who wrote the log line knew the rules existed.
processors:
transform/redact:
log_statements:
- context: log
statements:
- replace_pattern(body, "\\| user - .*$", "| user - <redacted>")
- replace_pattern(body, "\\?[^ ]*", "?<redacted>")
- replace_pattern(body, "[\\w.+-]+@[\\w-]+\\.[\\w.]+", "<email>")
- replace_pattern(body, "eyJ[A-Za-z0-9._-]+", "<jwt>")
- replace_pattern(body,
"(?i)(code|otp|password|token|secret)\"?\\s*[:=]\\s*\\S+",
"$$1: <redacted>")
- context: span
statements:
- delete_key(attributes, "db.statement")
- delete_key(attributes, "enduser.id")
- replace_pattern(attributes["url.full"], "\\?.*$", "?<redacted>")Some of those are guards for attributes we do not currently emit — db.statement only appears if someone turns on enhanced database reporting. A guard costs nothing and survives the upgrade that quietly enables it.
One rule is deliberately lossy. The validation library dumps the entire rejected request body into the log as a multi-line object, so the rule drops every continuation line of an object dump. It also removes harmless multi-line output, and we accepted that rather than try to parse the shape of a prose log.
The rules are tested the only way worth trusting. Send a request whose query string carries a unique marker, request a sign-in code, submit an invalid body containing a second marker — then search the dataset for every marker, the address, the subject line and eyJ. Zero hits, in logs and in spans, or the rule is not done.
Write the probe before the rule. A redaction regex that has never been shown to fire against a live request is indistinguishable from a comment saying you intended to write one.
Zero-Code Instrumentation, and Two Things That Bite
The application was not modified. Instrumentation is entirely environment variables on the task definition.
NODE_OPTIONS="--experimental-loader=@opentelemetry/instrumentation/hook.mjs \
--import @opentelemetry/auto-instrumentations-node/register"
OTEL_NODE_ENABLED_INSTRUMENTATIONS="http,express,mongoose,mongodb,aws-sdk,undici"
OTEL_NODE_RESOURCE_DETECTORS="env,host,os,process"
OTEL_EXPORTER_OTLP_ENDPOINT="http://127.0.0.1:4318"The first bite is ESM. On a CommonJS app, --import alone is enough. On an ESM app it silently misses express — you get HTTP server spans with no route attribute and no handler spans underneath, which looks like working instrumentation until you try to group by endpoint. The loader hook is the documented ESM path and it is not optional.
Pinning the resource detectors matters too. The default set probes GCP and Azure metadata endpoints on every start, which on ECS is a guaranteed wasted round trip to a host that does not exist. The ECS resource attributes come from the collector's own resourcedetection processor instead, where they belong.
The second bite is the one nobody warns you about. The auto-instrumentation entrypoint installs its own SIGTERM handler so it can flush spans. Once a handler exists, Node no longer exits on SIGTERM by default — so an application that used to stop instantly now waits out the full 30-second ECS grace period before SIGKILL, on every single deploy.
The stopgap is stopTimeout: 10 on the application container, which caps the wait. The real fix is a graceful shutdown in the application that drains connections and then exits, and it belongs on the hardening list rather than in the cutover.
What It Costs
Three processes in a task that used to run one. The staging task went from 256 CPU / 512 MB to 512 / 1024, roughly €10 a month. That is the entire infrastructure cost of the design, and it is the number to quote when someone asks whether a sidecar per task is wasteful.
Retention is 30 days for spans and logs and 13 months for metrics on the vendor side, with the CloudWatch copy raised to 30 days to match. Sampling is 100% with no sampler configured, which is correct while volume is small and is the first thing to revisit against measured production traffic rather than a guess.
The collector config ships as --config=env:OTELCOL_CONFIG, rendered by Terraform straight into the task definition. Nothing to host, no parameter store entry, no bucket, no custom image. A config change becomes a new task definition revision, which means the pipeline that reviews infrastructure changes also reviews the redaction rules.
Scope Tokens by Dataset, Because Permissions Won't
Three tokens, each doing one job: an ingest token for the collector, a read-only token for the assistant's MCP connection, and a full-access token used by Terraform at apply time. The permission model offers exactly three levels — All, Ingesting, Reading — with no configure-only tier, so a token that manages alert rules is necessarily an All token.
Where the permission model stops, the dataset scope continues. Every token is locked to a single environment's dataset, so the worst case for a leaked staging Terraform token is a staging blast radius. Scope on the axis the vendor gives you rather than the axis you wish it gave you.
The ingest token lives in Secrets Manager and is injected into the collector container and one Lambda. It is deliberately not given to the application container, which holds no observability credential at all — it exports plaintext OTLP to localhost and knows nothing about the backend.
The Alert We Deleted
We shipped a rule that fired when no spans arrived from the service for fifteen minutes. It sounds like exactly the alert a new observability stack should have: it catches the collector dying, the token expiring, the exporter misconfigured.
It fired constantly, and it was right to. Health-check spans are filtered out as noise at the collector, so on a staging environment with no traffic overnight there genuinely are no spans. The rule was measuring traffic and reporting it as a telemetry outage.
We removed it rather than tuning the window. The signal it was reaching for is already covered from outside: a synthetic check hits the health endpoint every minute, and because the collector is essential, a dead collector replaces the task and shows up there. An alert that fires nightly on a healthy system trains people to ignore the channel, which costs more than the gap it covers.
The latency rule needed its own carve-out. Two endpoints call a language model and legitimately take 43 to 73 seconds, so a p95 threshold across all operations is meaningless. They are excluded by name in the expression, with the exclusion written down next to the reason.
What We'd Change
Structured logging in the application. Every awkward part of the redaction config exists because the rules are regexes over English prose. A JSON logger turns them into field-level deletions, and the lossy object-dump rule stops needing to exist at all. It is the highest-value follow-up on the list.
Task restarts also deserve a direct signal. Today they are inferred from synthetic-check failures, which works but is indirect; an ECS task state event feed into the same pipeline would close it.
None of this is exotic. It is one extra container, a config file under review, and a refusal to let raw log lines leave the account — and the whole design falls out of taking that last requirement literally rather than treating it as a preference.
FAQ
A sidecar for a handful of services, a central collector once you are running many. The sidecar costs roughly €10 a month in extra task size and gives each task an in-process OTLP endpoint with no network hop and no shared failure domain. A central collector amortises better at fleet scale, but it becomes a service you have to run, scale and page on.
Only because of the components we needed. ADOT does not register the transform processor, which is where OTTL redaction statements run, and it has no fluentforward receiver for FireLens logs. If your pipeline needs neither, ADOT is the easier choice — it is mirrored in public ECR and supported by AWS. Check your required components against the distribution before you commit.
Yes. Set firelensConfiguration.type to fluentbit on the collector container, enable the fluentforward receiver on unix:///var/run/fluent.sock, and run the container as user 0 so it can bind the socket. ECS does not verify that the router is really Fluent Bit, and the wire protocol is the same. This removes the separate Fluent Bit container and its config hosting entirely.
In the collector, inside your own account, before any exporter runs. Vendor-side filtering means the raw data crossed the internet and was stored before it was removed. Doing it in a transform processor also means one set of rules covers logs, spans and any duplicate copy you write to CloudWatch.
The auto-instrumentation entrypoint registers a SIGTERM handler to flush telemetry, and Node stops exiting on SIGTERM once any handler is installed. The container then waits the full ECS grace period before SIGKILL on every deploy. Cap it with stopTimeout on the container and add a real graceful shutdown in the application.
Only if the service actually emits telemetry continuously. If you filter health-check spans as noise, a quiet environment produces no spans and the rule fires on healthy systems every night. Prefer an external synthetic check plus an essential collector container, so a dead collector replaces the task and surfaces as a real failure.
Usually not at low volume. A NAT path over TLS satisfies most data-residency requirements as long as both regions sit in the right jurisdiction. An interface endpoint is an always-on hourly charge plus per-GB, which rarely pays for itself until egress volume is substantial or a policy forbids internet egress outright.

We design and build OpenTelemetry pipelines on AWS that redact before egress, cost what you expect, and alert on things that are actually broken.
CEO & Founder @ Perfsys | Serverless architect with 10+ years of hands-on experience designing cloud-native architectures on AWS, backed by multiple AWS certifications. His writing bridges deep technical expertise with real-world business strategy, covering topics from AWS best practices to scaling tech-driven organizations.
AWS Experts, On-Demand
Need to move fast? Our cloud team is ready to scale, secure, and optimize your systems. Get serverless expertise, 24/7 support, and seamless CI/CD pipelines when you need it most.
Please accept cookies to load the booking widget.
