Back in July, we encountered a tricky observability bug while running an asynchronous compute pipeline inside AWS Lambda processing data from Amazon S3.

We had instrumented our workload with OpenTelemetry (OTel) to gain granular insights into execution bottlenecks, network latency, and compute time. But when we opened our tracing dashboard, instead of a clean, cohesive flamegraph showing our execution flow, we were greeted with total chaos: span explosion, orphaned traces, and completely disconnected parent-child spans.

Here is what went wrong, why our concurrency model broke OpenTelemetry’s default behavior, and how we solved it using distributed tracing context propagation.


1. Context & Architecture

Our serverless pipeline processes large objects retrieved from S3:

  1. Lambda Invocation: An S3 object notification triggers the Lambda handler on the main runtime thread.
  2. Concurrent Compute: Because each S3 payload contains hundreds of independent tasks, the handler dispatches these compute tasks across a pool of background worker threads to leverage multi-core Lambda compute efficiently.
  3. Task Completion & Sinks: Each worker performs its computation, writes results, and signals completion.
[S3 Event] ──> [Lambda Handler (Main Thread)]
                     │
                     ├── Dispatch ──> [Worker Thread 1: Process Task A]
                     ├── Dispatch ──> [Worker Thread 2: Process Task B]
                     └── Dispatch ──> [Worker Thread 3: Process Task C]

To monitor performance, the Lambda handler initiated a root trace (lambda-s3-compute), and each worker task was supposed to create a child span (process-task) linked directly to that parent.


2. The Symptoms: The “Broken Trace” Mystery

When inspecting traces in our telemetry backend (Jaeger / Honeycomb), two bizarre problems showed up:

Issue A: Disconnected Orphan Traces

Instead of seeing one single trace ID containing the root handler span and multiple nested child spans, we saw dozens of separate root traces:

WHAT WE EXPECTED:
└── [Trace ID: 4bf92f3577b34da6a3ce929d0e0e4736]
    └── lambda-s3-compute (Parent Span)
        ├── process-task (Worker 1)
        ├── process-task (Worker 2)
        └── process-task (Worker 3)

WHAT ACTUALLY HAPPENED:
├── [Trace ID: 4bf92f3577b34da6] -> lambda-s3-compute (Orphaned Parent)
├── [Trace ID: 8a1b92c431d87e01] -> process-task (Worker 1 as a new Root!)
├── [Trace ID: 5f2a1b9c84e12019] -> process-task (Worker 2 as a new Root!)
└── [Trace ID: 3e9d81a4b5c77204] -> process-task (Worker 3 as a new Root!)

Issue B: Span Explosion & Redundant Telemetry

Because child tasks were not tethered to their parent, tasks that spawned sub-tasks or HTTP requests generated runaway spans with distinct IDs. Telemetry ingestion costs spiked, and searching for the trace of a single S3 file execution was virtually impossible because none of the spans shared a common trace context.


3. The Investigation & Root Cause

Where was the context getting lost?

In the main Lambda handler, the span was active:

// Main handler thread
Tracer tracer = openTelemetry.getTracer("s3-compute");
Span parentSpan = tracer.spanBuilder("lambda-s3-compute").startSpan();
try (Scope scope = parentSpan.makeCurrent()) {
    // Context is active HERE on the main thread
    dispatchTasksConcurrently(tasks);
} finally {
    parentSpan.end();
}

Inside the worker threads, we were starting child spans like this:

// Worker thread (inside ExecutorService / Runnable)
public void run() {
    // We expected OTel to automatically know the parent!
    Span childSpan = tracer.spanBuilder("process-task").startSpan();
    try {
        processItem();
    } finally {
        childSpan.end();
    }
}

The “Aha!” Moment: Thread Boundaries Break ThreadLocal

By default in most OpenTelemetry SDKs (such as Java and Python), the active trace context is stored in ThreadLocal memory.

  • When parentSpan.makeCurrent() is called, the trace context is bound strictly to the Main Thread.
  • When worker tasks are dispatched to an ExecutorService or worker threads, the child threads do not inherit the parent thread’s ThreadLocal context.
  • When tracer.spanBuilder("process-task").startSpan() executes inside the worker thread, OpenTelemetry inspects Context.current(), finds nothing, and assumes:

    “There is no active span on this thread. I must create a brand-new Trace ID and treat this as a brand-new Root Span!”

This explained why all our worker spans were severed from the parent and generating arbitrary new trace IDs!


4. The Fix: Explicit Context Propagation Across Threads

To fix this, we applied the fundamental rule of distributed tracing: When crossing an execution boundary (network or thread), you must explicitly propagate the Context.

The Solution: Wrapping Runnables with OTel Context

OpenTelemetry provides built-in utilities to capture the active context from the parent thread and attach it to the worker execution.

Before (Broken):

// Parent thread dispatches without context:
executorService.submit(() -> {
    // Context is NULL here!
    doWork();
});

After (Fixed with Context.current().wrap()):

// 1. Capture the current context on the main thread
Context parentContext = Context.current();

for (Task task : tasks) {
    // 2. Wrap the Runnable so it carries the parent context into the worker thread
    Runnable wrappedTask = parentContext.wrap(() -> {
        // 3. Inside worker: Context is now restored!
        Span childSpan = tracer.spanBuilder("process-task")
            .startSpan();
        try (Scope scope = childSpan.makeCurrent()) {
            executeTask(task);
        } finally {
            childSpan.end();
        }
    });

    executorService.submit(wrappedTask);
}

Alternative: Using OpenTelemetry Context in Custom Handlers

If you’re passing state objects between threads, you can also pass the parent Context explicitly:

public void processInWorker(Context parentContext, Task task) {
    // Explicitly bind the parent context to the span builder
    Span childSpan = tracer.spanBuilder("process-task")
        .setParent(parentContext)
        .startSpan();

    try (Scope scope = childSpan.makeCurrent()) {
        execute(task);
    } finally {
        childSpan.end();
    }
}

5. The Outcome

Once we deployed the fix:

  1. Unified Traces: Every S3 invocation produced exactly one Trace ID.
  2. Proper Tree Hierarchy: All worker threads appeared as neat, concurrent parallel child bars under the main handler span in the flamegraph.
  3. Accurate Bottleneck Detection: We could instantly identify which specific S3 task worker was bottlenecked on I/O vs. compute.
  4. Clean Telemetry Bills: Zero runaway root spans or detached noise.

6. Key Takeaways

  1. ThreadLocal Never Crosses Thread Pools Automatically: Whenever you use ExecutorService, asynchronous callbacks, or worker pools, never assume the logging or tracing context will travel with the task.
  2. Use Context.current().wrap(): OpenTelemetry’s Context.wrap() is the cleanest, idiomatic way to pass trace propagation across concurrency boundaries.
  3. Always End Spans in finally Blocks: Ensure span.end() is guaranteed to execute, even on unhandled exceptions, to prevent memory leaks and dangling spans.
  4. Inspect Flamegraphs Early: Don’t just test if telemetry is arriving; verify that the Trace ID and Parent Span ID relationships are mathematically connected.