Just last week, we faced a sudden production failure in one of our core serverless data ingestion pipelines.

The pipeline runs on AWS Lambda with the Python 3.12 runtime. Its responsibility is straightforward: securely fetch an encrypted file from an external source, decrypt it, parse its hierarchical data structures (specifically large Excel spreadsheets with nested object trees), and stream the extracted records to downstream stores.

We had provisioned the Lambda with 1 vCPU of compute power (~1,769 MB of RAM)—what we initially assumed was generous headroom, bordering on “overkill.” Yet, during a batch run, the function abruptly died with an unceremonious Out-of-Memory (OOM) crash, leaving files half-processed and transactions incomplete.

Here is the breakdown of why our memory ballooned, how we isolated the bottleneck, and why switching to disk-backed ephemeral storage resolved the issue.


1. The Architecture & The Pipeline

The ingestion flow handles sensitive files that must be decrypted and validated on the fly:

[External Source / S3] 
       │ (Encrypted Ciphertext)
       ▼
[AWS Lambda (Python 3.12)]
   ├── 1. Fetch KMS / STS Secrets
   ├── 2. Decrypt Payload
   ├── 3. Parse Document Trees (Excel / XML Sheets)
   └── 4. OpenTelemetry Spans & S3 Upload

Before processing, the function executes peripheral tasks:

  • Assumes roles via AWS STS and fetches encryption keys.
  • Instruments spans using OpenTelemetry for end-to-end tracing.
  • Interacts with Amazon S3 using Boto3.

Each of these peripheral tasks consumes negligible CPU and memory (typically under 20–30 MB). So why were we exhausting over 1.7 GB of memory?


2. The Symptoms: Sudden Death in Production

During execution with larger files (30MB–80MB .xlsx spreadsheets), the Lambda invocation did not fail gracefully with an application-level exception. Instead, AWS CloudWatch reported:

REPORT RequestId: 8e4b3c92-11a5-4f7e-bc91-23d91f28014e
Duration: 18420.12 ms   Billed Duration: 18421 ms   Memory Size: 1769 MB   Max Memory Used: 1769 MB
Status: error
Error: Runtime.ExitError (Signal: killed)

The Linux kernel’s Out-Of-Memory killer had sent SIGKILL to the Python process. The file execution was aborted mid-stream, resulting in missing records and inconsistent state downstream.


3. The Investigation & Root Cause

Ruling Out the Red Herrings

We initially checked:

  1. OpenTelemetry Telemetry Buffers: Was the OTel SDK buffering too many spans? No—span batching was flushed frequently and used less than 15 MB.
  2. Boto3 & STS Client Leaks: Were connection pools leaking memory? No—connections were properly reused.

The Real Culprit: In-Memory Buffering + DOM Tree Amplification

Looking closely at the ingestion code, we discovered that the entire pipeline was being operated strictly in RAM:

# The Flawed In-Memory Approach:
response = s3_client.get_object(Bucket=bucket, Key=key)
ciphertext_bytes = response['Body'].read()  # Stored in RAM (e.g. 60 MB)

# In-memory decryption buffer:
decrypted_stream = io.BytesIO()
cipher.decrypt(ciphertext_bytes, decrypted_stream)
plaintext_bytes = decrypted_stream.getvalue()  # Another 60 MB in RAM!

# In-memory Excel parsing:
workbook = openpyxl.load_workbook(io.BytesIO(plaintext_bytes))

Notice what happened here:

  1. Double Buffering: At any given moment, the ciphertext byte string, the decryption output buffer, and the plaintext byte string were coexisting simultaneously in Python’s heap.

  2. The “Excel XML Expansion” Multiplier: An .xlsx file is actually a compressed ZIP archive containing raw XML files. A 50 MB compressed spreadsheet can expand to 400 MB–800 MB of raw XML.

  3. Python Object Overhead: When a library like openpyxl parses an XML document into Python objects (rows, cells, styles, formulas), each cell is represented as a full Python PyObject. In Python, a single integer or string has significant memory overhead compared to raw C primitives.

    A 50 MB file in memory rapidly expanded into 1.8+ GB of active Python heap memory, triggering an instant OOM crash.

[Compressed File: 50 MB] 
       │
       ▼ (Decryption in RAM)
[Plaintext Bytes: 50 MB]
       │
       ▼ (Unzipped XML: 500 MB)
[Python DOM Objects in Heap: > 1.8 GB] ──> 🔥 OOM KILL

4. The Fix: Offloading from RAM to Ephemeral /tmp Storage

Instead of treating Lambda’s expensive memory as a hard drive, we realized that AWS Lambda provides ephemeral disk storage (/tmp)—ranging from 512 MB up to 10 GB of high-speed NVMe storage.

Rather than buffering byte streams in memory, we redesigned the pipeline to stream directly to disk and parse in read-only chunks.

Step 1: Stream Decryption Directly to Disk

We replaced io.BytesIO with tempfile.NamedTemporaryFile backed by Lambda’s /tmp directory:

import tempfile
import openpyxl

def process_encrypted_file(s3_client, bucket, key):
    # Create temporary files on Lambda's local NVMe disk
    with tempfile.NamedTemporaryFile(dir="/tmp", delete=True) as enc_file, \
         tempfile.NamedTemporaryFile(dir="/tmp", delete=True) as dec_file:
        
        # 1. Stream download to disk in chunks (keeps RAM flat)
        s3_client.download_fileobj(bucket, key, enc_file)
        enc_file.flush()
        enc_file.seek(0)

        # 2. Decrypt in streaming chunks from file to file
        decrypt_file_stream(enc_file, dec_file)
        dec_file.flush()
        dec_file.seek(0)

        # 3. Stream-parse without loading the entire DOM into memory
        workbook = openpyxl.load_workbook(
            filename=dec_file.name, 
            read_only=True,     # Stream parser: does NOT load all cells at once
            data_only=True
        )

        for sheet in workbook:
            for row in sheet.iter_rows(values_only=True):
                process_row(row)

        workbook.close()

Step 2: Read-Only Streaming Mode

By passing read_only=True to load_workbook, openpyxl switches from building a massive in-memory DOM tree to an iterative XML parser (iterparse). It reads and yields rows one by one, dropping processed rows from memory immediately.


5. Results & Metrics

Metric In-Memory (Before) Disk-Backed Streaming (After)
Peak Memory Consumption > 1,769 MB (OOM Crash) ~140 MB (Flat & Predictable)
Max Processable File Size ~40 MB 500 MB+
Failure Rate ~18% on large batches 0%
Lambda Memory Sizing Required 2 GB+ (failing) Can safely run at 512 MB (Cost Savings!)

6. Key Takeaways

  1. Don’t Treat RAM as a File System: In serverless environments, RAM is the most expensive resource. If you’re holding full payloads in BytesIO buffers, you’re paying for memory you don’t need.
  2. Beware the Spreadsheet Amplification Factor: Spreadsheets and XML documents are deceptively compressed on disk. In-memory object trees can expand by 10x to 30x the file’s raw size.
  3. Lambda /tmp Storage is Fast NVMe: Lambda’s ephemeral storage is local, fast, and gives you at least 512 MB for free (expandable up to 10 GB). Use it to spool, decrypt, and parse large binary streams.
  4. Use Iterative Stream Parsers: Whenever parsing tabular or structured files, prefer streaming APIs (read_only=True, SAX, or iterparse) over loading the full DOM tree into memory.