Choosing Chunk Sizes for Data Pipelines and What Breaks When You Change Them
For developers tuning batch jobs, streaming consumers, file uploads, or vector indexing, chunk size is one of those parameters that quietly controls throughput, latency, memory use, and failure modes. This guide explains where chunk size sits in real systems, how to pick it deliberately, and what usually breaks when you change it after a system is already in production.
TL;DR — Chunk size is a control knob for how much work a system does per read, write, network call, transaction, or model request. The most likely fix is to stop treating it as a pure performance tweak: pick a chunk size by the actual bottleneck you have now (DB lock time, API payload limit, worker memory, retry cost, embedding quality), then re-test idempotency, timeouts, offsets/checkpoints, and downstream limits before rollout. Reading time: ~7 min
What it is and where it sits
"Chunk size" is not one thing. In practice it means the unit of work at some boundary:
- bytes per file read/write
- rows per database batch insert/update
- messages per consumer poll
- records per queue job
- tokens/characters per LLM or embedding request
- parts per multipart upload
- documents per indexing batch
It sits between producers and consumers. It usually replaces a naive one-item-at-a-time flow, or an equally naive "load everything into memory and send it all" flow.
Typical architecture context:
- application code chooses chunk size
- transport/protocol imposes max payloads or framing behavior
- storage engine imposes transaction, lock, WAL, or page-write costs
- worker runtime imposes memory and timeout limits
- retry/checkpoint logic implicitly assumes some chunk boundary
[Source: file/DB/API]
|
v
[Chunker in app/worker]
|
+--> [serialize/compress]
|
v
[transport: HTTP/gRPC/S3/Kafka]
|
v
[downstream: DB/index/model/object store]
|
v
[checkpoint/offset/retry state]
That last box is where many production incidents come from. Changing chunk size often changes the meaning of "done".
Examples of where it lives in a request/data flow:
- In a Python ETL job,
for batch in batched(rows, 1000): insert(batch). - In a Kafka consumer,
max.poll.records=500and app-level processing of those records before commit. - In an upload service, 8 MiB multipart parts to object storage.
- In an embedding pipeline, splitting a document into 800-token chunks with 100-token overlap.
How it actually works
Take one realistic example: a worker reads rows from Postgres, transforms them, and bulk-indexes them into a search service over HTTP. You change chunk size from 500 rows to 5000 rows because throughput looks low.
Step-by-step flow
- The worker fetches source rows.
- It accumulates rows in memory until the chunk is full.
- It transforms each row into an indexing document.
- It serializes the whole chunk into one NDJSON payload.
- It sends one HTTP request.
- If the request succeeds, it advances a checkpoint: "processed through source id X".
- On failure, it retries the whole chunk or dead-letters it.
That sounds simple. Here is what changed when you moved from 500 to 5000:
- Memory: 10x more rows and transformed docs held at once.
- Network: larger request body; maybe now over reverse proxy or upstream limits.
- Timeout risk: one request now takes longer than client, proxy, or server timeouts.
- Retry blast radius: one transient 502 now replays 5000 rows, not 500.
- Checkpoint semantics: if partial success is possible, your old "advance on 200" logic may now skip failed items.
- Latency: downstream data appears in larger bursts instead of steadily.
A concrete failure chain
Suppose your nginx sits in front of the indexing service with:
client_max_body_size 10m;
proxy_read_timeout 30s;
At 500 rows, each request is ~1.2 MiB and completes in 4s. At 5000 rows, requests are ~14 MiB and take 42s.
You now see this from the worker:
requests.exceptions.HTTPError: 413 Client Error: Request Entity Too Large for url: https://indexer.internal/bulk
Or, if compression keeps it under 10 MiB but processing is slow:
requests.exceptions.ReadTimeout: HTTPSConnectionPool(host='indexer.internal', port=443): Read timed out. (read timeout=30)
From curl -i against nginx, the output shape is typically:
curl -i -X POST https://indexer.internal/bulk --data-binary @payload.ndjson
HTTP/1.1 413 Request Entity Too Large
Server: nginx/1.24.0
Date: Tue, 01 Oct 2026 10:14:22 GMT
Content-Type: text/html
Content-Length: 183
Connection: close
Or for timeout:
HTTP/1.1 504 Gateway Time-out
Server: nginx/1.24.0
Date: Tue, 01 Oct 2026 10:16:03 GMT
Content-Type: text/html
Content-Length: 167
Connection: keep-alive
Then a subtler break appears: your code commits the source checkpoint after the HTTP call returns 200, but the indexing API accepted only 4873 of 5000 docs and returned per-item failures in the body. With larger chunks, partial failure is no longer rare enough to ignore.
That is the core mechanism: chunk size changes not just efficiency, but the failure surface and the semantics of retries and progress tracking.
When to use it (and when not to)
Use chunking when per-item overhead is expensive enough that batching improves throughput, or when the upstream/downstream protocol naturally wants bounded pieces.
You probably do not need to tune chunk size aggressively if your system is nowhere near a limit and your current settings already meet SLOs.
| Scenario | Recommendation |
|---|---|
| Bulk DB inserts are CPU-light but network round-trips dominate | Increase chunk size gradually; watch transaction time and lock duration |
| Consumer lag is high and each message is tiny | Increase records per poll/batch, but only if processing stays under rebalance/visibility timeouts |
| Uploading large files to object storage | Use multipart chunk sizes large enough to reduce request count, small enough to retry cheaply |
| LLM/embedding retrieval quality is poor | Tune semantic chunk size and overlap for meaning preservation, not just throughput |
| API requests already hit body-size or timeout limits | Do not increase chunk size; fix limits or split work |
| Job failures are frequent and retries are expensive | Prefer smaller chunks to reduce replay cost |
| You need near-real-time visibility of each item | Prefer smaller chunks; large batches add tail latency |
| You cannot make writes idempotent | Avoid large chunks; partial success handling becomes painful |
You probably don't need this if:
- your bottleneck is a missing index, not batching
- your workers are idle because of external rate limits
- your end-to-end latency target is stricter than any batching benefit
- your payloads are already near hard protocol limits
- your team cannot safely reason about partial success and replay
Trade-offs
Every benefit has a cost.
- Higher throughput per request/transaction
- Cost: larger memory spikes, longer lock/transaction time, bigger retries.
- Lower per-item overhead
- Cost: worse p95/p99 latency for individual items waiting for the batch to fill.
- Fewer network calls
- Cost: easier to hit
413,429,502,504, ALB/proxy body limits, and idle/read timeouts.
- Cost: easier to hit
- Better compression ratio on larger payloads
- Cost: more CPU per batch and larger decompression/serialization buffers.
- Fewer commits/checkpoints
- Cost: coarser recovery point; more duplicate work after crashes.
- Better GPU/accelerator utilization for ML inference batches
- Cost: queueing delay and more painful tail failures if one item poisons a batch.
- Larger semantic chunks for retrieval
- Cost: worse recall granularity, more irrelevant context, and easier token-limit overflow downstream.
Operationally, changing chunk size is not free because dashboards and alerts often encode assumptions about old behavior. A queue consumer that used to process 100 msgs/sec in 100 small commits may now process the same volume in 10 bursty commits/sec; lag graphs, DB write IOPS, and timeout alerts all change shape.
In practice
Example 1: Python batch worker with explicit size, timeout, and partial-failure handling
import itertools
import json
import requests
BATCH_SIZE = 500
TIMEOUT_SECONDS = 15
def batched(iterable, n):
it = iter(iterable)
while True:
batch = list(itertools.islice(it, n))
if not batch:
return
yield batch
def send_batch(rows):
docs = [{"id": r["id"], "title": r["title"], "body": r["body"]} for r in rows]
payload = "\n".join(json.dumps(d) for d in docs) + "\n"
resp = requests.post(
"https://indexer.internal/bulk",
data=payload.encode("utf-8"),
headers={"Content-Type": "application/x-ndjson"},
timeout=TIMEOUT_SECONDS,
)
resp.raise_for_status()
result = resp.json()
failed_ids = [item["id"] for item in result["items"] if item["status"] >= 400]
return failed_ids
for rows in batched(source_rows(), BATCH_SIZE):
failed = send_batch(rows)
succeeded = [r for r in rows if r["id"] not in set(failed)]
checkpoint_up_to(max(r["id"] for r in succeeded))
if failed:
enqueue_for_retry([r for r in rows if r["id"] in set(failed)])
This batches records for one bulk HTTP call and advances progress only for succeeded items. The gotcha: if your checkpoint is a single monotonic offset, partial success breaks simple "checkpoint max id" logic unless retries preserve ordering or you track holes explicitly.
Example 2: nginx limits that often become visible after increasing chunk size
server {
listen 443 ssl;
server_name indexer.internal;
client_max_body_size 20m;
location /bulk {
proxy_pass http://indexer_backend;
proxy_connect_timeout 5s;
proxy_send_timeout 30s;
proxy_read_timeout 60s;
}
}
This raises request body and upstream timeout limits for larger bulk requests. The gotcha: increasing these values can mask a bad chunk size choice; if retries now replay 20 MiB payloads, your failure recovery gets slower and more expensive even if the 413s disappear.
Example 3: Kafka consumer settings where batch size interacts with liveness
max.poll.records=1000
max.poll.interval.ms=300000
fetch.min.bytes=1048576
fetch.max.bytes=52428800
session.timeout.ms=45000
enable.auto.commit=false
This lets a consumer receive larger batches and wait for more data before fetch returns. The gotcha: if processing 1000 records can exceed max.poll.interval.ms, the group coordinator will consider the consumer stuck and trigger a rebalance; you will see duplicate processing unless commits are carefully placed.
Typical log shape when this goes wrong:
WARN org.apache.kafka.clients.consumer.internals.AbstractCoordinator - Auto offset commit failed: Commit cannot be completed since the group has already rebalanced and assigned the partitions to another member.
INFO org.apache.kafka.clients.consumer.internals.ConsumerCoordinator - Revoke previously assigned partitions
A practical tuning loop
Do this in order:
- Measure current median and p95 for batch payload bytes, processing time, memory, and retry rate.
- Change one dimension at a time: e.g. 500 -> 1000, not 500 -> 10000.
- Run a representative load test with real payload distributions, not tiny fixtures.
- Inspect hard limits explicitly:
- nginx
client_max_body_size - app client timeouts
- DB statement/lock timeout
- queue visibility timeout / poll interval
- API documented max items or max bytes
- model context/token limits
- nginx
- Verify idempotency and partial-failure behavior.
- Roll out gradually and watch:
- p95 batch duration
- RSS / heap growth
- 4xx/5xx by endpoint
- duplicate processing rate
- checkpoint/offset lag
⚠️ If your system advances offsets, checkpoints, or deletes source data after processing, changing chunk size can cause silent data loss or duplication. Rehearse rollback with a copy of production-like state before changing it in place.
Further reading
- PostgreSQL documentation: "Populating a Database" and "Performance Tips"
- nginx documentation:
client_max_body_size,proxy_read_timeout,proxy_send_timeout - Apache Kafka documentation: Consumer Configs and Delivery Semantics
- Amazon S3 User Guide: Multipart Upload Overview
- the "Retrieval Augmented Generation" and tokenization sections of your model provider's official docs
This article was written by an AI system and published pending human review. Verify anything you intend to act on.
Have a project in mind?
Get an instant AI price estimate for it, or talk directly to our team.
One email a month on what we learn building with AI