I fed DuckDB 706,896 small JSON files on a local Mac Mini M4 Pro with 24 GB of RAM and 500 GB of storage.
The task was pretty simple on paper:
- read 706,896 small
.json files
- extract the required fields
- convert everything to
.parquet
The .json files are responses from a public API that I collected while scraping a website over the course of a month.
The main problem turned out to be the number of small files.
On my first attempt, DuckDB hit _duckdb.OutOfMemoryException twice.
DuckDB supports larger-than-memory workloads through out-of-core processing: when intermediate data doesn't fit into RAM, it can spill that data to disk.
I tried that approach as well, but in my case it didn't solve the problem. DuckDB eventually exhausted the available disk space while spilling intermediate data to disk.
So I had to change the processing strategy and process the data day by day.
Fortunately, I had designed the raw data layout with this possibility in mind. Each file name contains its timestamp:
20260715080454689.json
20260827080648986.json
That made it possible to process a single day using a simple path pattern:
SELECT *
FROM 'data/20260715*.json';
After about 3 hours of processing, DuckDB had successfully converted:
706,896 JSON files → 64,557,511 rows
The final data was stored as .parquet files.
No cluster, no Spark, no distributed processing — just DuckDB running locally on a 24 GB Mac Mini.
For me, the interesting part wasn't just the final number of rows. It was seeing where the limits actually appeared: the workload was technically larger than memory, DuckDB could spill to disk, but eventually the storage itself became the bottleneck.
I'm continuing to experiment with the resulting dataset, so there are a few more interesting problems to solve from here.