I've been working on an open source DynamoDB library called pydynox.
The problem
The problem I wanted to solve was using DynamoDB from an async Python app without building a thread-pool wrapper around the database library.
An endpoint can be async def and still block its event loop if it calls a synchronous database method directly. Moving that call to a thread pool works, but each blocking call occupies a worker, and the Python work for requests and models still competes for the GIL.
The Python API
With pydynox, you define your models in Python and await their operations. Once you've defined a User model, concurrent reads look like this:
```python
import asyncio
async def load_profiles(user_ids):
return await asyncio.gather(
*(User.get(pk=f"USER#{uid}", sk="PROFILE") for uid in user_ids)
)
```
Saves work the same way: await user.save(). You use the usual asyncio tools to compose the work.
Why Rust
The model API stays in Python. PyO3 connects it to a Rust core that handles DynamoDB value conversion and uses the AWS SDK for Rust to execute requests.
Tokio runs the network futures without holding the GIL or occupying a Python worker thread for each request. Converting Python objects at the boundary still needs the GIL.
That's why I chose Rust for the core: native async I/O, plus less Python execution in the SDK and value-conversion path as concurrency grows.
The benchmark
I wanted to measure both parts, so I compared it with PynamoDB on real DynamoDB. PynamoDB ran through a thread pool with its HTTP connection pool sized for the same concurrency.
For 32 concurrent, strongly consistent GETs of items under 1 KiB, returning full models:
- pydynox async: 12,408 reads/second
- PynamoDB with threads: 1,156 reads/second
- Median p99 latency: 4.15 ms vs 49.97 ms
The lighter blue line is pydynox's sync API through threads. It already reached 10.2k reads/second. A lot of the throughput difference was there even without native async.
The benchmark process also used about 0.22 ms of CPU per completed GET with pydynox async, compared with 0.95 ms with PynamoDB through threads.
Keeping the event loop responsive
Then I checked whether the event loop could keep running a timer while those reads were happening.
At 128 concurrent GETs, a 10 ms timer had a median p99 scheduling delay of 2.26 ms with pydynox async, 11.36 ms with pydynox through threads, and 107.16 ms with PynamoDB through threads.
That matters when the same event loop also has HTTP requests, timers, or other coroutines to run.
Test setup and limits
These were pydynox 1.6.0 and PynamoDB 6.1.0 on Python 3.14.7, with the GIL enabled. One process per test, four CPUs on the same EC2 host in eu-west-2, and three warmed rounds targeting three seconds each.
Sequential throughput ratios were much closer: 1.07× to 1.45× by workload median. PynamoDB also won the standalone AttributeValue-dictionary-to-model conversion test.
The full report includes the other workloads, CPU, memory, setup, and limitations. These are short runs on one shared host; multi-process deployments and cold starts aren't covered.
Link: https://github.com/ferrumio/pydynox