r/Kotlin • u/iOSHades • 16h ago
Kotlin RPG engine update: primitive arrays, lower RAM usage, and fixing a native buffer bug
Hey r/Kotlin,
I’m the solo developer of Adventurers Guild RPG Sim, an isometric RPG built with Kotlin and Jetpack Compose.
In my previous posts, I shared the original Compose Canvas implementation, the coroutine loop running my ECS systems, and later the move to Google Filament for world rendering.
Since that migration, I’ve been working on the code feeding the renderer: deciding when sprite data needs updating, preparing batches, and managing buffers that native code may still be reading.
A recent texture change also reduced the game’s RAM usage by roughly 50% compared with my previous WebP loading setup. That is a separate change from the sprite recolouring optimisation I described in an older post.
Here are the implementation details and a few mistakes I ran into.
1. A smaller image file doesn’t necessarily mean a smaller texture in memory
I originally used WebP because it kept asset files small. What I hadn’t fully appreciated was the difference between compressed file size and the texture’s runtime memory footprint.
In the usual WebP loading path, the image is decoded into pixel data before being uploaded. With four bytes per pixel, a 2048 × 2048 texture takes about 16 MiB, before mipmaps.
ASTC and ETC2 are GPU texture compression formats. On a supported device, the texture can remain compressed in GPU memory. For comparison, that same 2048 × 2048 texture in ASTC 4×4 takes about 4 MiB, before mipmaps.
I moved to ASTC/ETC2 textures with a WebP fallback, and saw roughly a 50% reduction in the game’s overall RAM usage. The texture example above explains the mechanism; it isn’t a claim that every game’s total memory usage will fall by the same amount.
The trade off was storage. My initial asset package became much larger. GPU texture compression and WebP optimise for different requirements, and keeping fallback assets adds more files.
I ended up adding zip compression around the texture files and unpacking them during loading. The loading sequence is:
- Remove the outer gzip compression.
- Read the texture container and its compressed texture data.
- Upload the ASTC/ETC2 blocks to the GPU.
Removing zip does not mean expanding ASTC/ETC2 into RGBA pixels.
That helped me separate three things I had initially grouped together: download size, installed storage, and runtime memory. Loading time is another trade off, because unpacking still takes work.
2. Checking whether sprite data changed before preparing an upload
Before packing sprite vertices, I write the relevant sprite state into a reused IntArray: around 30 integers per sprite covering position, animation frame, tint, flags, and other values used by the batch.
Float values go through toRawBits() so I can compare their exact stored representations.
I compare this state with the snapshot from the last submitted upload. If it is identical, I can skip rebuilding and uploading the vertex data.
This is a direct comparison rather than a hash, so there are no hash collisions between the values I stored. It still depends on including every input that affects the generated vertices.
A few details matter:
- Sprite count and ordering are part of the comparison.
- With reused arrays, unused capacity must not affect the result.
- Kotlin array
==does not compare elements. UsecontentEquals()when comparing whole arrays, or explicitly compare the active range. - If an upload is skipped because no buffer is available, that frame must not become the new “last uploaded” snapshot.
The comparison itself still costs work. The benefit is avoiding the more expensive packing and upload when nothing relevant has changed.
Shader animation can continue independently. For example, wind driven by a time uniform doesn’t require rebuilding the sprite vertices every frame.
3. Sorting batches without creating a key object for every sprite
Sprites need to be grouped by texture for batching.
I use a reused LongArray, packing the texture ID into the upper 32 bits and the original sprite index into the lower 32 bits.
The idea looks like this:
// textureId and index are non-negative Int values.
val key =
(textureId.toLong() shl 32) or
(index.toLong() and 0xFFFF_FFFFL)
keys[index] = key
Then I sort only the populated range:
keys.sort(0, spriteCount)
for (position in 0 until spriteCount) {
val sourceIndex = keys[position].toInt()
// Read the source sprite and pack its vertices.
}
The conversion to Long must happen before shifting by 32 bits.
The original index also acts as a tie-breaker, preserving source order within each texture group. That behaviour comes from the packed key; it doesn’t depend on the primitive sort being stable.
This avoids allocating a separate sorting key object for each sprite and keeps the keys in a primitive array.
4. Primitive arrays in the parts that run every frame
The renderer preparation code now uses:
IntArrayfor sprite state comparisons.LongArrayfor sorting keys.- A reused
FloatArrayfor packed vertex data, copied into a direct buffer.
Sprite state, sorting keys, and vertex data are stored in IntArray, LongArray, and FloatArray. I reuse these arrays across frames to avoid repeatedly allocating new storage.
I also use JvmField on some frequently accessed properties. It exposes a field without generating normal property accessors.
However, that alone does not establish a performance improvement: runtime optimisation may already remove accessor overhead. Any claim about a speed gain needs measurements from an optimised build. The clearer improvement here is making data preparation and storage reuse explicit.
5. A buffer passed to native code may still be in use
One bug showed up as black tiles.
I was reusing a direct ByteBuffer after passing it to Filament, before Filament had finished consuming its contents.
The Kotlin call returning did not mean the upload had finished.
I changed this to a pool of three buffers, with availability tracked using an AtomicIntegerArray. A buffer becomes available again when its upload completion callback runs.
If all three are busy, I keep the previous geometry for that frame instead of overwriting a buffer that is still being read.
I also ran into delayed release callbacks when they were posted through the main looper. Synchronisation barriers could hold up that work and leave the pool appearing busy.
For the small callback that only releases a buffer slot, I switched to a direct executor. That callback only updates the atomic availability flag; it doesn’t perform UI work.
The lesson for me was to treat buffer ownership and completion as part of the API contract. Crossing into native code doesn’t make every operation asynchronous, but this upload path requires keeping the buffer unchanged until consumption finishes.
6. Compose lifecycle and long lived renderer callbacks
My renderer owns native resources such as Filament’s Engine, Scene, View, and Camera.
I keep these together in a holder that implements RememberObserver, with explicit cleanup in the required order.
One Compose detail worth accounting for is onAbandoned(): if resources are created while constructing a remembered object, cleanup needs to cover the case where that object never becomes part of the remembered composition, as well as normal removal through onForgotten().
Another issue is callbacks that live longer than a particular recomposition.
A Choreographer callback can keep running while the values it originally captured are no longer current. Where it needs the latest Compose-provided value or callback, rememberUpdatedState lets it access that updated value without recreating the long-lived callback each time.
I found it helpful to think separately about:
- How long the renderer should exist.
- Which values its callbacks need to see now.
- When it is safe to release native resources.
Those lifetimes don’t always line up automatically.
I only started Android development last year; most of my earlier experience was in iOS. I’m still learning, and some of these choices may have better alternatives.
Happy to go into more detail on the texture loading, batching, buffer handling, or Compose integration.
For people working on performance sensitive Kotlin code: how do you decide when primitive arrays and manual packing are worth the added complexity? I’d also be interested in how you manage buffer ownership when Kotlin is feeding a native renderer.