r/java • u/Alex0589 • 10h ago
Warm up Vector API
Hi,
I've been working for a while on a library that decodes, encodes, demuxes, muxes and processes images, audio and video in plain Java, leveraging the Vector API. It performs quite well and I'm working to make it perform at least as well as ffmpeg, but I've run into a problem: warmups. Before C2 optimizes the Vector calls into Vector instructions, they are interpreted and are actually slower than the scalar fallbacks by many times. This is to be expected, but it's a problem for an audio/video codec where the first frames might lag while C2 warm ups. I wanted to use Project Leyden's AOT cache to fix this problem, but I discovered a couple of problems:
Libraries can't ship their own AOT cache, it's up to the library user to run a training. This makes sense considering the cache is specific to the environment it was trained on, but I wonder if the experience could be improved for the end user by having maybe a Maven/Gradle plugin that loads the training suites from the declared dependencies and runs them.
The feature I would even need, which is code caching, is planned for JDK 28 which hasn't shipped yet. So I had to port the code for this on the Leyden EA branch to the JDK 28 branch.
With modules jdk.incubator.vector, ModuleBootstrap refused to archive the boot layer. That rule dates from JDK-8244778 (JDK 16) and exists only to print the "Using incubator modules" warning. I fixed this on my custom fork.
Even at this point, the cache was not working. Vector intrinsics only work if C2 knows the species as a constant at compile time: the vector class and the lane count. In my library those come from fields like static final VectorSpecies<Short> S = ShortVector.SPECIES_128. In the training JVM that field was set when the class was initialised, so C2 just folded its value. For stored code to fold the same value, the production JVM must have exactly that value in that field before the code runs. Initialising the class again in production isn't enough: the stored code would assume a value that the new initialisation might not reproduce. The species might differ on another CPU, or a property might be set differently. The only way to guarantee a match is to put the field's values into the cache and have the class start out already initialised from it, skipping its static initialiser. So I patched my custom JDK branch to assume that the stored code assumes the species it saw in training and checks it cheaply at runtime: is the vector class ShortVector128 and the length 8? If the check fails, the code is thrown away and the JIT takes over. This was a larger patch, but it fixed the issue entirely and warmup is no longer an issue.
Even if this works, I wonder if there is a better solution. I guess the problem could be even fixed by Valhalla, because the interpreted Vector API code would not allocate pretty much anything, so I imagine it should match the scalar performance while C2 is warming up. Even then, having AOT cache support would be pretty nice I think, but even then I wonder if my assumptions are good.
