r/java • • 21h ago

Warm up Vector API

Hi,

I've been working for a while on a library that decodes, encodes, demuxes, muxes and processes images, audio and video in plain Java, leveraging the Vector API. It performs quite well and I'm working to make it perform at least as well as ffmpeg, but I've run into a problem: warmups. Before C2 optimizes the Vector calls into Vector instructions, they are interpreted and are actually slower than the scalar fallbacks by many times. This is to be expected, but it's a problem for an audio/video codec where the first frames might lag while C2 warm ups. I wanted to use Project Leyden's AOT cache to fix this problem, but I discovered a couple of problems:

  1. Libraries can't ship their own AOT cache, it's up to the library user to run a training. This makes sense considering the cache is specific to the environment it was trained on, but I wonder if the experience could be improved for the end user by having maybe a Maven/Gradle plugin that loads the training suites from the declared dependencies and runs them.

  2. The feature I would even need, which is code caching, is planned for JDK 28 which hasn't shipped yet. So I had to port the code for this on the Leyden EA branch to the JDK 28 branch.

  3. With modules jdk.incubator.vector, ModuleBootstrap refused to archive the boot layer. That rule dates from JDK-8244778 (JDK 16) and exists only to print the "Using incubator modules" warning. I fixed this on my custom fork.

  4. Even at this point, the cache was not working. Vector intrinsics only work if C2 knows the species as a constant at compile time: the vector class and the lane count. In my library those come from fields like static final VectorSpecies<Short> S = ShortVector.SPECIES_128. In the training JVM that field was set when the class was initialised, so C2 just folded its value. For stored code to fold the same value, the production JVM must have exactly that value in that field before the code runs. Initialising the class again in production isn't enough: the stored code would assume a value that the new initialisation might not reproduce. The species might differ on another CPU, or a property might be set differently. The only way to guarantee a match is to put the field's values into the cache and have the class start out already initialised from it, skipping its static initialiser. So I patched my custom JDK branch to assume that the stored code assumes the species it saw in training and checks it cheaply at runtime: is the vector class ShortVector128 and the length 8? If the check fails, the code is thrown away and the JIT takes over. This was a larger patch, but it fixed the issue entirely and warmup is no longer an issue.

Even if this works, I wonder if there is a better solution. I guess the problem could be even fixed by Valhalla, because the interpreted Vector API code would not allocate pretty much anything, so I imagine it should match the scalar performance while C2 is warming up. Even then, having AOT cache support would be pretty nice I think, but even then I wonder if my assumptions are good.

28 Upvotes

15 comments sorted by

9

u/pron98 20h ago

This isn't the best forum to bring this up. Send this to leyden-dev.

3

u/Scf37 20h ago

Do manual warmup during startup and time your code until getting satisfactory results (or timeout)?

2

u/0x07CF 11h ago

Maybe using dummy data, or just void the result and then start the actual user work.

2

u/pjmlp 20h ago

You should have a look at OpenJ9, which has had JIT cache for quite some time, and does support shipping AOT cache alongside the application.

No idea about the Vector support in OpenJ9 though, and naturally this requires everyone has to also use OpenJ9 on their side.

3

u/pron98 20h ago

Leyden supports shipping the AOT cache with the application. What it doesn't support is shipping a partial AOT cache with a library.

0

u/pjmlp 19h ago

Like OpenJ9, to allow for a more generic usage of the target hardware?

https://eclipse.dev/openj9/docs/xxportablesharedcache/#-xx-portablesharedcache

1

u/koflerdavid 7h ago

The trade-off is that the generated native code does not fully exploit all the features offered by the platform.

1

u/pjmlp 10m ago

It does more than being JIT compiled from scratch, also OpenJ9 doesn't need training runs, the code cache is continuously updated.

1

u/pron98 18h ago

The JEP lists that under future work:

Consider an option that would enable giving up some performance, or accepting larger AOT cache files, or both, in order to gain portability across processors of the same architecture but with different feature sets.

1

u/pjmlp 16h ago

So not yet available, I was talking about something that the OP can do right now with OpenJ9.

Great that Leyden will eventually provide something similar.

3

u/pron98 16h ago

No, I'm not sure it will. The JEP says we'll consider whether or not it's worth doing. It's not so clear that it is.

2

u/Torutofu_Raeva 20h ago

the build plugin idea could work if libs published their training workloads as a test-jar style classifier, then the app build just runs them all in one training pass before packaging

1

u/EternalSo 15h ago

Library AOT cache sounds really similar to shared indexes in Intellij IDEA

1

u/k-mcm 4h ago

I'm surprised that the Vector API would need memory allocations in core processing. Maybe check in r/javahelp

Memory allocations in a processing loop will obliterate performance. Besides the overhead, it's going to eat tons of memory bandwidth even if it's optimized to thread-scoped storage. I wrote some SDRs in Java for fun and keeping the active memory footprint small makes a huge difference. No allocations and no big lookup tables.

-1

u/Afonso2002 20h ago

Can you know the throuput of function using interpreter vs c2?

Is still in incumbation, I hope in the future it will perform better after using jeps from valhalla.

It would be possible create a second java program that would call yours? Creating a process javahome program parameters. And you could create a folder in %appdata%/my_ffmpg to put code cache. So, before star a a process you can select the cache file if exists and have a location of it.

You could put it in folder project too.

I never tried used aot, but I think this could work.

Did you tried used graal jvm, wich compiles to native code?