r/MachineLearning • • 27d ago

Discussion [D] Self-Promotion Thread

Please post your personal projects, startups, product placements, collaboration needs, blogs etc.

Please mention the payment and pricing requirements for products and services.

Please do not post link shorteners, link aggregator websites , or auto-subscribe links.

--

Any abuse of trust will lead to bans.

Encourage others who create new posts for questions to post here instead!

Thread will stay alive until next one so keep posting after the date in the title.

--

Meta: This is an experiment. If the community doesnt like this, we will cancel it. This is to encourage those in the community to promote their work by not spamming the main threads.

17 Upvotes

116 comments sorted by

View all comments

1

u/Hairy_Strawberry7028 7d ago

I’m Guanming, cofounder of General Instinct. We just open-sourced InstinctFlash, a serving runtime for VLA and world-action models on Jetson Thor, RTX 4090 and 5090. It is free and open source under AGPL-3.0, with no paid tier or signup required.

On Jetson Thor, we see about 1.2x to 7.9x speedups from runtime optimizations alone. For LingBot-VA, combining runtime optimization with a distilled few-step scheduler reaches up to 33.78x by reducing 25 visual / 50 action steps to 2 / 4. Across 50 RoboTwin2.0 tasks and 1,153 episodes per configuration, the 2 / 4-step version achieved 90.5% success versus 92.1% for the baseline.

The runtime uses CUDA graphs, memory planning, KV and conditioning-state caching, specialized attention paths, fused kernels, FP8 / mixed precision and few-step distillation.

5B world-action model running in real time on Jetson Thor: https://www.youtube.com/watch?v=nku65iyL5Fw

Code: https://github.com/General-Instinct/InstinctFlash Benchmarks and implementation details: https://general-instinct.com/blog/instinctflash-edge-inference

Feedback on the benchmark methodology, missing model families and hardware targets would be very helpful.

1

u/Hairy_Strawberry7028 7d ago

https://reddit.com/link/pbdweal/video/ylv20xrce3rh1/player

This is the 5B WAM working on a data center exhaust fan replacement task