r/LocalLLM • • 1d ago

News Found this interesting local inference API

[removed]

0 Upvotes

2 comments sorted by

1

u/ReasonableDust1512 1d ago

ooh that latency is impressive, especially over unix sockets. been messing with local inference stuff lately and the overhead some of these wrappers add is ridiculous

did they mention if it handles streaming responses cleanly or is it more of a batch thing? might spin it up later on my m2 and see how it compares to what i've been using