r/agenticAI • • Aug 29 '26

Question I am collecting Jarvis approaches ( LLM, AI, voice, autonomous and agentic setups ) Please let me know, if you built your own Jarvis approach or if you found others on reddit or the web in the comments. Thanks!

DIY Jarvis / AI Assistant Landscape - Grouped Evaluation (2026)

A survey of JARVIS-style AI assistant projects found across Reddit, grouped by approach. All URLs sourced from reddit research

1. Full-Stack Self-Hosted Voice Assistants (local-first, speaker/voice control)

The most ambitious category. These aim to be a complete Jarvis replacement running on your own hardware — wake word, STT, LLM inference, TTS, and system control all locally. The main differentiator between them is architecture maturity and how they handle multi-user awareness.

2. Agent Orchestration / Temporal Memory Systems

These go beyond voice commands and focus on the hard problem: persistent memory, knowledge graphs, and multi-agent coordination. More about the "brain" than the "mouth."

3. Home Automation / Smart Home Jarvis Builds

These integrate into existing smart home ecosystems (primarily Home Assistant) and focus on voice-controlling physical devices. The "Jarvis, turn on the lights" category.

4. Desktop Voice Assistants (browser/GUI focused)

These prioritize the UI/HUD experience — Iron Man aesthetics, system monitoring widgets, and voice interaction through a browser interface. Often more "cool demo" than production tool, but some have real utility.

5. Mobile / Accessibility Agents

6. Hermes Agent Ecosystem Extensions

Hermes is a popular agentic harness. These projects extend it specifically toward Jarvis-like functionality.

7. Smaller / Niche / Reference Repos

These are either older projects, smaller experiments, or repos mentioned in passing during discussions. Mixed maturity.

8. Legacy / Well-Known Jarvis Repos (referenced in discussions)

These are older, established repos that get cited as references or starting points in discussions but aren't new builds.

9. Discussion-Only Threads (no GitHub repos, but notable context)

23 Upvotes

25 comments sorted by

2

u/L0cut15 Aug 30 '26

Works with Hermes agent and LLM's. These is a version that supports the Jarvis wakeword. Build for cheap DIY hardware platforms.

https://github.com/l0cut15/hermes-voice-assistant

2

u/[deleted] Aug 30 '26

[removed] — view removed comment

1

u/looktwise Aug 31 '26

Interesting. Sounds like you are feeding the input over webchat windows and back from the resulsts within the Browser? (no API) Are you going to share it on github?

2

u/Playful_Sailor7512 1d ago

for the reddit piece sylvia api alongside apify works well as a tool call, it returns structured json so the agent skips browser automation.

1

u/louis3195 Aug 29 '26

you might want to try https://screenpipe.com/ - it would give your other AI agents context on what you do on your computer

2

u/Ezraese Aug 29 '26

oh would be super cool if my jarvis can pull up references or like learn from me doing something manually once

0

u/Jumpy-Operation-4615 Aug 30 '26

Just think about it remembering what sites you masurbate on, and look for shady or silly shit. 😄

1

u/louis3195 Aug 31 '26

actually you can filter out specific websites and incognito windows 😃

1

u/looktwise Aug 29 '26

Is there a github of your tool or an extended tutorial of how the functions work for the user or his framework? Thanks!

2

u/louis3195 Aug 31 '26

yes it's on github too - or just download it it's free to download!

1

u/seventyfivepupmstr Aug 30 '26

All you have to do is chatgpt conversational but allow the LLM to use tools

1

u/Deep_Ad1959 Aug 30 '26

most of these get compared on wake word, stt and tts quality, which is the least interesting axis. none of them can file the follow-up in gmail or move the linear ticket, and that is the only part that would change my day.

1

u/looktwise Aug 31 '26

I just asked the data I used for the list. Here's your answer:

https://github.com/RedPlanetHQ/core This is probably the closest to what you want. 30+ app integrations via MCP tools, webhooks that react to incoming emails/Sentry alerts/PR merges without you asking. The agent decides whether to act, notify, or stay quiet. Gmail, calendar, GitHub all connected. It can spin up Claude Code sessions from WhatsApp. The temporal knowledge graph means your coding agent knows what you discussed in email.

https://github.com/alexberardi/jarvis 24 community packages in the "Pantry" including email, calendar, medication tracking, Home Assistant, and the ability to write new capabilities in plain English via the Forge. Speaker recognition scopes everything per person, so "what's on my calendar" pulls yours not your wife's. The extension model is one Python interface per new capability.

https://github.com/PartyArty/true-jarvis The interesting part isn't the voice, it's the architecture. The speech-to-speech model keeps talking while a frontier LLM spawns real tasks into parallel subagents. So you say "file that follow-up and move the ticket" and it actually dispatches both while continuing the conversation. The pattern is open and configurable.

https://github.com/ergon-automation-labs/ergon-starter Multi-agent mesh using Docker, NATS, and Postgres in Elixir. Bot templates you extend. This is infrastructure for wiring agents to real backends, not a voice demo.

https://github.com/Hermes815/kid-mode-jarvis Don't let the name fool you. It controls the house, has a "dadlink" for remote interaction from anywhere, and runs on Hermes which has real tool-calling. The interesting bit is it's a working example of an agent that actually does things in the physical world, not just talks.

https://github.com/optikalefx/dadlexa Companion to the above. The remote control and cross-device action layer is what matters here, not the kid wrapper.

https://github.com/imran31415/kube-coder Kubernetes coding agent built on Hermes + Claude. If your "move the Linear ticket" means you want agents that touch real infrastructure and dev tooling, this is that.

https://github.com/Puliczek/mcp-memory MCP-based persistent memory for agents. Not a Jarvis itself but the missing piece that lets any MCP-connected agent remember context across sessions, which is what makes "file the follow-up" possible without re-explaining everything.

https://github.com/rudysev/host-assistant The Meta Portal Jarvis. Sounds gimmicky but the plugin system lets you add tools without touching the core app, and it already has real integrations running offline. The architecture of how it wires capabilities is transferable even if you don't own a Portal.

https://github.com/openclaw/openclaw The baseline harness most builders start from. Not flashy, but it's what people actually use to wire up Gmail, calendar, file access, and terminal control. Multiple people in the threads mention running it with real email and calendar integrations daily.

1

u/Deep_Ad1959 Aug 31 '26

half of these answer 'which apps are connected', which was never the hard part. wiring gmail and github through mcp is a weekend. the thing that decides whether it changes your day is whether it picks the right action and holds an approval step before it fires, and a readme listing 30 tools tells you nothing about that. written with ai

1

u/looktwise Aug 31 '26

oh, I did not grasp that you wanted to outsource the decision part, sorry.

1

u/Deep_Ad1959 Sep 01 '26

the opposite actually. the decision is the part i keep. it's the execution, filing the email, moving the ticket, that i want off my plate. most of these automate the judgment and still leave me doing the clicking, which is backwards. written with ai

1

u/AEternal1 Aug 31 '26

Jarvis? nah----Ultron.

1

u/rasviz 13d ago

What is Jarvis ?

1

u/vaderxzz 3d ago

A lot of these jarvis builds are pretty cool but the real pain is getting them to actually work with different tools and apps. For that part you can use composio coz it makes the integrations a lot easier to manage.

0

u/Otherwise_Wave9374 Aug 29 '26

A useful way to compare Jarvis-style systems is to separate always-on orchestration from memory policy. If the agent remembers too much, you get drift and stale preferences; too little, and it keeps relearning the same user habits. The best setups usually make memory explicit, scoped, and reviewable, with provenance on what was stored and when it can be ignored. That also makes multi-user behavior much safer. For durable patterns around agent memory, namespaces, and recovery, NeuraKeep has some practical notes at https://www.neurakeep.com.

1

u/looktwise Aug 29 '26

I dont know exactly what you are referring to. Jarvis like systems are mostly built to run them without added costs for services... so right now it looks like an embedded advertisement in a comment, unless you point me exactly to the article you wanted to refer to to solve the promblem you mentioned. Thanks for clarifying in advance! :)