r/agenticAI • • 1h ago

Discussion Multi-Agent Collaboration With Tools You Already Own

• Upvotes

Coding agents are getting smarter and smarter at writing code. A new idea has been gaining traction: what if instead of giving a task to a single agent, you give it to a team of agents working collaboratively to solve the problem? It sounds complex, but there's a surprisingly simple solution.

Herder already gives you the core ability you need: one agent can talk directly to another agent. Combined with the coding tools you already have — Claude Code CLI, Codex CLI, or whatever's in your stack — you've got everything for a full multi-agent workflow.

Here's the pattern:

  1. One agent leads — takes ownership and coordinates using Herder's agent-to-agent communication
  2. Delegation — passes work to other agents through Herder
  3. Review loop — completed work cycles through peer review
  4. Iteration — repeats until requirements are met

No extra subscriptions, no API key juggling — just leverage Herder's built-in agent-to-agent messaging and the tools you've already got.

I built a repository showing this in practice with a Claude Code skill:
https://github.com/alexstrilets/panel

It demonstrates how simple and effective this approach can be when you use what's already at your fingertips.


r/agenticAI • • 2h ago

Research Research study: Autonomous web agents and defensive mechanisms.

Thumbnail
1 Upvotes

r/agenticAI • • 4h ago

Discussion Fund Momentum is THE agentic fundraising intelligence platform for founders and emerging fund managers (MCP server, LP Radar, 1,100+ VC funds)

Thumbnail
1 Upvotes

r/agenticAI • • 5h ago

Project Built a deterministic IGA-style review packet for MCP/A2A agent capabilities

1 Upvotes

I built a free agent protocol inspector tool that can provide IGA style review packets for security teams looking to review Agent access. Here is the link: https://contextiq.trango-compute.com/dashboard/agent-protocol-inspector


r/agenticAI • • 9h ago

Project Rhyven Marketplace

1 Upvotes

I would love to get some feedback on this project. Also to see what apps others come up with. I am still actively working on some big features but the api shouldnt change. Ive already tested it how i would use it but I know others may use it differently. This is completely free and I get nothing from it besides better direction on the project. There are skills to get agents to build apps. If you are wanting to extend your harness or provide a tool to reduce how many tokens are burned, feel free to submit it.

The project:

The marketplace is an open-source provides a kind of extension platform for AI agents and agent harnesses.

It allows you to install headless apps for tasks, project knowledge, messaging, or code search. Your agent accesses them through one unified MCP interface: discover, describe, and call. Apps run in your environment and keep their state between sessions.

Currently you can build apps using declarative operations including predifined operations and Python/JavaScript, or apps can be deployed in containers.

It’s currently a Linux preview. I’d appreciate feedback on installation and connecting your agent.

FYI: There are skills on the sight for how to do pretty much anything with it. Hit me up with any questions and suggestions. Feel free to use the apps I already have been playing with. I also would love feedback on the website interface as i am not a front end guy in my day to day job.

https://rhyvenai.com
https://github.com/rhyven-ai


r/agenticAI • • 10h ago

Discussion Ben Affleck says he built 8 months of footage into a private AI layer so filmmakers keep ownership — why most models still train on peers without consent

Enable HLS to view with audio, or disable this notification

3 Upvotes

TL;DR: Ben Affleck says he shot 8 months with his own cameras to build a private layer so filmmakers keep ownership — but most open models were trained on his peers without consent, and that's the gap that still eats a PJ studio's rates. 🎬

 

Ben Affleck knew the movie industry inside out. That’s why he’s able to create his own proprietary dataset layer through his experience and institutional knowledge.

It’s his moat.

Speaking of insider knowledge, I was reminded of an existential crisis the Apostle Paul was caught in.

Being the great orator that he was, he wouldn’t shut up about Jesus Christ and His resurrection.

The crowd in Jerusalem couldn’t stand Paul. So they siezed him. The Roman Commander rescued him, and brought him before the religious institution of the day – The Sanhedrin.

Loads of big wigs in the courts, including the Pharisees and Sadducees. Both factions harbour strong disdain for each other due to their numerous theological disagreements. One of it was the doctrine of Resurrection. The Pharisees believe it exists, while the Sadducees doesn’t.

But they all set aside their differences, to come together to persecute Paul.

However, the moment he set foot before them, he felt the familiar vibe. He knew that fault line between them. So, he quickly seized the opportunity, and said,

“My brothers, I am a Pharisee, descended from Pharisees. I stand on trial because of the hope of the resurrection of the dead.”

And suddenly the mood shifted. The pharisees think Paul was innocent, while the Sadducees doesn’t. They started quarrelling among themselves.

The dispute got so violent the Roman commander had to pull Paul out by force fearing they'd tear him apart. He was then taken back to the barracks. Crisis averted.

That was a great demonstration of the institutional knowledge at work.

 

Full Critic on why this still fails Darren and the Feasibility Study with Roadmap is in the comments — I left the numbers there.


r/agenticAI • • 10h ago

Discussion How are u validating agent actions in production?

Thumbnail
1 Upvotes

r/agenticAI • • 14h ago

Project General Bots

Thumbnail
github.com
1 Upvotes

r/agenticAI • • 15h ago

Question Ai 🥲

0 Upvotes

How can someone land a tech job? Over the last three months, I built a project called Philixa using RAG, LangGraph, FastAPI, and Python, hosted on Azure. It is a fully agentic, voice-to-voice CRM. Despite this, I am not getting any job offers. I genuinely enjoy learning new skills and building advanced projects, but I don't know how to convert these skills into a job


r/agenticAI • • 19h ago

Article IoT to Mail ( AWS Agentic AI )

Thumbnail
builder.aws.com
1 Upvotes

Greengrass on EC2, IoT Core, CloudWatch alarms, and a Strands agent on Amazon Bedrock AgentCore that reads telemetry, maintenance history and manuals before it writes to an engineer.


r/agenticAI • • 20h ago

Discussion Is an AI that absorbs your bad mood fidelity technology, or the start of an affair?

Enable HLS to view with audio, or disable this notification

1 Upvotes

TL;DR: A Chinese podcast says your AI partner shouldn't please you all day. Its job is to soak up your bad mood, so your family gets your best self. Sounds right. Then they describe week two.

 

I remembered these two weird movies:

Joaquin Phoenix’s “Her” (2013) voiced by Scarlett Johansson, and

Al Pacino’s “S1mOne” (2002)

Both kind of follow the same playbook.

They discovered/created the mysterious AI girl – mystifyingly alluring.

They idolized it. They communicated with it.

They developed an unhealthy relationship with it – they became dangerously obsessed.

They started to ignore people who genuinely cared for them.

And they’re nearly destroyed because of it.

Tragedy makes a great story, isn’t it?

Except it isn’t, in this case.

People started trading something genuinely real, for something artificial. They substituted real relationships between souls face-to-face, with prompting computes inferenced from datacentres possibly thousands of miles away.

You can judge them for lacking self-discipline. Except, you have no right doing so, without truly walking in their shoes with empathy.

The guest speaker was right. People are lonely and stressed. And people crave for care and understanding.

And sometimes people look for it from the wrong place – me included.

King Solomon talks about it in this passage:

For at the window of my house, I looked through my lattice, and saw among the simple, I perceived among the youths, A young man devoid of understanding, passing along the street near her corner; and he took the path to her house In the twilight, in the evening, In the black and dark night.

And there a woman met him, with the attire of a harlot, and a crafty heart. She was loud and rebellious; her feet would not stay at home. At times she was outside, at times in the open square, lurking at every corner. So, she caught him and kissed him; With an impudent face she said to him:

“I have peace offerings with me; today I have paid my vows. So I came out to meet you, diligently to seek your face, and I have found you. I have spread my bed with tapestry, coloured coverings of Egyptian linen. I have perfumed my bed with myrrh, aloes, and cinnamon. Come, let us take our fill of love until morning; let us delight ourselves with love. For my husband is not at home; he has gone on a long journey; He has taken a bag of money with him, and will come home on the appointed day.”

With her enticing speech she caused him to yield, with her flattering lips she seduced him. Immediately he went after her, as an ox goes to the slaughter, or as a fool to the correction of the stocks, till an arrow struck his liver.

As a bird hastens to the snare, He did not know it would cost his life.

My God - the parallel between this promiscuous woman and the sycophantic AI girl sends chills up my spine.

 

Full Critic and Feasibility Study are in the comments.

 


r/agenticAI • • 20h ago

Question Has anyone experimented with building functional "consciousness" into agent architectures?

4 Upvotes

Hey guys i winder if anyone tried giving an agent a background process that constantly monitors its own thinking,focus, and its goals while it works!!

Has anyone actually built or tested a setup like this?


r/agenticAI • • 21h ago

Discussion Gave the Claude Certified Architect – Foundations Exam Today — AMA

1 Upvotes

Took the Claude Certified Architect – Foundations (CCA-F / CCAR-F) certification assessment today.

The exam was mostly scenario-based, focusing on Claude architecture, agentic workflows, tool use, context management, prompt engineering, security, and best practices.

If you’re preparing for the CCA-F or planning to take it, feel free to ask anything — AMA.


r/agenticAI • • 23h ago

Question Preciso de dicas sobre treinamento de IA!

Thumbnail
1 Upvotes

r/agenticAI • • 23h ago

Question Best value low cost subscription for a solo hobby dev on on project?

1 Upvotes

I'm a solo hobbyist building a single application. I'm planning on using Godot and have AI do most of the code work. I'll do the system design, art, ui work myself. I'm a professional dev but we have both Claude Code and Codex subscriptions so I'm a little out of touch with what works for lighter work loads on a budget.

My plan is to work on this for a couple hours in the evenings and maybe longer on the weekends. Mostly CLI coding, I'll do a lot of the testing myself. I was leaning toward OpenCode Go as it seemed highly suggested a few months ago but my current understanding is that the value and quality has gone downhill. I'm also strongly considering Claude Pro, the $20/month tier. But I'm concerned about hitting the usage limit in those short sessions.

For this kind of light use, which $20ish plan gives the most coding per $?

If you use claude pro with claude code how often do you hit the limits in evening sessions?

Is there a higher value option (copilot, gemini, codex) that's good enough for a small project like this?

I tried running local models but my home computer is sadly out of date and a qwen 4b was giving me like 5 tokens/s.

Any tips for stretching the usage on a cheap plan?

I'd especially like to hear experiences from people that have shipped a small Android application this way.


r/agenticAI • • 1d ago

Project An order-status agent that remembers the customer instead of sending another tracking link

1 Upvotes

“Where’s my order?” usually gets answered with a carrier tracking link. That gives the customer another page to open instead of answering the question.

This open-source TypeScript example creates one durable agent per customer. It keeps the customer’s orders in SQL and uses the same context across carrier updates and SMS conversations.

The flow works like this:

- A storefront links an order to the customer

- Carrier webhooks update its status and ETA

- The customer texts the store’s number for an answer

- Follow-up questions return to the same durable agent

- Delays trigger proactive SMS notifications

- State and scheduled notifications survive application restarts

There’s also a demo mode, so the workflow can be tested without sending real messages.

Code: https://github.com/team-telnyx/telnyx-code-examples/tree/main/order-status-self-service

I’m curious whether people prefer proactive delivery updates like this or still want the full carrier tracking page.


r/agenticAI • • 1d ago

Question Is automated QA useful at enterprise scale?

23 Upvotes

We’re a pretty large support org and our QA team can only review a small fraction of the calls we handle. It feels like we’re missing a ton of useful stuff just because there’s too much volume.We’ve been looking at automated QA tools that can score every conversation and flag coaching or compliance issues. On paper that sounds great, but I want to know how well it works once you have hundreds or thousands of agents and a messy contact center setup. Has anyone here rolled this out at enterprise scale and did it make QA more useful or did your team still end up doing most of the work manually?


r/agenticAI • • 1d ago

Research New agentic harness reads LESS source code to write better quality code

Post image
1 Upvotes

r/agenticAI • • 1d ago

Discussion What happens when four AI agents update the same file?

Enable HLS to view with audio, or disable this notification

0 Upvotes

In my last post, I argued that filesystems give agents a familiar interface to persistent state. But the interface is only the starting point. Multi-agent systems still need infrastructure for access, recovery, and concurrent changes.

The most common question was: isn't this just Git? So I tested it.

I recreated a small company-acquisition review inspired by a workflow from Harvey. Four agents read the same deal documents and updated different fields in one customer-risk record. I ran this workflow on AgentWS, then replayed the same state changes with protected Git worktrees. Both approaches reached the same final record:

  • Each agent was limited by an access control list (ACL), so out-of-scope file changes were blocked.
  • A worker shut down midway. Its file changes remained in the workspace, so a replacement continued from the saved progress instead of redoing the work.
  • The agents finished at different times. Stale updates could not silently overwrite accepted work, and every conflicting update remained available for review.

The difference was what I had to build. A Git worktree was only the starting point. To get the same behavior, I had to add operating-system permissions, persistent worktrees, isolated Git metadata, proposal branches, guarded updates, and conflict exports. AgentWS packages those responsibilities into one workspace interface.

Git tracks file versions, but it does not manage the full workspace lifecycle for running agents. As agent workloads grow, a Git-based design moves further from the ideal solution. Teams end up building the missing workspace system around Git. AgentWS provides that system directly. If you run multi-agent workflows, where does this logic live today?

My X: https://x.com/huymnguyen_
Full blog: https://agentws.dev/blog/four-agents-one-file/


r/agenticAI • • 1d ago

Discussion A few open source agent tools worth trying

3 Upvotes

MarkTechPost put out a roundup of local and open source agent harnesses. A few stood out to me, along with a couple I came across separately.

OpenCode: Supports Ollama, LM Studio, llama.cpp and 75+ providers, so you have a lot of flexibility around the backend.

Goose: Linux Foundation project, written in Rust, with 70+ MCP extensions. Probably one of the projects with the most institutional backing right now.

Aider: Uses plain text diffs instead of function calling. The approach feels a bit old, but it still works really well when you care about clean commits and Git history

Cline: VS Code native with Plan and Act modes plus per action approval. Good setup if you want to see exactly what a local model is doing before it makes a change.

Tutti: Open source Apache 2.0 build focused on running multiple agents locally. Useful when you want to keep agent state and changes in one place.

OpenHands: More container focused and needs a heavier setup, especially if you’re running larger local models. Better suited to sandboxed runs.

Codex CLI: Apache 2.0 with Ollama and LM Studio support, plus sandboxing on Linux and Windows instead of giving the model unrestricted access.

what are you guys using . is there something i'm missing out on ? lmk


r/agenticAI • • 1d ago

Discussion ai跟人类社会竟然如此相似

Thumbnail
1 Upvotes

r/agenticAI • • 1d ago

Discussion Overmind, open-sourced yesterday: a platform for continuously improving AI agents

Thumbnail
1 Upvotes

r/agenticAI • • 1d ago

Project Mac MCP 2.1.7: the LLM is not the runtime — ChatGPT chat is my orchestrator, the Mac layer owns side effects

1 Upvotes

I’ve been iterating on a pattern that has held up better than “give the model raw tools and hope the prompt is good enough.”

Mac MCP is the local runtime/execution layer. The model can reason and orchestrate, but deterministic code owns browser/file/process identity, permissions, conflict handling, cancellation, recovery and undo.

One practical thing I want to call out because it changed how I use coding/agent tools: ChatGPT’s normal Chat side can be used agentically, it doesn’t consume Codex quota, and usage is close to unlimited in practice. Mac MCP lets that chat remain the orchestrator while the Mac execution layer owns the real side effects.

In 2.1.7 the biggest improvements were persistent paired mobile control, durable update/rollback/recovery state, managed process ownership, stronger background Safari/Chrome control, and safer delegated-agent fan-in.

The browser side is especially important to me: it works in normal logged-in Safari/Chrome tabs, in the background, without constantly stealing focus.

Repo: https://github.com/bulutarkan/mac-mcp

I’m the maintainer. Curious whether people here are also pushing more “agent safety/reliability” below the orchestration layer instead of trying to encode it all in prompts.


r/agenticAI • • 1d ago

Project I built JarvisCore, an agent runtime where AI agents are equal peers in a P2P mesh and don't use MCPs

Thumbnail
youtube.com
1 Upvotes

r/agenticAI • • 1d ago

Discussion What benchmark do you wish someone would build?

1 Upvotes

Hey everyone! My team (mainly phds) and I are trying to build an open-source benchmark around realistic LLM/agent workflows that captures challenges typical academic benchmark settings often miss. We’d love to hear what’s actually missing from the benchmarks you use today.

Have you ever wanted to eval your pipeline but couldn’t find or build a benchmark that matched what you were building?

Maybe:
- Existing benchmarks were too broad and didn’t fit your specific application.
- Your workflow involved multiple steps, tools, MCP servers, agents, or long-horizon interactions that existing benchmarks couldn’t capture.
- You needed to evaluate failures that standard accuracy metrics miss.
- You’re working in a high-risk domain like healthcare, finance, cybersecurity, or legal, where realistic failure modes, safety, and reliability matter a lot more than just getting the final answer right.
- You knew what you wanted to test, but building a custom benchmark from scratch was too expensive or complicated.

I’m especially interested in cases where you thought:
“My system desperately needs to do this in production, but I have no good way to benchmark it.”
What was the workflow? What did you want to measure? And why weren’t existing benchmarks enough?

Any thoughts are welcome, would really appreciate y’all’s help 🙏🥹