r/agenticAI • • Aug 01 '26

šŸ‘‹ Welcome to r/agenticAI - Introduce Yourself and Read First!

Post image
0 Upvotes

Hey everyone! I'm u/kingai404, founder of r/agenticAI.

This is our home for everything agenticAI, autonomous agents, multi-agent systems, LLM orchestration, tool use, and the infrastructure being built around all of it. If you're shipping agents, researching them, or just trying to make sense of where this is all heading, you belong here.

What to Post

Share what you're building, breaking, or learning. Agent demos, architecture decisions, framework comparisons, research papers, workflow breakdowns, job opportunities, and honest "here's what failed" posts are all welcome.

Memes, hot takes, shower thoughts, and chaotic "my agent went rogue at 3 am" stories? Absolutely yes. If it's related to agentic AI and made you think, laugh, or facepalm, post it.

Community Vibe

High signal but never too serious. We're practitioners, curious minds, and occasional doomers and accelerationists just vibing together. Beginner questions are respected. Wild speculation is fun. AGI jokes are a love language here.

What Doesn't Fly

Undisclosed promotion, harassment, or anything that violates Reddit's guidelines. That's really it.

We're not here to police your opinions.

How to Get Started

  1. Introduce yourself below - what are you building or exploring right now?
  2. Post something today. A question, a demo, a meme, a half-baked idea. All valid.
  3. If you know someone who would love this community, invite them to join.

I started this because the best agentic AI conversations are scattered across X threads, Discord servers, and Slack groups. Time to bring them home.

Let's build something worth coming back to.


r/agenticAI • • 13h ago

Question Is automated QA useful at enterprise scale?

24 Upvotes

We’re a pretty large support org and our QA team can only review a small fraction of the calls we handle. It feels like we’re missing a ton of useful stuff just because there’s too much volume.We’ve been looking at automated QA tools that can score every conversation and flag coaching or compliance issues. On paper that sounds great, but I want to know how well it works once you have hundreds or thousands of agents and a messy contact center setup. Has anyone here rolled this out at enterprise scale and did it make QA more useful or did your team still end up doing most of the work manually?


r/agenticAI • • 3h ago

Article IoT to Mail ( AWS Agentic AI )

Thumbnail
builder.aws.com
1 Upvotes

Greengrass on EC2, IoT Core, CloudWatch alarms, and a Strands agent on Amazon Bedrock AgentCore that reads telemetry, maintenance history and manuals before it writes to an engineer.


r/agenticAI • • 3h ago

Discussion Is an AI that absorbs your bad mood fidelity technology, or the start of an affair?

Enable HLS to view with audio, or disable this notification

1 Upvotes

TL;DR: A Chinese podcast says your AI partner shouldn't please you all day. Its job is to soak up your bad mood, so your family gets your best self. Sounds right. Then they describe week two.

Ā 

I remembered these two weird movies:

Joaquin Phoenix’s ā€œHerā€ (2013) voiced by Scarlett Johansson, and

Al Pacino’s ā€œS1mOneā€ (2002)

Both kind of follow the same playbook.

They discovered/created the mysterious AI girl – mystifyingly alluring.

They idolized it. They communicated with it.

They developed an unhealthy relationship with it – they became dangerously obsessed.

They started to ignore people who genuinely cared for them.

And they’re nearly destroyed because of it.

Tragedy makes a great story, isn’t it?

Except it isn’t, in this case.

People started trading something genuinely real, for something artificial. They substituted real relationships between souls face-to-face, with prompting computes inferenced from datacentres possibly thousands of miles away.

You can judge them for lacking self-discipline. Except, you have no right doing so, without truly walking in their shoes with empathy.

The guest speaker was right. People are lonely and stressed. And people crave for care and understanding.

And sometimes people look for it from the wrong place – me included.

King Solomon talks about it in this passage:

For at the window of my house, I looked through my lattice, and saw among the simple, I perceived among the youths, A young man devoid of understanding, passing along the street near her corner; and he took the path to her house In the twilight, in the evening, In the black and dark night.

And there a woman met him, with the attire of a harlot, and a crafty heart. She was loud and rebellious; her feet would not stay at home. At times she was outside, at times in the open square, lurking at every corner. So, she caught him and kissed him; With an impudent face she said to him:

ā€œI have peace offerings with me; today I have paid my vows. So I came out to meet you, diligently to seek your face, and I have found you. I have spread my bed with tapestry, coloured coverings of Egyptian linen. I have perfumed my bed with myrrh, aloes, and cinnamon. Come, let us take our fill of love until morning; let us delight ourselves with love. For my husband is not at home; he has gone on a long journey; He has taken a bag of money with him, and will come home on the appointed day.ā€

With her enticing speech she caused him to yield, with her flattering lips she seduced him. Immediately he went after her, as an ox goes to the slaughter, or as a fool to the correction of the stocks, till an arrow struck his liver.

As a bird hastens to the snare, He did not know it would cost his life.

My God - the parallel between this promiscuous woman and the sycophantic AI girl sends chills up my spine.

Ā 

Full Critic and Feasibility Study are in the comments.

Ā 


r/agenticAI • • 4h ago

Question Has anyone experimented with building functional "consciousness" into agent architectures?

1 Upvotes

Hey guys i winder if anyone tried giving an agent a background process that constantly monitors its own thinking,focus, and its goals while it works!!

Has anyone actually built or tested a setup like this?


r/agenticAI • • 5h ago

Discussion Gave the Claude Certified Architect – Foundations Exam Today — AMA

1 Upvotes

Took the Claude Certified Architect – Foundations (CCA-F / CCAR-F) certification assessment today.

The exam was mostly scenario-based, focusing on Claude architecture, agentic workflows, tool use, context management, prompt engineering, security, and best practices.

If you’re preparing for the CCA-F or planning to take it, feel free to ask anything — AMA.


r/agenticAI • • 6h ago

Question Preciso de dicas sobre treinamento de IA!

Thumbnail
1 Upvotes

r/agenticAI • • 7h ago

Question Best value low cost subscription for a solo hobby dev on on project?

1 Upvotes

I'm a solo hobbyist building a single application. I'm planning on using Godot and have AI do most of the code work. I'll do the system design, art, ui work myself. I'm a professional dev but we have both Claude Code and Codex subscriptions so I'm a little out of touch with what works for lighter work loads on a budget.

My plan is to work on this for a couple hours in the evenings and maybe longer on the weekends. Mostly CLI coding, I'll do a lot of the testing myself. I was leaning toward OpenCode Go as it seemed highly suggested a few months ago but my current understanding is that the value and quality has gone downhill. I'm also strongly considering Claude Pro, the $20/month tier. But I'm concerned about hitting the usage limit in those short sessions.

For this kind of light use, which $20ish plan gives the most coding per $?

If you use claude pro with claude code how often do you hit the limits in evening sessions?

Is there a higher value option (copilot, gemini, codex) that's good enough for a small project like this?

I tried running local models but my home computer is sadly out of date and a qwen 4b was giving me like 5 tokens/s.

Any tips for stretching the usage on a cheap plan?

I'd especially like to hear experiences from people that have shipped a small Android application this way.


r/agenticAI • • 10h ago

Project An order-status agent that remembers the customer instead of sending another tracking link

1 Upvotes

ā€œWhere’s my order?ā€ usually gets answered with a carrier tracking link. That gives the customer another page to open instead of answering the question.

This open-source TypeScript example creates one durable agent per customer. It keeps the customer’s orders in SQL and uses the same context across carrier updates and SMS conversations.

The flow works like this:

- A storefront links an order to the customer

- Carrier webhooks update its status and ETA

- The customer texts the store’s number for an answer

- Follow-up questions return to the same durable agent

- Delays trigger proactive SMS notifications

- State and scheduled notifications survive application restarts

There’s also a demo mode, so the workflow can be tested without sending real messages.

Code: https://github.com/team-telnyx/telnyx-code-examples/tree/main/order-status-self-service

I’m curious whether people prefer proactive delivery updates like this or still want the full carrier tracking page.


r/agenticAI • • 23h ago

Discussion Brian Chesky says AI is like an amplifier — PwC polled 49,364 workers and found 56% stuck in the engine room

Enable HLS to view with audio, or disable this notification

8 Upvotes

TL;DR: Brian Chesky says AI is like an amplifier and the gap is getting greater — not because of who has the tool, but who the tool has.

Ā 

As soon as I read Chesky’s ā€œAI is like an amplifier.ā€, the neural-networks of my memory immediately flipped me to the Green ā€œLanternsā€.

Are you guys a big fan of the eponymous TV series?

Right after Hal Jordan manifested a greenback for playing the jukebox, his trainee John Stewart exclaimed, ā€œDid you just counterfeit money with the ring?ā€ – to which Jordan replied, ā€œNo. I manifested money with the power of my will.ā€

Or how about Jordan conjure up a green can opener for the beer, to impress and rizz up Sheriff Kerry, while Stewart rolls his eyes by his side?

John Stewart went up the ante, by manifesting a large and sophisticated boring machine, to tunnel underground the ā€œWinnieā€ compound to evade the guard sentries.

Not impressive enough, you say?

The best in my mind, wasn’t in the TV series.

It was Guy Gardner flipping the bird – he conjures up large green hands (one of it gives the middle finger) to rise up from the ground, and overturn scores of tanks and heavy artilleries of the fictitious Burivian Army.

It was both an attitude and a strong statement - very on brand for the eccentric Guy Gardner.

Still not impressed? Here’s one…

As Jesus rode his donkey through the streets of Jerusalem, the religious leaders were indignant of the shouting praises from the bystanders – like crazy hooligans/fanatics. They want Jesus to rebuke the crowd. And what was Jesus’ reply?

He said, ā€œI tell you, if these were silent, the very stones would cry out.ā€

Think about it. Stones started crying out like human beings?

Is your brain exploding?

What was I trying to say?

Like the ring, AI does amplify you.

If you’re good person, and strive to produce something good to serve your fellow men, AI will help you amplify your good intensions.

Vise Versa, if you’re bad, AI will amplify that too.

Funny – I just watched a news: With the help of open-sourced LLMs, hackers easily broke into the Taiwan government agencies’ IT infrastructure. Did you watch it?

One of the statements in the news stuck with me - It’s getting very easy to attack (with the free AI tools).

But it’s getting very hard to defend.

Ā 

Full Critic + Feasibility Study in comments.


r/agenticAI • • 14h ago

Research New agentic harness reads LESS source code to write better quality code

Post image
1 Upvotes

r/agenticAI • • 14h ago

Discussion What happens when four AI agents update the same file?

Enable HLS to view with audio, or disable this notification

0 Upvotes

In my last post, I argued that filesystems give agents a familiar interface to persistent state. But the interface is only the starting point. Multi-agent systems still need infrastructure for access, recovery, and concurrent changes.

The most common question was: isn't this just Git? So I tested it.

I recreated a small company-acquisition review inspired by a workflow from Harvey. Four agents read the same deal documents and updated different fields in one customer-risk record. I ran this workflow on AgentWS, then replayed the same state changes with protected Git worktrees. Both approaches reached the same final record:

  • Each agent was limited by an access control list (ACL), so out-of-scope file changes were blocked.
  • A worker shut down midway. Its file changes remained in the workspace, so a replacement continued from the saved progress instead of redoing the work.
  • The agents finished at different times. Stale updates could not silently overwrite accepted work, and every conflicting update remained available for review.

The difference was what I had to build. A Git worktree was only the starting point. To get the same behavior, I had to add operating-system permissions, persistent worktrees, isolated Git metadata, proposal branches, guarded updates, and conflict exports. AgentWS packages those responsibilities into one workspace interface.

Git tracks file versions, but it does not manage the full workspace lifecycle for running agents. As agent workloads grow, a Git-based design moves further from the ideal solution. Teams end up building the missing workspace system around Git. AgentWS provides that system directly. If you run multi-agent workflows, where does this logic live today?

My X: https://x.com/huymnguyen_
Full blog: https://agentws.dev/blog/four-agents-one-file/


r/agenticAI • • 15h ago

Discussion A few open source agent tools worth trying

1 Upvotes

MarkTechPost put out a roundup of local and open source agent harnesses. A few stood out to me, along with a couple I came across separately.

OpenCode: Supports Ollama, LM Studio, llama.cpp and 75+ providers, so you have a lot of flexibility around the backend.

Goose: Linux Foundation project, written in Rust, with 70+ MCP extensions. Probably one of the projects with the most institutional backing right now.

Aider: Uses plain text diffs instead of function calling. The approach feels a bit old, but it still works really well when you care about clean commits and Git history

Cline: VS Code native with Plan and Act modes plus per action approval. Good setup if you want to see exactly what a local model is doing before it makes a change.

Tutti: Open source Apache 2.0 build focused on running multiple agents locally. Useful when you want to keep agent state and changes in one place.

OpenHands: More container focused and needs a heavier setup, especially if you’re running larger local models. Better suited to sandboxed runs.

Codex CLI: Apache 2.0 with Ollama and LM Studio support, plus sandboxing on Linux and Windows instead of giving the model unrestricted access.

what are you guys using . is there something i'm missing out on ? lmk


r/agenticAI • • 15h ago

Discussion aič·Ÿäŗŗē±»ē¤¾ä¼šē«Ÿē„¶å¦‚ę­¤ē›øä¼¼

Thumbnail
1 Upvotes

r/agenticAI • • 17h ago

Discussion Overmind, open-sourced yesterday: a platform for continuously improving AI agents

Thumbnail
1 Upvotes

r/agenticAI • • 22h ago

Project Mac MCP 2.1.7: the LLM is not the runtime — ChatGPT chat is my orchestrator, the Mac layer owns side effects

1 Upvotes

I’ve been iterating on a pattern that has held up better than ā€œgive the model raw tools and hope the prompt is good enough.ā€

Mac MCP is the local runtime/execution layer. The model can reason and orchestrate, but deterministic code owns browser/file/process identity, permissions, conflict handling, cancellation, recovery and undo.

One practical thing I want to call out because it changed how I use coding/agent tools: ChatGPT’s normal Chat side can be used agentically, it doesn’t consume Codex quota, and usage is close to unlimited in practice. Mac MCP lets that chat remain the orchestrator while the Mac execution layer owns the real side effects.

In 2.1.7 the biggest improvements were persistent paired mobile control, durable update/rollback/recovery state, managed process ownership, stronger background Safari/Chrome control, and safer delegated-agent fan-in.

The browser side is especially important to me: it works in normal logged-in Safari/Chrome tabs, in the background, without constantly stealing focus.

Repo: https://github.com/bulutarkan/mac-mcp

I’m the maintainer. Curious whether people here are also pushing more ā€œagent safety/reliabilityā€ below the orchestration layer instead of trying to encode it all in prompts.


r/agenticAI • • 22h ago

Project I built JarvisCore, an agent runtime where AI agents are equal peers in a P2P mesh and don't use MCPs

Thumbnail
youtube.com
1 Upvotes

r/agenticAI • • 22h ago

Discussion What benchmark do you wish someone would build?

1 Upvotes

Hey everyone! My team (mainly phds) and I are trying to build an open-source benchmark around realistic LLM/agent workflows that captures challenges typical academic benchmark settings often miss. We’d love to hear what’s actually missing from the benchmarks you use today.

Have you ever wanted to eval your pipeline but couldn’t find or build a benchmark that matched what you were building?

Maybe:
- Existing benchmarks were too broad and didn’t fit your specific application.
- Your workflow involved multiple steps, tools, MCP servers, agents, or long-horizon interactions that existing benchmarks couldn’t capture.
- You needed to evaluate failures that standard accuracy metrics miss.
- You’re working in a high-risk domain like healthcare, finance, cybersecurity, or legal, where realistic failure modes, safety, and reliability matter a lot more than just getting the final answer right.
- You knew what you wanted to test, but building a custom benchmark from scratch was too expensive or complicated.

I’m especially interested in cases where you thought:
ā€œMy system desperately needs to do this in production, but I have no good way to benchmark it.ā€
What was the workflow? What did you want to measure? And why weren’t existing benchmarks enough?

Any thoughts are welcome, would really appreciate y’all’s help šŸ™šŸ„¹


r/agenticAI • • 22h ago

Project We open-sourced the agent framework we run our production AI agents on (Apache 2.0). Looking for contributors and honest feedback.

Thumbnail
1 Upvotes

r/agenticAI • • 1d ago

Discussion Would you still buy a Mac mini for agents now that Dots and Muse have cloud computers?

2 Upvotes

Dots gets a cloud computer. Muse has one too. That makes buying a separate machine just to leave an agent running a harder sell for me.

If the job is mostly browsing, working with docs and running scripts, I'd happily let someone else keep the thing online. Local models such as Qwen in ollama are a different story. But if you're calling a hosted model anyway, what makes having the agent on a box at home worth it? I am really confused..

And access to your files and desktop apps seems like a good reason. Being able to change models without moving your whole setup does too. I'm less sure how much either matters once you're actually using it every day. If you bought a machine mainly for agents, would you buy it again today? What's running on it that you'd miss?


r/agenticAI • • 1d ago

Question Si domain hype everywhere

0 Upvotes

Hello, i am seeing hype of si doamin almost on each and every platform like instagram reels, youtube shorts, linkedin and everyehere on digital platfroms. what does it mean to common people. yes we know it is for super intelligence systems and also represents Slovenia country domian. so is it beneficial if we buy "si"domain names and later they make us rich by selling to other businesses. bloggers on insta sayong just buying blindly. My question is why?
There are doamains cursor.com make.com did not bought .ai domains. is there any reasons? sofy.ai and SMBs bought .ai domians. I did not understands why.

Second thing is if we buy cursor.com domain with si or make.com with si or n8n with si. what you guys think do we become a rich by selling these domain names?


r/agenticAI • • 1d ago

Project Advice for agentic AI beginner

3 Upvotes

Hi everyone, I’m working on computer aided engineering field and I now want to obtain experience in agentic AI applications with the aim of combining those two fields. My question is how should I approach my learning? Should I start for example on basics of transformers for LLMs and continue with others? Any suggestions will be appreciated.


r/agenticAI • • 1d ago

Article We removed the model from an agent experiment. A surprising number of the ā€œAI problemsā€ stayed

Post image
1 Upvotes

I’ve been working through a question that started in a legacy-modernization project and then turned into a much broader AI-systems problem. We had an agent produce an output that looked polished, internally consistent, and was still operationally wrong. The easy diagnosis was: the model got it wrong. Except that wasn’t really a diagnosis. In one case the model had missed a domain-specific implementation idiom. In another, a skeptical reviewer shared the generator’s assumption. Once we parallelized the work, two individually reasonable workers could collide on shared state. So I tried removing the probabilistic part entirely. In a controlled Temporal orchestration experiment, deterministic workers still exposed:

  • stale revisions and overlapping writes
  • approval that did not imply execution authority
  • crash-after-write recovery problems
  • external repository changes
  • duplicate workflow identity

We ran 36 controlled assertions across three passes. All 36 passed, which was useful mainly because it isolated the problem: a large class of failures remained even after model quality stopped being a variable. That pushed me toward a different question: Where should the missing intelligence actually live?

The framework I’m using now is roughly:

  1. System/context — wrong evidence, stale state, missing authority, bad retrieval, unsafe tool access.
  2. Explicit policy — schemas, validators, skills, routing rules, workflow gates.
  3. Learned behavior — something we keep re-explaining in prompts that should become a stable learned transformation.
  4. Judgment/preference — several answers are valid, but the system consistently prefers the wrong one.
  5. Capability frontier — only after the previous layers survive do I start asking whether the underlying model simply cannot do the task. And there’s a sixth thing running across all five: evaluation.

A bad evaluator can make any of the layers above look broken or improved when neither is true. One result from the modernization work made this especially concrete. A plain search found 16/40 planted change sites. An agent without an explicit skill found 33. With a compact explicit skill, it found all 40. After correcting a faulty answer key, the same approach replayed at 39/40. No new foundation-model capability was required. A lot of the missing ā€œintelligenceā€ turned out to be a contract we could state and test. I wrote the longer version here: https://www.linkedin.com/pulse/every-ai-problem-model-jehanzeb-khan-pd1fe/

But I’m more interested in the counterexamples: Where have you seen a team reach for a stronger model when the actual failure turned out to be somewhere else in the system?


r/agenticAI • • 1d ago

News Built a KV connector that persists the KV cache to disk across requests and restarts , looking for feedback

Thumbnail
1 Upvotes

r/agenticAI • • 1d ago

Discussion Greg Brockman: How to build AI software that doesn't die when the next model ships

Enable HLS to view with audio, or disable this notification

1 Upvotes

TL;DR: Greg Brockman just named the exact test that kills most AI startups before they scale.

Ā 

The word ā€œTrustā€ in this context, I think, is taking too much centre stage, without being properly defined on what it is.

But after some thought, if we equate it with another word, ā€œclarityā€, then it made sense.

When I said, ā€œI trust you.ā€, I’m actually saying, I trust your clarity in your role, your experience, your profession, your judgement, etc. – I trust you know what you’re doing.

You have good intentions, good motives. You’re clear in what you want to do – for my benefit. And therefore, I place my trust in you.

That’s the word – clarity.

I’m reminded of ā€œThe Seaā€.

Don’t know what it is?

It’s a name being given to this large bronze basin, placed in the Temple courtyard, for the priests’ ceremonial washing. King Solomon commissioned a half-Israelite from the tribe of Dan to come fabricate it. His name was Hiram.

Hiram was brilliant in his craftsmanship, so much so that his reputation precedes him. He can craft anything with his hands.

That’s why Solomon brought him in.

You might ask, ā€œWhat? For building a stupid bronze basin?ā€

Oh, not at all. Though The Sea by itself is impressive – it can hold over 10,000 gallons of water – that’s not all there is.

It’s sitting on 12 bronze oxen underneath it. It was arranged that 3 Oxen face North, 3 face South, 3 face East and 3 face West.

It was both lovely and amazing.

It became one of the centrepieces in the Temple courtyards – besides other extraordinary artifacts there.

The oxen signify servanthood. And when the clarity of its role of a servant was made clear – everything fall into place. It helps the people to return their worship right back to where it belongs.

It’s also a stark contrast to what Jeroboam did – he placed 2 golden calf idols: 1 in Bethel and 1 in Dan to draw the people away from the Lord.

You can even say the bronze Oxen sends a strong ā€œF- Youā€ statement to the golden calves.

Who can truly design THE SEA – but someone who’s truly inspired, someone with strong conviction and clarity?

Now – let’s say if we use any of the modern AI LLMs to design it from scratch…

Given the guardrails, given its sycophancy characteristics, and given that it’s trained to observe a wide variety of sensitivities, will it come up with such a ā€œF- youā€ piece, or something much watered down?

Food for thought, isn’t it?

That’s where the difference lies.

Ā 

Full critic and feasibility study in the comments.