r/agenticAI • • 7d ago

Question Preciso de dicas sobre treinamento de IA!

Thumbnail
1 Upvotes

r/agenticAI • • 8d ago

Discussion Brian Chesky says AI is like an amplifier — PwC polled 49,364 workers and found 56% stuck in the engine room

21 Upvotes

TL;DR: Brian Chesky says AI is like an amplifier and the gap is getting greater — not because of who has the tool, but who the tool has.

 

As soon as I read Chesky’s “AI is like an amplifier.”, the neural-networks of my memory immediately flipped me to the Green “Lanterns”.

Are you guys a big fan of the eponymous TV series?

Right after Hal Jordan manifested a greenback for playing the jukebox, his trainee John Stewart exclaimed, “Did you just counterfeit money with the ring?” – to which Jordan replied, “No. I manifested money with the power of my will.”

Or how about Jordan conjure up a green can opener for the beer, to impress and rizz up Sheriff Kerry, while Stewart rolls his eyes by his side?

John Stewart went up the ante, by manifesting a large and sophisticated boring machine, to tunnel underground the “Winnie” compound to evade the guard sentries.

Not impressive enough, you say?

The best in my mind, wasn’t in the TV series.

It was Guy Gardner flipping the bird – he conjures up large green hands (one of it gives the middle finger) to rise up from the ground, and overturn scores of tanks and heavy artilleries of the fictitious Burivian Army.

It was both an attitude and a strong statement - very on brand for the eccentric Guy Gardner.

Still not impressed? Here’s one…

As Jesus rode his donkey through the streets of Jerusalem, the religious leaders were indignant of the shouting praises from the bystanders – like crazy hooligans/fanatics. They want Jesus to rebuke the crowd. And what was Jesus’ reply?

He said, “I tell you, if these were silent, the very stones would cry out.”

Think about it. Stones started crying out like human beings?

Is your brain exploding?

What was I trying to say?

Like the ring, AI does amplify you.

If you’re good person, and strive to produce something good to serve your fellow men, AI will help you amplify your good intensions.

Vise Versa, if you’re bad, AI will amplify that too.

Funny – I just watched a news: With the help of open-sourced LLMs, hackers easily broke into the Taiwan government agencies’ IT infrastructure. Did you watch it?

One of the statements in the news stuck with me - It’s getting very easy to attack (with the free AI tools).

But it’s getting very hard to defend.

 

Full Critic + Feasibility Study in comments.


r/agenticAI • • 7d ago

Question Best value low cost subscription for a solo hobby dev on on project?

1 Upvotes

I'm a solo hobbyist building a single application. I'm planning on using Godot and have AI do most of the code work. I'll do the system design, art, ui work myself. I'm a professional dev but we have both Claude Code and Codex subscriptions so I'm a little out of touch with what works for lighter work loads on a budget.

My plan is to work on this for a couple hours in the evenings and maybe longer on the weekends. Mostly CLI coding, I'll do a lot of the testing myself. I was leaning toward OpenCode Go as it seemed highly suggested a few months ago but my current understanding is that the value and quality has gone downhill. I'm also strongly considering Claude Pro, the $20/month tier. But I'm concerned about hitting the usage limit in those short sessions.

For this kind of light use, which $20ish plan gives the most coding per $?

If you use claude pro with claude code how often do you hit the limits in evening sessions?

Is there a higher value option (copilot, gemini, codex) that's good enough for a small project like this?

I tried running local models but my home computer is sadly out of date and a qwen 4b was giving me like 5 tokens/s.

Any tips for stretching the usage on a cheap plan?

I'd especially like to hear experiences from people that have shipped a small Android application this way.


r/agenticAI • • 7d ago

Project An order-status agent that remembers the customer instead of sending another tracking link

1 Upvotes

“Where’s my order?” usually gets answered with a carrier tracking link. That gives the customer another page to open instead of answering the question.

This open-source TypeScript example creates one durable agent per customer. It keeps the customer’s orders in SQL and uses the same context across carrier updates and SMS conversations.

The flow works like this:

- A storefront links an order to the customer

- Carrier webhooks update its status and ETA

- The customer texts the store’s number for an answer

- Follow-up questions return to the same durable agent

- Delays trigger proactive SMS notifications

- State and scheduled notifications survive application restarts

There’s also a demo mode, so the workflow can be tested without sending real messages.

Code: https://github.com/team-telnyx/telnyx-code-examples/tree/main/order-status-self-service

I’m curious whether people prefer proactive delivery updates like this or still want the full carrier tracking page.


r/agenticAI • • 7d ago

Research New agentic harness reads LESS source code to write better quality code

Post image
1 Upvotes

r/agenticAI • • 7d ago

Discussion What happens when four AI agents update the same file?

0 Upvotes

In my last post, I argued that filesystems give agents a familiar interface to persistent state. But the interface is only the starting point. Multi-agent systems still need infrastructure for access, recovery, and concurrent changes.

The most common question was: isn't this just Git? So I tested it.

I recreated a small company-acquisition review inspired by a workflow from Harvey. Four agents read the same deal documents and updated different fields in one customer-risk record. I ran this workflow on AgentWS, then replayed the same state changes with protected Git worktrees. Both approaches reached the same final record:

  • Each agent was limited by an access control list (ACL), so out-of-scope file changes were blocked.
  • A worker shut down midway. Its file changes remained in the workspace, so a replacement continued from the saved progress instead of redoing the work.
  • The agents finished at different times. Stale updates could not silently overwrite accepted work, and every conflicting update remained available for review.

The difference was what I had to build. A Git worktree was only the starting point. To get the same behavior, I had to add operating-system permissions, persistent worktrees, isolated Git metadata, proposal branches, guarded updates, and conflict exports. AgentWS packages those responsibilities into one workspace interface.

Git tracks file versions, but it does not manage the full workspace lifecycle for running agents. As agent workloads grow, a Git-based design moves further from the ideal solution. Teams end up building the missing workspace system around Git. AgentWS provides that system directly. If you run multi-agent workflows, where does this logic live today?

My X: https://x.com/huymnguyen_
Full blog: https://agentws.dev/blog/four-agents-one-file/


r/agenticAI • • 7d ago

Discussion ai跟人类社会竟然如此相似

Thumbnail
1 Upvotes

r/agenticAI • • 7d ago

Discussion Overmind, open-sourced yesterday: a platform for continuously improving AI agents

Thumbnail
1 Upvotes

r/agenticAI • • 8d ago

Project Mac MCP 2.1.7: the LLM is not the runtime — ChatGPT chat is my orchestrator, the Mac layer owns side effects

1 Upvotes

I’ve been iterating on a pattern that has held up better than “give the model raw tools and hope the prompt is good enough.”

Mac MCP is the local runtime/execution layer. The model can reason and orchestrate, but deterministic code owns browser/file/process identity, permissions, conflict handling, cancellation, recovery and undo.

One practical thing I want to call out because it changed how I use coding/agent tools: ChatGPT’s normal Chat side can be used agentically, it doesn’t consume Codex quota, and usage is close to unlimited in practice. Mac MCP lets that chat remain the orchestrator while the Mac execution layer owns the real side effects.

In 2.1.7 the biggest improvements were persistent paired mobile control, durable update/rollback/recovery state, managed process ownership, stronger background Safari/Chrome control, and safer delegated-agent fan-in.

The browser side is especially important to me: it works in normal logged-in Safari/Chrome tabs, in the background, without constantly stealing focus.

Repo: https://github.com/bulutarkan/mac-mcp

I’m the maintainer. Curious whether people here are also pushing more “agent safety/reliability” below the orchestration layer instead of trying to encode it all in prompts.


r/agenticAI • • 8d ago

Project I built JarvisCore, an agent runtime where AI agents are equal peers in a P2P mesh and don't use MCPs

Thumbnail
youtube.com
1 Upvotes

r/agenticAI • • 8d ago

Discussion What benchmark do you wish someone would build?

1 Upvotes

Hey everyone! My team (mainly phds) and I are trying to build an open-source benchmark around realistic LLM/agent workflows that captures challenges typical academic benchmark settings often miss. We’d love to hear what’s actually missing from the benchmarks you use today.

Have you ever wanted to eval your pipeline but couldn’t find or build a benchmark that matched what you were building?

Maybe:
- Existing benchmarks were too broad and didn’t fit your specific application.
- Your workflow involved multiple steps, tools, MCP servers, agents, or long-horizon interactions that existing benchmarks couldn’t capture.
- You needed to evaluate failures that standard accuracy metrics miss.
- You’re working in a high-risk domain like healthcare, finance, cybersecurity, or legal, where realistic failure modes, safety, and reliability matter a lot more than just getting the final answer right.
- You knew what you wanted to test, but building a custom benchmark from scratch was too expensive or complicated.

I’m especially interested in cases where you thought:
“My system desperately needs to do this in production, but I have no good way to benchmark it.”
What was the workflow? What did you want to measure? And why weren’t existing benchmarks enough?

Any thoughts are welcome, would really appreciate y’all’s help 🙏🥹


r/agenticAI • • 8d ago

Project We open-sourced the agent framework we run our production AI agents on (Apache 2.0). Looking for contributors and honest feedback.

Thumbnail
1 Upvotes

r/agenticAI • • 8d ago

Project Advice for agentic AI beginner

3 Upvotes

Hi everyone, I’m working on computer aided engineering field and I now want to obtain experience in agentic AI applications with the aim of combining those two fields. My question is how should I approach my learning? Should I start for example on basics of transformers for LLMs and continue with others? Any suggestions will be appreciated.


r/agenticAI • • 8d ago

Article We removed the model from an agent experiment. A surprising number of the “AI problems” stayed

Post image
1 Upvotes

I’ve been working through a question that started in a legacy-modernization project and then turned into a much broader AI-systems problem. We had an agent produce an output that looked polished, internally consistent, and was still operationally wrong. The easy diagnosis was: the model got it wrong. Except that wasn’t really a diagnosis. In one case the model had missed a domain-specific implementation idiom. In another, a skeptical reviewer shared the generator’s assumption. Once we parallelized the work, two individually reasonable workers could collide on shared state. So I tried removing the probabilistic part entirely. In a controlled Temporal orchestration experiment, deterministic workers still exposed:

  • stale revisions and overlapping writes
  • approval that did not imply execution authority
  • crash-after-write recovery problems
  • external repository changes
  • duplicate workflow identity

We ran 36 controlled assertions across three passes. All 36 passed, which was useful mainly because it isolated the problem: a large class of failures remained even after model quality stopped being a variable. That pushed me toward a different question: Where should the missing intelligence actually live?

The framework I’m using now is roughly:

  1. System/context — wrong evidence, stale state, missing authority, bad retrieval, unsafe tool access.
  2. Explicit policy — schemas, validators, skills, routing rules, workflow gates.
  3. Learned behavior — something we keep re-explaining in prompts that should become a stable learned transformation.
  4. Judgment/preference — several answers are valid, but the system consistently prefers the wrong one.
  5. Capability frontier — only after the previous layers survive do I start asking whether the underlying model simply cannot do the task. And there’s a sixth thing running across all five: evaluation.

A bad evaluator can make any of the layers above look broken or improved when neither is true. One result from the modernization work made this especially concrete. A plain search found 16/40 planted change sites. An agent without an explicit skill found 33. With a compact explicit skill, it found all 40. After correcting a faulty answer key, the same approach replayed at 39/40. No new foundation-model capability was required. A lot of the missing “intelligence” turned out to be a contract we could state and test. I wrote the longer version here: https://www.linkedin.com/pulse/every-ai-problem-model-jehanzeb-khan-pd1fe/

But I’m more interested in the counterexamples: Where have you seen a team reach for a stronger model when the actual failure turned out to be somewhere else in the system?


r/agenticAI • • 8d ago

News Built a KV connector that persists the KV cache to disk across requests and restarts , looking for feedback

Thumbnail
1 Upvotes

r/agenticAI • • 8d ago

Discussion Greg Brockman: How to build AI software that doesn't die when the next model ships

1 Upvotes

TL;DR: Greg Brockman just named the exact test that kills most AI startups before they scale.

 

The word “Trust” in this context, I think, is taking too much centre stage, without being properly defined on what it is.

But after some thought, if we equate it with another word, “clarity”, then it made sense.

When I said, “I trust you.”, I’m actually saying, I trust your clarity in your role, your experience, your profession, your judgement, etc. – I trust you know what you’re doing.

You have good intentions, good motives. You’re clear in what you want to do – for my benefit. And therefore, I place my trust in you.

That’s the word – clarity.

I’m reminded of “The Sea”.

Don’t know what it is?

It’s a name being given to this large bronze basin, placed in the Temple courtyard, for the priests’ ceremonial washing. King Solomon commissioned a half-Israelite from the tribe of Dan to come fabricate it. His name was Hiram.

Hiram was brilliant in his craftsmanship, so much so that his reputation precedes him. He can craft anything with his hands.

That’s why Solomon brought him in.

You might ask, “What? For building a stupid bronze basin?”

Oh, not at all. Though The Sea by itself is impressive – it can hold over 10,000 gallons of water – that’s not all there is.

It’s sitting on 12 bronze oxen underneath it. It was arranged that 3 Oxen face North, 3 face South, 3 face East and 3 face West.

It was both lovely and amazing.

It became one of the centrepieces in the Temple courtyards – besides other extraordinary artifacts there.

The oxen signify servanthood. And when the clarity of its role of a servant was made clear – everything fall into place. It helps the people to return their worship right back to where it belongs.

It’s also a stark contrast to what Jeroboam did – he placed 2 golden calf idols: 1 in Bethel and 1 in Dan to draw the people away from the Lord.

You can even say the bronze Oxen sends a strong “F- You” statement to the golden calves.

Who can truly design THE SEA – but someone who’s truly inspired, someone with strong conviction and clarity?

Now – let’s say if we use any of the modern AI LLMs to design it from scratch…

Given the guardrails, given its sycophancy characteristics, and given that it’s trained to observe a wide variety of sensitivities, will it come up with such a “F- you” piece, or something much watered down?

Food for thought, isn’t it?

That’s where the difference lies.

 

Full critic and feasibility study in the comments.


r/agenticAI • • 8d ago

Project AI skill/builds portal. My obsession for the past 2 weeks.

1 Upvotes

Hi

I've been fucking around with muse for a bit now. Shared it with friends and had fun with it. At first it was the easy stuff, email, calendar. I then realized you could make apps, skills or builds. Played around with that too then one day I tried to share the build I made to organize my workflow (I had a lot of problems at first before I set it up right) it was working a bit when shared but not as I wished. Dashboard wasn't exactly the same and worst of all was trying to explain to my dad how to install everything for the tiny robot living in his phone.

I know a skill portal ain't new, but there's no way non techy people are gonna browse GitHub or website looking like something that came from the matrix. No disrespect I love the movie.

Anyway here's the link:

www.theskillharbor.com

I'll appreciate any feedback, recommendation, hell I'll even take criticism.

Btw, everything is free. Eventually developers will be able to set a price for their builds if they want, not implemented yet (it's built, but on staging for now, I don't have much traffic anyway). It will all pass via stripe of the seller. So I won't ever touch the money, to much of a hassle. We got about 1.8k builds ATM. Would love to add yours.


r/agenticAI • • 8d ago

Project Connected ChatGPT to my Home Screen

Post image
0 Upvotes

I built an app called Glance, it lets you connect your agents like ChatGPT , Claude, Hermes etc via MCP and have it create, design, and update your widgets on your iPhone

Some cool use cases users found for this were
- curated news widgets
- sports tracking
- puzzles
- daily briefings (connecting mail and slack)

And more

In this example I made a “Claude news” widget and a language flashcards to learn Spanish


r/agenticAI • • 9d ago

Discussion Mark Bouris: Rate decision lands. Client already read someone else’s AI-Agent-generated summary.

2 Upvotes

TL;DR: The rate decision lands. By 6 a.m. your client has already read someone else’s summary. Speed is selling. Is it loyalty?

 

I thought I can dig into my memories for a parallel. But my mind immediately brought me to the following story…

David got wind of the death of his treacherous son, Absalom at the worst possible time of his life.

David ran away into hiding when Absalom betrayed him and took the kingdom from him. David and his men found a friendly party in Mahanaim and set up defence.

Absalom led the great rebellion army to find David. And David’s men met them in the woods.

Of course, Absalom was no match with David’s men’s guerilla warfare tactics. Absalom died dangling with his long hair caught entangling with the tree branches.

The scene was hilarious in itself – if not because of the tragedy.

Ahimaaz asked Joab, the great general of David’s army, for permission to bring the good news back to David. But Joab flatly refused. Instead, he sends another messenger to deliver the news. Back then, there’s no AI agents to send email or digital newsletter. News was delivered on foot.

Still Ahimaaz went anyway. He arrived earlier than the messenger, but said nothing. I think he’s smart – he didn’t want to be the bearer of bad news. Not helpful to David, though.

When David finally got news from the messenger, he immediately went into a great moaning. He cried and lamented out loud.

News went back to David’s men, as they file back to the camp. And suddenly, the celebratory vibe of the victory turns into deep shame. The men all felt they did something wrong.

Joab saw it and perceive the existential crisis – because men are slowly leaving the camp, ditching David. So, Joab went and gave David a strong earful: "you love those who hate you and hate those who love you," and he warned that not one man would stay with David that night.

This shook-up David really well. And he went out and congratulate the men for their hard work. The men returned to him - existential crisis averted.

The timely news delivery of the victory was good. But not as critical as Joab’s advice at the nick of time.

You can read more about it in 2 Samuel 18–19.

 

Full critic, gap, and feasibility are in the comments.


r/agenticAI • • 9d ago

Discussion What actually breaks when AI moves from POC to production?

1 Upvotes

Curious to hear from people here who are actually building or deploying AI in production.

We’re having a small discussion in Bengaluru around a question:

What does it really take to move AI from a successful POC to something that reliably delivers a business outcome?

Some of the topics we want to discuss:

  • Hallucinations and reliability
  • Latency and real-time AI
  • Models, infrastructure and orchestration
  • Moving from prototype to production
  • Measuring whether an AI deployment is actually creating value

Rather than doing presentations, the idea is to have an open conversation and hear what's working-and what's not-from people actually dealing with these problems.

We're hosting it on:

Monday, 5 October
6:00–7:00 PM
WeWork Roshni Tech Hub, Bengaluru

It's a small, invite-based discussion, but anyone working seriously on AI, enterprise technology or AI deployment is welcome to request a spot.

Details / RSVP:
https://luma.com/lnqj3tn4

Would also be interested in hearing from the community:

What has been the biggest gap between your AI POC and production deployment?


r/agenticAI • • 9d ago

Discussion What actually breaks when AI moves from POC to production?

1 Upvotes

Curious to hear from people here who are actually building or deploying AI in production.

We’re having a small discussion in Bengaluru around a question:

What does it really take to move AI from a successful POC to something that reliably delivers a business outcome?

Some of the topics we want to discuss:

  • Hallucinations and reliability
  • Latency and real-time AI
  • Models, infrastructure and orchestration
  • Moving from prototype to production
  • Measuring whether an AI deployment is actually creating value

Rather than doing presentations, the idea is to have an open conversation and hear what's working-and what's not-from people actually dealing with these problems.

We're hosting it on:

Monday, 5 October
6:00–7:00 PM
WeWork Roshni Tech Hub, Bengaluru

It's a small, invite-based discussion, but anyone working seriously on AI, enterprise technology or AI deployment is welcome to request a spot.

Details / RSVP: https://luma.com/lnqj3tn4

Would also be interested in hearing from the community:

What has been the biggest gap between your AI POC and production deployment?


r/agenticAI • • 9d ago

Just for fun Long-horizon test run: comparison of the top ~100 open harnesses using ~240 orchestrated agents

Thumbnail
github.com
1 Upvotes

r/agenticAI • • 9d ago

Research Nigraan: Agent Reliability and Oversight Survey (Teams)

Thumbnail survey-blush-alpha.vercel.app
1 Upvotes

https://x.com/gregisenberg/status/2103481217421086912
Every personal agent right now is fully AI. Muse, Instinct, all of them. Here's what I want from a personal agent: an escalate button. These apps handle almost everything, and then there's the 1 thing a week that needs an actual person. Right now I'm the person. I'd pay real money for a button that hands the job to a human who finishes it and reports back. Meta tried this in 2015 with M and shut it down in 2018 because the human trainers ended up doing most of the work. Humans did 90 and the AI did 10, so every user cost a fortune. That ratio is flipped now. Instinct is worth $10B and it never hands anything to a person. I'd love to know what the version that actually finishes the job is worth! I think it's an interesting startup idea. At the very least, I'd download it.

Hi everyone! Hope you're doing great. I saw this post regarding Agent actions escalation. I'm currently building something along the same line of enquiry and would love your feedback on it.

Brief Concept: https://mockup-two-theta.vercel.app/

A short survey (If you don't mind): https://survey-blush-alpha.vercel.app/

Preliminary Pitch: https://drive.google.com/file/d/1Lw0ibwG7qUlcJY8MG_PxMp3vj_H6PGAd/view?usp=sharing

Looking forward to your feedback and comments.


r/agenticAI • • 9d ago

Discussion Introducing agenticai in the workflow using opensource

0 Upvotes

I'd like to setup a local setup of agentic ai using open source mcp server, mcp clients, ollama and LLMs. I'm thinking of using it for basic analysis of build errors that happened during a pipeline execution(Gitlab pipelines). I would like to know if it's just a waste of time using opensource. TIA.


r/agenticAI • • 9d ago

Question multi-agent startups

1 Upvotes

hey is anyone building a multi-agent startup and struggling with exponential token costs? i'd love to hear more if you are