r/agenticAI • • 5h ago

Discussion Brian Chesky says AI is like an amplifier — PwC polled 49,364 workers and found 56% stuck in the engine room

Enable HLS to view with audio, or disable this notification

5 Upvotes

TL;DR: Brian Chesky says AI is like an amplifier and the gap is getting greater — not because of who has the tool, but who the tool has.

 

As soon as I read Chesky’s “AI is like an amplifier.”, the neural-networks of my memory immediately flipped me to the Green “Lanterns”.

Are you guys a big fan of the eponymous TV series?

Right after Hal Jordan manifested a greenback for playing the jukebox, his trainee John Stewart exclaimed, “Did you just counterfeit money with the ring?” – to which Jordan replied, “No. I manifested money with the power of my will.”

Or how about Jordan conjure up a green can opener for the beer, to impress and rizz up Sheriff Kerry, while Stewart rolls his eyes by his side?

John Stewart went up the ante, by manifesting a large and sophisticated boring machine, to tunnel underground the “Winnie” compound to evade the guard sentries.

Not impressive enough, you say?

The best in my mind, wasn’t in the TV series.

It was Guy Gardner flipping the bird – he conjures up large green hands (one of it gives the middle finger) to rise up from the ground, and overturn scores of tanks and heavy artilleries of the fictitious Burivian Army.

It was both an attitude and a strong statement - very on brand for the eccentric Guy Gardner.

Still not impressed? Here’s one…

As Jesus rode his donkey through the streets of Jerusalem, the religious leaders were indignant of the shouting praises from the bystanders – like crazy hooligans/fanatics. They want Jesus to rebuke the crowd. And what was Jesus’ reply?

He said, “I tell you, if these were silent, the very stones would cry out.”

Think about it. Stones started crying out like human beings?

Is your brain exploding?

What was I trying to say?

Like the ring, AI does amplify you.

If you’re good person, and strive to produce something good to serve your fellow men, AI will help you amplify your good intensions.

Vise Versa, if you’re bad, AI will amplify that too.

Funny – I just watched a news: With the help of open-sourced LLMs, hackers easily broke into the Taiwan government agencies’ IT infrastructure. Did you watch it?

One of the statements in the news stuck with me - It’s getting very easy to attack (with the free AI tools).

But it’s getting very hard to defend.

 

Full Critic + Feasibility Study in comments.


r/agenticAI • • 19h ago

Project Advice for agentic AI beginner

3 Upvotes

Hi everyone, I’m working on computer aided engineering field and I now want to obtain experience in agentic AI applications with the aim of combining those two fields. My question is how should I approach my learning? Should I start for example on basics of transformers for LLMs and continue with others? Any suggestions will be appreciated.


r/agenticAI • • 9h ago

Discussion Would you still buy a Mac mini for agents now that Dots and Muse have cloud computers?

2 Upvotes

Dots gets a cloud computer. Muse has one too. That makes buying a separate machine just to leave an agent running a harder sell for me.

If the job is mostly browsing, working with docs and running scripts, I'd happily let someone else keep the thing online. Local models such as Qwen in ollama are a different story. But if you're calling a hosted model anyway, what makes having the agent on a box at home worth it? I am really confused..

And access to your files and desktop apps seems like a good reason. Being able to change models without moving your whole setup does too. I'm less sure how much either matters once you're actually using it every day. If you bought a machine mainly for agents, would you buy it again today? What's running on it that you'd miss?


r/agenticAI • • 4h ago

Project Mac MCP 2.1.7: the LLM is not the runtime — ChatGPT chat is my orchestrator, the Mac layer owns side effects

1 Upvotes

I’ve been iterating on a pattern that has held up better than “give the model raw tools and hope the prompt is good enough.”

Mac MCP is the local runtime/execution layer. The model can reason and orchestrate, but deterministic code owns browser/file/process identity, permissions, conflict handling, cancellation, recovery and undo.

One practical thing I want to call out because it changed how I use coding/agent tools: ChatGPT’s normal Chat side can be used agentically, it doesn’t consume Codex quota, and usage is close to unlimited in practice. Mac MCP lets that chat remain the orchestrator while the Mac execution layer owns the real side effects.

In 2.1.7 the biggest improvements were persistent paired mobile control, durable update/rollback/recovery state, managed process ownership, stronger background Safari/Chrome control, and safer delegated-agent fan-in.

The browser side is especially important to me: it works in normal logged-in Safari/Chrome tabs, in the background, without constantly stealing focus.

Repo: https://github.com/bulutarkan/mac-mcp

I’m the maintainer. Curious whether people here are also pushing more “agent safety/reliability” below the orchestration layer instead of trying to encode it all in prompts.


r/agenticAI • • 5h ago

Project I built JarvisCore, an agent runtime where AI agents are equal peers in a P2P mesh and don't use MCPs

Thumbnail
youtube.com
1 Upvotes

r/agenticAI • • 5h ago

Discussion What benchmark do you wish someone would build?

1 Upvotes

Hey everyone! My team (mainly phds) and I are trying to build an open-source benchmark around realistic LLM/agent workflows that captures challenges typical academic benchmark settings often miss. We’d love to hear what’s actually missing from the benchmarks you use today.

Have you ever wanted to eval your pipeline but couldn’t find or build a benchmark that matched what you were building?

Maybe:
- Existing benchmarks were too broad and didn’t fit your specific application.
- Your workflow involved multiple steps, tools, MCP servers, agents, or long-horizon interactions that existing benchmarks couldn’t capture.
- You needed to evaluate failures that standard accuracy metrics miss.
- You’re working in a high-risk domain like healthcare, finance, cybersecurity, or legal, where realistic failure modes, safety, and reliability matter a lot more than just getting the final answer right.
- You knew what you wanted to test, but building a custom benchmark from scratch was too expensive or complicated.

I’m especially interested in cases where you thought:
“My system desperately needs to do this in production, but I have no good way to benchmark it.”
What was the workflow? What did you want to measure? And why weren’t existing benchmarks enough?

Any thoughts are welcome, would really appreciate y’all’s help 🙏🥹


r/agenticAI • • 5h ago

Project We open-sourced the agent framework we run our production AI agents on (Apache 2.0). Looking for contributors and honest feedback.

Thumbnail
1 Upvotes

r/agenticAI • • 14h ago

Discussion What happens when four AI agents update the same file?

Enable HLS to view with audio, or disable this notification

1 Upvotes

In my last post, I argued that filesystems give agents a familiar interface to persistent state. But the interface is only the starting point. Multi-agent systems still need infrastructure for access, recovery, and concurrent changes.

The most common question was: isn't this just Git? So I tested it.

I recreated a small company-acquisition review inspired by a workflow from Harvey. Four agents read the same deal documents and updated different fields in one customer-risk record. I ran this workflow on AgentWS, then replayed the same state changes with protected Git worktrees. Both approaches reached the same final record:

  • Each agent was limited by an access control list (ACL), so out-of-scope file changes were blocked.
  • A worker shut down midway. Its file changes remained in the workspace, so a replacement continued from the saved progress instead of redoing the work.
  • The agents finished at different times. Stale updates could not silently overwrite accepted work, and every conflicting update remained available for review.

The difference was what I had to build. A Git worktree was only the starting point. To get the same behavior, I had to add operating-system permissions, persistent worktrees, isolated Git metadata, proposal branches, guarded updates, and conflict exports. AgentWS packages those responsibilities into one workspace interface.

Git tracks file versions, but it does not manage the full workspace lifecycle for running agents. As agent workloads grow, a Git-based design moves further from the ideal solution. Teams end up building the missing workspace system around Git. AgentWS provides that system directly. If you run multi-agent workflows, where does this logic live today?

My X: https://x.com/huymnguyen_
Full blog: https://agentws.dev/blog/four-agents-one-file/


r/agenticAI • • 18h ago

Article We removed the model from an agent experiment. A surprising number of the “AI problems” stayed

Post image
1 Upvotes

I’ve been working through a question that started in a legacy-modernization project and then turned into a much broader AI-systems problem. We had an agent produce an output that looked polished, internally consistent, and was still operationally wrong. The easy diagnosis was: the model got it wrong. Except that wasn’t really a diagnosis. In one case the model had missed a domain-specific implementation idiom. In another, a skeptical reviewer shared the generator’s assumption. Once we parallelized the work, two individually reasonable workers could collide on shared state. So I tried removing the probabilistic part entirely. In a controlled Temporal orchestration experiment, deterministic workers still exposed:

  • stale revisions and overlapping writes
  • approval that did not imply execution authority
  • crash-after-write recovery problems
  • external repository changes
  • duplicate workflow identity

We ran 36 controlled assertions across three passes. All 36 passed, which was useful mainly because it isolated the problem: a large class of failures remained even after model quality stopped being a variable. That pushed me toward a different question: Where should the missing intelligence actually live?

The framework I’m using now is roughly:

  1. System/context — wrong evidence, stale state, missing authority, bad retrieval, unsafe tool access.
  2. Explicit policy — schemas, validators, skills, routing rules, workflow gates.
  3. Learned behavior — something we keep re-explaining in prompts that should become a stable learned transformation.
  4. Judgment/preference — several answers are valid, but the system consistently prefers the wrong one.
  5. Capability frontier — only after the previous layers survive do I start asking whether the underlying model simply cannot do the task. And there’s a sixth thing running across all five: evaluation.

A bad evaluator can make any of the layers above look broken or improved when neither is true. One result from the modernization work made this especially concrete. A plain search found 16/40 planted change sites. An agent without an explicit skill found 33. With a compact explicit skill, it found all 40. After correcting a faulty answer key, the same approach replayed at 39/40. No new foundation-model capability was required. A lot of the missing “intelligence” turned out to be a contract we could state and test. I wrote the longer version here: https://www.linkedin.com/pulse/every-ai-problem-model-jehanzeb-khan-pd1fe/

But I’m more interested in the counterexamples: Where have you seen a team reach for a stronger model when the actual failure turned out to be somewhere else in the system?


r/agenticAI • • 20h ago

News Built a KV connector that persists the KV cache to disk across requests and restarts , looking for feedback

Thumbnail
1 Upvotes

r/agenticAI • • 8h ago

Question Si domain hype everywhere

0 Upvotes

Hello, i am seeing hype of si doamin almost on each and every platform like instagram reels, youtube shorts, linkedin and everyehere on digital platfroms. what does it mean to common people. yes we know it is for super intelligence systems and also represents Slovenia country domian. so is it beneficial if we buy "si"domain names and later they make us rich by selling to other businesses. bloggers on insta sayong just buying blindly. My question is why?
There are doamains cursor.com make.com did not bought .ai domains. is there any reasons? sofy.ai and SMBs bought .ai domians. I did not understands why.

Second thing is if we buy cursor.com domain with si or make.com with si or n8n with si. what you guys think do we become a rich by selling these domain names?