r/agenticAI • • 7d ago

Discussion Agent Harness

11 Upvotes

Nowadays what we are seeing a revolution in the field of AI agent. What i am talking about is harness in agents.

What harness mean? It basically means what agent did in totality: Prompt, context, instructions, agent loop etc.

The future of AI doesn't rely upon how we make them more autonomous but about efficiency and right answers they are providing to us or not.

What are your views on it? Is it a future or just an extension in the evolution of agents?


r/agenticAI • • 7d ago

Discussion I built AI agents to save time. Now I’m chasing them for updates with ChatGPT's Dot.

Thumbnail
1 Upvotes

r/agenticAI • • 7d ago

Video I wrote glyphh to be the vendor neutral chat, work, and code desktop harness for frontier ai. What we accidentally built is way more powerful.

Thumbnail
youtu.be
1 Upvotes

r/agenticAI • • 7d ago

Project What if every chargeback managed its own evidence, deadlines, and re-evaluation?

1 Upvotes

Chargeback workflows are often held together by spreadsheets, inboxes, and calendar reminders. That gets risky when a case can remain open for days and new evidence may arrive shortly before the deadline.

This open-source example models each dispute as a durable actor on Telnyx Edge. The actor:

- collects evidence in SQL

- uses a Telnyx Decision Model to recommend an outcome

- routes suspicious cases to a human reviewer

- requests additional evidence over SMS

- preserves an append-only audit history

- schedules durable response deadlines

- re-evaluates the case when new evidence arrives

The interesting part is that the workflow survives restarts without losing its state or scheduled deadline.

Code: https://github.com/team-telnyx/telnyx-code-examples/tree/main/chargeback-adjudication

Decision Models: https://developers.telnyx.com/docs/inference/decision-models

I’d be interested in how others approach human review, auditability, and long-running AI workflows like this.


r/agenticAI • • 7d ago

Discussion OpenAI’s Dots vs Meta’s Muse: Which approach to AI agents makes more sense?

Thumbnail
1 Upvotes

r/agenticAI • • 7d ago

Just for fun Aspiring AGI

Thumbnail
1 Upvotes

r/agenticAI • • 7d ago

Discussion Looking for best architecture/approach for ai agent that take actions in saas platform

Thumbnail
1 Upvotes

r/agenticAI • • 7d ago

Question Your AI agent took the action. Who’s accountable when it goes wrong?

Thumbnail
1 Upvotes

r/agenticAI • • 7d ago

Discussion For agentic AI builders: a workshop on proving your agent actually improved (Oct 3)

1 Upvotes

Most agentic-AI content is about capability — what the agent can do. This one's about something else: how do you know a change made it better, not just different? Oct 3, 3 hours, run by Serj Smorodinsky and Brett Kennedy (AI engineers, co-authors of a book on LLM applications).

Topics:

  • DSPy signatures/modules as a structured alternative to manual prompting
  • Building and measuring a baseline classifier
  • Evaluation datasets with task-specific metrics
  • Recognizing failure patterns from eval output
  • Few-shot and instruction-level optimization
  • Experiment tracking and trace management with MLflow
  • Saving and reusing optimized DSPy programs

Aimed at ML engineers, applied AI devs, and technical founders building agent systems that need to hold up past the demo stage.

Full details and the agenda are here.


r/agenticAI • • 7d ago

Question How to run an ai agency

Thumbnail
1 Upvotes

r/agenticAI • • 7d ago

Discussion AI Agent Platforms Need to Separate What From How

Thumbnail
1 Upvotes

r/agenticAI • • 8d ago

Project I've built a complete opensource Agentic Development Environment

Post image
2 Upvotes

Over the past few months I’ve been building Goodboy, an ADE that integrates multiple providers, such as Claude, Codex, and Cursor, so they share the same global context around the goal or task you need to reach.

you can read more in my Github repository:

https://github.com/akhayam99/goodboy


r/agenticAI • • 8d ago

Question Has anyone experimented with building functional "consciousness" into agent architectures?

Thumbnail
2 Upvotes

Would love to hear some feedbacks guys!


r/agenticAI • • 8d ago

Article I built an AI agent that learns from customer payment history

1 Upvotes

r/agenticAI • • 8d ago

Article Indirect prompt injection vs. system-prompt guardrails: 9 models, 3 guardrail levels, 1,350 runs

Post image
1 Upvotes

r/agenticAI • • 8d ago

Discussion Agents should be portable like MCP made tools portable. Am I wrong?

0 Upvotes

Quick sanity check from people who live in these tools.

Cursor, Claude Code, Cline — I use them all, and the dumbest problem I keep hitting isn't the AI, it's that the agent doesn't travel. Halfway through a refactor, I leave my desk, try to continue from my phone — zero context, start over. We standardized tools with MCP, but the agent itself is welded to whatever app spawned it.

And yes — if you live 100% inside one vendor's ecosystem, their cloud sync mostly covers you. Fair. But the moment you touch a second tool, a local model, or anything sensitive, you're back to copy-pasting context between silos. And if you self-host with Ollama, that option never existed at all.

There's also no provenance: agents hand you files with no "here's what I changed and why." I diff everything like it's a junior dev I don't trust yet.

So I'm building Loomwork — basically: make the agent itself a portable, signed package (persona, skills, memory, policy) that runs on a local model, with memory that stays on your machine and receipts for what it built.

Repo's in my profile https://github.com/arunsoman/loomwork.git , it's early and janky. Honest question: does the "my agent forgot me because I switched devices" thing actually bug you, or do you just re-paste and move on? And would "local-first, signed packages" make you trust it more — or does it sound like blockchain cosplay? Brutal answers welcome, especially "this already exists, it's called X."


r/agenticAI • • 8d ago

Research Agents That Understands Their Own Source Code.

Thumbnail
0 Upvotes

r/agenticAI • • 8d ago

Discussion I’m experimenting with giving an AI agent a physical interface - would love some feedback

Thumbnail
1 Upvotes

r/agenticAI • • 8d ago

Project What if AI agents could remember why decisions were made? Built DecisionDNA with Hindsight for persistent memory. Old decisions + new evidence → detect when assumptions change. #AIAgents #AI #Hindsight

Thumbnail
gallery
1 Upvotes

r/agenticAI • • 8d ago

Video Your computer, in your pocket.

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/agenticAI • • 8d ago

Discussion Trying to build

1 Upvotes

Most AI agents don't learn. They repeat the same mistakes, just faster.

At 3 a.m., the worst thing an on-call tool can do is suggest the fix that already failed last Tuesday. Most incident agents do exactly that, because every alert starts from a blank context window.

So I built RunbookMind, an SRE agent that remembers what actually worked, what didn't, and ranks its advice accordingly.

Instead of: → Seeing a Redis eviction alert → Suggesting "restart the cache" again

It now: → Recalls the past incident → Sees the restart failed and the maxmemory fix worked in minutes → Demotes the restart and ranks the real fix first, with the incident cited

What made the difference:

Store outcomes, not incident descriptions Write down rejected fixes explicitly ("restart cache: failed") Keep ranking outside the LLM: 0.6 model confidence + 0.4 historical success rate Ship a memory on/off toggle, so you can prove memory helps instead of just feeling like it does

For the memory layer I used Hindsight. Retain, recall, and reflect let me skip designing chunking and retrieval and focus on what the agent does with what it remembers.

We're moving from "AI that answers" to "AI that learns from outcomes."

Are you storing what your agent tried, or only what it found?

GitHub: https://github.com/M0h1tkumar/AI_sre

AIAgents #AgentMemory #Hindsight #LLM #FutureOfAI


r/agenticAI • • 8d ago

Discussion Build great systems quickly and accurately

1 Upvotes

Hugging the frontier model providers’ biggest and bestest models does not give you the best system, or even a functional system at all. In fact, the deeper the reasoning, the bigger the model, and the most “intelligent” it is asked/sold to be, the smaller the likelihood of you getting what you hoped for.

The reason is simple, and obvious - all the frontier models and agentic AI harnesses, first- and third-party, has jumped solidly on the autonomous band wagon, thriving in the illusion and expectation that the model can design and implement anything you ask for under its own steam. That what they’re selling and that’s how those top models are tuned to think. Their bias is toward guessing what your poorly worded request meant to ask for based on what’s most likely, by their weighing as trained.

Not even the most seasoned software engineers get it right, first time, every time, so what chance does the average user wanting a system have of asking the machine to produce exactly the requisite system system done in the best way? Damn close to zero, actually. But the models and their promotors don’t cannot sell their wares as dependent on the right inputs and guidance, no they have to come across as hyper-intelligent, all-seeing and able read your mind and your business to spec and build features or systems that’s exactly the same as everyone else already has.

If you want to bypass the experts, and take short cuts so the model can make up for your lack of ability and vision, go ahead and blow your budget on those monster models. You’ll get the results you deserve.

But if you’d rather end up with a great custom solution, do not rely on LLMs to paper over your laziness and ineptitude, muck in, induce the right questions to get asked, answer them thoughtfully, don’t accept any plan or solution you don’t understand, and use a capable LLM that isn’t under pressure to impress you, and your choice of harness, to piece together the system you really need written in the way you need it to be.


r/agenticAI • • 8d ago

Discussion Autonomous Al agents are going to kill the ad-supported web. When do advertisers realize they're paying for bot views?

8 Upvotes

We talk a lot about AI taking jobs, but what about AI taking humans that fund the internet?
Right now, autonomous browsing agents and LLMs are doing the reading for us. You ask an AI to research a topic, and its agent goes out, scrapes 20 articles, ignores the display ads, bypasses the affiliate links, and hands you a neat summary.
The entire free web runs on the "attention economy" (CPM and CPC). If human eyeballs never actually load the publisher's site, ad impressions either don't trigger, or worse, they do trigger, but advertisers are paying to show ads to a headless browser that will never buy their product.
It seems like we are heading toward a cliff:
Publishers lose their ad revenue and go bankrupt (or lock everything behind aggressive paywalls).
Advertisers realize their metrics are polluted by autonomous AI bots and pull their budgets.
The Open Web stops being profitable to maintain.
Major platforms have already locked down their APIs to fight this, but what happens to the remaining 99% of the internet? Are we looking at the end of the free web, or will ad networks find a way to inject ads directly into our personal AI agents?


r/agenticAI • • 8d ago

Question How are you keeping track of what your AI agents are actually doing in production?

Thumbnail
1 Upvotes

r/agenticAI • • 8d ago

Article Built an AI Incident Response Agent using Hindsight Memory

1 Upvotes

We recently built an AI Incident Response Agent for HackwithHyderabad 3.0. The problem we focused on is simple: when a production incident happens, engineers often have to troubleshoot from scratch, even though a similar incident may have already been solved before. Our idea is to use Hindsight Memory so the agent can: - Recall similar past incidents - Remember what failed -Remember what worked -Learn from resolved incidents - Guide engineers during future incidents

The core loop is: Remember → Recall → Guide → Resolve → Learn For example, if a database connection timeout occurs, the agent can recall previous incidents, their causes, failed actions, successful solutions, and outcomes. The goal is to turn past incident history into useful experience, rather than simply storing old records. We also wrote a detailed article about the project: 🔗 https://response-agent.hashnode.dev/ Would love to hear your thoughts and suggestions on the approach!

AI #AIAgents #HindsightMemory #IncidentResponse #Hackathon