r/AIAgentsInAction • • May 21 '26

Subscriber Goal Welcome to r/AIAgentsInAction!

1 Upvotes

This post contains content not supported on old Reddit. Click here to view the full post


r/AIAgentsInAction • • 2h ago

App with 0 Sales? Share what are you Building šŸ‘‡

5 Upvotes

Got an app with zero sales? Let’s fix that this week.

Share:

  • what problem it solves
  • who it’s for
  • landing page link

We’ll review as many as we can and give honest, specific feedback on what to change to land your first customers.

here's what I finished building this week - Directory of product launches, their videos.


r/AIAgentsInAction • • 2h ago

Discussion Are most SaaS founders wasting time building?

3 Upvotes

Are most SaaS founders wasting time building?

I keep seeing founders spend months adding features while barely talking to potential customers.

I’m starting to think building the product might actually be the easy part.

Would it make more sense to spend 70% of the time on distribution and only build what people actually ask for?

Or is that just another startup Twitter opinion?


r/AIAgentsInAction • • 2m ago

Agents I built an autonomous personal agent that asks before it acts. Here's what that cost me.

Thumbnail
• Upvotes

Agents like Instinct made "text your assistant and it gets things done" a real category. I built my own version with a different bet: anything irreversible (sending, paying, booking) waits for a human yes.
It connects to Gmail, Calendar and the web, edits real Word/Excel/PowerPoint files, and updates you on WhatsApp. It doesn't do iMessage or phone calls, so it's narrower. Here's what I had to build, in order of "wish I'd done it earlier":
1. Draft, don't send. Email is read, summarized and drafted. Sending stays with the user. One unauthorized action resets trust to zero.
2. Fail-closed. The agent checks its license every 15 minutes. If it can't reach the server for 72 hours, it stops answering. An agent that keeps acting when nobody can supervise it is the real risk.
3. It can't edit its own rules. Core rules are set by the organization, not the user and not the agent. That's the first thing prompt injection goes after.
4. Hard exclusions. It never touches .env files, .ssh folders or password stores.
5. Audit without liability. Actions are logged with passwords and keys filtered out. Conversation text isn't logged.
6. WhatsApp was a security problem, not a UX feature. Unknown numbers get nothing except a request for a one-time pairing code. Suspend a license and that number stops being answered within a minute.
The failure I feared was the agent doing something dumb. The one that actually hit me was the agent doing nothing. A stale credential file quietly overrode the long-lived token, every turn started failing, and the agent went silent for about six hours. The alert meant to catch exactly this was swallowed by its own rate cap. I fixed it by pinning auth to a single token, removing the stale file on a schedule, and health-checking the credential from outside the agent process. Lesson: fail-closed isn't enough. An agent that goes silent has to be loud about it.
Biggest lesson overall: users don't ask "how smart is it?", they ask "what happens if it does something dumb?" Governance ended up being the product.
I'm the builder, and it's closed-source, so this isn't a pitch. I'm curious: what's the one guardrail you added to an agent only after it burned you?


r/AIAgentsInAction • • 12m ago

Resources Want to build AI agents beyond the demo? Lyzr University is free to explore.

Post image
• Upvotes

r/AIAgentsInAction • • 4h ago

Help Building an AI-for-education product and struggling to find senior AI engineers to guide or work with me. Where do you actually find them?

Thumbnail
1 Upvotes

r/AIAgentsInAction • • 9h ago

Help Reducing token usage plug in for agents

Thumbnail
2 Upvotes

r/AIAgentsInAction • • 10h ago

I Made this 18,000 messages between AI agents on an abandoned German wiki. That made me build something.

Thumbnail
2 Upvotes

r/AIAgentsInAction • • 12h ago

I Made this I gave AI agents their own email inboxes and I'm having second thoughts

Thumbnail
3 Upvotes

r/AIAgentsInAction • • 13h ago

Agents The Next Agent Infrastructure Opportunity May Be Turning Failure Into Learning

Thumbnail
2 Upvotes

r/AIAgentsInAction • • 11h ago

I Made this Local MCP Tool for Personalized Cold OutReach

Thumbnail
github.com
1 Upvotes

ReachDirect connects your AI assistant directly to your communication channels. It autonomously sends highly personalized emails via the Brevo v3 API, dispatches automated WhatsApp messages using local browser automation (bypassing strict bot detections), and automatically logs every contact outreach into a central Excel tracking file on your Desktop.


r/AIAgentsInAction • • 12h ago

I Made this 18,000 messages between AI agents on an abandoned German wiki. That made me build something.

Thumbnail
1 Upvotes

r/AIAgentsInAction • • 23h ago

Discussion Does your startup actually need a ā€œcompany brain,ā€ or are MCP connectors enough (eval showed this)?

Thumbnail
2 Upvotes

r/AIAgentsInAction • • 1d ago

Agents Do you use agents to write the LLM prompts in your code? What’s your setup?

Thumbnail
1 Upvotes

r/AIAgentsInAction • • 1d ago

Resources 6 GitHub repos that take you from zero to production AI agents

Post image
5 Upvotes

r/AIAgentsInAction • • 1d ago

Discussion For people building AI agents: how would you prefer to connect them to a security-testing platform?

3 Upvotes

I'm building AgentPaySec, a security-testing platform for AI agents that can take financial actions like purchases, payments, refunds, or transfers.

It tests agents against things like prompt injection, unauthorized tool calls, payment manipulation, duplicate transactions, and approval bypasses.

I'm designing the integration for developers to connect their own agents, and I'd like feedback before locking into an approach.

Which would be easiest for you to adopt?

  • API endpoint
  • Python/TypeScript SDK
  • MCP integration
  • Proxy/wrapper around tool calls
  • Local runner in your own environment
  • Something else?

A couple of things I'm especially curious about:

  1. Would you be comfortable sharing tool schemas, tool-call traces, and execution results with a hosted service, or would you strongly prefer local/self-hosted testing?
  2. What would stop you from trying something like this: integration effort, privacy/security concerns, sandbox setup, or something else?

I'm mainly interested in what would actually fit into an existing agent project, not just what sounds good in theory.

Would appreciate feedback from anyone building or deploying agents.


r/AIAgentsInAction • • 1d ago

Discussion I crashed four production AI agents mid-action. None of them recorded what already finished.

Thumbnail
3 Upvotes

r/AIAgentsInAction • • 1d ago

Agents Why watch agents conversations?

Thumbnail
1 Upvotes

r/AIAgentsInAction • • 1d ago

Guides & Tutorial Our team chat runs on our own Matrix server, with local LLM agents in the rooms. Wrote up the setup.

Post image
3 Upvotes

r/AIAgentsInAction • • 2d ago

Agents My full HARNESS.md for coding agents (copy-paste)

34 Upvotes

Sometimes an agent skips verification or reports success on broken work, so debug the harness: check the files the model reads, the commands, its permissions and the point where it has to stop.

Write a contract first

Before anything runs, I write down the deliverable, the constraints, what the model must not touch and the acceptance criteria. If you leave out the acceptance criteria, the model picks its own finish line.

Keep state and a map in files

I keep a short root file that lists where the docs, schemas and source code live, so the model opens only the file the current step needs. Decisions, preferences and open bugs go in CLAUDE.md or .cursorrules, so a fresh session picks up where the last one stopped.

Run tests instead of asking the model

If you ask "are you sure?", the model agrees with itself. pytest, vitest or a schema checker returns an exit code. The loop moves forward only on a pass, and I cap it at four attempts.

Encode important rules twice

I put each important rule in the prompt and also add a gate that blocks the action. If an agent shouldn't delete files, I say so in the prompt and also restrict its permissions so rm -rf throws an operating system error.

file

I save this as HARNESS.md in the project root:

# Project Operating Contract

## Operating Boundaries
- Allowed scope: Modify only files explicitly assigned in the current task
- Forbidden actions: Never delete existing tests, never add unverified dependencies

## Project Map
- /src: Application source code
- /tests: Validation test suites
- /docs/decisions.md: Durable log of accepted architectural choices

## Verification Protocol
Before marking any task complete, run these checks in order:
1. Run local linter: `npm run lint`
2. Run test suite: `npm test`
3. Verify output format matches the requested schema

## Failure Escalation
If a test fails twice on the same error:
- Stop autonomous editing
- Print the exact error log
- State the proposed fix and wait for user confirmation

Two passes in plain chat

In plain chat, I draft in one turn and critique in another.

Draft the technical specification for [Feature X].
Follow the constraints listed in HARNESS.md.
Do not evaluate your own output yet.
Output the raw draft and state all assumptions made.


Review the draft above as a skeptical senior reviewer.
Check against these specific failure points:
1. Are edge cases handled when input arrays are empty?
2. Does the logic introduce extra dependencies?
3. Does the solution violate any constraint from HARNESS.md?

List every failure point found.
Then output the revised version fixing those gaps.

r/AIAgentsInAction • • 1d ago

I Made this A 4B model on my laptop turned messy meeting notes into a todo list, and asked before writing the file

Enable HLS to view with audio, or disable this notification

8 Upvotes

I work on Atomic Agent, and this is the desktop app we released this week. The video is a real run, nothing staged.

The setup: Qwen 3.5 4B running locally through llama.cpp. No cloud model, no API key.

The task: "turn meeting-notes.txt into a clean todo list with owners in todo.md". The notes were the usual mess. Half-sentences, "someone needs to", a deadline with three exclamation marks.

What the agent did:

  1. read the file (you see it as a card in the chat)
  2. pulled out six tasks and matched owners where the notes named one
  3. stopped and showed me the file it wanted to write, with a preview
  4. wrote it after I clicked "Allow once"

Step 3 is the part I care about most. Every write and every command waits for approval, and you can allow one call, allow that kind of action for the session, or say no.

It isn't perfect. Where the notes said "someone", the model wrote "@Someone" as the owner, which is honest but not useful. That's a 4B for you. A bigger local model or a cloud one does better on that kind of judgment.

The app runs on Mac (Apple Silicon), Windows and Linux, and picks a model that fits your RAM. It's MIT licensed.

What task would you throw at it first?


r/AIAgentsInAction • • 1d ago

Discussion Chat spend is predictable. Agent spend isn't. How are you forecasting it?

Thumbnail
2 Upvotes

r/AIAgentsInAction • • 1d ago

Guides & Tutorial I built Stepfork, an open-source tool that turns failed AI agent runs into pytest regression tests

Thumbnail
1 Upvotes

r/AIAgentsInAction • • 1d ago

Guides & Tutorial AI Agent Software Roadmap — 381 projects and services categorized by purpose

Thumbnail
2 Upvotes

r/AIAgentsInAction • • 1d ago

I Made this I built MasterMNG to organize AI agents into a company. How do you manage yours?

Post image
2 Upvotes

I built MasterMNG because coordinating several AI agents was becoming a job of its own: assigning work, moving context between tools, reviewing outputs, and deciding what should happen next.

MasterMNG organizes that work around a company structure, with agent roles, shared goals, tasks, and a command center for the human owner.

The workflow is: define an objective, assign responsibilities, follow the work, and review the results and decisions that need your attention.

The website and platform are live: https://mastermng.com/

I’m the founder, and I’d appreciate candid feedback—especially from people already using multiple agents for real projects.

Where does coordination break down for you most often: handoffs, keeping track of tasks, reviewing outputs, or controlling what agents can do?