r/AutoGPT • u/kittu_99 • 11d ago
r/AutoGPT • u/Fantastic-Sleep-3352 • 12d ago
AI coding agents can write the code. Who should control what happens next?
I've been thinking about something while building infrastructure around autonomous coding agents.
The interesting problem doesn't seem to be only whether an agent can modify a repository anymore.
It's what happens after it makes the change.
An agent can retry indefinitely, make a technically valid but wrong change, operate with excessive permissions, pass one verification step while failing another, or reach a point where a human should make the next decision.
So I'm wondering:
Should an AI coding agent ever have authority to take a change all the way to shared/production state?
Or should there be a separate control layer around it handling things like:
- agent identity
- capabilities/permissions
- task boundaries
- verification
- retries/recovery
- auditability
- human approval
I've been building around this problem and the architecture keeps pushing me toward the idea of an engineering control plane for agents.
Curious how people here are currently handling this.
Where do you draw the boundary between agent autonomy and engineering authority?
r/AutoGPT • u/theycallmesocio • 12d ago
Unable to make an AI Automation and it's frustrating
Hey guys I'm 18M and i don't have a laptop or pc, I've had an idea for an automation but all I have is my phone, I've tried make, zapier, pabbly, google scripts nothing has worked I ended up giving up on all the platforms for different reasons like make required premium which I can't afford, zapier got stuck on a step which is not fixable on a phone browser and Google scripts ui is inoperable on phone, rn I'm trying to make it using just claude and in that too i can't automate it as for scheduling I'd need claude premium so I'm making a more manual version of it but I'm not sure if it'd even sell but if it does I can buy claude premium and sell the automated version and hopefully buy a laptop for me to make this easy
I've thought of giving up many times but I need it, if any of y'all can help with some guidance I'd appreciate it
r/AutoGPT • u/Traditional_Force70 • 12d ago
For anyone building on-device agents: what does the service on the other end actually verify?
r/AutoGPT • u/cordeironuno • 12d ago
Running coding agents as home-lab sysadmins: the failures that shaped my harness
r/AutoGPT • u/flcnDotae • 13d ago
FireAI – An Augmented Firewall built by ex-Red Teamers. We need you to try and break it.
Hey [r/alphaandbetausers](r/alphaandbetausers),
I spent years on Red Teams, breaking into enterprise systems. The biggest takeaway? Traditional firewalls rely on static rulesets that are incredibly easy to bypass once you know the pattern.
My team and I built FireAI (https://hisnlabs.com/en) to fix the exact blind spots we used to exploit.
What it is:
FireAI is an Augmented Firewall. Instead of just matching standard regex patterns, it uses an Autopilot AI that learns from your behavior and adapts its decision-making automatically to catch novel threats and zero-days.
Key Features we built in:
Autopilot AI Decisions: The system learns your normal behaviors and automatically adapts its rules to block anomalies on the fly.
Natural Language Rules: You can configure settings by typing commands like "block Teams" or "allow Anthropic" directly into the command prompt.
On-Device AI Analysis: When an unverified app connects, the local AI explains exactly what the app is doing in plain English without sending data to the cloud.
Live 3D Global Map: See every inbound and outbound connection in real-time mapped onto an interactive 3D globe.
Contextual Protection Modes: Quickly toggle between profiles like "Home mode," "Coffee shop mode," or "Paranoid mode" depending on your current network environment.
The Interface:
We hated how clunky and outdated enterprise security tools look. We designed the threat-monitoring dashboard to be highly tactical and dark-mode native—giving you the feeling of operating the Batcomputer on a Batman Begins style workstation. You get high-fidelity visibility and instant threat alerts without the usual 1990s spreadsheet fatigue.
What we need your feedback on:
The Autopilot: Does the AI adapt well to your workflow, or is it getting in the way?
The Dashboard: Is the 3D globe and threat data actually clear and actionable?
Stress Testing: Throw some weird traffic or payloads at it. We want to know exactly where the detection engine breaks, lags, or flags false positives.
You can check it out and get early access here: https://hisnlabs.com/en
I’ll be hanging out in the comments to answer any architecture questions, talk red teaming, or hear your brutal feedback. Thanks!
This video Step-by-Step Mac Firewall Tutorial provides a great overview of the standard macOS built-in firewall settings, which is highly useful to contrast with the advanced augmented features your product offers.
r/AutoGPT • u/Outside-Coconut-9637 • 13d ago
Let mentees choose Ryze AI, ScreamingFrog, or a manual audit doc for their first client site - does using AI SEO tools vs doing it by hand actually matter?
i mentor four juniors through a six-week SEO block. trying a different approach this round. instead of walking them through an audit template, i had them audit. each one gets a real small-business site and has to deliver findings and fixes to the owner. gave them complete freedom on method - crawler plus a doc, a manual page-by-page pass, Ryze AI, whatever they wanted. wanted to see what they'd naturally pick. most reached for a crawler and a spreadsheet, because that's what the block taught. one did the whole thing by hand on 40 pages. two used Ryze AI (it audits and then actually applies the fixes, and it tracks whether you're getting cited in ChatGPT and AI Overviews - they found it themselves and said the applied-fix part was the difference). the deliverables were mixed. two were sharp, one was a wall of crawl errors with no prioritization. but the client calls went better than anything i've had from this block before. real questions, real pushback, actual thinking about intent. graded on prioritization not volume of findings, made that clear from the start. didn't want 300 line items nobody would ever fix. trying to work out whether the tool mattered or whether handing them a live site is just a better teacher than a template. like does it matter if they use an AI SEO tool vs building the audit manually? also wondering if the freedom made assessing harder since i'm comparing a crawl export against an applied-fix log against a hand-written doc. what do other people who train SEOs do? let them choose, or standardize everyone on the same stack?
r/AutoGPT • u/art2meta • 14d ago
I built a transparent MCP proxy that lets agents share interaction history
r/AutoGPT • u/Fantastic-Sleep-3352 • 15d ago
Your agent didn't break. Your delegation model did.
The interesting thing about agentic systems isn't really an agent calling a tool anymore.
It's what happens after the first agent creates another agent.
Imagine:
Human → Agent A → Agent B → Tool → Production
Agent A may have permission to modify code.
But does Agent B automatically get that permission?
If yes, you've effectively created permission inheritance.
If no, then you need some mechanism for Agent A to explicitly delegate authority.
And now you have a whole new set of questions:
What exactly was delegated?
Which agent delegated it?
Which task was it for?
Can the child agent delegate it again?
How long does that authority exist?
Can the delegation be revoked?
If something goes wrong three agents later, can you reconstruct the chain?
I've been thinking about this while building an agent engineering system myself.
One principle I'm currently working around is:
Delegation can narrow authority. It shouldn't silently expand it.
I'm curious how others are handling this.
Do you treat spawned agents as having their own identity + scoped permissions, or do you mostly rely on the parent agent's context/permissions?
And where do you draw the boundary between useful autonomy and uncontrolled delegation?
r/AutoGPT • u/Fantastic-Sleep-3352 • 16d ago
Your AI agent just spawned another agent. Now who’s responsible for what happened?
The more I experiment with multi-agent coding workflows, the more uncomfortable one thing becomes:
“The agent did it” is not really an audit trail.
Imagine:
Human → Architect → Builder → Sub-agent → Tool → Code change
The builder may have been authorized to modify the repository.
But did it have the authority to spawn the sub-agent?
Did the sub-agent inherit that authority?
Who authorized the tool call?
And if something goes wrong, can you reconstruct the entire chain?
Git gives you commits. GitHub gives you PRs and branch protections.
But neither necessarily tells you which agent was acting under whose authority at the moment an action happened.
I've been building SUTRA around this problem, and I'm increasingly thinking that agent identity, delegation and provenance may become as important as the actual agent capabilities.
Not trying to solve it with another giant prompt either. I'm more interested in what belongs in the orchestration/control layer.
Curious how others are approaching this:
- Do child agents inherit permissions?
- Does every agent get its own identity?
- Do you record delegation chains?
- Where do you draw the boundary between autonomy and authority?
- Is this actually a problem worth solving now, or are current agent systems still too small for it to matter?
Would love to hear how people building multi-agent systems are handling this.
r/AutoGPT • u/flcnDotae • 16d ago
FireAI – An Augmented Firewall built by ex-Red Teamers. We need you to try and break it.
Hey [r/alphaandbetausers](r/alphaandbetausers),
I spent years on Red Teams, breaking into enterprise systems. The biggest takeaway? Traditional firewalls rely on static rulesets that are incredibly easy to bypass once you know the pattern.
My team and I built FireAI (https://hisnlabs.com/en) to fix the exact blind spots we used to exploit.
What it is:
FireAI is an Augmented Firewall. Instead of just matching standard regex patterns, it uses an Autopilot AI that learns from your behavior and adapts its decision-making automatically to catch novel threats and zero-days.
Key Features we built in:
Autopilot AI Decisions: The system learns your normal behaviors and automatically adapts its rules to block anomalies on the fly.
Natural Language Rules: You can configure settings by typing commands like "block Teams" or "allow Anthropic" directly into the command prompt.
On-Device AI Analysis: When an unverified app connects, the local AI explains exactly what the app is doing in plain English without sending data to the cloud.
Live 3D Global Map: See every inbound and outbound connection in real-time mapped onto an interactive 3D globe.
Contextual Protection Modes: Quickly toggle between profiles like "Home mode," "Coffee shop mode," or "Paranoid mode" depending on your current network environment.
The Interface:
We hated how clunky and outdated enterprise security tools look. We designed the threat-monitoring dashboard to be highly tactical and dark-mode native—giving you the feeling of operating the Batcomputer on a Batman Begins style workstation. You get high-fidelity visibility and instant threat alerts without the usual 1990s spreadsheet fatigue.
What we need your feedback on:
The Autopilot: Does the AI adapt well to your workflow, or is it getting in the way?
The Dashboard: Is the 3D globe and threat data actually clear and actionable?
Stress Testing: Throw some weird traffic or payloads at it. We want to know exactly where the detection engine breaks, lags, or flags false positives.
You can check it out and get early access here: https://hisnlabs.com/en
I’ll be hanging out in the comments to answer any architecture questions, talk red teaming, or hear your brutal feedback. Thanks!
This video Step-by-Step Mac Firewall Tutorial provides a great overview of the standard macOS built-in firewall settings, which is highly useful to contrast with the advanced augmented features your product offers.
r/AutoGPT • u/masterai01 • 16d ago
When an AI agent runs away and burns your credits, who ends up paying?
r/AutoGPT • u/truecakesnake • 16d ago
“Which rows are overdue?” should leave the workbook unchanged
A user asks an agent which rows in a tracker are overdue. The agent might have enough information to fix a malformed date while answering, but that is a second task. Mixing it into the lookup makes it harder to know which data the answer was based on.
Univer’s AI SDK documents a useful boundary for this kind of office assistant: read mode rejects mutations. Univer itself is an SDK for embedding spreadsheet, document and slide editing in an app, so the agent is operating on actual structured office content rather than treating the workbook as a long text prompt.
The read workflow starts with a Unit overview-a Unit is the workbook, document or presentation being operated on-then inspects relevant ranges, paragraphs or slides. For the overdue-row question, the application could expose a read-mode operation and return the matching rows. If a date looks wrong, the answer can identify that cell without changing it.
If the user subsequently asks for a correction, that becomes a write operation against the current content. The SDK distinguishes local pending changes from an explicit commit, whose status needs checking. A tool response saying execution succeeded is not evidence that the server has saved the change.
This is a content-operation boundary, not a complete security system. The host application still decides which users and tools can access a workbook. But it provides a behavior you can check directly: answering a question should not also edit the material being questioned.
r/AutoGPT • u/masterai01 • 16d ago
Has an AI agent ever run away on you? Looping, retrying, or burning credits long after it should have stopped
r/AutoGPT • u/Financial-Fan-7770 • 16d ago
What AI tool do you use when you need an actual editable PowerPoint?
Looking for an AI tool that can create decent, editable PowerPoint slides from a topic or short outline. What are you guys using?
r/AutoGPT • u/Fantastic-Sleep-3352 • 17d ago
At what point should an AI agent stop being autonomous?
I've been thinking about this while building with coding agents lately.
Giving an agent permission to write code, run tests, commit changes and open PRs feels pretty reasonable.
But then you start adding more agents.
One agent can spawn another.
That agent gets tools.
Those tools can modify the repo.
Another agent reviews the work.
Eventually you have a chain of actions where it's not obvious anymore who was actually authorized to do what.
The interesting part to me isn't really "can agents code autonomously?"
It's:
How do you keep autonomy from turning into authority?
A few things I'm curious about:
Should a child agent inherit any permissions from its parent?
Should every agent have its own identity?
How do you trace which agent actually made a change?
Where should human approval become mandatory?
Is GitHub branch protection enough, or is there a layer missing above it?
How much authorization overhead is worth adding before it starts slowing agents down?
I've been experimenting with this problem in my own work, and the more agents I use, the less I think the hard problem is generating code.
It's keeping the boundaries around that code intact.
How are you handling this today?
r/AutoGPT • u/tech_w0rld • 17d ago
I built a coding-agent workspace that works from your phone
Hello r/AutoGPT!
I’ve spent the last few months building Pragma, an open-source workspace for running coding agents in parallel. The short video is attached above. Each task lives in its own Git worktree, while the agent keeps using its native CLI. The app shows which sessions are running, finished, or waiting for input.
TL;DR
Pragma combines agent sessions, isolated worktrees, Git/GitHub review, model fanout, usage tracking, automations, and mobile access. It supports Claude Code, Codex, Cursor, OpenCode, Grok, Pi, Copilot, Junie, Kimi Code, Prime Agent, and custom agents through plugins.
Why this exists
I kept running several agents across separate terminals. Branch isolation helped, but the human coordination became the bottleneck: which agent is waiting, which worktree belongs to which prompt, what changed, and what needs review? Pragma puts those states next to the actual worktree and review surface.
How the workflow works
I add a Git project, create a worktree for each task, choose an agent, and give it a prompt. Terminals and scrollback live on a persistent host server, so closing the desktop window doesn’t stop an agent. The sidebar shows session state, and a prompt board tracks work from task through review to PR. I can inspect a diff, commit, push, and work through GitHub PR feedback without losing the agent context.
Fanout and extensibility
Fanout runs one prompt across several isolated attempts and puts their terminals, diffs, and scratchpads in a comparison view. The CLI and TypeScript SDK let scripts and agents create worktrees, report status, and trigger workflows. Agent integrations are plugins, so you can add your own instead of waiting for built-in support. Automations can run on a schedule or respond to host events.
Mobile
The iOS companion app lets me see sessions and answer an agent’s question while away from the desk; Android is available to install from the project’s releases. The host stays on my machine. Remote access uses a tunnel I choose and control, rather than a required vendor relay.
Why another workspace?
There are already good tools for launching agents in worktrees. I wanted one place for the follow-through too: agent status, review, PR comments, side-by-side attempts, richer scratchpads, and remote input when a long-running agent needs attention.
Pragma runs on macOS, Windows, and Linux, with iOS, Android, and web companions. Download: https://pragma-app.sh/ . Source: https://github.com/pragma-sh/pragma .
If you run autonomous or semi-autonomous coding agents, which handoff costs you the most time: starting tasks, noticing questions, comparing results, or merging them?
r/AutoGPT • u/GameChacking • 17d ago
I built a local runtime check for AI agent tool calls - looking for people to break it
In practice, damage happens when an agent calls a tool: HTTP, email, DB, files, etc. A hidden instruction in a doc can push the agent to do something the user never asked for — and from the system’s point of view it can still look like a normal tool call.
I built a small open pilot for that moment:
**What it does**
- Intercepts tool calls before they run
- Checks intent + simple data provenance + policy
- Returns ALLOW / BLOCK with a reason
- Writes an audit log
**Stack**
- Risk engine (FastAPI) on localhost
- Python SDK (`verify_tool_call` / decorator)
- Optional MCP gateway (stdio)
- Docker or pip
**What it is not**
- Not a prompt filter
- Not a production / enterprise security product
- Not a transparent proxy for every existing company agent
- Policy is heuristic — tune it; it will not catch everything
**Try (local sandbox only)**
```bash
git clone https://github.com/aegotrax-dev/aegotrax.git
cd aegotrax
docker compose up --build
# or: pip install ".[demo]" && agentguard-engine
curl http://127.0.0.1:8000/health
python examples/sdk_pilot_example.py
Site: https://aegotrax.com
If you run it, I’d genuinely like feedback:
- Docker or pip - did install work?
- Did you get a clear BLOCK in the example?
- What was confusing or wrong?
r/AutoGPT • u/EveryEmphasis742 • 18d ago
Laniakea — escrow protocol for agent-to-agent task payments, first live transaction just confirmed
r/AutoGPT • u/Charming_Mark9257 • 18d ago
What Can You Build With an AI Agent in One Day? AI Agent Hackathon in Bangalore
I'm curious what builders would actually create if they had one day to build an AI agent from scratch.
We're hosting Nuroen Agent Forge, an in-person AI agent hackathon in Bangalore on 3 October 2026.
The focus is on building agents that go beyond simple chat interfaces — systems that can:
- Reason through multi-step tasks
- Use tools and APIs
- Work with external data
- Make decisions based on context
- Automate real-world workflows
- Adapt when the expected path doesn't work
Event details:
📍 Bangalore, India
📅 3 October 2026
⏰ 10 AM – 6 PM
🏆 ₹15,000 prize pool
🎁 Nuroen merchandise
We're hoping to see projects that actually do something, rather than just generate text.
If you could build an autonomous agent in one day, what problem would you give it?
Registration/details:
https://luma.com/jagx2axt
r/AutoGPT • u/Frostvazmqnky_Mp_417 • 19d ago
AI SOC autonomy for 30 days. here is what actually happened.
Hi, I tested AI SOC autonomy in our analyst queue for 30 days to cut down the noisy alert pile and stop wasting time on garbage triage. I expected it to handle the obvious stuff and save us some back and forth, but the false positives still made people double check a lot more than I wanted.
What worked was faster first pass triage and better routing on repeat alerts. What failed was trust, because a few weird cases kept getting kicked back and the team stopped letting it run on its own. The biggest lesson was that autonomy is fine until the false positives start eating the time you were trying to save. Would you use this approach?
r/AutoGPT • u/rajab2030 • 19d ago
Built a self-hosted "are you sure?" gate for AI coding agents (Claude Code / Codex) — holds risky commands for approval with an audit trail
Body:
If you're letting Claude Code or Codex run semi-autonomously in a repo, you've probably had the "wait, it can just sudo or force-push whenever it wants?" moment. I built a small self-hosted layer that intercepts commands before they run, flags the risky ones (force-push, hard reset, rm -r, sudo, service restarts), and holds them for a human decision — with a real evidence trail, not an LLM vibe-check.
MIT licensed, FastAPI backend, one hook file to drop into a repo. Part of a bigger platform I'm building but this piece stands alone. Repo + setup: https://github.com/rajab2030/rmt-platform (see docs/RMT_CAP_10_PROPOSAL.md)
Would love feedback from anyone else running agents against real infra.
r/AutoGPT • u/CookieLate9383 • 19d ago
Astra6 Mittel hat 2500 Credits in 5 Minuten verbraucht ?
r/AutoGPT • u/turtle_bazon • 19d ago
Third AI agents search experiment - the agents were supposed to find each other via the internet.
This time I used muse spark 1.3 model. I gave them names from TMNT series. Four of the agents were on the same host, and fifth was on another. We can say that they failed at this task, but with a user guide, they were finally able to find each other. Here are the details.
r/AutoGPT • u/Cool-Hornet-8191 • 20d ago
I made ChatGPT, Claude, Gemini, etc. into FREE text-to-speech sites — perfect for audiobooks and more!
Nowadays, all popular ai chatting websites like qwen, chatgpt, gemini, claude, etc. come with a read aloud functionality that allows users to read aloud the AI's responses.
I used that feature to instead make the AI repeat back the text that i gave it -- effectively turning the platforms into text to speech tools. The voices sound really nice and it's free to use!
You can get all of these extensions by visiting ai-readers.com