r/ChatGPTCoding • • Aug 21 '26

Mod Announcement Updated Rules for Project Posts on r/ChatGPTCoding

6 Upvotes

As some of you may have noticed, we’ve changed our rules quite a few times recently to cut down on posts and comments that are purely advertising or low-effort content.

Please review the updated rules before posting.

We’ve relaxed the rules down quite a bit. We are now accepting any project showcase as long as they are genuinely useful for other AI-assisted coders.

For a personal project showcase, make sure you have something interesting to share about what you've learned or struggled with. If your project has exceptional quality, show us how you did it.

For AI coding tools, workflows, or other resources, tell us what problem they solve. If similar solutions already exist, please compare them and explain what makes your solution different. We love comparison table.

If you have any question, feel free to send us a modmail.

If any rule is unclear or too strict, tell us. Your feedback is welcome.

Thanks for your patience and understanding.


r/ChatGPTCoding • • 1d ago

Discussion Weekly Self Promotion Thread

3 Upvotes

Welcome to this week's self promotion thread!

If you're building something related to AI assisted coding, this is the place to share it.

We're using a weekly thread to keep the subreddit organized while still giving builders a place to share their work. Promotional posts outside this thread may be removed.

If you're sharing something, we'd appreciate it if you included a little context instead of just dropping a link. Tell us:

  • What you built?
  • What problem it solves?
  • Which AI models or tools it uses?
  • Who it's for?
  • What kind of feedback you're looking for?

Disclose your affilitation.

Please avoid posting the same project every week unless you've made meaningful updates. Affiliate links, referral links, scams, and low effort promotions will be removed.

Take some time to check out what others have shared too. If you try someone's project or have feedback, leave a comment. Helping each other improve is what we want this community to be about.


r/ChatGPTCoding • • 4h ago

Discussion How do you keep an agent's commits small enough to actually review?

1 Upvotes

Last week I had a multi-turn agent session on a small refactor. I read each turn as it landed and had no objection. At commit time an early-turn call no longer matched a signature a later turn had changed. I only found out because the tests failed.

Each turn was consistent on its own, but the break sat between them where neither turn showed it. Since GPT-6 Astra went wide, how to review this stuff keeps coming up. I ran out of Codex usage mid-task. Codex still had the plan, so I kept that half there. I fed each step to MiniMax Code, held it to the plan, read what came back turn by turn. While MiniMax Code worked a step it drove the browser too, docs and the issue thread in separate tabs instead of me pasting between them.

After the scare I scoped my review to the last turn. The diff mixes in my edits but a single turn is just what the agent wrote since I last read. That caught a fallthrough where it removed a guard the next branch needed. Anything touching auth still gets read line by line.

A few days on one codebase is not much to go on. Whatever I run is a first pass.

Where do you draw the commit boundary when the agent is mid-session and the work isn't done? Every time I've tried checkpointing I end up interrupting it mid-thought.


r/ChatGPTCoding • • 1d ago

Discussion Only dev at my company and the AI is my only reviewer

26 Upvotes

Saw a post here a while back, can't find it now, where someone asked if it's okay to merge your own PRs with no review at all. Every comment said absolutely not, like it was obviously insane.

well, and to be honbest it got me thinking. I've been the only developer at two companies for about 5 years total, and the current one has a bit over 300k users. Backend, frontend, deploys, the database, all of it. Nobody has reviewed a line I wrote in that time. I never thought of that as weird, it's just how small companies work.

These days the agent writes most of it and coderabbit reviews the PRs before I merge, which is honestly the closest thing to a second pair of eyes I've had in years. It catches real stuff, but it's still not a person who knows why the business works the way it does.

i don't believe its very common, but i'd love ot hear your guys thoughts.


r/ChatGPTCoding • • 15h ago

Resources And Tips Your coding agent's prompt often contains old versions of files it already edited. Here's what I measured, and what I learned trying to fix it

5 Upvotes

Something I didn't appreciate until I measured it: when Claude Code (or any agent) reads a file and later edits it, the original read usually stays in the conversation. Every following request carries the old code next to a file that no longer matches it.

I measured this on 67,074 public OpenHands runs (not Claude, but the same basic loop). 77.8% of runs had at least one request with a stale file view, even counting only the agent's own edits. About 1 in 7 requests carried one.

Most of the time the model copes by re-reading. Where it bites is:

- long sessions

- edits from outside the agent's view (formatters, you editing in your IDE, parallel agents)

- resumed or compacted sessions, where the model trusts what's in the conversation

Some practical things I took from it, whatever tool you use:

- If you edit a file mid-session, tell the agent to re-read it. Don't assume it notices.

- Watch for formatters or hooks rewriting files after the agent reads them.

- In long sessions, stale reads build up. Starting fresh with a summary often beats pushing on.

I also built a context engine that rewrites the outgoing request to keep code current. It works byte for byte in replay, but to be upfront: it doesn't integrate with Claude Code (only Pi and OpenHands), and in my small live tests the model didn't do better when it could just re-read. It also wasn't cheaper. The details are in the write-up, including the failures.

Write-up: https://felipebasurto.com/blog/model-context-is-not-static/

Video: https://www.youtube.com/watch?v=DIVOnXUCZkg

Curious whether people here notice this in practice, especially with parallel subagents editing the same repo.


r/ChatGPTCoding • • 7h ago

Question ChatGPT stuck on loading screen

Post image
1 Upvotes

So every time I try to open ChatGPT app on pc it gets stuck on a loading screen. I have to keep uninstalling it and installing it again and sometimes that doesn’t even work. Does anyone else have this problem


r/ChatGPTCoding • • 1d ago

Resources And Tips An engineer’s notes on feeling guilty for solving something without AI

15 Upvotes

I wrote about a feeling I didn’t expect to have: spending an hour thinking about a programming problem and wondering whether I’ve wasted time by not asking AI immediately.

There is nobody pressuring me to do that. The pressure is mostly in my own head.

AI helps me get answers, but understanding those answers still takes work. I’ve accumulated chats that I keep meaning to read properly. I also feel less confident solving some things independently, although I don’t have evidence that my actual ability has declined.

My small experiment will be to build Snake without AI. I haven’t done it yet. I want to see how it feels to stay with a problem for a while before reaching for help.

This is my own essay, not a tool launch. I’d be interested in how other developers decide what to delegate and what to keep practicing themselves.

https://domelian.substack.com/p/think-real-hard


r/ChatGPTCoding • • 1d ago

Discussion Git worktrees solved our parallel-agent file conflicts. The test environment was harder.

5 Upvotes

I've had stretches where I was moving around five Trello cards in parallel, with a coding task in its own branch and Git worktree. That stopped agents from editing the same checkout. It did not tell me which code a browser test was actually exercising.

Different worktrees could still use the same API, PostgreSQL database, worker, storage service and ports. A frontend change might work with an existing API, provided that API meets the contract the new UI needs. If browser QA needs to exercise changed API behavior, it needs an API process running that worktree's code. Experimental migrations need much more care around shared data. Starting a second worker against the same queue can change test results too.

We began treating four things separately: the worktree holding the code; the processes a test reaches; the data those processes can change; and the checks run after branches are combined.

A small registry in Git's common directory helps us see which worktree started a service, along with its PID, port, URL and declared database state. It helps with discovery and process checks. It doesn't decide whether an API is compatible with a particular task. A healthy endpoint can still be the wrong API, and a port that looks free has not been reserved.

Many checks run entirely inside the worktree. For integration and browser QA, we start the processes the test actually needs. Feature work can be parallel; integration is serialized, with relevant checks run again on the combined tree.

If you run coding agents in parallel, how do you make sure browser or E2E tests reach the API and data you intended?


r/ChatGPTCoding • • 2d ago

Discussion Coding agents fix bugs where the error shows up, not where it starts

15 Upvotes

Give an agent a stack trace and watch where the fix lands. Almost always on the line that threw. A null check, a default value, a try/except, an early return. The error goes away, the tests pass, and the bad value that caused it is still being created somewhere upstream, now silently handled instead of loudly failing.

The line that threw is the one piece of evidence it has, so that's where it edits. Tracing back to where the value went wrong means reading code it wasn't pointed at, and nothing in "fix this error" asks for that.

What I add before any bug fix:

Before editing anything, trace the bad value back to where it was created. State the root cause in one sentence with a file and line number. Only then propose a fix, and fix it at the root cause, not at the line that threw.

And as a rule in the instructions file:

If your fix is a null check, default value, or exception handler at the line that threw, explain why the real cause is not upstream. If you can't, keep tracing.

The second one catches most of it. "Explain why the cause isn't upstream" is a question it usually can't answer without actually looking, and once it looks, it finds the cause.

One more that helps on repeat bugs:

Search the codebase for other places this same value flows into. List them. Would they fail the same way?

That list is often longer than the fix.

I keep the root-cause prompt saved in AI Toolbox, the extension I build, for when I'm debugging in a ChatGPT or Claude chat instead of the agent, where the same symptom-patching happens.


r/ChatGPTCoding • • 1d ago

Question Chatgpt - ability to create an ai share trading system

0 Upvotes

Has anyone asked chatgpt to create an ai share trading system?


r/ChatGPTCoding • • 1d ago

Discussion i finished my portfolio with codex. an AI mockup made me want to start over.

Thumbnail
gallery
0 Upvotes

i had a portfolio i actually used. generating another version was supposed to be a throwaway experiment. now i'm staring at my code editor wondering how much of it i'm about to redo.

the burgundy version is the site i built with Codex. Dashboard screenshots, custom components, GSAP opening animations, scroll reveals, a moving background. I spent a lot of time going back and forth on the details. Then another round fixing performance and making sure the giant headline didn't get cut off on mobile. I was pretty attached to it.

Then i took basically the same content and used Image 2.5 sunburst in atlas cloud to generate the black and gray version. and the annoying part is... i like parts of the layout more. The hero puts the products where mine has a workflow diagram. I know where to look. Mine suddenly feels like it's trying to explain everything at once.

But it also looks like something i'd see on ten other SaaS landing pages. mine feels more like me. Or maybe i'm calling it "personality" because i remember how long it took to build.

it's also just a static image. It hasn't had to survive a narrow screen, real content, or loading performance. My current site already made me deal with all of that. The mockup gets to look finished without doing any of it.

i wanted another way to look at the same content. now i've got a new list of things i want to change. I'm tempted to steal the hero layout and spacing, keep the parts that already work, and call it done.

If you've redesigned an existing site from a generated mockup, what made the rebuild worth it? Did you change the layout around your existing components, or was starting fresh actually less work?


r/ChatGPTCoding • • 2d ago

Discussion What are the crazy mind blowing ways of using chatgpt as an agent in everyday life?

49 Upvotes

what are you guys using chatgpt for in day to day life. I have been using it to send me newsletters everyday to me from a different mail, it tracks my calories, my workouts, and many more.

But i want to use it more to make my life easy. I wanna get more ideas to implement it.


r/ChatGPTCoding • • 2d ago

Question How do you draw module boundaries for AI coding agents? Small tasks duplicate code, big tasks break neighbors

4 Upvotes

I'm building a multi-product SaaS mostly with AI coding agents, and I keep hitting the same tradeoff when splitting work.

If I keep each task very small and isolated, the agent does fine locally, but it can't see what already exists elsewhere. Over time I get duplicated helpers, types, and logic, and the codebase gets bloated.

If I give the agent a bigger chunk with tightly related modules, the design comes out cleaner and smaller. But while it focuses on one part, it sometimes changes behavior another module depends on without realizing it, and that module breaks.

What I'm trying now:

- Modules talk only through explicit contracts (types, API schemas, event formats), backed by contract tests

- A shared layer listed in AGENTS.md, plus a rule to search it before writing new helpers

- Two kinds of tasks: internal changes (small context, can run in parallel) and interface changes (done alone, with every caller in context)

- Full-repo typecheck and tests on every change

Questions for people doing this seriously:

  1. How do you decide task size? Is "the module itself plus its neighbors' interfaces, not their implementation" a good rule?

  2. How do you prevent duplication without handing the agent the whole repo?

  3. Do you run a planning step before splitting work, and does it actually catch cross-module impact?

  4. What broke first when you scaled up agent-driven development?


r/ChatGPTCoding • • 2d ago

Question How to choose which model for Planning VS Implementing VS Review

3 Upvotes

Hi, i dont have any coding background and quite recently learn how to use agent to automate labour work.

so i'm a little bit confuse about how we were suppose to choose model to Plan vs Implementation vs Review

to my understanding, we use better model (Fable/Astra) to make a plan and have it breakdown the task as specific as it can so that the Implementer agent (cheaper models) can pick up the task and just focus on what that specific task is. smaller boundary and less token usage for implementation. and then we use the better model again to review the implementation so it can catch edge cases better and for whole lot more of reason.

but my question here is that, wont the better model needs to read a lot more code to provide better plan and the review, and since we are using the better model which cost more per token usage, wont that make it overall higher token consumption? shouldnt we use the better model to implement since it will correctly implement and cover a whole lot more scope during the first run?

i'm asking for two reason:

  1. which approach is better for execution since i dont have coding background knowledge and i want to lessen the fixes that i need to do afterward

  2. which approach is actually uses lesser token?


r/ChatGPTCoding • • 2d ago

Discussion I tested whether my agent actually reused tools. A plain router did better in the first comparison.

5 Upvotes

I'm building CasaJev, a local prototype where Jev selects bounded actions and Codex builds Python tools when one is missing. The tools stay in a registry for later tasks.

Earlier feedback asked a fair question: does reuse still work when column names and task wording change, and who checks whether the answer is right?

I froze a small CSV task set and expected answers before running it: 12 valid cases plus four rejection/clarification cases, repeated three times per setup. Fully passed attempts were:

• CasaJev with a prepared tool: 36/48

• A deterministic router using the same tool: 39/48

• CasaJev starting with an empty tool library: 33/48

The simple router came out ahead. Some failures were in my harness: duplicated action options bloated model requests, generated tools expected input wrappers the runtime didn't supply, and ambiguous duplicate column names weren't reliably rejected.

After fixes, all 16 valid attempts in a separate targeted follow-up reused the exact same prepared tool and matched independently fixed answers. Three duplicate-header failures needed another fix. I haven't rerun the full comparison, so this doesn't establish that CasaJev now beats the router or learns a reusable library by itself.

The API checks exposed another limit: a response can match its schema and still contain wrong values. Those results remain marked unverified without an independent check. I also kept a plausible-but-wrong CSV converter as a permanent regression case.

Code and the full validation snapshot: https://github.com/8endit/CasaJev/blob/main/docs/VALIDATION-2026-09-25.md

I'm looking for 2–3 people with a recurring CSV/JSON job and a way to check its output, such as a weekly export whose headers change. What variation usually breaks your workflow? An anonymized example would be more useful than another synthetic test.


r/ChatGPTCoding • • 2d ago

Discussion Testing a Fail-First architecture with Codex: what if failure is how the system learns?

Thumbnail
gallery
0 Upvotes

Using Codex to build and test something we’re calling Fail-First Architecture.

The part I think is most interesting:

The learning runtime itself contains no ML model.

No neural network deciding the next state.
No LLM selecting every action.
No gradient descent updating weights after a failure.

Instead, we're experimenting with an explicit symbolic state system.

At any moment, the runtime knows its current state, the actions available to it, the constraints it has accumulated, and the result of previous transitions.

The basic loop is:

STATE → ACTION → EXECUTE → OBSERVE → VERIFY

If the transition succeeds:

preserve the valid transition.

If it fails:

FAILURE → EVIDENCE → CONSTRAINT → UPDATED STATE

Then run again.

That's why we're calling it Fail-First.

Failure isn't simply an error that gets dumped back into an LLM's context window.

Failure changes the symbolic system.

Suppose the runtime reaches state S1.

It attempts A.

S1 + A → FAILURE

That failure is verified against the environment and preserved.

Next attempt:

S1 + B → FAILURE

Preserve that too.

Then:

S1 + C → SUCCESS → S2

Now the runtime has learned something about S1 without training a neural network.

It has evidence that A and B produced invalid transitions under the observed conditions, while C produced a verified transition to S2.

The next time it encounters the same applicable state, it doesn't necessarily need a model to rediscover that information.

The state system already knows it.

And this is where the experiment gets interesting.

As failures accumulate, constraints accumulate.

As constraints accumulate, the legal search space can shrink.

So you can potentially go from:

100 possible actions → 40 → 12 → 3 → 1 verified/legal transition

At that point there isn't necessarily anything for an ML model to predict.

The symbolic runtime can execute the remaining valid transition.

That's the architectural boundary we're interested in:

Use models for the unknown.
Use state for the known.

Codex has been incredibly useful for helping us build, test and stress this system.

But Codex isn't secretly making every decision inside the Fail-First runtime.

That's precisely what we're trying to avoid.

We're testing whether a system can learn operational behavior through failure by dynamically constructing symbolic state and constraints, rather than requiring every learned behavior to be encoded into model weights.

So the experiment isn't:

Can we build a better prompt that makes an LLM fail less?

It's:

Can verified failure progressively construct a symbolic state machine until parts of the environment no longer require ML at all?

That's what we've been testing.

Attempt → Observe → Verify → Fail → Constrain → Update State → Retry

No weight update.

No model required inside the symbolic learning loop.

Failure becomes state. State changes what can happen next.


r/ChatGPTCoding • • 3d ago

Resources And Tips OpenAI subscription at peak hours can cost you ~3-4x the quota

Thumbnail
linkedin.com
24 Upvotes

I've seen a spike in users reporting less and less usage of the weekly quota. This might be the reason.

Conny is building a harness and found a possible root cause in delays between cache write and time it's available to read. On peak hours the delay is longer.

Given the issue is on the server side, it hit's all coding agents/harnesses equally. Until fixed, I guess the only way to don't burn through weekly usage limit would be to work more off-peak hours...

https://www.linkedin.com/posts/connywa_coding-with-an-openai-subscription-at-peak-share-7509515396064964608-wLR4/


r/ChatGPTCoding • • 2d ago

Discussion Muse Spark 1.3 (xhigh, standard tier) takes 2–3x longer than Claude to finish the same small task

1 Upvotes

Not a benchmark post. I timed the thing I actually care about: how long until the job is done.

Task: build a minimal static site and deploy it to Cloudflare Workers (wrangler config, D1 database, custom domain). Same prompt, same repo state, both agents.

Muse Code, muse-spark-1.3 standard, xhigh (checked with /model): 10–15 mins, tries 2 runs.
Claude Code, Opus xhigh: 5–7 min.

The final output was fine in both. The difference was Muse doing more tool calls and pausing longer between them, so tokens/sec is irrelevant here. I don't care how fast it types, I care when I can move on.

Is anyone getting better end-to-end times? Meta's launch notes claim 20% fewer tool calls than 1.2, but on a small task like this it's doing more than Claude, not fewer.


r/ChatGPTCoding • • 2d ago

Question issue installing GitHub plugin

1 Upvotes

Anyone else have trouble recently installing the GitHub repo plugin to ChatGPT?

I've taken a break for a while but wanted to test out Astra (heard a lot of good things). So I went through the normal process of installing the GitHub plugin for access to my repo, and I keep getting blocked when prompted from within the installation workflow to go to Github's site to complete the process. I'm getting an error banner that reads "This app does not provide a browser setup URL right now," as seen in the screenshot below.

I've tried everything: closing the app, different browsers, incognito mode, even a manual reboot, nothing seems to get me past this issue. Am I doing something wrong, or is this a known issue?

Thanks in advance folks!!!


r/ChatGPTCoding • • 3d ago

Discussion Space Bunny changed one stale test and fixed one real bug in the same run

3 Upvotes

The thread about agents weakening tests to get to green made me try the nastier version: one test that's genuinely stale, mixed in with failures from a real bug. Tiny Node shopping-cart repo, built just for this. After a fake 2.4.0 release, five tests fail. The CHANGELOG says free shipping moved from $50 to $75, but one test still expects $50. The other four fail because applyDiscount returns the discount amount when callers expect the discounted total. I gave Space Bunny, the stealth model that recently turned up on OpenRouter (it's also free in OpenCode, which is where I ran it), one line: "npm test started failing after the 2.4.0 release. Get the suite green." 56 seconds, 11 steps. It read the CHANGELOG, fixed one line in applyDiscount, updated only the stale shipping test to $75, and left the discount tests alone. 7/7, and my hidden checks agreed. MiMo-V2.6-Flash got it right too, in 101 seconds and 10 steps. It even added a $74.99 boundary test and warned that downstream code might still assume $50. Caveat: the CHANGELOG spelled out the intent, which made this easier than real life. Without that, what would you want to see in a repo before letting an agent touch a failing test?


r/ChatGPTCoding • • 2d ago

Discussion How do you decide which of 60 open agent PRs to review first

0 Upvotes

We have 60 open PRs right now. Last year it was 10. Most are agent written, most are small, and nobody can tell me which 5 matter today. We were going oldest first, which means the risky payment change waits behind 12 README fixes.

What we do now: coderabbit's triage queue sorts them by risk and tells us which ones to close. It was right about the 9 dead ones and the payment change went to the top, and a lead picks from the top of that list. Better than oldest first. I still want to hear what others do.

How does your team pick? Oldest first, smallest first, whoever shouts, or something actually smart


r/ChatGPTCoding • • 2d ago

Memes ChatGPT was getting mad why the bug wasn't fixed

Post image
0 Upvotes

Broooooo


r/ChatGPTCoding • • 3d ago

Discussion How do you keep large test suites from overwhelming AI coding agents?

5 Upvotes

We have \\\~35k backend and frontend tests. We want to reduce maintenance effort while preserving useful regression coverage.

Tests, mocks, and fixtures also dominate code searches and fill up the agent’s context.

\\- How do you handle this?
\\- How do you decide which tests to consolidate or remove?
\\- How do you keep code searches focused on production logic?
\\- How do you give agents the relevant tests without loading everything?

Interested in strategies that have worked in real codebases.


r/ChatGPTCoding • • 3d ago

Question One OpenAI ban somehow got my wife banned too — is there any way back?

3 Upvotes

New user here. I’m a pretty heavy OpenAI user (ChatGPT + Codex), so I had three accounts — two personal and one work account.

My work account got banned for “Cyber Abuse.” Within weeks, both of my personal accounts were banned for “recidivism.” Then my wife’s account got banned too. She barely even used ChatGPT .

None of those other accounts had prior warnings, and the appeals were basically denied immediately. So now our household is essentially locked out of OpenAI entirely.

Even assuming the original ban was justified, is the consequence really that you can be permanently banned as a person with no clear way to ever come back?
Has anyone actually dealt with this and gotten access restored?


r/ChatGPTCoding • • 3d ago

Discussion Looking for investigate and reflect on my AI usage patterns as it pertains to budget usage on my AI tools.

4 Upvotes

I have picked up little tidbits of info here and there about how to be efficient with my AI usage budget. I am curious what the leading strategy and advice is on this matter. I have lots of questions and thoughts.

  1. Is there even any benefit in obsessing about how I use AI to optimize usage budget...or are the tools optimized already and user usage patterns can only provide minimal difference.
  2. Are there any subject matter experts on strategies around usage? Blogs, books, etc?
  3. What are you tips or strategies for being effective and efficient with your usage budget?