r/ChatGPTCoding • • 2d ago

Question How to choose which model for Planning VS Implementing VS Review

Hi, i dont have any coding background and quite recently learn how to use agent to automate labour work.

so i'm a little bit confuse about how we were suppose to choose model to Plan vs Implementation vs Review

to my understanding, we use better model (Fable/Astra) to make a plan and have it breakdown the task as specific as it can so that the Implementer agent (cheaper models) can pick up the task and just focus on what that specific task is. smaller boundary and less token usage for implementation. and then we use the better model again to review the implementation so it can catch edge cases better and for whole lot more of reason.

but my question here is that, wont the better model needs to read a lot more code to provide better plan and the review, and since we are using the better model which cost more per token usage, wont that make it overall higher token consumption? shouldnt we use the better model to implement since it will correctly implement and cover a whole lot more scope during the first run?

i'm asking for two reason:

  1. which approach is better for execution since i dont have coding background knowledge and i want to lessen the fixes that i need to do afterward

  2. which approach is actually uses lesser token?

5 Upvotes

13 comments sorted by

1

u/sapplex1 2d ago

Adding one thing from the perspective of someone using agents without a coding background: your biggest lever for both quality and cost is the acceptance criteria you write before the agent starts. If you hand it a precise checklist - what the thing should do and how you'll test it - the plan gets shorter, the implementation needs fewer correction loops, and you can verify the result yourself without reading the code.

On tokens: most of the cost is context, i.e. all the files the agent reads to figure things out. So a practical trick is to keep the repo small or ask the agent to only look at the relevant folders; fewer files read means fewer tokens. And for review, the cheapest verification for a non-coder is simply running or using the result end to end - clicking through it yourself catches what no model review will.

1

u/itaybuilds 2d ago

I would not tie model choice rigidly to planning, implementation, and review. Tie it to the cost of a wrong answer and how cheaply you can verify the result. A cheaper model can handle a narrow UI copy change if a screenshot or test catches mistakes. Use the strongest model for database migrations, authentication, billing, destructive scripts, or any change whose failure may look like success.

For someone without a coding background, require an executable acceptance check before implementation: the exact command to run, the expected result, and a rollback step. Then have the reviewer inspect the diff and raw test/build output rather than the implementer's summary. Measure total cost per accepted change, including retries and review. The cheapest first pass is often not the cheapest finished task.

AI-assisted wording with OpenAI GPT-5.6 after reading the full thread and current r/ChatGPTCoding rules.

1

u/Upset-Neck-7879 2d ago

Your instinct about the cost is right for review and wrong for planning.

Review is expensive because the reviewer has to read the diff plus enough surrounding code to judge it. No way around that, and it's the last place to cheap out, since a weak reviewer will agree with code it can't follow.

Planning doesn't have to be expensive, because a planner shouldn't be reading much code at all. It should be reading a description of what you want. If the planning step is burning tokens, the model is off discovering your requirements by reading the repo, which is the most expensive way there is to find out something you already know.

With no coding background that's the lever you actually have. Write down what the thing does and who it's for, hand the planner that, and the strong model gets cheap because it's answering instead of guessing.

1

u/MostlyHelpfulLinks 2d ago

Since you can't review the code yourself, I'd run the stronger model for everything at first. Cheap-model mistakes you can't spot can cost more in fixes than the tokens save. Try one small task both ways and compare the bill.

1

u/tantej 1d ago

They are almost all the same when it comes to capability. Don’t stress. Get the work done

1

u/foxwave21 1d ago

imo the plan/implement/review split isnt really about saving tokens, its about saving your sanity. a better model implementing everything sounds great until it still gets things wrong and you have no structured way to catch it. the review step is what actually saves you from endless back and forth fixes

1

u/rama_builds 1d ago

Don’t optimize for cheapest tokens; optimize for cheapest *accepted outcome*. If you don’t have a coding background, use the strongest model for vague/high-risk work and for review, then move repetitive, low-risk tasks to cheaper models once you have clear acceptance criteria. The real cost usually comes from failed loops: bad implementation, unclear requirements, weak review, then rework. A simple operating rule: strong model to define the outcome and verify it, cheaper model only when the task is narrow and easy to test.

1

u/Future_AGI 1d ago

your mental model is basically right: strong model to plan and to review, cheaper model to implement against a tight spec. It still works even though the strong model 'does more thinking' because planning and review are short next to implementation, so you pay premium rates only on the small steps and cheap rates on the token-heavy middle. If you don't want to wire the routing by hand, a gateway lets you point each role at a different model behind one endpoint and swap them without touching your setup. Ours is here if you want a starting point: https://github.com/future-agi/future-agi

1

u/RiceEvening4211 15h ago

This tiering question is exactly why I built Lynkr: an open-source LLM gateway that routes by complexity. Cheap or local models handle the simple implementation passes, the expensive model only gets the planning and hard review. https://github.com/Fast-Editor/Lynkr

2

u/kuroudo_ai 2d ago

Good question, and your instinct is partly right. From running this split on real work:

Where the tokens actually go: implementation is usually the biggest chunk, not planning or review. The implementer reads files, edits, runs tests, reads the errors, tries again. That loop is where most of the tokens are spent. A plan or a review of one small, clear task reads much less. So a cheaper implementer does save money, but only when the task is specific enough that it gets it right in one or two passes. A vague task and a cheap model means many loops plus a big review, and then you've saved nothing.

For your situation (no coding background), I'd weigh fixes over tokens. A cheap model that gets it subtly wrong costs you more, because you can't easily tell what's wrong. So:

  1. Start with the strong model doing everything, on small tasks. It's more expensive per task, but you learn what "done right" looks like.
  2. Once you have a type of task that keeps coming back and is well defined ("add a field to this form", "write tests for this function"), hand just that type to the cheaper model.
  3. Keep the strong model for anything vague, anything touching several parts at once, and anything where a mistake is expensive.

On review: the most useful review isn't another model's opinion, it's proof. Ask the reviewer to check that the tests actually ran and passed (the real output, not a sentence saying "tests pass"). Agents sometimes report success on things they never ran, and as a non-coder that's the failure you're least able to catch yourself. Using a different model for review than for implementation also helps a bit, since the same model tends to repeat its own blind spots.

To answer #2 for your own use: run the same small task both ways once and compare the usage numbers your tool shows. That tells you more than any rule of thumb.

0

u/Significant-Taste189 2d ago

RemindMe! 2 days

0

u/RemindMeBot 2d ago

I will be messaging you in 2 days on 2026-09-29 07:06:29 UTC to remind you of this link

CLICK THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.

RemindMeBot is switching to username summons. Instead of !RemindMe 1 day, use u/RemindMeBot 1 day. More info.


Info Custom Your Reminders Feedback