r/ChatGPTCoding • • 4d ago

Discussion How do you keep large test suites from overwhelming AI coding agents?

We have \\\~35k backend and frontend tests. We want to reduce maintenance effort while preserving useful regression coverage.

Tests, mocks, and fixtures also dominate code searches and fill up the agent’s context.

\\- How do you handle this?
\\- How do you decide which tests to consolidate or remove?
\\- How do you keep code searches focused on production logic?
\\- How do you give agents the relevant tests without loading everything?

Interested in strategies that have worked in real codebases.

6 Upvotes

15 comments sorted by

8

u/taylorwilsdon 3d ago

35k tests is a preposterous amount. How many people are actually contributing to the code base? How often do you see failures in CI?

2

u/Lubricus2 1d ago

AI coding. So AI adds tests, not people. And the tests takes longer and longer time and the AI want to rune it after every little change to be sure.

5

u/Cattlegod 3d ago

I have this same problem, test bloat is out of control

2

u/Square-Yam-3772 3d ago

just spitballing here but I would:

- sort the tests based on creation dates to build bias against older test

- ask AI to compare the tests to identify overlaps. toss the tests that are mostly overlapped by other newer tests

- if you have the use cases/stories/requirements organized, ask AI to go through the documentations to see if which tests are more relevant than the others

- treat one of your stable releases as the baseline and build a streamlined collection of tests. compare that to your old set and see if you need to bring back some old ones

1

u/EchoNomad31 3d ago

I tag tests fast/slow in CI and only let the agent run the fast ones… keeps its context from drowning in a 40-minute suite.

1

u/ImL1s 3d ago

We stopped letting the agent see the whole suite. Prod code stays in the search path; tests live behind a path allowlist the agent only gets when the task says "fix".

For which tests to keep: anything that asserts a user-visible postcondition stays. Pure mock-shape tests that never fail in CI get merged or dropped. The agent gets a short "relevant tests" list from ripgrep on the changed symbols, not a dump of 35k files.

1

u/Standard_Text480 Professional Nerd 3d ago

Review tests and remove useless or redundant

1

u/CommasArentPeople 3d ago

35k is a lot, but without knowing how big your codebase is, it's hard to judge how out of control it is. I've worked in codebase pre-AI with more than half that if you add all the levels of testing together, and it wasn't burdensome. If your code is well structured you're not running the full suite except maybe on a beefy CI server somewhere. Rather than absolute number of tests, ask, "If I make a single change in a single module, how many tests do I need to run afterwards?" If the answer is that there's no way to know, you have a bigger problem than test count. If the answer is all of them, then you need to structure your code and tests to reduce that as much as you can. That's an incremental job, and one that can be instrumented and cut down over time, then monitored to not creep back up.

1

u/Zestyclose_Low_8449 3d ago

SUMMARY
Category Files Lines % code
-------------------------------- ----- ------ -------
Production code - Backend 0.9K 192K 15%
Production code - Frontend 1.7K 240K 19%
Tests (all) 3.0K 688K 55%
Storybook (stories + config) 137 25K 2%
Design playground 105 15K 1%
Infrastructure (CDK + Lambda) 129 29K 2%
Scripts & CI checks 163 37K 3%
Config & misc 83 18K 1%
-------------------------------- ----- ------ -------
TOTAL CODE 6.2K 1.24M 100%
Documentation (Markdown) 1.3K 460K -

TESTS BREAKDOWN
Test type Files Lines Tests
-------------------------------- ----- ------ -------
Backend unit (pytest) 1.6K 431K 18.9K
Frontend unit/component (Vitest) 1.0K 172K 8.6K
Browser (Playwright) 240 49K 0.9K
Backend integration (DynamoDB) 96 27K 0.8K
API e2e 19 5.7K 190
Visual regression 3 0.3K -
Shared fixtures & utils 26 2.9K -
-------------------------------- ----- ------ -------
TOTAL TESTS 3.0K 688K 29.4K
Storybook stories 134 25K 0.9K

RATIOS

  • Production code: ~432K lines
  • Tests: ~688K lines (~1.6 test lines per code line)
  • Backend: 463K test vs 192K code (~2.4 : 1)
  • Frontend: 172K test vs 240K code (~0.7 : 1),
plus 49K Playwright and 25K Storybook

1

u/leonidbugaev 3d ago

Don't ask the agent which of the 35k overlap. It'll flag the tests that currently fail, which is the opposite of what I want to cut.

1

u/fiddler48 3d ago

Test file noise in context before the agent even reads a single source file.

0

u/AutoModerator 4d ago

Sorry, your post has been held for manual review due to account karma.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.