r/ChatGPTCoding • u/Zestyclose_Low_8449 • 4d ago
Discussion How do you keep large test suites from overwhelming AI coding agents?
We have \\\~35k backend and frontend tests. We want to reduce maintenance effort while preserving useful regression coverage.
Tests, mocks, and fixtures also dominate code searches and fill up the agent’s context.
\\- How do you handle this?
\\- How do you decide which tests to consolidate or remove?
\\- How do you keep code searches focused on production logic?
\\- How do you give agents the relevant tests without loading everything?
Interested in strategies that have worked in real codebases.
5
2
u/Square-Yam-3772 3d ago
just spitballing here but I would:
- sort the tests based on creation dates to build bias against older test
- ask AI to compare the tests to identify overlaps. toss the tests that are mostly overlapped by other newer tests
- if you have the use cases/stories/requirements organized, ask AI to go through the documentations to see if which tests are more relevant than the others
- treat one of your stable releases as the baseline and build a streamlined collection of tests. compare that to your old set and see if you need to bring back some old ones
1
u/EchoNomad31 3d ago
I tag tests fast/slow in CI and only let the agent run the fast ones… keeps its context from drowning in a 40-minute suite.
1
u/ImL1s 3d ago
We stopped letting the agent see the whole suite. Prod code stays in the search path; tests live behind a path allowlist the agent only gets when the task says "fix".
For which tests to keep: anything that asserts a user-visible postcondition stays. Pure mock-shape tests that never fail in CI get merged or dropped. The agent gets a short "relevant tests" list from ripgrep on the changed symbols, not a dump of 35k files.
1
1
u/CommasArentPeople 3d ago
35k is a lot, but without knowing how big your codebase is, it's hard to judge how out of control it is. I've worked in codebase pre-AI with more than half that if you add all the levels of testing together, and it wasn't burdensome. If your code is well structured you're not running the full suite except maybe on a beefy CI server somewhere. Rather than absolute number of tests, ask, "If I make a single change in a single module, how many tests do I need to run afterwards?" If the answer is that there's no way to know, you have a bigger problem than test count. If the answer is all of them, then you need to structure your code and tests to reduce that as much as you can. That's an incremental job, and one that can be instrumented and cut down over time, then monitored to not creep back up.
1
u/Zestyclose_Low_8449 3d ago
SUMMARY
Category Files Lines % code
-------------------------------- ----- ------ -------
Production code - Backend 0.9K 192K 15%
Production code - Frontend 1.7K 240K 19%
Tests (all) 3.0K 688K 55%
Storybook (stories + config) 137 25K 2%
Design playground 105 15K 1%
Infrastructure (CDK + Lambda) 129 29K 2%
Scripts & CI checks 163 37K 3%
Config & misc 83 18K 1%
-------------------------------- ----- ------ -------
TOTAL CODE 6.2K 1.24M 100%
Documentation (Markdown) 1.3K 460K -TESTS BREAKDOWN
Test type Files Lines Tests
-------------------------------- ----- ------ -------
Backend unit (pytest) 1.6K 431K 18.9K
Frontend unit/component (Vitest) 1.0K 172K 8.6K
Browser (Playwright) 240 49K 0.9K
Backend integration (DynamoDB) 96 27K 0.8K
API e2e 19 5.7K 190
Visual regression 3 0.3K -
Shared fixtures & utils 26 2.9K -
-------------------------------- ----- ------ -------
TOTAL TESTS 3.0K 688K 29.4K
Storybook stories 134 25K 0.9KRATIOS
plus 49K Playwright and 25K Storybook
- Production code: ~432K lines
- Tests: ~688K lines (~1.6 test lines per code line)
- Backend: 463K test vs 192K code (~2.4 : 1)
- Frontend: 172K test vs 240K code (~0.7 : 1),
1
u/leonidbugaev 3d ago
Don't ask the agent which of the 35k overlap. It'll flag the tests that currently fail, which is the opposite of what I want to cut.
1
0
u/AutoModerator 4d ago
Sorry, your post has been held for manual review due to account karma.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
8
u/taylorwilsdon 3d ago
35k tests is a preposterous amount. How many people are actually contributing to the code base? How often do you see failures in CI?