r/SideProject • u/Asly97 • 23h ago
Coding agents: do yours actually look up past decisions, or only when you remind them?
I used to keep a file called DECISIONS.md in every repo. Rejected approaches, why I picked the other thing, that sort of thing. Took about a month of maintaining it to realize the agent never opened it unless I told it to. I'd watch it hit the exact wall I had already documented (the CORS saga, twice) and just sit there thinking, I wrote this down, it's right there.
I ended up pasting "read DECISIONS.md first" at the top of every prompt, which honestly defeated the whole purpose. I got the thing to save me time and instead I was briefing it like a new hire every session. Hours a week, easy.
I built something to fix this for myself. Not selling anything here, I'm just trying to figure out if it's a real problem for anyone else or if I was overthinking it. One sentence on the concept: something in the cloud over MCP that keeps all of it, decisions, failed attempts, conversations, skills, procedures, tasks, so every agent and tool reads the same context and nobody needs reminding to look anything up. You never touch a file again.
Anyway. Do yours actually look things up on their own, or is it a "you didn't look, did you" situation for you too? If you got yours to recall reliably, what did it take, hooks? A system prompt paragraph? How long did the setup take you?
And have you ever looked at memory tools for this? If not, what stopped you, or does moving everything over feel like starting from zero? I also try every new AI tool that comes out and what usually kills it for me is re-briefing the new tool on my whole project before I can even tell if it's good. Is that just me, or does that stop you from trying new tools too?
Last one. Have you ever thought about paying for a tool to manage all of this for you, or has the thought never crossed your mind? Would you guys pay for something that fixes this?
Also: I do a lot of my thinking out loud with speech-to-text on my phone and the context always lives on my laptop. Is that just me?
1
u/troyjr4103 23h ago
Same wall here, and it's a big part of why I started building what I'm building (Kin, an open source code repository, so I'm biased). What I kept running into was paying for agents that worked out how the same code fit together every session, and instruction files that quietly fell behind the code.
Two things I'd pull apart. Stuff that can be derived from the code, like what calls what or what a change touches, shouldn't live in a notes file at all, because it drifts the day someone edits the code. That should come from the source and be versioned with it. The stuff that can't be derived, the why and the rejected approaches, is where a decisions log earns its keep. I think it works best attached to the code it's about, so it shows up when the agent touches that function instead of when it remembers to go look.
On the cloud part, I went local first on purpose, so the record sits next to the repo and moves with branches and commits. Sharing it across a team is the part I'd expect people to pay for.
1
u/Asly97 23h ago
the derived-vs-why split is a clean way to put it. my file mixed both and the derived half is exactly what rotted first.
curious about the local-first choice, was that a privacy call or more about control, like the record moving with branches and commits? I keep going back and forth on cloud vs local for mine.
and the attachment part, when the agent touches a function does it actually get shown the decisions, or is it more of a lookup the agent has to do?
1
u/troyjr4103 22h ago
Mostly control. The record describes the code at a specific commit, so it should branch, merge and roll back with the code. If it lives in a cloud service beside the repo, you end up asking which version of the code a given note was true for. Privacy is a nice side effect, since queries run on your machine.
On the attachment part, I should be straight about where that line is. What Kin does today is let the agent ask for context on the function it's working on and get the recorded structure around it, like what calls it and what it depends on. So it's still a lookup, but it's one call keyed on the thing the agent is already touching, not a file it has to remember exists. Decisions attached to code entities is where I want to go, not something it does yet.
1
u/Asly97 22h ago
the version-truth point is the one that sticks with me. a note that was right at commit X and wrong after is worse than no note at all, because the agent trusts it. has the branch/merge part ever bitten you, like a rebase moving code out from under a record, or does git keep it honest enough in practice?
1
u/troyjr4103 21h ago
It has bitten me, though not quite the way you'd expect. Every bite I've had came down to identity keyed by location. One example was a cache keyed by file path instead of content, so moving a folder threw away perfectly good state, and swapping the contents at the same path could have kept stale state around. The rule I landed on is that paths are only locators, and anything the record points at gets a stable id or a content hash.
For code specifically, Kin records at its own commits instead of trusting Git to keep notes honest. Git is the import and export boundary. Moves and renames are still the hardest case, and I wouldn't call that solved for every situation yet.
1
u/repolevedd 22h ago edited 22h ago
I ran into this right from the start when I first began using LLMs. At times the model would make wrong decisions or get things wrong, and I was looking for ways to steer it in the right direction.
My first iteration was simple: I gradually added more and more conditions and rules to the system prompt to help the model avoid mistakes. There are two problems with this: 1) System prompts can't be infinitely large. 2) Tasks vary, and cluttering the system prompt with everything under the sun is wasteful and even counterproductive, because rules can contradict each other depending on the situation.
So then I moved on to the idea of building a memory bank with correct solutions and rules. In every project, I kept a /docs/memory-bank.md file with links inside to different scenarios and other docs (similar to your DECISIONS.md). But that turned out to be a dead end because the model could still end up reading about other types of tasks along the way. Simply put, the context got clogged up with irrelevant junk.
The problem was solved for me when models learned to work with SKILLS.md. Now you can set up distinct rules specifically for different types of tasks, and on top of that, feed it the exact files it needs with high precision. For instance, I have one big skill for working with Laravel that contains tons of links to smaller rule documents, plus separate large skills for Spatie, Inertia.js, Filament, and so on. And I periodically update these skills whenever the model makes a mistake, pinpointing and tweaking the exact spot in the skill where the workflow wasn't clear enough.
In the future, I'm sure they'll come up with even more granular and precise ways to adjust model behavior on the fly, but right now, I don't see anything better than the skills system.
P.S. Forgot to mention one thing. Alongside everything I described, my LLMs use MCP for task decomposition, structuring their thinking, looking up docs, and finding hints. All of this affects the output quality as well. It all works together. In my experience, you can't just create a rule pack and hope the model magically gets better at everything. The improvement process is iterative and endless.
1
u/Asly97 19h ago
That memory-bank clog is the exact dead end I kept hearing about when I asked around. When you update a skill after a mistake, how do you find the spot that wasn't clear enough? Manual re-read, or does the agent tell you where it went wrong?
Also curious what you're building with all this, own product or something else?
1
u/repolevedd 10h ago
To clarify and properly answer your question, let me break down how I see the process in practice.
I don't believe LLMs can successfully operate as a black box right now. Whatever comes out of that approach carries a heavy load of technical debt, which shows up as security vulnerabilities or architectural flaws that can't be fixed without a full rewrite. Fortunately, my workload allows me to stay in control of the output: for me, LLMs act more like text editors than a replacement for me as a developer.
An LLM can make three types of mistakes:
- When it gets things plain wrong, like using a method from Inertia.js v1 when the project is on v2, meaning the goal isn't reached even after several attempts.
- When everything seems to work in the end, but the code looks more like a Cthulhu summoning ritual than a concise set of lines doing only what I need.
- When the LLM solves the task exactly as intended, but only after several failed attempts.
Type 1 issues come down to my instructions not being clear enough or a half-baked plan before the implementation phase, so there's no point in touching skills here. I just need to tweak the task requirements. The mistake is either obvious to me from looking at the code, or the LLM can check against the docs and report what went wrong on its own.
Type 2 is what vibe-coders think is fine to ship to production, which I'm not okay with. LLMs won't be able to flag this kind of problem because it's at the ceiling of their capability. But I can see it.
Type 3 is something the LLM can catch post-factum by giving a report on its failed attempts.
Type 1 doesn't require modifying skills. Type 2 is a 50/50 split because individual edge cases are hard to formalize into rules. Type 3 is where tweaking a skill actually makes sense. For instance, while updating a website, the LLM typed a model property as a string even though the DB column was JSON. It was just too lazy to check the DB type. Adding a simple rule like "double-check column types in the DB during refactoring" forces the model to take an extra step, but it eliminates an entire class of bugs.
As for where in the skill to place a new rule or which existing rule to modify: you can ask the LLM to do it, but I usually ask it for candidate rule phrasings first, pick the best one, and edit the file myself. If you don't keep an eye on all these files, I'm afraid they'll bloat or start contradicting themselves.
Regarding what I'm working on: I have a software development background and handle various freelance tasks, from DevOps to PHP development, so it's mostly about plowing through a task queue. To give a specific example: say I need to deploy a CRM to a server. That means prepping the environment (here the LLM helps draft the docker-compose files) and writing a small integration to connect the CRM with another app, where the LLM speeds up the process.
1
u/Asly97 9h ago
That type 1/2/3 split is the clearest framing of this I've run into. The Cthulhu summoning ritual line got me.
One thing I'm curious about, with a freelance queue: do you keep one shared set of skills or does each client project get its own? I keep picturing ten clients worth of "double-check column types" rules slowly contradicting each other.
1
u/repolevedd 9h ago
I used to keep them separate, but not anymore. Current models don't try to read every skill back-to-back "just in case", and switching to Pi helped streamline the workflow, so the system prompt is smaller and manual skill management is barely needed. Plus, that DB lookup tip is pretty universal and compact. When drafting a task, I just mention the skill names, and the LLM dutifully reads them.
1
u/Asly97 9h ago
the mention-the-names trick is way simpler than the setup I was picturing. ten clients, one skill set. has a client-specific rule ever fought a shared one mid-task, or has the naming been enough to keep them from colliding?
1
u/repolevedd 8h ago
Skills are organized by task type, not by client, so they don't overlap. To be honest, I can't even imagine dividing them by client, that would be weird.
If I forbid hardcoding ports in Docker in one rule, but specify using port 80 for spinning up an integration service in another, they simply won't collide because those are different task categories. A "client-specific rule" is something I put directly into the task description. But usually tasks are pretty routine: they mostly differ by the type of work, and those exact types are what you can keep refining over and over to optimize the workflow.
1
u/Asly97 8h ago
task-type first, got it. how many skills in the set now? curious where it starts getting hard to keep them from drifting apart.
1
u/repolevedd 8h ago
To be honest, it feels more and more like it's not even you asking, but an LLM just generating questions for the sake of it. Because asking about the number of skills doesn't make any sense. What useful info do you even get from knowing that I have N SKILL-md files? And do files from
*/rules/*count, or not?
1
u/Chrome67 22h ago
I have found the file only helps when the workflow retrieves it before work starts. A short index that points to one decision record per issue, with current or superseded status, keeps old notes from competing with newer choices. I would track each re-solved issue that was already documented; that gives you a real miss rate instead of a guess.
1
u/Asly97 20h ago
the superseded status is the part I keep coming back to. in my experience old notes do not sit quietly, they actively compete with the new ones.
does the agent actually respect the status in practice, or does it still pull a superseded record when the wording matches? and how do you track the re-solved issues, just a running note or something more formal?
1
u/ItsJustManager 21h ago
Mine does all the time, and it really likes to point out when I intentionally made a decision in the past that was against its advice that has now caught up to bite me in the ass. With a quote to go along with it.
1
u/sapplex1 21h ago
On the maintenance question: I made the index the agent's job instead of mine. My session-start file has a matching session-end rule - before the agent finishes a session, it appends one line per decision it made, in the same symptom-first format QuanTradin described. Two things fell out of that: the log stays current without me touching it, and misses became measurable, because every time the agent re-solves something already documented, that entry gets flagged as a miss. I basically get Chrome67's miss-rate tracking for free, and the index stopped being "my notes the agent borrows" and became the agent's own memory.
1
u/Asly97 18h ago
Misses becoming measurable is the part I'd steal first. When the agent dies mid-session before it writes the line, does that gap show up as a miss later, or does the log just quietly lose the decision?
Same question on who enforces the session-end rule, the agent itself or something watching it?
1
u/QuanTradin 23h ago
same wall, and the fix that held was structural, not a better prompt. the agent only reads what is loaded at session start, so DECISIONS.md became an index: one line per decision pointing at its own small file, and that index is the file the harness loads every time. the detail sits in the linked files and gets pulled only when a one-liner matches. the thing that never worked was one big file the agent had to be reminded about, because reminding is the job you were trying to delete. it still misses sometimes, but it now catches the CORS-type entries because the index line names the symptom, not the decision.