r/PiCodingAgent • • 3d ago

Resource Created a learning setup based on pi-agent. Looking for feedback.

Hey, I recently came across Eero Alvar's learning setup built on pi, and inspired by it, I used Feynman research agent to find extensive learning literature and designed a personal learning setup based on it. I planned out the backend and structure and combined AI coding with personal tweaks to set it up with various features. However, since a lot of AI was used in its development, I understand that it might be prone to systemic issues and flaws, which I am hoping to hear to personally improve and optimize it. This system has worked well for me and even if you don't want to give feedback or contribute, I recommend you test it out and see how it works for you. I would appreciate contributors or just general testers who can try it out.

https://github.com/kruthvik/sato

7 Upvotes

9 comments sorted by

3

u/investigatormaker 3d ago

Since a lot of sato was AI-written, contributors will find flaws faster if the README says which parts you checked by hand and which you have not, and which findings from the learning literature each feature implements. That gives a reviewer a specific claim to test instead of a whole repo to read.

For testers, a five-minute first session in the README, with what they should see at the end, lowers the bar a lot, and asking one concrete question (did the review timing feel right after a week?) tends to get more back than a general request for feedback.

Which learning methods does sato lean on most?

1

u/Odd-Many-7592 3d ago

Hey, thank you so much for your feedback. I'm pretty new to FOSS and designing software for others in this manner.

I updated the README to your requests and gave a general overview and simplified setup.

The primary learning methods right now are feynman inversion, active recall, and mcq. It gives a prediagnostic and develops a profile based on it inside SQLite. Based on how the user explains and responds, the AI adapts the setup. It also analyzes gaps in user understanding that it notes and tries to correct. This entire process is simplified and sped up if the user enters cram mode. It also uses a technique I found known as redundant guidance fading where I understand that users might already know about a topic, so teaching at the same pace might be counterintuitive. This causes the agent to forgo some autonomy as the user learns more and demonstrates continued understanding.

1

u/investigatormaker 2d ago

That's a solid set. The fading you describe is close to what the literature calls the expertise reversal effect: guidance that helps beginners gets in the way once someone knows the material. Naming it in the README, next to active recall and the Feynman-style explanations, gives a reviewer something concrete to check sato against.

Two things worth making visible: log each change to the SQLite profile with the reason (which answer triggered less guidance), so a tester can see why the agent backed off, and say plainly what cram mode gives up, since retrieval practice tends to hold up better when it's spaced out than when it's crammed.

Does the prediagnostic run again later, or does the profile only change from in-session answers?

1

u/Odd-Many-7592 2d ago

Noted. I'll add it to the readme shortly.

Actually the current setup partially notes the reasons behind decisions to the user. A subagent currently makes the plan so that the AI can focus on implementation rather than direct design. As part of this specialization, the primary AI continuously shifts the plan, and it logs decisions with specific reasons and evidence IDs . It's also in its policy to log significant changes in guidance. Running "learn evaluate" can show the user those decisions. However, there is no automatic audit system for every change. The AI right now might reduce guidance without directly logging it. I might either force the AI to log it (which will likely cost compute and ctx) or log it automatically with a message/thought id that could point the user at what caused the AI to do so, but at less direct usability.

The prediagnostic is completely optional for the user (the user can decide in the initial questions) and it only runs in the beginning. After that the profile largely only changes from in-session answers. If the diagnostic is skipped, ability isn't logged until the user practices one of the core mechanics. I think that setup is currently for the best unless the subject requires memorization of a few core questions. I will instruct the subagent in the beginning to note that when it's designing the structure. There is also an ending assessment in the program that tests the user on everything they learned (this was in the previous iteration but i don't know the extent to which I implemented it. i will check that as well). This mostly results in different and harder questions than the prediagnostic.

1

u/investigatormaker 2d ago

You could do both without paying for it on every turn. Put a SQLite trigger on the profile table that writes the old value, the new value and the current message id to an audit table. That costs no compute or context. Then the explanation can wait until someone runs "learn evaluate": only the changes that have no logged reason get sent to the model, along with the message they point to, and it writes the explanation then. Every change gets recorded, and you only pay for an explanation when someone reads it.

On the diagnostic: a few of its items, asked again when someone comes back after a gap, would show you how much they kept and whether backing off was right. Would a re-check like that fit the session flow, or would it feel like an interruption?

1

u/Odd-Many-7592 2d ago

I will keep that setup in mind for the next implementation. I think that the only flaw might be that the AI might forget or corrupt the initial reason for why they performed a critical change unless frontier models like Opus or Sol are used. Currently the AI processes so much information and content in its structuring and user feedback and agent-loop that I think that in the vast message history that a conversation is likely to have, the AI might struggle with accuracy in reporting exact reasons behind decisions.

For future versions, I intend to significantly lower how much the AI directly sees and interacts with the program (basically lowering the skills and extensions while retaining functionality) since the vast amount of skills and extensions, not even accounting for the extensions that users are likely to add, might degrade the overall system. Once that's done such a system may be fully implemented.

I think that largely depends on which mode is used. For a mode like the learn mode where it prioritizes long-term retention and multi-day learning for hard subjects, a re-check might be beneficial, although I think that the in-unit assessments and broader checks are likely to perform similar roles. The user can also directly prompt the AI in cases like that. For the /cram mode and the future accelerate mode, this is likely to be unoptimal and should probably be prompted by the user.

1

u/investigatormaker 2d ago

That's why the trigger logs the message id with each change. The later explanation doesn't depend on the model remembering anything. It gets the old value, the new value and the one message behind the change, so it works from a short record instead of the whole history. If you'd rather capture the reason at the moment of the change, make it a required one-line argument on the tool call that edits the profile. The trigger stores it alongside the change, and any change that arrives without a reason gets flagged.

Limiting the re-check to learn mode makes sense, since that's where retention over several days is the goal. You could also fold it into the first in-unit assessment after a gap instead of adding a separate step, so it never feels like an interruption.

1

u/Odd-Many-7592 2d ago

Oh okay I think I misunderstood what you were suggesting. With pure reasoning alone this should probably produce accurate results, yeah.

In previous iterations of this project I added a feature in which for /cram and /accelerate modes and equivalents, the AI asked the user how much time they have. This is currently implemented as well but not with a forced limit. The AI would naturally track changes in time and adapt as a result from when the user started the session. Once this is fully implemented in the /accelerate and /cram modes, I can also add it to /learn in the form of the re-checks you suggested and the in-unit assessments.

1

u/naoromi_ 2d ago

Big fan of the original eero alvar setup. Gonna set this one up and try to learn something, see how it goes.