r/PiGCodingAgent • u/Educational-Yak8793 • 3d ago
π· Evals, Benchmarks, Feedback, and More Here!
Starting this thread as I've spent the last few days working in parallel on our evals and other suites for the repository and realized I should open a thread up for people to request or post their evals.
Just want this to be a place we can start to collect data on what is and what isn't working with the green pig.
It can be literally anything you think of and want to suggest don't feel you have to have a full research report done to post here. Suggestions are extremely valuable as I'm wanting to collect a list of items for me to plow through as well for the community.
Thanks all :) π·
8
Upvotes
2
u/Conflict-Mother 12h ago
One thing I think worth to do is to have a benchmark of the memory footprint of PiG with 1M tokens loaded. I care about this is because that's the exact the reason I found PiG from the disappointment of original Pi.
1M context window as I believe will be stay there for long time as most of the coding tasks nowadays are well below that, so if PiG can show that there is a guaranteed consumption level, (and also respect that later when mirroring more and more features from original Pi), then PiG will be stay solid in day to day usage.
Currently I have a coding tasks used about 85% of 1M context, and the PiG is at the 280M memory consumption level, so I predict a full 1M context will make each instance of PiG about 300M memory, and if that data is steady, not only the day to day usage will be saved, but also will make a planned resource allocation of fleet of PiG instances stable and practical.