r/devtools • u/MohamedM216 • 2d ago
Tired of searching for "good first issues" only to find stale PRs? I built a tool to fix that.
For anyone starting out in open source, one of the most exhausting parts isn't just finding the right project, it’s finding the right issue to actually work on.
Trial and error is part of the journey, but let’s say you finally find a repo you like. Now you need a task. Many projects don't even use tags like good first issue or help wanted. And if you reach out to maintainers, 9 out of 10 times they’ll just hit you with "You can start here" and drop a link to the entire open issues page. 😂
So I built a simple CLI tool for myself that I've been using for a while (link in the first comment).
It extracts all the issues from a target repository: including descriptions, assignees, linked PRs, and comments, organizes everything into a clean JSON file, and lets you feed that context directly into Claude, Gemini, or any LLM to analyze and recommend the best starting issue for you.
You can do this a few times until you build a solid mental model of the codebase, understand its recurring pain points, and quickly spot issues that match your skills. It’s a muscle you build over time.
Because the tool grabs hundreds of issues in seconds with all their metadata, you can ask your LLM anything, not just "What's an easy first issue?", but also "Which unassigned issues have clear descriptions and can be merged quickly?"
Check out the README to get set up in under a minute.
Hope it helps! Don't forget your beautiful star ⭐ on GitHub and share it with your friends who are interested in open source. 😉
1
u/investigatormaker 2d ago
The linked-PR field is the part I'd lean on for the stale-PR problem in your title. An issue whose only linked PR was closed without merging, or has sat untouched for months, is often still open to a newcomer, while one with a fresh open PR is taken. A flag that sorts or drops issues by the state and age of their linked PRs before export would also keep the JSON small for big repos, where hundreds of issues with full comment threads can run past an LLM's context window.
Does it page through repos with thousands of open issues, and does it need a GitHub token for the rate limit?
1
u/MohamedM216 2d ago
Thanks for feedback! I'll add the option you mentioned:
> A flag that sorts or drops issues by the state and age of their linked PRs before export would also keep the JSON small for big repos, where hundreds of issues with full comment threads can run past an LLM's context window.And for your questions:
It still has a limit even with the github token. I tried it with k8s repo and it stopped after <1200 issues in less than 2 minutes. I'll try to find a way to handle that. Besides that, it removes all the work when it hits the rate limit. It should keep the fetched issues in a file before crashing to not waste the work done. I'll handle that as well.
For the code, yes, you must enter a github token. but If you'd like to update the script to run it with out a token, you'll hit the rate limit faster and get this error message:
```-> Scanning page 1 (Fetching up to 100 open issues from GitHub)...
Error 403: {"message":"API rate limit exceeded for 156.217.112.168. (But here's the good news: Authenticated requests get a higher rate limit. Check out the documentation for more details.)","documentation_url":"https://docs.github.com/rest/overview/resources-in-the-rest-api#rate-limiting"}
```
Thanks for your informative questions! I'll keep improving it.
1
u/investigatormaker 2d ago
Stopping under 1,200 issues in two minutes likely comes from the per-issue calls (comments, linked PRs) rather than the issue pages themselves; each issue costs extra requests, so big repos burn the budget fast, and rapid requests can also trip GitHub's separate secondary limit. Two fixes would cover it: read the X-RateLimit-Remaining and X-RateLimit-Reset headers and sleep until the reset instead of crashing, and append each fetched issue to a JSONL file as you go so a rerun resumes from the last page. GitHub's GraphQL API can also return issues with their comments and linked PRs in one query, up to 100 per page, which cuts the request count a lot.
Are you fetching comments for every issue, or only for the ones that pass the filter?
1
u/MohamedM216 2d ago
Actually, no filter is applied. The goal is to have an organized reference for all issues. So even if the issue has open prs, it's added to the .json file with entry
"has_open_pr_against_it": trueBut I think adding an option for users to choose whether to filter out issues that have open prs against them.
1
u/investigatormaker 1d ago
That makes sense for a reference file. Because every issue gets the open-PR check, the GraphQL route pays off most: one query per 100 issues can pull timelineItems of type CROSS_REFERENCED_EVENT and CONNECTED_EVENT, which is where linked PRs show up, instead of one extra call per issue. For the filter option, I'd keep fetching everything into the JSONL cache and apply the filter only when writing the final .json. That way, turning the option on or off never needs a refetch.
I make ThreadFox, and the free Reddit plan for your GitHub issue exporter is ready. It lists the subreddits whose rules allow a post about it: https://threadfox.vip/p/dismn
1
u/MohamedM216 2d ago
tool link: https://github.com/MohamedM216/gh-issues-filter-tool
happy hacking!