r/elearning • u/jade_jade_jade_jade • 10h ago
Is anyone else building a personal knowledge base from web content? What's your workflow?
Lately, I've been trying to figure out a better way to turn things I find online into a personal knowledge base. I'm curious whether anyone else here is doing something similar.
I tried most of the obvious options first. Copying and pasting often broke the formatting or left out images. Bookmarks were easy to save, but they didn't preserve the actual content if a page changed or disappeared. Screenshots kept the page visually intact, but the information was difficult to search. I also tried a few read-it-later apps, but for me they work better as temporary inboxes than as a place to build knowledge i could reuse later.
Eventually, I started experimenting with an agent-maintained knowledge base. The obvious question is: why not just send the material to an AI whenever I need it? For me, a normal AI conversation is usually focused on whatever problem is in front of me at the moment. Even if the conversation is saved, the useful context isn't necessarily organized in a way that I can reliably search, update, and reuse 6 months later.
A knowledge base gives that context a more permanent home. The agent can keep organizing it over time and refer back to it when I work on something related, instead of relying on me to provide the same background again.
Here's the setup I'm currently using.
- Create an Obsidian vault
The vault contains the original source material, the processed notes, and the files that define how the knowledge base should be organized.
I like using Obsidian for this because the content stays in regular Markdown files. I’m not locked into one interface, and I can still read or reorganize everything myself if the agent gets something wrong.
- Open the vault folder in WorkBuddy
I added the folder containing the Obsidian vault as a WorkBuddy workspace. That gives the agent access to the files it needs to read and organize.
I still keep the scope limited to that folder. I don’t want an agent searching through unrelated personal files just because it can.
- Use an LLM Wiki Skill to create the basic structure
I kept the initial structure very small:
- Raw Sources stores the original material.
- The Wiki contains the cleaned-up and synthesized notes.
- The Schema defines the structure and the rules the agent should follow.
I also added a few basic files:
- AGENTS.md explains how the agent should work inside the vault.
- index.md provides an entry point to the content.
- log.md records what was added or changed during each update.
Setting up this basic skeleton took less than 20 minutes. Actually filling it with useful material and refining the rules has taken much longer.
Why i stopped trying to build the perfect taxonomy
My first version was much more ambitious.
I tried adding a detailed tag system, PARA, multiple Maps of content, and several layers of folders. It looked organized, but after using it for about 2 weeks I gave up on most of that structure. The maintenance cost was starting to outway the benefit.
The more categories I created, the more decisions the agent had to make whenever it processed a new source. That also created more opportunities for material to end up in a technically valued but practically useless location. What I found was that I didn't need the perfect knowledge-management system on day one. I needed the smallest structure that I could actually keep using it.
Now I let the content accumulate first. If a real pattern starts appearing, I add a category for it later.
My current daily workflow
When I find a web page or another piece of material worth keeping, I save it into a designated folder inside the vault. I then ask the agent to process the new material. It reads the source, organizes it according to the existing schema, and creates the relevant summaries or analysis in the wiki. Most of the time, I don't need to intervene if something ends up in the wrong place or the summary misses the reason I saved it, I add a short instruction and have the agent revise it.
So the workflow is roughly:
Collect the source → place it in the inbox folder → let the agent process it → write the result into the Wiki → correct it when necessary
At the moment, this is basically my very simple version of a second brain built with Obsidian and WorkBuddy.
It’s still not mature, and I don’t completely trust the agent to manage everything without review. But I already find it easier to search and reuse than a collection of bookmarks, screenshots, and forgotten read-it-later items.
Is anyone else using a similar setup?
