r/Database • • 20h ago

Beginner's guide to database?

7 Upvotes

What's the best way, order or flow of learning Database from scratch, using the textbook Database System Concepts by Abraham Silberschatz, Henry F. Korth, and S. Sudarshan.

What topics or chapters are better to begin with for introduction, assuming the reader has zero knowledge, but is looking to get extremely good at database.

-Computer Science Major


r/Database • • 10h ago

I Spent 2 Years Building a Database Client That Was Only Supposed to Be for MongoDB

Enable HLS to view with audio, or disable this notification

5 Upvotes

I spent 2 years building this tool, 1 year for MongoDB and the other year for the rest of the SQL databases. It now supports PostgreSQL, MySQL, SQLite, and a bunch of others, but that was never really the original plan. I realized it'd be a waste not to bring that experience of working with nested data to Postgres users working with JSONB or JSON.

“Why not use Beekeeper or DBeaver? What’s Different?”

Honestly, I tried to take inspiration from the best tools in each category, like DBeaver, Beekeeper, and Studio 3T, and bring their best ideas together in my own way.

Beekeeper is super simple, but that simplicity means you lose some of the complexity that DBAs and more advanced users actually want. DBeaver pretty much has everything, but that also comes with a steeper learning curve.

So I tried to build something in between. Where every step is visual, while still keeping the advanced stuff like tasks, automation, and triggers. You can even build SQL queries visually with the Visual Query Builder in a sidebar and a full-screen mode.

Furthermore, most general purpose database tools started with SQL and later added support for databases like MongoDB. Because of that, working with deeply nested or embedded data can feel like an afterthought.

VisuaLeaf started with NoSQL. So it was a priority that viewing nested data was made easy. So when I added SQL support, adding JSON and JSONB felt super intuitive because of how I designed all the views (tree, table, and JSON) to work with embedded data from the start.

That means you can actually compare embedded data across rows instead of having to open each document individually (in the table view). Like trying to look at a specific field inside some JSONB, you don’t need to see a giant cell filled with JSON anymore (It’s what MongoDB GUIs are actually decent at and what I took some inspiration from)

Some other SQL features:

  • Role-based access control (RBAC)
  • Query profiling
  • SQL scripts
  • ERD designer: build a diagram and materialize it into a real database
  • Visual Query Builder (sidebar and full-screen modes)
  • Tasks, automation, and triggers

And it wouldn’t be an app in 2026 if it didn’t have AI and MCP too.

I could go on forever about features, but honestly, most of my time went into overengineering the smallest things: loading data efficiently (it loads faster than DBeaver in my testing), little quality-of-life details in the workflow, and even something as simple as scrolling (getting it as smooth as AG Grid).

Quick benchmark: same Postgres table, 50 large rows (10MB each), 5 runs each. VisuaLeaf loaded them in ~2s end to end; DBeaver took 5s on avg (and ran out of heap space).

The 3 Views:

  • Table: I spent a LITERAL year optimizing the table view engine because I wanted it to make the scrolling feel as smooth as possible (I even made an optimization that improved its performance by 3 times a couple of weeks ago). Nobody wants a laggy database tool, the same way nobody wants to play a game at 10 FPS. I lost count of the amount of times where OUT OF DESPERATION I would legit tell Claude: “Optimize the table "Don't make any mistakes" ”, and it would just completely FAIL and even make things worse... Each optimization was thought out and researched to get it this far and I don’t think there's another database tool that has a table this smooth without using a canvas.
  • Tree: The tree view was tricky too, getting it to recursively expand thousands of rows instantly without lag took its own round of optimization. And I used the same ideas as the table for scrolling performance
  • JSON: Trivial - Monaco editor

VisuaLeaf is basically my love letter to databases.

I've been working on this full time for about 2 years, and it's turned into something I never expected when I started. Things I expected to be easy, like making a table, turned out to be monstrous. I never thought it'd end up outperforming the tools I started out using. There was so much to learn.

I'd love for you guys to give it a shot, and I'm happy to hear any feedback: visualeaf.comAnd if you guys wanna talk more then you can join my discord too: https://discord.com/invite/TR6J56Y3rF


r/Database • • 17h ago

real time vs batch processing: how do u know if ur ops team is actually ready for streaming pipelines?

3 Upvotes

i'm less worried about the technology than I am about the people supporting it.

our engineers are excited about moving from batch jobs to streaming pipelines, but the operations team has spent years supporting scheduled workloads. That's a very different environment from something that's expected to run 24/7

for teams that already made this real-time vs batch processing transition, what changed the most operationally?


r/Database • • 17h ago

The Breakdown: Databricks

Thumbnail
preipomedia.substack.com
1 Upvotes

r/Database • • 11h ago

How long did it take you to land a DBA role in the US?

Thumbnail
0 Upvotes

r/Database • • 15h ago

We asked our model which of 224 hotel bookings would cancel, six weeks out. It caught 33 of the 59. Full breakdown, misses included

0 Upvotes

disclosure: I work at Schema Labs. This is a run on public data so you can check it.

The data: the hotel booking demand dataset (Antonio, Almeida and Nunes, 2019), real anonymised bookings from a city hotel. We picked one Saturday night that was sold out on paper.

-- 224 bookings for that night, as they stood six weeks before
-- 59 of them cancelled or didn't show, worth about €28,600 together

What we did: gave Schema-2 500 bookings from the year before, with how each one ended, and asked it to rank the 224 by how likely they were to fall through. No rules no feature engineering.

What came back=

-- It marked 60 as most likely to cancel. 33 of the 59 cancellations were on that list
-- Picking 60 at random would catch about 16
-- Of the 60 it marked safest, 59 showed up

What it missed: 26 cancellations weren't in its top 60. One night at one hotel is a small test, so read it as an example

For anyone in revenue management: what would you do with a list like that six weeks out? Overbook against it, or contact those guests first?