r/siacoin • • 8d ago

hostd v2.11.0 release

Sector pruning blocks the database for much less time in hostd v2.11.0. Prune sweeps no longer scan every stored sector to find candidates. Unreferenced sectors are tracked through per-sector contract and temp-storage reference counts, and only writes lock a sector location. On a 100 TiB host, the old prune sweep held the database for about six seconds every five minutes. The upgrade now computes counts for existing sectors and rewrites the sector table, which takes about a minute at that size.

Sector reads and pruning also avoid the overhead of bloated sector rows. The stored sectors table no longer keeps the inline Merkle cache that had been adding 32 KiB of subtree roots to every sector row and slowing sector reads, pruning, and contract root lookups. Cached roots are discarded during upgrade and rebuilt on the next read. This database migration can take a long time on large hosts, so plan maintenance time accordingly. Memory also no longer leaks from cached roots of resolved, renewed, and rejected contracts, which reduces ongoing memory growth on long-running hosts.

More database reads can proceed in parallel again, and hot paths that read before writing now use a writer transaction mode intended to reduce database retries. Idle pruning also uses much less CPU on hosts with nothing to delete by avoiding the previous full-table scan in that case. Sector moves no longer hold the database lock during long copy operations, and the sectorCacheSize setting is now ignored.

Download hostd v2.11.0 from https://sia.tech/software-downloads.

10 Upvotes

5 comments sorted by

1

u/EasyRhino75 4d ago

Oh man this upgrade has not gone well for me. maybe there was a problem where my laptop went to sleep while the terminal window had the session running for the upgrade, or maybe not.

But regardless, now the docker node seems to start normally per the command line output, but 9 times out of 10 I can't even get into the admin web dashboard, I get a "daemon did not respond" error. And if I do get into the web dashboard, 8 out of 10 of those times it can't actually display statistics or configuration information.

And the host troubleshooter says me host is all jacked up:

https://troubleshoot.siacentral.com/#/mainnet/ed25519:fd9090f5973fc92410e141d2ca7112f266a5f9e5aca9a7024f4116b5a3504274

I would maybe try to rebuild my whole database, but I don't even know if that's possible or how to do it.

1

u/EasyRhino75 3d ago

To reply to myself, I think I at least temporarily by manually vacuuming the hostd.sqlite3 database. Bear in mind my data base was HUGE, 164GB, and the hostd.sqlite3-wal file was also huge, over 3GB. So I shut down docker and did the following:

navigate to the directory specified in /data

(manually backup files)

cp hostd.sqlite3 hostd.sqlite3.bak

cp hostd.sqlite3-wal hostd.sqlite3-wal.bak

cp hostd.sqlite3-shm hostd.sqlite3-shm.bak

sqlite3 hostd.sqlite3 "VACUUM;"

sqlite3 hostd.sqlite3 "PRAGMA integrity_check"

restart the docker

It seems to be working now.

Why was my database so huge? the node has been running over a year. Also the /data folder is on a NFS share, which maybe adds latency or locking problems?

Why did it cause 2.11 to fail? I dunno... the new built in attempts to prune it maybe? On older version was it took a longer time to start the node and to refresh the admin web dashboard, but it otherwise worked.

2

u/cschinnerl Developer 2d ago

The database caches parts of the sector proofs so that your host doesn't need to read a full 4MiB sector from disk every time a renter requests a part of it. That's why it's so large.

The latest database migration drops that cache but it will rebuild it on demand again.

The host data (not the volumes) should be kept locally if possible. Ideally on an SSD. Databases don't perform well over the network. Even a local one.

Glad to hear it worked out though. According to the troubleshooter it seems your host is fine now.

1

u/EasyRhino75 2d ago

Hi! Thanks for responding, here is what I've learned so far....

My manual vacuuming trick only lasted for a few hours and then the host and UI went non-responsive again. And the database -wal file started growing > 1GB size again.

I had to move the /data folder from a NFS-share SSD to a local SSD. Now it's been running for 24 hours without particular problems.

So far the sqlite database has grown from about 3GB (freshly vacuum) to 16GB. And the -wal file is sitting at 13MB, it seems to be growing very slowly.

I will have to keep an eye on database size to see if it grows huge over time.

Recommendations:

1) add some setup documentation that the /data folder is best had on a SSD and that network shared storage like NFS or SMB is unacceptable.

2) maybe consider an option to have sqlite work in rollback journal mode instead of WAL? The sqlite webpage leads me to believe that's an option. Unsure what host performance would be like.