User Details
- User Since
- Nov 2 2014, 11:35 PM (611 w, 1 d)
- Availability
- Available
- IRC Nick
- andrewbogott
- LDAP User
- Unknown
- MediaWiki User
- Andrewbogott [ Global Accounts ]
Yesterday
I have granted the quota increase. Do please create a subticket about reverting, or follow up here when you're done with the migration.
+1 lgtm
ok, I'm caught up now! I hadn't realized that we had reproducible uuids for new hosts, I was still just taking nova's word for it :(
The stated quota increase is to make room for an identical VM to the main one we use now, which is of special flavour g4.cores24.ram122.disk20
Our tool cannot run in Toolforge
I made another pass at rebalancing but 1004 is still taking a bit longer than 24 hours (24 hours and 6 minutes) to complete so this needs yet one more pass.
Fri, Jul 17
There are lots of reasons for a host to become unreachable, but falling off the network is one of them. This dashboard is a good place to start:
I think I'm behind on something. When you say...
Thu, Jul 16
This is now done. I did wind up doing a staggered reboot of all servers because NFS clients are not great at handling change.
Wed, Jul 15
Typically I would restart nfs clients after an nfs migration to make sure they're mounting the new server properly. So the most straightforward thing is for me to just shut down all the -app hosts in the cluster, flip the nfs switch, and then start them up again. That's also the best way to avoid any weird in-between states; does that sound ok? I'm assuming the above jobs run on the app hosts and not the apache host, please lmk if that's wrong.
Thank you @jsn.sherman
Thanks! I have put this host back in service and we'll keep an eye on it.
Wed, Jul 8
I have a start on this
Tue, Jul 7
this needs more testing before I ship it to eqiad1, but the specific blubber update is done.
this is done. The old server dumps-nfs-1 is shutdown and can be deleted after a few days of good health.
Oh, also, to be clear: This is entirely a storage-based solution so far. Ideally we would be blocking and deleting users that cross the line but at the moment we're not doing anything to prevent users from re-creating their notebooks every 24 hours.
The current running code regards 'freeroot' as a strike against a notebook but not an automatic block. If that gets file usage under control then that may be good enough but I'll keep an eye on it.
Let's do a memory scan on this server too, while it's out of service. Thanks!
Oh -- it's out of service now, so y'all can switch it off whenever.
For now: this host is in the maintenance aggregate and drained of all user VMs.
And, just like T410470, a totally useless syslog during the downtime:
Mon, Jul 6
find . -name README.md -exec grep "Foxytoux" {} \; | wc
3401 32322 244072
A bit of documentation added here: https://wikitech.wikimedia.org/w/index.php?title=PAWS%2FGetting_started_with_PAWS&diff=2433391&oldid=2254420
Thu, Jul 2
Here is a snippet of the kind of thing the attached patch would do. Might be too aggressive...
lgtm
Wed, Jul 1
Tue, Jun 30
I'm pretty sure that the cleanup script that copies things into /srv/paws/files-to-remove/ was lost when I rebuilt the NFS server. Going to wait 24 hours to confirm that that's right -- I certainly can't find a job scheduled anyplace.
I'll go ahead and run this update next week if I don't hear anything here in the meantime.
This is almost certainly the result of a recent upstream change which limits app credential delegation. In theory we've already activated all the workarounds but I'll see if there's another one for this exact case. In the meantime: painful as this is you can likely do this via Horizon which allows access as a human user.
Mon, Jun 29
Belinda -- assigning back to you to close or follow up as needed.
These numbers are never full stable due to ongoing decoms and new hardware, but here's a snapshot of cloud resources:
Fri, Jun 26
Here are the two nfs mounts we will be replacing:
we're pitching NFS overboard as part of T402054: twl: Replace deprecated Bullseye VMs in Cloud VPS this sprint.
Thu, Jun 25
Hello @jsn.sherman! I have not yet started work on this, but twl is one of the projects that needs an NFS update sometime soon (see T429793). Do you have any idea what went wrong with the NFS mount? And, is this a project that still requires nfs for sharing files between hosts?
I'm not sure why this is happening. Most likely it's due to a different cookbook (maybe wmcs.openstack.cloudvirt.safe_reboot) failing and leaving things in an inconsistent state.
+1
Tue, Jun 23
I built fresh base images for Bookworm and Trixie which should resolve this.
Mon, Jun 22
- Add a runbook to the alert -- if it's generic, make an index page for all nfs full alerts
- Add a daily timer to delete all files older than 2 weeks from /srv/paws/files-to-remove/ on the paws NFS server
There are some very brief docs about this here: https://wikitech.wikimedia.org/wiki/PAWS/Admin#Deleting_user_data_in_case_of_spam_or_credential_leaks
Jun 18 2026
Indeed, the alert seems to have cleared. Thank you!
