[08:11:00] federico3: I was checking the read-only cookbook and I saw: [08:11:04] usage: cookbook [GLOBAL_ARGS] sre.mysql.global-read-only [-h] [--ignore-dirty-dbctl] [-t TASK_ID] [-r REASON] [--sections SECTIONS] [08:11:04] {set-ro,unset-ro,set-rw} [08:11:12] What's the difference between unset-ro and set-rw? [08:28:57] FIRING: SystemdUnitFailed: swift_rclone_sync.service on ms-be1069:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:46:08] marostegui: it would appear that one is effectively an alias, so the code is "if read_only else". We probably shouldn't have aliases where argparse doesn't show them as aliases of each other - the help message set for the argument attempts to do this, but it isn't clearly displayed [08:46:55] yeah, we should have just RO and not RO (whether we want it to be unset or set rw, I don't mind) [08:50:00] FIRING: [2x] SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter.service on db1180:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:50:06] Given the tight scope of this cookbook, it may have been clearer to use the equivalent of --add / --remove as mutually exclusive boolean flags , there should only ever be these 2 options [08:50:37] Should the DEFAULT_SECTIONS be restricted to s? [08:50:49] rclone issue is the same 4 objects as last week, need some more time to investigate :/ [08:52:29] DEFAULT_SECTION usually comes from MW which is s3 now [08:53:57] FIRING: [2x] SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter.service on db1180:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:53:57] I mean "global-read-only" is not global when only a subset of clusters are read-only, that was more what I was pointing to [08:54:19] ah sorry I thought you were talking about MW default_section [08:54:51] cezmunsta: yes, default sections should be s*, x1, x3 (and starting this week x4 too) [08:55:07] But probably also the writtable es* [08:55:13] pc and ms can be excluded [08:55:39] Basically the same thing we do during the DC switchover when we go RO [08:56:16] | | |-- sre.switchdc.mediawiki.02-set-readonly [ServiceOps] [08:56:24] sorry: | | |-- sre.switchdc.mediawiki.03-set-db-readonly [ServiceOps] [09:05:44] I can remove one of them, e. g. we keep set-rw? [09:06:09] +1 [09:06:14] that works for now yeah, thanks [09:33:57] RESOLVED: SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter.service on db1180:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:56:22] are the testcommonswiki links tables going to be on x1 or x4 long-term? [11:56:44] Amir1: ^ [11:56:54] x1 [11:57:05] I can move it if it's too much hassle [11:57:42] just means that they're not goint to be replicated then, since there's no x1 replica? [11:58:20] i suspect we can live with that since there probably aren't many tools relying on those, just means I need to revert some docs I just updated assuming they would both be on x4 [11:58:22] dhinus: ^ [12:00:33] I can set it up there [12:00:53] we will soon-ish want to set up x1 too [12:01:15] I made some changes about it too weeks ago [12:18:06] yeah if they can be in x4 it would be great, but I agree it's not a big issue if not [12:53:49] Emperor: o/ I've read your updates for the aqs nodes, it is weird that the firmware upgrade didn't work as expected. On cassandra-dev I reimaged all the three nodes multiple times without any issues.. [12:54:18] anyway, just to confirm - are you doing the idrac upgrade in two bumps? 5->6->7? And also, are you running provisioning afterwards? [12:54:34] it shouldn't be 100% needed but I'd run it just to be sure [12:56:32] zabe: at what time will you be around tomorrow for s4/x4 split? [12:59:42] elukey: doing the f/w upgrade in 1 go, then a reprovision. [12:59:54] super [13:00:11] I'm almost entirely sure the problem is that the preseeding is relying on the default grub boot device (sda) [13:00:57] I would be curious to know if using the trixie's installer changes anything [13:01:05] [because I had one this morning where I ran "chroot /target ; grub-install /dev/sde" and that fixed it] [13:02:24] elukey: we can't use trixie on these nodes (java / cassandra version incompatbility) [13:02:55] yep yep I know [13:18:43] dhinus: I'll set it up to read from x4, no worries [13:20:16] Amir1: thx <3 [13:59:48] marostegui, cezmunsta: currently the silences and pooling metrics report hostname,instance_group,section and the first one also has role. We don't have a naming scheme to distinguish the sanitarium hosts; I could add a flag for hosts on port 3306 or not... but maybe we can think of a better naming? [14:04:35] I would exclude any host that is not on 3306 [14:04:43] As all the ones in dbctl are on 3306 [14:06:23] for the pooled metric yes, but for the silence metric we might want to show it also for hosts not on 3306 [14:07:01] marostegui: does 9am UTC work for you? [14:07:10] zabe: it does! [14:07:36] Alright:) [14:07:41] federico3: what is supposed to populate the puppet_X tables? [14:07:45] federico3: We'd have to check because I think all the hosts with a port other than 3306 have notifications disabled? [14:08:47] marostegui: I don't think that that backup replicas do, do they? [14:09:08] cezmunsta: ah indeed, I think backups do alert (not pag3) [14:10:56] federico3: can you include the host_meta data, which should tell you that a host is in the sanitarium? [14:11:43] cezmunsta: https://gitlab.wikimedia.org/repos/data_persistence/zarcillo/-/blob/main/zarcillo/webapp/datasync.py?ref_type=heads#L601 currently not in use - we should not scrape VCSes, there's a related task [14:14:57] federico3: that question was superseded ... see host_meta [14:20:20] ok I filtered out the non-poolable hosts [14:25:24] federico3: what is populating that table? It looks to be static data and I don't see it getting updated when adding a new instance, did I miss something? [14:27:12] cezmunsta: https://gitlab.wikimedia.org/repos/data_persistence/zarcillo/-/blob/main/zarcillo/webapp/ids.py?ref_type=heads#L357 - this is for ancillary data that we never gathered elsewhere [14:27:48] luckily it's a small amount of data for now [14:28:15] ok the metrics are disappearing -> https://grafana.wikimedia.org/d/fc7sbw8/mariadb-pooling-and-silences-overview?from=now-15m&to=now&timezone=utc [14:33:36] federico3: tomorrow x4 will be a new section, do we have to update cookbooks/zarcillo to support it? [14:33:39] it is like any other section [14:36:13] there shouldn't be updates needed on cookbooks; on zarcillo if new hosts are added as x4 they will show up. The only bit to tweak is the section<-> port mapping and I can do it now [14:36:43] (but it's not blocking anything) [14:37:34] I am wondering about the DC switchover cookbook, let me check [14:46:43] https://gerrit.wikimedia.org/r/plugins/gitiles/operations/software/spicerack/+/refs/heads/master/spicerack/mysql.py#29 [14:46:47] I will add x4 there tomorrow [14:47:48] rolling_restart has hardcoded sections too [14:47:59] ok I will modify that too [14:48:00] thanks [15:01:04] marostegui: on zarcillo https://zarcillo.wikimedia.org/ui/hosts "Add instance" uses existing rows to populate the Section radio box so for the first host to add it will have to be done using clone.py or SQL [15:13:11] not sure what you mean federico? [15:15:42] federico3: section names are hard-coded, aren't they? [15:16:02] ... yet there is a sections table [15:16:36] Yeah [15:16:46] The hosts are already in zarcillo under x4 [15:16:48] All that is done [15:24:13] yes, indeed they are hardcoded so I'm going to updated it, however there's a tentative TODO to remove the hardcoding in theory [15:24:50] (yet we add new sections very infrequently so might as well keep it hardcoded and expose them in an API for cookbook/scripts) [15:27:17] Why hardcoded? Also, why isn't the sections table referenced? It would seem to make more sense to pull the sections from the sections table and not from the section_instances one, which then allows "add section" to be added and exposed later [15:34:04] for speed and safety. Ok, I updated zarcillo and x4 is showing in https://zarcillo.wikimedia.org/ui/hosts [15:42:15] You had to do a deployment for that though, right? [15:42:37] yes [15:43:47] Reading from the database would have been speedier and safer, wouldn't it? I mean, that is what it is doing for most data, just not all of the data [23:22:48] FIRING: PuppetFailure: Puppet has failed on ms-be2089:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure