[08:25:02] federico3: do you happen to know anything about "MediaWiki periodic job db-lag-stats-reporter failed" ? [08:25:22] uhm no... any context? [08:25:58] https://alerts.wikimedia.org/?q=%40state%3Dactive&q=team%3Ddata-persistence [08:36:29] Hmm it seems to be mentioned here: https://wikitech.wikimedia.org/wiki/Maintenance_server [08:37:14] I will ping Hugh [09:19:17] that's a mediawiki periodic job that is failing for whatever reason [09:19:30] not sure if that's right to assign to your team [09:21:30] weird etcd error was the source of failure, and it has run successfully since [09:27:02] hnowlan: thanks for the info - as for which team, yes not sure - we can wait until Manuel is back in case he would prefer it to remain with DP as the owner [09:28:38] cezmunsta: I'd give serviceops a shout and see if they have any ideas, they did the migration of those jobs [09:29:04] I say they, *my* name is on the git blame for that alert [09:29:27] but generally I think a failure of that job is of interest to you - I think that in this case it's weird external failure that's causing it to alert [09:30:06] keep it around if you want to ask, but if you're not worried I think you can delete the stale job and move on [09:31:57] thanks [10:21:17] that failure was 41h ago btw, there are looooads of successful passes since [10:35:43] hnowlan: yes, is there a reason that it doesn't clear itself? [10:37:03] Failing jobs stick around for analysis in most cases I believe, I think there's a setting to clear them up after a period [12:19:01] cezmunsta: unit tests are passing and runs locally, I'm trying to also fix CI and deploy https://gitlab.wikimedia.org/repos/data_persistence/zarcillo/-/merge_requests/1 [12:35:05] federico3: ack - there seem to be a number of pipelines showing as stuck and running, do they need cleaning up? [12:35:30] they need a good bunch of fixing :D [14:44:55] I just deployed the new version of zarcillo with split api/web ui [14:47:41] most stuff should work, the only thing is the etcd client is failing to connect so the status of candidates is stale. In case of rollback the previous version was v0.0.1-b335-7457560f [14:48:57] federico3: is it a good idea to leave it like that given that cookbooks use it? [14:50:32] Do you know why it can't connect to etcd? [14:50:44] it showing a certificated error, oddly [14:51:04] but it's difficult to replicate outside of prod [14:52:15] Do you have the specific error? [14:53:55] that's some verbose error msg https://phabricator.wikimedia.org/P96541 [14:54:10] but maybe the cert issue is a red herring and it failed to connect or something else [15:05:22] I would say that it is better to roll it back for now. unless you are confident that is the only issue and that it doesn't affect anything in cookbooks, or scripts [15:07:11] As the container seems to have a shell, can't you kubectl exec to test to see if you can get a better idea? I mean, perhaps rolling back might also show the same issue? [15:07:31] yes I updated the paste with it [15:07:58] no, the previous version was not showing this issue and it was on Bookworm [15:09:32] yep I'm rolling back while investigating [15:13:08] The output suggests that the one failing is using urllib3, may be relevant