[03:53:32] I am a bit too sleepy to start deleting toolforge-nfs files. I will look in the morning but will also not complain if someone else cleans up while I sleep! [15:45:53] good morning! We have CI jobs being slower since July 25th and I wonder whether maybe something might have changed on the WMCS OpenStack infrastructure? [15:45:57] there are some graphs at https://phabricator.wikimedia.org/T434024#12195653 [15:53:29] hashar: July 25th was a Saturday so it seems unlikely there was any intentional change applied on that day [15:53:58] yeah I haven't dig into the details. Maybe it was a Debian package being upgraded or a kernel or whatever [15:54:38] at least it is good to know that was not an entire overall of the WMCS infra around that time :] [15:57:36] i also don't see anything particularly relevant in puppet that day or the day before [15:57:55] hashar: are you seeing this on all worker VMs or only some? [15:58:21] I haven't looked at that yet :) [15:58:30] I reached out here just in case something big happened at that time :] [15:58:57] the graphs on the task are a daily moving average, so certainly the slow down happened earlier [15:58:59] but I will find [17:27:18] so fun things, there is elevated latency in Ceph since July 24th [17:29:23] hm... we upgraded in mid-june, but that doesn't line up [17:29:32] hashar: can you point me to your data? [17:30:37] the only thing that corresponds to july 24th that I know of is https://phabricator.wikimedia.org/T429387#12154524 [17:30:49] yeah I am copy pasting on the task [17:30:52] will link once done [17:37:09] https://phabricator.wikimedia.org/T434024#12196170 [17:37:43] looks like cloudcephosd1045 is at 100% disk utilization https://grafana.wikimedia.org/d/rtOg0AiWz/wmcs-ceph-eqiad-osd-host-details?orgId=1&from=now-30d&to=now&timezone=utc&var-datasource=000000006&var-ceph_hosts=$__all&viewPanel=cloudcephosd1045$panel-5 [17:38:09] (it might be nice to have a dashboard with a breakdown per metrics instead of per host, so we could instantly see the disk utilization for all hosts) [17:39:11] andrewbogott: ^ :) [19:38:25] hashar: any chance everything is better now? (I know the graph is better but wondering if you can tell about CI jobs) [21:07:37] andrewbogott: I guess we will find out as builds are happening on CI [21:07:50] if the disk utilization has been resolved and the latency is down, i guess that explains it [21:08:06] thank you for the very complete bug report! I definitely steered me towards a serious problem, we will find out if it was also your problem. [21:08:13] *it definitely [21:08:50] larssandergreen in #wikimedia-fundraising would be able to confirm [21:09:03] fr-tech reported the task due to their build being slow/timing out [21:09:35] and if all fine possibly close the tasks and get congratulations from greg-g who poked me about it earlier today :) [21:10:16] I'll poke him :) [21:12:45] there are a couple mysteries left on that task but if the builds are fixed I will save the mysteries for next week [21:18:39] yeah sounds good [21:18:44] thanks for the quick fix up [21:18:53] Lars is going to trigger a few builds and will report back [21:26:52] great! I'll be here [21:43:14] * greg-g peaks in and smiles [21:43:17] thanks both!