[10:01:32] lunch [11:27:39] Hi, I was referred here by Leila Zia from Wikimedia Research regarding an issue with two French Wikipedia articles: [11:27:40] • Aimé Nkanu — https://fr.wikipedia.org/wiki/Aim%C3%A9_Nkanu [11:27:40] • Junior Mpiana — https://fr.wikipedia.org/wiki/Junior_Mpiana [11:27:41] After publication, both articles appeared in Google Search. They then disappeared, later appeared again, and then disappeared again. [11:27:41] The last time I personally saw them appearing in Google was around September 27, 2026. Since then, they have remained absent from Google Search, although both articles are still normally available on Wikipedia. [11:27:42] This repeated pattern — appearing, disappearing, reappearing, and then disappearing again — is the reason I am contacting you. [11:27:42] Leila Zia advised me to contact the Wikimedia Search Platform team about this issue. [11:27:43] Could someone please help me understand what is causing this and, if possible, help resolve the problem so that these articles can appear normally in Google Search again? [11:27:43] Thank you very much for your help. [13:16:15] o/ [13:44:54] \o [13:51:51] .o/ [14:04:29] o/ [14:08:09] o/ [14:17:07] dcausse: WE has released structured contents for Japanese. Would you have time to add that to the parsing/embedding extraction DAG? [14:17:29] pfischer: sure, will do that now, do we need a ticket? [14:19:15] dcausse: not yet, I’ll create one. [14:19:25] thanks! [14:20:17] T440327 [14:20:17] T440327: Create embeddings for Japanese - https://phabricator.wikimedia.org/T440327 [14:23:55] up for review: https://phabricator.wikimedia.org/T438984#12401140 (numbers and a quick conclusion of the perf test done yesterday) [14:26:28] wiki_fr243: Hi, thanks for the report and the links. The Search team doesn’t control whether individual Wikipedia articles appear in Google Search, but we can have a quick look for anything obviously unusual on the Wikimedia side. If everything looks normal there, this would likely be something on Google’s end. [14:28:21] they left, but quickly looked and those pages were created very recently (mid-september) so I would not be surprised that it'd take a bit more time to have stable search results [14:56:07] !log [15:02:09] Hi, I was referred here by Leila Zia from Wikimedia Research regarding an issue with two French Wikipedia articles: [15:02:10] • Aimé Nkanu — https://fr.wikipedia.org/wiki/Aim%C3%A9_Nkanu [15:02:10] • Junior Mpiana — https://fr.wikipedia.org/wiki/Junior_Mpiana [15:02:11] After publication, both articles appeared in Google Search. They then disappeared, later appeared again, and then disappeared again. [15:02:11] The last time I personally saw them appearing in Google was around September 27, 2026. Since then, they have remained absent from Google Search, although both articles are still normally available on Wikipedia. [15:02:12] This repeated pattern — appearing, disappearing, reappearing, and then disappearing again — is the reason I am contacting you. [15:02:12] Leila Zia advised me to contact the Wikimedia Search Platform team about this issue. [15:02:13] Could someone please help me understand what is causing this and, if possible, help resolve the problem so that these articles can appear normally in Google Search again? [15:02:13] Thank you very much for your help. [15:05:50] wiki_fr243: unfortunately this is on googles end, we only manage the on-site search. We have almost no visibility into what google does [15:06:34] at a general level, my memory is that google does not index new pages immediately, they probably have some way to decide if it gets included based on age and other metrics [15:10:15] Thank you for your reply. The unusual part is that these pages were already indexed by Google after publication. [15:10:16] Both articles appeared normally in Google Search, then disappeared. They later appeared again, and then disappeared again. They have now remained absent for several days. [15:10:16] Google's Rich Results Test can currently fetch the Junior Mpiana article successfully: HTTP 200, crawling allowed and indexing allowed. [15:10:17] So the issue is not simply that Google has never indexed the new pages. They were indexed and then repeatedly dropped from Google Search. [15:10:17] Do you know if there is another Wikimedia team that deals with external search engines/SEO, or someone who has access to Wikimedia's Google Search Console data who could check this? [15:12:50] wiki_fr243: we don't have an SEO team, I've heard there are some people in analytics/research that have access to the search console but not sure who, i've never looked at it [15:13:12] i've seen some data from it before and it wasn't very useful, the top queries was just a list of page titles [15:14:39] Thank you, that's very helpful. I'll try to find the appropriate person in Analytics/Research who has access to Wikimedia's Google Search Console. [15:32:20] dcausse, ebernhardson: thanks for T438984, that is a lot of data to chew on. [15:32:21] T438984: Exploration: Route queries based on query type and/or target wiki and/or query cost - https://phabricator.wikimedia.org/T438984 [16:26:28] looked closer into the single REST handler limitation, turns out security plugin uses it. So we can't do that. There is a way with ActionHandlers, but that doesn't really get everything. I think the envoy proxy probably is better than an opensearch plugin [16:28:46] oh right, we still use tlsproxy in puppet [16:36:17] We can look at migrating to envoy if it would help. I think I already have a ticket open for it [16:36:39] yup T398070 [16:36:39] T398070: Migrate CirrusSearch TLS termination from nginx to envoy - https://phabricator.wikimedia.org/T398070 [16:38:25] dunno about envoy but it should be fairly trivial to add a header to nginx running on the opensearch hosts? [16:42:15] yea it's pretty trivial for both, i put together a tlsproxy patch and a patch for deployment-charts which should add X-Cluster-Name everywhere [16:43:33] perhaps harder on the cirrus side to propagate the header down to the caller, might need to rely on the global state? [16:45:42] dcausse: yea i'm not sure yet how I will get that out of the transport and down to the cirrus layer, [16:59:39] heh, dumb idea: transport has access to the connetion object, and we can `$conn->setParam( $k, $v )` to pass data around [17:07:03] why not? client has a kind of getLastRequest, why not a getLastRequestHeaders from connection? [17:08:52] just read about the schema limitations for array of struct in T259924 [17:08:52] T259924: HiveExtensions.convertToSchema does not properly convert arrays of structs - https://phabricator.wikimedia.org/T259924 [17:09:18] sounds like a limitation of our own tooling but not really a limitation of hive/iceberg? [17:10:28] dcausse: we did have a problem in the past where we added a field, and then spark querying it out put data in the wrong fields, like it was indexing into the object with numbered indexes and not names [17:11:07] ouch [17:14:15] that's annoying tho... hopefully we won't analyze snippets all that often... [17:14:29] yea it will be annoying to extract [17:15:30] even we ran a manual "alter table" prior to shipping the schema change, this will mess-up spark? [17:19:47] will possibly try this tomorrow in my own hive db just out of curiosity, but feel free to ship the schema change as-is [17:20:27] hmm, maybe a manual alter table would work? Does seem worth testing, i don't really want to deal with this wart for years [17:20:55] sounds good, will do some testing tomorrow [17:20:58] dinner