[10:17:35] lunch [12:39:39] pfischer: japanese should be available, but there's a problem regarding enterprise formatting, it's not new but possibly a lot more annoying for japanese, enterprise html->text conversion generally adds extra space, while this is not always visible in languages with spaces it might be more frequent in japanese esp. for links where an extra space is added before and after the link [12:40:50] in short when you read https://ja.wikipedia.org/wiki/%E9%9D%92#%E8%87%AA%E7%84%B6%E7%95%8C%E3%81%AE%E9%9D%92 there will be an extra space around all blue links [13:17:09] o/ [13:57:26] \o [14:03:37] o/ [14:04:59] ran some analysis simulating what "perfect" more_like caching would do, if we never lost data to TTL. Really we aren't losing much there, given the 1 day TTL we expect a hit rate of about 68% and the logs report 58-65% [14:08:48] does kinda confirm the geo-distributed bit thogh, codfw sees 3-4M unique pages per day, eqiad sees 8-10M, so anytime we switch to codfw it's unlikely to have half the pages cached [14:29:52] makes sense [14:31:20] .o/ [15:51:28] took frever to calculate, but it looks like if we ran 7-day ttl's hit rate could increase to ~86-89%. Would be perhaps half as many more like queries landing on the cluster [15:52:47] err, no thats too much. actual/optimal gives ~74%, so it would perhaps be 25% i think [15:54:15] no it's half, actual/optimal gives a different result. Right now ~70% are cached, we could get to 85%, so thats half of the req's [16:31:48] this is keying only by page_id? does it take into account variations of the limit requested? [16:44:07] dcausse: hashing the query params, hopefully close enough