[02:21:49] Yes I think we need guidance from the team on whether the HTML will get any processing when displaying on an embedded wiki. IMO the outcome described by Mateo is the ideal, but that requires a difference between the display on AW and the language wiki. (re @Al: Where did we end up with this? I’m thinking we could always link to the Wikipedia page for the view [02:21:49] language, if it has a wikili...) [02:26:34] Al's schemes are roughly the best we can do without some kind of html conversion, but IMO they sound a bit ugly and they won't be trivial to build (especially for the link-injection/removal functions). (re @u99of9: Yes I think we need guidance from the team on whether the HTML will get any processing when displaying on an embedded wiki. IMO ...) [02:29:11] The difficulty of link injection is one reason I'm still unsure about whether language configs should be in HTML or monolingual text. The decision here could sway our whole best practice. (re @u99of9: Al's schemes are roughly the best we can do without some kind of html conversion, but IMO they sound a bit ugly and they won't b...) [02:31:20] (link injection will get much messier in languages which routinely make morphological changes that don't correspond with the label/lemma) [02:58:32] I've been thinking about this more, and although it's sometimes useful and possibly to split compound concepts, I don't think it will always work, and not should this be forced on AW as a constraint. For example, we will need articles about compound concepts (e.g. "capital city", "unincorporated town", "secondary school", "rainforest", "ice cream") or they will be [02:58:33] useful for oth [02:58:33] er items. If I'm right, and we don't want to make compound lexemes for them all, then we will need a way to navigate from the compound concept item QID to the head concept item QID, and the separately to the modifier QIDs. Unless there is already one in place I think this would need new item-typed properties on Wikidata. (re @Jan_ainali: Another conundrum is complex [02:58:33] concepts lik [02:58:34] e Q1538989. Perhaps I am not just good enough in linguistics, but I would think that in...) [03:29:48] At least four of your examples are common nouns that have entries in Swedish dictionaries and already have lexemes. I guess rainforest and ice cream will have that in English too, right? (re @u99of9: I've been thinking about this more, and although it's sometimes useful and possible to split compound concepts in manual fragmen...) [03:36:59] That will require more than just CSS, though. If someone could provide the JavaScript, I can suggest it. (re @u99of9: Are you willing to consider a non-error call to action for missing lexemes? (Like we have for missing labels/configs etc) That i...) [03:38:51] I'm deliberately choosing examples that you and I think of as single concepts that are actually composites. (re @Jan_ainali: At least four of your examples are common nouns that have entries in Swedish dictionaries and already have lexemes. I guess rain...) [03:39:47] In Swedish glass (ice cream) is not a composite 😃 (re @u99of9: I'm deliberately choosing examples that you and I think of as single concepts that are actually composites.) [03:42:14] But my point is that composite lexemes are okay. [03:43:01] Okay perfect, so maybe that's an even better way to show that different languages will think differently about whether a lexeme is needed and therefore whether they should decompose to head+modifier. In some languages "Czech municipality with town privileges" may not be a compound, but we will still want an automated way to decompose it. [03:44:33] Sure. And that decomposition needs to happen for all languages so that we always render lexemes and not labels. (re @u99of9: Okay perfect, so maybe that's an even better way to show that different languages will think differently about whether a lexeme ...) [03:44:50] Sure but you (quite fairly) don't want to make one for every compound, e.g. megadiverse country. (re @Jan_ainali: But my point is that composite lexemes are okay.) [03:46:54] I agree we need a way to do it, but I don't agree with your reason. Nevertheless, since we both want to be able to do it automatically, what do we need for that? Am I right that it would need new item-item properties on Wikidata? (re @Jan_ainali: Sure. And that decomposition needs to happen for all languages so that we always render lexemes and not labels.) [03:47:24] Another reason to do that is that it will be more predictable. Labels change quite often. (re @Jan_ainali: Sure. And that decomposition needs to happen for all languages so that we always render lexemes and not labels.) [03:49:08] You mean like a qualifier on P279, (re @u99of9: I agree we need a way to do it, but I don't agree with your reason. Nevertheless, since we both want to be able to do it automat...) [03:50:01] Lexemes have their issues too. I'm not anti-lexeme, but I'm pro-fallback, and pro whatever-works-for-a-language. [03:50:56] Maybe for the head concept. How about for the modifiers? (re @Jan_ainali: You mean like a qualifier on P279,) [03:51:26] On that I agree. But currently it's not used as fallback but as intended main method. (re @u99of9: Lexemes have their issues too. I'm not anti-lexeme, but I'm pro-fallback, and pro whatever-works-for-a-language.) [03:52:05] Do we have an appropriate qualifier? (re @Jan_ainali: You mean like a qualifier on P279,) [03:53:28] If you mean in English implementations, feel free to add tests where it fails. (re @Jan_ainali: On that I agree. But currently it's not used as fallback but as intended main method.) [03:54:47] It _fails_, and fallback, if a lexeme is not used. [03:56:21] So if you never try to get the lexeme, the implementation is flawed even if tests pass. [03:57:42] But the non-sandbox template is protected, need someone with the required user rights to help editing it. (re @Al: Yes, just go ahead boldly.) [04:03:22] If every conceivable test passes, it is good enough for English for now, certainly compared to delivering error messages. IMO we need to demonstrate the potential to our language communities with whatever we've got. [04:03:44] Okay, I've just put it into the main page. Do you mind testing it out? (re @Winston_Sung: But the non-sandbox template is protected, need someone with the required user rights to help editing it.) [04:08:14] While you're here, do you have a verdict on this? I guess it would help your language community among others that implement lexemes already. (re @u99of9: Are you willing to consider a non-error call to action for missing lexemes? (Like we have for missing labels/configs etc) That i...) [04:11:54] Yeah, it works. But to make it appears on the main page we need to mark it as ready for translation. (re @u99of9: Okay, I've just put it into the main page. Do you mind testing it out?) [04:22:54] I think I've also now done this. Try again? (re @Winston_Sung: Yeah, it works. But to make it appears on the main page we need to mark it as ready for translation.) [06:26:12] Yes, https://t.me/Wikifunctions/35778 (re @u99of9: While you're here, do you have a verdict on this? I guess it would help your language community among others that implement lexe...) [06:38:01] Tested and confirmed it worked. (re @u99of9: I think I've also now done this. Try again?) [06:38:43] Sure. But it can only be called a fallback if something else is tried, don't work, and then it falls back, right? If it's the only method it's by definition not a fallback. [06:38:43] A way we could highlight this is for any label we render that we show that doesn't have a P31 statement (and therefore could be a proper noun) we put the background in orange. (re @u99of9: If every conceivable test passes, it is good enough for English for now, certainly compared to delivering error messages. IMO we...) [06:42:48] Sorry I missed it. I must be thinking of something different to you. My hope would be to output HTML with some symbol like ✏️ (https://www.wikidata.org/wiki/Q4501797) or 🔧 (https://www.wikifunctions.org/wiki/Z36194) that links to a resource/qid where lexeme fixing is desirable. I haven't thought through the whole scheme. Can you elaborate on your JS>CSS idea? (re [06:42:48] @Jan_aina [06:42:49] li: Yes, https://t.me/Wikifunctions/35778) [06:45:23] I was talking about that people are skeptic about sitelinks to Abstract articles. So for now, we are likely to hide them with CSS for non-logged in users. While we were discussing this, we saw that frwiki hides them for everyone (https://fr.wikipedia.org/w/index.php?title=MediaWiki:Common.css&diff=prev&oldid=236930046). (re @u99of9: Sorry I missed it. I must be [06:45:23] thinking of someth [06:45:24] ing different to you. My hope would be to output HTML with some symbol like ✏️ or...) [06:45:32] I don't call the current English ones fallbacks. I know you don't like them, but those that work in 100% of conceivable test cases are perfect implementations ;-) (re @Jan_ainali: Sure. But it can only be called a fallback if something else is tried, don't work, and then it falls back, right? If it's the on...) [06:48:48] I’m pretty sure it should be neither. The question for me is how close to HTML it should be. An approach like [[Wikifunctions:Type proposals/HTML fragment structure]] is about as close to HTML as we should go. The specific details aren’t too important; it’s just about ensuring that there is a language-neutral stage at the end of the pipeline. The same input to the final [06:48:48] sta [06:48:48] ge, which is the output from language-specific functions, can then be used in different contexts to provide plain text, HTML or other marked-up text (including monolingual text). (re @u99of9: The difficulty of link injection is one reason I'm still unsure about whether language configs should be in HTML or monolingual ...) [06:49:55] Ahh, I see, you want the call to action near to the parallel page in sv-wiki? Okay, I have no issue with that. I mean one inside AW when a page is rendered. So, like the two types we have so far for missing labels and missing configurations, but for a sentence that is missing important lexeme info: : https://tools-static.wmflabs.org/bridgebot/053ea26a/file_83278.jpg [06:51:38] I said we're about to hide the link. So it's more like _you_ want a call to action :) (re @u99of9: Ahh, I see, you want the call to action near to the parallel page in sv-wiki? Okay, I have no issue with that. I mean one inside...) [06:53:36] But it must be highly unlikely that a non-logged in user clicks the Abstract link, sees a broken page, learns about Wikidata and lexemes, and that's the way the get started editing. But the possibility is perhaps enough? (re @Jan_ainali: I said we're about to hide the link. So it's more like you want a call to action :)) [06:53:48] I have no opinion on hiding links from other wikis. But within AW I want sv users like others to see articles with best-available defaults (including CTAs) in preference to errors. (re @Jan_ainali: I said we're about to hide the link. So it's more like you want a call to action :)) [06:56:08] Maybe? I mean, the red box is better than the sentence before it. : https://tools-static.wmflabs.org/bridgebot/993290e3/file_83279.jpg [07:02:14] I accept that we may not have the right type yet. Your comment helps me understand the point of your type proposal. With that proposal structure in mind, how would an AW editor specify which list items should be linked? (re @Al: I’m pretty sure it should be neither. The question for me is how close to HTML it should be. An approach like [[Wikifunctions:Ty...) [07:08:26] True, but that's because you still have remnants of the English in Z37736 : https://tools-static.wmflabs.org/bridgebot/a18a6d3f/file_83280.jpg [07:09:45] if we take out that English superlative wrapper, and just return the ltd fallback or QID, it will look more like the Japanese. (re @u99of9: True, but that's because you still have remnants of the English in Z37736) [07:11:04] (except since there's now a lexeme connected to flatness I'll need to demonstrate elsewhere...) [07:12:16] <[[smlckz]]> What can I read to better understand the linguistics side of things you've been doing here, lexemes and such? [07:13:39] done https://www.wikifunctions.org/w/index.php?title=Z37736&diff=302746&oldid=292720 (re @u99of9: if we take out that English superlative wrapper, and just return the ltd fallback or QID, it will look more like the Japanese.) [07:14:54] On another note, some sentences just sound weird in some languages 😆 : https://tools-static.wmflabs.org/bridgebot/e80461a7/file_83281.jpg [07:16:02] Maybe here: https://www.wikidata.org/wiki/Wikidata:Lexicographical_data/How_to_help (re @wmtelegram_bot: <[[smlckz]]> What can I read to better understand the linguistics side of things you've been doing here, lexemes and such?) [07:17:08] Rewrite how you would want that written, and make it a test. Satisfying the tests will be harder in some languages than others. (re @Jan_ainali: On another note, some sentences just sound weird in some languages 😆) [07:19:19] It sounds weird because the Swedish word for continent is "world part". So stating that it's "in the world" becomes redundant in a silly way. (re @u99of9: Rewrite how you would want that written, and make it a test. Satisfying the tests will be harder in some languages than others.) [07:20:46] In Swedish, we preferably would use another function and skip the location completely. [07:21:17] Is it needed in English? [07:22:09] Would "on Earth" be better? I'm also uncomfortable with using the QID for "world", because it's concept is about a term rather than a world [07:22:11] If it is the flattest continent, wouldn't that be precisely as clear? [07:23:29] Agreed! Maybe we even have an un-located superlative function already. Thanks. (re @Jan_ainali: If it is the flattest continent, wouldn't that be precisely as clear?) [07:46:22] It depends. I think in terms of link suppression, but an express requirement to suppress could be a specific wrapping function or a styling parameter. There may be other options. The ultimate effect would probably be to return a span like Douglas Adams (or not strong, as the case may be). (re @u99of9: I accept that we may [07:46:23] not have [07:46:24] the right type yet. Your comment helps me understand the point of your type proposal. With that pr...) [07:47:32] Z37738 is an interesting failing test and I am not sure what I would like to do. It fails because both L409186 and L35731 links to Q18204 and the function happens to pick one over the other. The shorter one (klass) is probably more used in non-bureaucratic Swedish. Perhaps some P6191 statement is needed... [08:01:23] Sounds plausible to me. (re @Jan_ainali: Z37738 is an interesting failing test and I am not sure what I would like to do. It fails because both L409186 and L35731 links ...) [08:06:54] Is it possible to reorder the listing of implementations? Or could we have automated sorting of them. Having to click to see the connected ones, as on Z27327, is tedious. [08:13:58] The display order on the function page depends on internal tables, so only staff could change that, I believe. Ordering or filtering by status sounds like a good idea to me, but I think it would be a new Feature request. (re @Jan_ainali: Is it possible to reorder the listing of implementations? Or could we have automated sorting of them. Having to click to see the...) [08:18:29] This is also an interesting case where we have two composition implementations that both passes all tests but by looking at the description, I can't tell why there are two different ones. (re @wikilinksbot: Z27327 – best lexeme for Wikidata item) [08:33:39] I've added a language style to one of them, and created a new test that passes 🎉 Z39919 (re @Jan_ainali: Z37738 is an interesting failing test and I am not sure what I would like to do. It fails because both L409186 and L35731 links ...) [08:34:17] How’s your Italian? [08:34:18] *Implementazioni [08:34:19] * [08:34:21] miglior lessema per elemento Wikidata, QID lingue (https://www.wikifunctions.org/view/it/Z36592) [08:34:22] miglior lessema per elemento, con 0 categorie less (https://www.wikifunctions.org/view/it/Z37525) [08:34:24] I guess “less” is a curtailment of “lessicale” rather than “lessema” (re @Jan_ainali: This is also an interesting case where we have two composition implementations that both passes all tests but by looking at thei...) [08:35:34] I would guess that it is nearly non-existent 😅 (re @Al: How’s your Italian? [08:35:34] Implementazioni [08:35:36] miglior lessema per elemento Wikidata, QID lingue [08:35:37] miglior lessema per elemento, con 0 cat...) [08:35:41] If both are acceptable, you can change the validation of the test to (something like) "is listed in" with a list of both options. In English both "Jupiter is the largest/biggest ..." seem acceptible. (re @Jan_ainali: Z37738 is an interesting failing test and I am not sure what I would like to do. It fails because both L409186 and L35731 links ...) [08:37:39] Oh! That sounds interesting! Do you know about some existing test using this method that I can look at and learn from? (re @u99of9: If both are acceptable, you can change the validation of the test to (something like) "is listed in" with a list of both options...) [08:39:48] Z13137 (re @Jan_ainali: Oh! That sounds interesting! Do you know about some existing test using this method that I can look at and learn from?) [08:42:28] Z34274 looks like a duplicate of Z27627? [08:42:42] Z31316 (re @u99of9: If both are acceptable, you can change the validation of the test to (something like) "is listed in" with a list of both options...) [08:43:38] That one was more similar, thanks! (re @Al: Z31316) [08:44:23] I can't find an ordinal/superlative without a location qualifier... 😔 (re @u99of9: Agreed! Maybe we even have an un-located superlative function already. Thanks.) [08:46:06] So I might eventually make "The Australian world part is the smallest world part in the world." (re @Jan_ainali: It sounds weird because the Swedish word for continent is "world part". So stating that it's "in the world" becomes redundant in...) [08:49:11] As of yesterday lexemes were giving me this for Z39865 : https://tools-static.wmflabs.org/bridgebot/a7614696/file_83285.jpg [08:49:37] Does this look correct? Z37738 (re @u99of9: If both are acceptable, you can change the validation of the test to (something like) "is listed in" with a list of both options...) [08:50:23] Yes. [08:59:15] Yes it's way easier to accurately remove links Z36835 than to accurately inject them Z37661. Strong could be toggled or switched in later Z37726. (re @Al: It depends. I think in terms of link suppression, but an express requirement to suppress could be a specific wrapping function o...) [09:12:16] "A canonical string representation is parsed to its structured form. This is a non-trivial operation, and initial implementations targeting an ad-hoc structure should be available before this proposal is finalised." Are there any decisions involved in this? (re @Al: I’m pretty sure it should be neither. The question for me is how close to HTML it should be. An [09:12:16] approach like [[W [09:12:18] ikifunctions:Ty...) [09:12:49] I think we need just one superlative function with optional qualifiers. But even the superlative could be a qualifier (essentially, part of the framing of the relation). How abstract is too abstract? (re @u99of9: I can't find an ordinal/superlative without a location qualifier... 😔) [09:28:24] We can wrap ultra-abstract for the AW writers, but it's hard on the language-composition functioneers. (re @Al: I think we need just one superlative function with optional qualifiers. But even the superlative could be a qualifier (essential...) [09:28:54] Broadly, we need some idea of what we want the HTML to look like. But I’m inclined to revisit that part and make final conversion to HTML a standalone function (with equality/equivalence still based on the resultant HTML, or (better?) a canonical structured representation that joins the string-lists). (re @u99of9: "A canonical string representation is parsed to its [09:28:54] structured f [09:28:55] orm. This is a non-trivial operation, and initial implementation...) [09:32:24] By "optional qualifiers" do you mean a list of some kind that can be empty, or are you waiting for optional args? Do you mean actual WD qualifier triples? (re @Al: I think we need just one superlative function with optional qualifiers. But even the superlative could be a qualifier (essential...) [09:34:28] I assume we aim for all of this to be made transparent to AW users? (re @Al: Broadly, we need some idea of what we want the HTML to look like. But I’m inclined to revisit that part and make final conversio...) [09:40:57] I’m not waiting for optional args. I think it’s a list of pairs: with QIDs representing “size”, “rank”, “context/location” etc. Not unlike snaks, of course. (re @u99of9: By "optional qualifiers" do you mean a list of some kind that can be empty, or are you waiting for optional args? Do you mean ac...) [09:47:09] Absolutely. AW only shows the rendered HTML fragments. However, if AW contributors need to construct HTML in situ, I believe it would be better if the construction were via the intermediate representation rather than by injecting strings directly into Z89s. (re @u99of9: I assume we aim for all of this to be made transparent to AW users?) [09:49:20] Looks like we need something similar to Scribunto mw.html library does. [10:04:53] Yes, it’s in the same problem space. My focus is not on the construction of the HTML but on deferring its construction, so that functions can operate on the structure without having to interpret actual HTML-fragments. (re @Winston_Sung: Looks like we need something similar to what Scribunto mw.html library does.) [10:05:04] <[[smlckz]]> I wonder why you need to parse and generate HTML like this anyway. Generate Wikitext, or an AST of Wikitext, and insert templates, for example, when you need something extra. As independent of natural languages, so independent of the output format/markup language as well. [10:11:24] Yes, I agree. But that’s not where we are today. I’m happy to back further away from HTML in the intermediate representation, so long as the generated text has adequate markup, particularly with regards to provenance (that this text is from some lexeme form or item label and, if modified, by which function). (re @wmtelegram_bot: <[[smlckz]]> I wonder why you need to parse [10:11:25] and [10:11:25] generate HTML like this anyway. Generate Wikitext, or an AST of Wikitext, and in...) [10:14:49] Do we have any chance of putting together a full list of markup requirements, or it it better to leave it extensible? (re @Al: Yes, I agree. But that’s not where we are today. I’m happy to back further away from HTML in the intermediate representation, so...) [10:16:36] "if modified by which function" sounds like a can of worms .. "by which tree of functions in which order with which parameters etc etc" [10:16:52] <[[smlckz]]> So you want metadata along with the generated output.. like debugging information in code generated by compiler, to make an analogy. [10:18:12] <[[smlckz]]> You can have different levels of depth of debugging information/metadata to choose etc. [10:18:49] Why do you choose QIDs rather than properties? (re @Al: I’m not waiting for optional args. I think it’s a list of pairs: with QIDs representing “size”, “rank”, “context/location” etc. ...) [10:23:32] Hah, yes… it is. But nothing breaks if it’s missing 😉 (re @u99of9: "if modified by which function" sounds like a can of worms .. "by which tree of functions in which order with which parameters e...) [10:26:48] They might be PIDs, but it’s a bit easier if they’re all the same kind reference 🤷‍♂️ (re @u99of9: Why do you choose QIDs rather than properties? (For the first member of each pair.)) [10:31:45] PIDs have the advantage that they are usually very well defined by community use. They are also one step closer to the source data if it comes from WD. On the other hand they are a limited set so we would need to know they had all the options we need. They are also one step further from lexemes. [10:34:23] QIDs are also not always clear in their directionality. Also, wouldn't it would be hard for functioneers to deal with infinite possible "predicate" QIDs? [10:35:00] Yes, I’m certainly not objecting to PIDs, but I think it would have to be PID/QID in general (or just QID). (re @u99of9: PIDs have the advantage that they are usually very well defined by community use. They are also one step closer to the source da...) [10:36:29] By "PID/QID" in general do you mean pick either and stick to it, or PIDs are incomplete so need QIDs mixed in? (re @Al: Yes, I’m certainly not objecting to PIDs, but I think it would have to be PID/QID in general (or just QID).) [10:36:45] The latter. (re @u99of9: By "PID/QID" in general do you mean pick either and stick to it, or PIDs are incomplete so need QIDs mixed in?) [10:37:05] Oh no, that sounds terrible! [10:38:07] So hard to handle if we don't know which type a thing is in advance! What are PIDs missing that you expect we will need? [10:38:30] It’s NLG 😏 (re @u99of9: Oh no, that sounds terrible!) [10:40:14] But that means that every functioneer has to consistently verify infinity QIDs into the same meaning across languages??? [10:40:32] *verbify [10:42:27] Are we even still talking about the superlative function, or have you generalised to everything? (re @Al: It’s NLG 😏) [10:57:46] This is the “too much abstraction” problem. In general, contextual qualifications are open-ended in natural language. I think it’s more acute with ranks and superlatives, because the context may be contrived (or particularised) so as to improve the ranking. For example, who would think of saying that Europe is a continent with relatively high flatness? (re @u99of9: Are we [10:57:46] e [10:57:46] ven still talking about the superlative function, or have you generalised to everything?) [11:03:18] In practice, [nth most] has the degenerate case where n=1, and the context may be implicit according to the placement of the assertion (or generally): “Paris is the capital of France and its largest city” [by most measures, if not all]. (re @u99of9: Are we even still talking about the superlative function, or have you generalised to everything?) [11:05:41] An Englishman with poor vision? (re @Al: This is the “too much abstraction” problem. In general, contextual qualifications are open-ended in natural language. I think it...) [11:07:14] Always a risk! 😎 (re @u99of9: An Englishman with poor vision?) [11:29:24] "Less" was for "Lessicale", i.e. "lexical" in English (re @Al: How’s your Italian? [11:29:25] Implementazioni [11:29:27] miglior lessema per elemento Wikidata, QID lingue [11:29:28] miglior lessema per elemento, con 0 cat...) [12:36:24] My stretch goal is for final HTML to be reversible to the intermediate structure and, ideally, to a plausible constructor for the intermediate structure. I suppose that’s analogous to reverse engineering plausible source code from object code (hard!). The “metadata” is a representation of the content that would otherwise be lost in conversion, like the Z11K1 as a lang [12:36:24] attri [12:36:25] bute when the Z11K2 becomes the text node. (re @wmtelegram_bot: <[[smlckz]]> So you want metadata along with the generated output.. like debugging information in code generated by compiler, to...) [13:21:30] 7233 [13:21:40] welcome Kieran [14:11:18] <[[smlckz]]> Al: "reverse engineering plausible source code from object code" that's why compared it with debugging info from compiling (otherwise how could you use debugger to find symbols to jump at to debug, etc.) if we are to do this well, even the underlying programming system should be revamped, for this is not really something that can be bolted on. How [14:11:18] <[[smlckz]]> willing are you/the team to make the overhaul needed? [14:21:50] <[[smlckz]]> My suggestion is to keep natural language generation, and output presentation seperate; like you say, delaying generating HTML (or other formats for that matter) as much as possible. Make a "debugging level" parameter, depending on it, the generator will attach metadata, the exact specificity depending on it. You'd inspect the format independent [14:21:50] <[[smlckz]]> AST of the result from generation stage.. okay this sounds more and more like stages of target-independent compiler, sigh [14:29:20] <[[smlckz]]> If you want to go the deep end, and must HTML, there's always RDFa 😿 https://en.wikipedia.org/wiki/RDFa [17:22:24] Yes, I think that’s the broad shape of the plan. I should try to articulate something on-wiki, but I wouldn’t want to appear to be constraining the new community. [17:22:25] At a super-high level we can relate to the [[:en:Natural language generation#Stages]]. I’d allocate Content determination and Document structuring to the editorial community on Abstract Wikipedia. So AW is essentially a KR repository, like Wikidata. Then we skip to the end of the pipeline, Realization, which is (almost currently) the semi-persistent HTML. [17:22:27] That leaves Aggregation, Lexical choice and Referring expression generation as the Wikifunctions domain, but we skip Aggregation, for the most part. We’d probably need to split it up between language-neutral and language specific, with language-neutral aggregation residing within AW or at the Wikifunctions interface. For example, uniting the concepts of capital city and [17:22:27] largest [17:22:28] city for Paris, is an editorial (AW) decision, but different languages may prefer different forms of expression (separate sentences, clauses or phrases, for example). [17:22:30] Referring expression generation (“REG”) is harder in theory than in practice (I hope) because we should conserve identity from the KR through to the lexicalized pre-HTML. This is where the current “single-shot” approach fails, because the scope of natural-language substitutions will typically extend to whole paragraphs or sections. It may be possible to limit the NLG pass [17:22:30] [17:22:31] es to two, assuming that (encyclopedic) content completion is performed once at the outset. [17:22:33] In summary, AW provides the persistent KR specification. Its language-neutral actualization is the input to language specific processing, and ought to be persisted (probably after Aggregation). Lexicalization is then per-language (NLG pass one) followed by article/section level REG and Realization (NLG pass two). The result is persisted as HTML (in the current rollout [17:22:33] plan). (re [17:22:34] @wmtelegram_bot: <[[smlckz]]> My suggestion is to keep natural language generation, and output presentation seperate; like you say, delaying gene...) [18:01:26] <[[smlckz]]> "I wouldn’t want to appear to be constraining the new community." I have a similar feeling when I want to talk about type theory and such. I have found little indication from the community about how to proceed in developing the programming system aspect further, in the long term. I'd otherwise be tempted to write a long article talking about the [18:01:26] <[[smlckz]]> theory and latest research I came across. [18:05:06] <[[smlckz]]> For example: https://github.com/oils-for-unix/oils/wiki/Perlis-Thompson-Principle [18:05:57] The Wikifunctions community is more mature than the Abstract Wikipedia community. Assuming you’re talking about implementing pure functions, I’d just go for it! (re @wmtelegram_bot: <[[smlckz]]> "I wouldn’t want to appear to be constraining the new community." I have a similar feeling when I want to talk abou...) [18:12:40] <[[smlckz]]> Or should we proceed towards compilation, requiring analysis that'd benefit from having multi-language semantics: https://dl.acm.org/doi/10.1145/3571220 [18:16:04] <[[smlckz]]> Al: okay, I'll then do a write-up. [18:21:57] <[[smlckz]]> "Assuming you’re talking about implementing pure functions" I'm talking about in this sense: https://www.wikifunctions.org/wiki/Wikifunctions:Project_chat#c-Smlckz-20260812113100-Wikifunctions_as_a_programming_system [18:27:22] <[[smlckz]]> I still await a reply from DVrandecic (or others) in this thread about clarification on the teams' understanding of the nature of the programming system and the direction they wish to develop. I wouldn't want to be presumtuous in this regard. [18:35:00] Right, but in Wikifunctions: [18:35:00] Everything is an object, specifically a ZObject with a JSON representation. [18:35:01] All functions are pure (unless they’re not, like the builtin fetches). [18:35:03] Apart from that, anything goes (in theory) but the project supports Wikimedia, so non-Wikimedia uses could extend WikiLambda but (I imagine) not enhance it. [18:35:04] In any event, it sounds like your interests are not off-topic for Wikifunctions, even if they are not an immediate priority. (Stress: if.) (re @wmtelegram_bot: <[[smlckz]]> "Assuming you’re talking about implementing pure functions" I'm talking about in this sense: https://www.wikifuncti...)