Commit Graph
100 Commits
Author SHA1 Message Date
aarneranta 0f4dd88f9e added a missing ParadigmsGer.mkPrep case to maintain bw compatiblity of the API 2026-09-09 09:47:27 +02:00
aarneranta f6576c4b36 added a type annotation to make ExtendGer compile with v. 3.12 type checker 2026-09-08 14:29:12 +02:00
aarnerantaandClaude Opus 4.8 363219730e polish: put the AdvVP adverbial after the verb
AdvVP built its result with setPrefix, so the adverbial landed before the
verb. Polish is neutrally SVO, and this matters beyond adverbs proper: an
object reached through PrepNP arrives at AdvVP as an Adv, so a two-place
verb came out as *"π zbiór pusty przecina" rather than "π przecina zbiór
pusty". Both orders are grammatical, but the preverbal one topicalizes.

"śpi tutaj" is as good as "tutaj śpi", so the RGL's own example is
unaffected.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-16 19:50:20 +02:00
aarnerantaandClaude Opus 4.8 8c4cd58bd1 polish: complete the RGL enough to build an application grammar
Adding Polish to informath (grammars/next) exposed a set of gaps. Unlike
Czech, StructuralPol was already complete; what was missing was a handful
of lins, a way to govern the nominative, and some outright mistakes.

Missing lins, all reached as qualified Grammar.X by the application's
functors and so not shimmable from the application side:

  AdjectivePol  AdvAP
  SentencePol   SSubjS
  IdiomPol      ImpP3
  StructuralPol that_Subj
  MarkupPol     new module (Markup is not part of the RGL build script,
                so it is compiled on demand, as MarkupCze is)
  ParadigmsPol  mkAdv (every other ParadigmsX has it; the application
                interfaces reach it as Paradigms.mkAdv)

ImpP3 has no third-person imperative to use, so it takes the standard
"niech" periphrasis. The copula takes the future ("niech x będzie grupą")
and every other verb the present ("niech x należy do A"): być is the only
Polish verb with a synthetic future, and imienne marks exactly the copular
VPs.

NomPrep: ComplCase had no nominative at all, so "jako", "niż" and
"zdefiniowany jako" -- all of which govern the nominative -- came out
locative ("mniejszy niż liczbie"). mkCompl now maps Nom to a new NomPrep,
and the dep tables of the pronouns, nounPN, mkPN and the structural NPs
cover it. LexiconNounPol is marked DO NOT EDIT, but paris_PN and john_PN
inline their own dep tables and so had to be extended too.

Two corrections:

  VerbPol.CompCN used the nominative for a predicative noun, giving
  "x jest grupa"; Polish uses the instrumental, as CompNP right below it
  already does.

  ExtendPol.ExistsNP inherited ExistNP from ExtendFunctor, giving
  "jest macierz". Polish distinguishes the two: "there exists" is istnieć,
  which is what mathematical prose uses.

ParadigmsPol.guess_paradigm_basic never matched -ość, the productive
feminine abstract suffix, so sprzeczność took a masculine declension
(*sprzeczności/a, *sprzecznościowi) despite being tagged feminine. It now
routes to the kość paradigm.

Also documents, without changing, why the 2-string mkN throws its genitive
away: guess_paradigm's 2-string table is unsound (its first branch matches
every noun in -a, and its branches disagree about whether mkNTable* takes
the nominative or the bare stem), so passing sggen to it turns
"liczba"/"liczby" into *liczbaa. The 1-string guesser is the sound path
until that table is repaired.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-16 19:22:24 +02:00
aarnerantaandClaude Opus 4.8 11f73acc28 czech: implement three-place verbs (V3)
Add a proper V3 category to the Czech RGL, previously absent (V3 fell
through to the {s:Str} default and Slash2V3/Slash3V3 were notYet stubs).

- CatCze: lincat V3 = VerbForms ** {c, c2 : ComplementCase}; extend
  VPSlash with a trailing `ind` field for the incorporated indirect
  object (mirrors the English VPSlash c2 field).
- VerbCze: implement Slash2V3 and Slash3V3; ComplSlash now renders the
  `ind` string after the object slot, and SlashV2a sets ind = []. This
  keeps the object before the indirect object under the shared
  ComplV3 v o d = ComplSlash (Slash3V3 v d) o.
- ParadigmsCze: add the mkV3 paradigm (default acc/dat, plus an
  explicit two-complement-case form).
- MissingCze: drop the now-implemented Slash2V3/Slash3V3 stubs.

V2 predication is unchanged (ind is empty for V2-derived VPSlash).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-16 17:51:41 +02:00
aarnerantaandClaude Opus 4.8 344342012d czech: fix noun paradigm selection, add loanword and adjectival types
declensionNounForms only tested hardConsonant (d t g h k n r), so the neutral
consonants (b f l m p s v), which take the hard endings, matched nothing and
fell through to the soft declSTROJ fallback: graf, atom, profil, přístup were
all declined as if soft. Same gap in guessNounForms. Add hardishConsonant
(hard + neutral + the foreign z, x) and use it in both.

The genitive passed to the three-argument mkN was only used to pick a
paradigm, then thrown away, so the oblique stem was re-derived from the
nominative by dropFleetingE. That guesses wrong in both directions: it fires
on člen -> člnu, and fails to fire on uzel -> uzelu, doplněk -> doplněku,
výpočet -> výpočetu. Give declHRAD, declHRADA and declMUZ stem-taking
variants and let declensionNounForms pass the stem it was given. The
one-argument guessNounForms keeps the old dropFleetingE behaviour.

New paradigms, all previously falling back to declSTROJ:
  declLATINUS   algoritmus - algoritmu, kosmos - kosmu (also -os)
  declLATINUSA  the same, masculine animate
  declLATINUM   kontinuum - kontinua
  declGREEKMA   schéma - schématu
  declHRADA     les - lesa, zákon - zákona
  declADJF      proměnná - proměnné
  declADJM      nultý - nultého
  declINVAR     indeclinable loans: bombé, tamari, software

Also match feminine consonant stems by genitive (-i -> kost, -e/-ě -> píseň),
masculine -e/-ě genitives (kužel - kužele, král - krále), and neuter -ě
(těžiště), none of which were reachable before.

Measured on the 1177 nouns of informath's WikidataWordsCze: 225 hit the
fallback before, 1 after (the acronym ADE).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 16:32:25 +02:00
aarnerantaandClaude Opus 4.8 a79202f667 czech: fill in the lins needed to build an application grammar
The Czech resource was too incomplete to compile a Syntax/Grammar client:
Sentence, Idiom and Question defined almost nothing, and many Structural
words were absent, so MissingCze's notYet stubs were reached at PMCFG
generation time.

Added:
  CatCze        lincat VV
  SentenceCze   AdvS, ExtAdvS, SSubjS, EmbedS, EmbedQS, ImpVP
  IdiomCze      ImpP3, ExistNP, ImpersCl, GenericCl
  VerbCze       CompCN, ComplVS, ComplVV, PassV2
  NounCze       SentCN, PredetNP
  AdverbCze     SubjS
  QuestionCze   QuestIAdv
  PhraseCze     UttImpSg, UttImpPl, UttImpPol
  StructuralCze all_Predet, both7and_DConj, between_Prep, by8means_Prep,
                can_VV, either7or_DConj, every_Det, if_Subj, no_Quant,
                on_Prep, someSg_Det, that_Subj, under_Prep, where_IAdv
  ExtendCze     ExistsNP no longer excluded

Three of these approximate, because VerbForms lacks the required forms:
ImpVP uses the 1st person plural present ("předpokládáme, že ...") as there
is no imperative; PassV2 uses the reflexive passive ("číslo se dělí") as
passpart is commented out in ResCze; ImpP3 uses "nechť", which suits
mathematical text more than the "let John walk" of the RGL example.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 15:29:55 +02:00
aarneranta 8a82c2ca53 added Cze.UttVP 2026-07-10 14:56:08 +02:00
aarnerantaandClaude Opus 4.8 e866940d65 add the Czech Markup module
MarkupNP marks s, clit and prep, since in Czech these are alternative
surface forms of a pronoun chosen by context, as in MarkupRomance.
The nominative clitic is left alone: it is the pro-drop subject, empty
for every personal pronoun, so marking it up gave "<b> </b> je starý".

Inherit MarkupCze in LangCze, as LangFin does; stringMark is excluded
because abstract Lang already excludes it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 14:16:34 +02:00
aarneranta 8f1b3eeece compiling SymbolicCze by default 2026-07-10 14:08:24 +02:00
aarnerantaandClaude Opus 4.8 b14c0f8492 complete the Czech Symbol module
Implement the remaining Symbol functions, which previously fell back to
GF's default linearization: IntPN, FloatPN, NumPN, CNNumNP, CNIntNP,
CNSymbNP, SymbS, SymbNum, SymbOrd, and the [Symb] list.

Symbols and numbers are indeclinable neuters, but a cardinal used as a
name still declines (NumPN). CNNumNP treats the numeral as an invariable
label, so the noun carries the case: "město / městu / městem pět".
CNSymbNP goes through numSizeForm and numSizeAgr, as DetCN does, so
SymbNum's Num5 size gives "n měst x a y je staré".

SymbOrd is only approximate: CatCze has no lincat Ord, so it defaults to
{s : Str} and the masculine "n-tý" cannot agree.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 12:29:30 +02:00
aarneranta 2e6a750c1a some fixes in ExtendFin to make it compile again 2026-07-10 11:28:00 +02:00
aarneranta 992e670c10 ParadigmsGre: case for mkPN Str Gender 2026-05-12 14:53:15 +02:00
Aarne Ranta f4aae6a61a Merge pull request #466 from hleiss/corrections
(Ger) Corrections in Cat, Irreg, Lexicon ; added QVP in Question and GenIP, GenModIP in Extend
2025-10-31 17:18:05 +01:00
Aarne Ranta f01d509376 word order in Ger CompCN and CompNP: RS in ext part, adv before it; agr for CompCN todo 2025-09-18 20:36:03 +02:00
Aarne Ranta 607d9240c2 Merge pull request #460 from hleiss/master
(Ger) Added missing linearizations pot21 to pot5plus in NumeralGer
2025-03-17 08:02:35 +01:00
Aarne Ranta ababe72fb8 started adding capital letters to Gre.FemAccFinalN; most letters still todo 2025-03-12 16:32:40 +01:00
Aarne Ranta 693cd88f7b heuristic rule for Gre mkPN 2025-02-13 10:43:17 +01:00
Aarne Ranta 6876b79d43 added MarkupFre to LangFre 2025-02-02 15:32:17 +01:00
aarneranta 8613a78929 changed a runtime + to BIND in Amharic 2024-12-20 11:17:11 +01:00
aarneranta e63a3c9d17 fixed ExtendSwe/VPS in accordance with Ger 2024-09-04 15:03:50 +02:00
Aarne Ranta d895f4ceeb MorphoDictSpa copied from DictSpa 2024-07-30 17:02:36 +02:00
Aarne Ranta ba2e4a4964 added Extend.SubjunctRelCN 2024-07-30 08:44:01 +02:00
Aarne Ranta 78f333e79e added a case to MakeStructuralFre.mkDet 2024-07-29 18:03:09 +02:00
Aarne Ranta 710e73688d in Formal.usePrec, removed variant parenthOpt for a better control of parentheses ; use Formal.parenth explicitly if you want to parse superfluous parentheses 2024-07-29 17:44:05 +02:00
Aarne Ranta 82075be37b corrected the formal subject of ExistsNP in Fre 2024-07-26 17:33:32 +02:00
Aarne Ranta 7d631fafa2 put RelVPS into RomanceFunctor and also inherited other VPS functions from there 2024-07-26 17:22:17 +02:00
Aarne Ranta 78fc625174 added VPS functions to Romance Extend modules 2024-07-26 16:47:43 +02:00
Aarne Ranta 911cbb06c1 MorphoDictPor 2024-07-25 12:01:19 +02:00
Aarne Ranta 71f3b2dc78 MakeStructural.mkDet in Ita and Por 2024-07-25 11:55:55 +02:00
Aarne Ranta 23ccacf222 mkStrongDet, mkWeakDet in MakeStructuralGer 2024-07-25 09:18:31 +02:00
Aarne Ranta d18b889df6 one more fix word order in ExtendGer.PredVPS ; not yet correct for Sub with ConjVPS 2024-07-22 23:59:55 +02:00
Aarne Ranta 6084aef91b trying to fix word order in ExtendGer.PredVPS ; not yet correct for Sub 2024-07-22 23:47:03 +02:00
Aarne Ranta 687f0cefc8 lost AP.ext in CompAP restored, Inv worder still to be fixed 2024-07-22 21:10:50 +02:00
Aarne Ranta 479fe7236f took away duplicates of Structural from MorphoDictEng in conformance with the other MorphoDicts 2024-07-19 18:15:31 +02:00
Aarne Ranta 15ecf3217c fixed import Cat in MorphoDictFinAbs 2024-07-18 19:17:12 +02:00
Aarne Ranta 1f98f642af DictFre copied to MorphoDict 2024-07-17 18:00:52 +02:00
Aarne Ranta 03e81a3fab ImpP3 in Ita, Swe ; order fixed in Fre (but not necessarily the only correct one) 2024-05-07 15:53:33 +02:00
aarneranta 726d04f00f made lincat LN = NP in Arabic, to deal with all names; also added mkLN Str Gender 2024-04-18 10:14:28 +02:00
aarneranta a251318dad Fin linref NP: Nom, not Acc 2024-04-11 14:11:58 +02:00
Aarne Ranta 2363da4b2f Croatian GenRP corrected 2024-04-03 18:47:59 +02:00
Aarne Ranta cd3a2e0ac4 unverified definitions of some Croatian functions 2024-04-03 17:44:00 +02:00
aarneranta bd75a0529c Ara.ExtAdvS 2024-03-22 13:51:20 +01:00
Aarne Ranta bb5107fa99 added GenRP to Extend Fin,Fre,Ger,Ita,Swe 2024-03-20 16:34:30 +01:00
Aarne Ranta 97b713520f commas in Ger subord clauses: things remaining to do 2024-03-20 09:27:47 +01:00
aarneranta f8d100edee ParadigmsAra.mkLN added 2024-02-08 11:55:14 +01:00
aarneranta 651fd743c7 added MorphoDictAra to support WordNetAra 2024-02-08 11:48:52 +01:00
aarneranta 5bdd45a8be ExtendAra: quick and dirty PassVPSlash and CompoundN, to be revisited 2024-02-08 10:39:30 +01:00
Aarne Ranta 82891f3c84 slightly improved adjective inflection in Ara.wmkA, more todo 2024-02-06 11:33:00 +01:00
Aarne Ranta d2eb7bd46d NumeralSwe: fem miljoner 2024-02-06 09:03:51 +01:00
aarneranta 1e55251235 DocumentationAra completed with Adj tables 2024-02-01 16:20:48 +01:00
Aarne Ranta 9c3e3985bc DocumentationAra verb tables (active) 2024-02-01 07:52:26 +01:00
aarneranta 4c9747a8b6 DocumentationAra: basic implementation, to be completed 2024-01-31 17:03:57 +01:00
aarneranta 04d6642afe fixed Fin heading Erisnimi 2024-01-31 16:28:53 +01:00
Aarne Ranta b746a425d5 Fin: default case for plural names is exterior 2024-01-30 16:40:22 +01:00
Aarne Ranta d98bb99fc1 added missing InLN and dayMonthYearAdv in Fin 2024-01-30 16:23:17 +01:00
Aarne Ranta 4a12ecd9d3 plural inflection of Fin LN 2024-01-30 16:00:59 +01:00
aarneranta 441c8305dc added NamesAra to GrammarAra 2023-12-08 14:04:38 +01:00
aarneranta af6edf2e6e saved WordNetAra for post-editing 2023-09-29 10:44:57 +02:00
Aarne Ranta 67d1e24761 producing a compilable WordNetAra.gf, with a lot of junk 2023-09-28 16:18:23 +02:00
Aarne Ranta f19dcc01f9 import of unicodedata in arabic_utilities 2023-09-27 12:23:43 +02:00
Aarne Ranta 1c355ce9dd factored out arabic_utilities.py as a separate file 2023-09-25 09:22:21 +02:00
Aarne Ranta 561a8c130d to_wordnet applied to a new format of data 2023-09-25 08:22:47 +02:00
Aarne Ranta aa1dff6702 added MoreAra.gf 2023-09-21 17:29:38 +02:00
aarneranta 7e383b746e moved wikt-specific paradigms to a separate file (for the moment) 2023-09-21 15:46:41 +02:00
aarneranta fdd7c9641e Ara: improving Adj inflection by identifying fcl patterns from concrete forms 2023-09-20 16:05:46 +02:00
Aarne Ranta 2419931105 some more paradigms for Arabic Wiktionary generation 2023-09-20 11:54:59 +02:00
Aarne Ranta abcb3a9f2a improving evaluation of wiktionary generated lexicon 2023-09-20 11:54:29 +02:00
Aarne Ranta 9e8c5eaad5 arabic/wiktionary: including root in the form list 2023-09-18 08:52:32 +02:00
aarneranta 73f0b8ef00 commented and refactored read_wiktionary.py 2023-09-15 14:48:23 +02:00
Aarne Ranta edecc3fe57 a quick way to extract wordnet morphology 2023-09-14 18:21:18 +02:00
Aarne Ranta 3e9be76e52 evaluation of generated lexicon 2023-09-14 15:19:05 +02:00
Aarne Ranta d5e6e7e389 Arabic Wiktionary: functions for normalization and evaluation 2023-09-14 12:21:48 +02:00
aarneranta 8e029bd8dd Arabic Wiktionary: started comparing evaluation 2023-09-13 17:24:21 +02:00
aarneranta 3c0adada11 new function in ParadigmsAra to deal with Wiktionary data; lots of untested guesses 2023-09-13 15:29:28 +02:00
Aarne Ranta afc84a61cb arabic/wiktionary using paradigms with records as arguments to cope with heterogeneous information 2023-09-13 09:06:02 +02:00
Aarne Ranta 8eceb53643 compilable MorphoDictAra generation except for V, not yet using all forms 2023-09-12 19:38:14 +02:00
Aarne Ranta 714d8abac0 GF abstract dict generation 2023-09-12 17:04:50 +02:00
Aarne Ranta ae1c7f0061 extracting Arabic from Wiktionary, next step GF generation 2023-09-12 16:35:21 +02:00
Aarne Ranta 6312624a5f preparing to read Arabic morpholex from Wiktionary 2023-09-12 12:08:31 +02:00
aarneranta 03d3a8cbc2 disambiguated Number in Slovenian 2023-09-08 13:11:10 +02:00
aarneranta 5f9322683c added NamesAra 2023-09-08 13:10:49 +02:00
Aarne Ranta 84a03ef78a added PiedPiping to Extend and corrected the comment on RelSlash in RelativeEng 2023-08-21 22:21:47 +03:00
Aarne Ranta 3ab3eab36b Merge branch 'master' of https://github.com/GrammaticalFramework/gf-rgl 2023-08-21 19:45:20 +03:00
Aarne Ranta 20dcab3208 some more infinitives and their interpretations in Fin 2023-08-21 19:44:20 +03:00
Aarne Ranta a00d41dcfa re-enabled Ger.possess_Prep 2023-08-21 16:08:52 +03:00
Aarne Ranta 0cbae9883d commented out Names from Grammar temporarily to avoid failure with notYet 2023-08-21 15:59:33 +03:00
Aarne Ranta f08c9c0c93 1st infinitive long 2023-08-21 15:58:57 +03:00
Aarne Ranta e729e4e707 Fin: some more infinitives interpreted 2023-08-18 08:40:33 +03:00
Aarne Ranta fccc0428fd a couple more infinitives interpreted 2023-08-17 23:04:41 +03:00
Aarne Ranta 7a75cf3fb3 interpreting infinitives 2023-08-17 17:08:33 +03:00
Aarne Ranta 146bc71a06 clean-up after Fin infinitives 2023-08-16 16:05:28 +03:00
Aarne Ranta 31b025bdf1 completed infinitive and participle structures of Finnish 2023-08-16 11:32:11 +03:00
Aarne Ranta a9408305df added SCAcc in Finnish and fixed the subject form in passives 2023-08-15 22:16:20 +03:00
Aarne Ranta 7d483b1539 uses of Finnish infinitive forms in syntax 2023-08-15 22:15:38 +03:00
Aarne Ranta 9c5b87e1c8 Lav: comma in please_Voc like in VocNP 2023-08-15 11:29:25 +03:00
Aarne Ranta edfe72514b fixed irreg verbs in DictFre; reject non-verbs in ParadigmsGer (check that infinitive ends -n) 2023-04-04 15:51:09 +02:00
aarneranta ac2c5c52ac ExtraSpa.UseComp_ser 2023-03-29 14:00:29 +02:00
aarneranta 9c96fc6653 added Extend.PassVPSlash to Ita, Spa 2023-03-20 13:02:12 +01:00
Aarne Ranta 1da671016b Merge pull request #424 from harisont/master
More cases where the definite article "lo" (rather than "il") is used
2023-03-13 09:34:40 +01:00