diff --git a/docs/arz/oracles.md b/docs/arz/oracles.md new file mode 100644 index 00000000..ec9ddc8c --- /dev/null +++ b/docs/arz/oracles.md @@ -0,0 +1,10 @@ +# Egyptian Arabic (arz): parked — thin, inconsistently voweled + +UniMorph `arz` has ~1,320 verb lemmas, but only **636** carry both a +3sg-masc perfective and imperfective (the principal parts the templatic +engine seeds from) — 48%, far below the 99.5% lemma-coverage bar. The +vocalisation is also inconsistent (كان unvoweled next to سِمِع partly +pointed), so a rule engine can't reliably match the target orthography. +kaikki.org has **no** Egyptian Arabic verb dump, so there is no second +oracle and no way to cross-check. Revisit if a fuller, consistently +voweled dialect lexicon (e.g. CALIMA-EGY output) becomes downloadable. diff --git a/docs/syc/oracles.md b/docs/syc/oracles.md new file mode 100644 index 00000000..2997ba57 --- /dev/null +++ b/docs/syc/oracles.md @@ -0,0 +1,10 @@ +# Classical Syriac (syc): parked — sparse seeds, participle-heavy + +UniMorph `syc` has ~755 verb lemmas / 17k rows, but after stripping the +Syriac vowel points only **269** lemmas (36%) carry both a 3sg perfect and +imperfect, so principal-part seeding covers nowhere near the 99.5% lemma +bar. The paradigm is unusually participle-heavy (the active/passive +participles inflect for person and dominate the rows) and the roots are +written defectively (2–3 consonants), which the templatic engine can't +disambiguate from the citation alone. kaikki has no usable Syriac verb +extraction. A dedicated Aramaic FST oracle would be needed; parked until then.