From c0167d6e1987bb95683594fc464c310ab6a1233e Mon Sep 17 00:00:00 2001 From: v4nn4 <1796093+v4nn4@users.noreply.github.com> Date: Mon, 24 Aug 2026 00:34:15 +0200 Subject: [PATCH] docs: park Egyptian Arabic + Syriac (data-quality blocked) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The Semitic expansion shipped the three Arabic-linked languages with usable oracle data — Arabic (Verified), Hebrew (Verified), Amharic (Beta). The remaining cousins are blocked by data quality, documented here alongside the existing Maltese note: - Egyptian Arabic: only 48% of lemmas seedable, inconsistent voweling, no 2nd oracle. - Classical Syriac: only 36% seedable, participle-heavy, defective roots, no 2nd oracle. - Maltese: kaikki expands 3 of 7,315 verbs (see docs/mlt/oracles.md). Co-Authored-By: Claude Opus 4.8 (1M context) --- docs/arz/oracles.md | 10 ++++++++++ docs/syc/oracles.md | 10 ++++++++++ 2 files changed, 20 insertions(+) create mode 100644 docs/arz/oracles.md create mode 100644 docs/syc/oracles.md diff --git a/docs/arz/oracles.md b/docs/arz/oracles.md new file mode 100644 index 00000000..ec9ddc8c --- /dev/null +++ b/docs/arz/oracles.md @@ -0,0 +1,10 @@ +# Egyptian Arabic (arz): parked — thin, inconsistently voweled + +UniMorph `arz` has ~1,320 verb lemmas, but only **636** carry both a +3sg-masc perfective and imperfective (the principal parts the templatic +engine seeds from) — 48%, far below the 99.5% lemma-coverage bar. The +vocalisation is also inconsistent (كان unvoweled next to سِمِع partly +pointed), so a rule engine can't reliably match the target orthography. +kaikki.org has **no** Egyptian Arabic verb dump, so there is no second +oracle and no way to cross-check. Revisit if a fuller, consistently +voweled dialect lexicon (e.g. CALIMA-EGY output) becomes downloadable. diff --git a/docs/syc/oracles.md b/docs/syc/oracles.md new file mode 100644 index 00000000..2997ba57 --- /dev/null +++ b/docs/syc/oracles.md @@ -0,0 +1,10 @@ +# Classical Syriac (syc): parked — sparse seeds, participle-heavy + +UniMorph `syc` has ~755 verb lemmas / 17k rows, but after stripping the +Syriac vowel points only **269** lemmas (36%) carry both a 3sg perfect and +imperfect, so principal-part seeding covers nowhere near the 99.5% lemma +bar. The paradigm is unusually participle-heavy (the active/passive +participles inflect for person and dominate the rows) and the roots are +written defectively (2–3 consonants), which the templatic engine can't +disambiguate from the citation alone. kaikki has no usable Syriac verb +extraction. A dedicated Aramaic FST oracle would be needed; parked until then.