Skip to content

Add Modern Standard Arabic — the first templatic engine (Verified) - #151

Merged
v4nn4 merged 1 commit into
mainfrom
rf/arabic-engine
Aug 23, 2026
Merged

v4nn4 merged 1 commit into
mainfrom
rf/arabic-engine

Conversation

@v4nn4

@v4nn4 v4nn4 commented Aug 23, 2026

Copy link
Copy Markdown
Collaborator

The headline of the Arabic-family goal, and ablaut's first root-and-pattern engine.

Approach

Arabic is cited by the unvoweled consonantal skeleton (كتب), which doesn't fix the vocalisation, so the paradigm can't be voweled from the lemma alone. The engine resolves each lemma to four fully-voweled principal parts (perfect/imperfect × active/passive 3sg-masc, mined into data/ara/parts.tsv) and derives the whole paradigm by regular person/number/gender/mood/voice affixation + weak-root morphophonemics — hollow, defective, doubled, assimilated, hamzated, generalised across the derived forms II–X. A small mined-override layer (data/ara/overrides.tsv, 1,112 cells) patches the irregular residue (doubly-weak, hamza-hollow).

Verification

Two independent oracles — kaikki.org (Wiktextract) ∩ UniMorph — normalised to a shared canonical feature vocabulary. The otiose-alif spelling of the -ū plural was the main systematic split; normalising it lifted agreement to ~94%, a ~48,300-cell gold across 640 lemmas.

Golden gate: forms 48288/48288 (100.00%), lemma coverage 640/640 (100.00%).
Split: 97.7% rule-derived, 2.3% overridden.

Registration

Lang::Ara, from_code ar/ara/arabic, Conjugation::Ara + build dispatch, a 9-group serde Table (perfect/imperfect/subjunctive/jussive/imperative × active+passive), reverse index, PyO3 + wasm bindings, Cargo includes, correctness.py, and the CI two-oracle gate. All green: 330 lib tests, clippy -D warnings, fmt --check, --features python and --features wasm.

Notes

  • The unvoweled lemma is inherently ambiguous for homographs (كتب resolves to the dataset's Form-II kattaba, not Form-I kataba); the engine faithfully reproduces whichever reading the oracles carry.
  • Participles and the verbal noun are nominal and out of the verb paradigm.
  • Web wiring (RTL, tense-group tiles) is a separate follow-up.

First of the Semitic set; Egyptian Arabic (data already banked), Hebrew, Amharic, Syriac, Maltese follow.

🤖 Generated with Claude Code

Arabic is root-and-pattern: cited by the unvoweled consonantal skeleton of the
3sg-masc perfect, whose vocalisation the skeleton doesn't fix. The engine
resolves each lemma to four fully-voweled principal parts (perfect/imperfect ×
active/passive, mined into data/ara/parts.tsv) and derives the whole paradigm
by regular person/number/gender/mood/voice affixation plus weak-root
morphophonemics — hollow, defective, doubled, assimilated and hamzated roots,
generalised across the derived forms II–X. A small mined-override layer
(data/ara/overrides.tsv, 1112 cells) patches the genuinely irregular residue
the rules don't reach (doubly-weak and hamza-hollow roots).

Verified against the agreement of two independent oracles — a kaikki.org
(Wiktextract) extraction crossed with UniMorph — normalised to a shared
canonical feature vocabulary. The otiose-alif spelling of the -ū plural was the
main systematic split; normalising it lifted agreement to ~94%, a ~48,300-cell
gold across 640 lemmas.

Golden gate: forms 48288/48288 (100.00%), lemma coverage 640/640 (100.00%),
97.7% rule-derived / 2.3% overridden. Full registration: Lang::Ara, from_code
ar/ara/arabic, Conjugation + build dispatch, reverse index, PyO3 + wasm
bindings, Cargo includes, correctness.py, and the CI two-oracle gate.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@v4nn4
v4nn4 merged commit 8724039 into main Aug 23, 2026
2 checks passed
@v4nn4
v4nn4 deleted the rf/arabic-engine branch August 23, 2026 21:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant