The near-Aleppo dataset

The near-Aleppo dataset is similar to MAM-parsed-plus (mpplus). This page describes what near-Aleppo changes relative to mpplus. For what is in common with mpplus, see the documentation for mpplus.

The mpplus dataset is a parsed form of the Wikitext sources for Miqra According to the Masorah (MAM). The Hebrew of near-Aleppo is nearer to the Aleppo Codex's body text than the Hebrew of mpplus.

By “body text,” we mean the contents of Aleppo's main, large-letter columns. There are usually two or three such columns per page. We exclude Aleppo's small-letter content, such as its column-margin notes (Masorah qetannah) and its top- and bottom-margin notes (Masorah gedolah).

The one type of small-letter content we do include is the information in column-margin qere notes. We include that information and supply the qere letters with the pointing that we judge to be implied by the pointed ketiv words in the body text. Unlike mpplus, near-Aleppo provides pointed ketiv words rather than only their letters.

The near-Aleppo dataset differs from mpplus in more than just its Unicode strings. It also uses a slightly different set of templates to structure the string data. The JSON reference describes those differences.

Most changes to MAM that near-Aleppo makes are based on information found in MAM itself. Some of that information is in the general policies laid out in MAM's Introduction. When the Introduction describes a systematic, invertible difference between MAM and the Aleppo Codex, near-Aleppo applies the inverse to MAM's text. For example, MAM supplies stress helpers for all accents, whereas the Introduction describes Aleppo as generally lacking helpers for accents other than pashta. So, near-Aleppo inverts (in this case undoes) MAM's work by omitting those helpers except where MAM's notes justify keeping them. For example, at Genesis 2:7, we can see that near-Aleppo removes MAM's telishah qetannah stress helper:

mpplus

וַיִּ֩יצֶר֩

Near-Aleppo

וַיִּיצֶר֩

Other changes that near-Aleppo makes to MAM are based on MAM's documentation notes. A MAM documentation note targets a particular word within a particular verse. When a note says that MAM diverges from Aleppo at a particular word, the note typically gives Aleppo's form, so in such cases near-Aleppo uses Aleppo's form of the word rather than MAM's form. For example, at Isaiah 27:5, MAM's note gives Aleppo's form without the dagesh on zayin, and near-Aleppo follows that form:

mpplus

בְּמָעוּזִּ֔י

Near-Aleppo

בְּמָעוּזִ֔י

The near-Aleppo dataset is not solely derived from mpplus and the Introduction to MAM; some external references were consulted as well. For instance, for tricky cases of ketiv pointing, we have consulted external references such as the Jerusalem Crown edition.

Near-Aleppo aims to provide a continuous text that moves seamlessly between an Aleppo-diplomatic edition where the codex survives and an Aleppo-flavored edition where it is missing. In missing sections, preserved testimony and photographs taken before the loss remain evidence of the codex’s text. Where that evidence is lacking, we apply editorial policies reflecting our best guess of the Tiberian Masoretic consensus. We do not present these editorial guesses as a proposed reconstruction of the codex.

Near-Aleppo is 24 JSON files in this repository's out/near-aleppo/plus/, one for each of MAM-parsed-plus's book files, in its layout and its serialization. Of its 23,202 verses, near-Aleppo's text differs from MAM-parsed-plus's at 13,877 (59.8%).

In addition to the JSON near-Aleppo dataset, we provide an example HTML edition to show one kind of human-readable edition that can be made from near-Aleppo.

Documentation