Almost errors

Introduction

This page documents the accent-grammar checker's “almost errors”: cantillation features that a naïve checker would flag, but that we do not flag — and, in each case, the reading we chose is a choice, not a forced move. Two kinds appear here.

First, the editorial charities: places where the checker silently normalizes away a genuine quirk of WLC — sometimes a real Leningrad Codex feature, sometimes an artifact introduced in BHS or WLC — and reads the text charitably rather than reporting an error. Here something at least questionable is being forgiven; the value is transparency about exactly what the checker quietly fixes, in which direction, and why.

Second, the masoretically-blessed oddities: features that look error-like — two accents crowding one letter or one word, or one divider written twice in a row — but that are 100% official masoretic tradition, attested in the standard manuscripts, not leniencies specific to LC, BHS, or WLC. Nothing here is forgiven; the checker accepts these, and where it must pick how to represent one for parsing — keep both marks as a sequence, fuse a pair into one token, carry a single mark, or collapse a repeated divider to one — that choice is, for the multi-accent cases, among readings that all parse cleanly (the telisha gedola exhibit below shows the alternatives). The headline case is Ezekiel 20:31’s mahapakh + qadma (mahapakh!qadma), the only word in Tanakh with two conjunctive accents on one letter.

Companion pages: the prose Goerwitz checker run and the poetic checker run list the verses the checker actually flags; this page is the inventory of what it deliberately does not.

Editorial charities

Each charity below names what the checker normalizes, the direction, and why, with a citation. Both reinterpret a mark the manuscript should not have here — a prose geresh muqdam read as a plain geresh, and a stray poetic geresh read as a geresh muqdam. The two signs are false friends — alike in shape, unrelated in grammatical role.

A third kind of charity, supplying a mark, arises only in detangling the dually-cantillated passages (the two Decalogues and Genesis 35:22): where WLC drops one reading’s accent on a word, that one mark is supplied from MAM so the reading parses. Because it adds a mark rather than rereading one, it is inventoried separately at Supplied and erased marks.

Geresh muqdam → geresh (prose)

Geresh muqdam (U+059D) is a poetic-only sign. In the 21 prose books WLC uses it just twice — Lev. 1:3 (alone) and 2 Kings 17:13 — as a typographic device standing in for a plain geresh. The checker reads it as a plain geresh, so the prose grammar (which has no geresh muqdam) sees the geresh it expects. Direction: poetic-looking sign → its prose false friend. tanach.us itself made the same correction in both verses — changes 2020.09.22-1 (Lev. 1:3) and 2020.09.22-2 (2 Kings 17:13), each described “Change geresh muqdam to geresh.”

The two verses differ in what happens next. In Lev. 1:3 the geresh muqdam stands alone, so the charity is the whole story. In 2 Kings 17:13 the converted geresh then sits on a word that also carries a telisha gedola — so once the charity has run, what remains is one of the telisha gedola + geresh oddities below (the only one of those five whose geresh reaches the checker by way of a charity).

Plain geresh → geresh muqdam (poetic): Psalms 124:4

A plain geresh in a poetic verse is otherwise a lexical error: the poetic grammar has no plain geresh. The sole charitable exception is Psalms 124:4, where a revia and a plain geresh share one letter. There the checker rereads the plain geresh as a geresh muqdam — its same-shape false friend — first normalizing the two same-letter marks into order (revia + geresh → geresh + revia). That reread is the whole charity; the pre-existing revia is left untouched. The ordinary geresh muqdam + revia → revia mugrash rule then consumes both as a single revia mugrash, the established poetic compound.

This within-letter order normalization is legal precisely because it stays within a single letter — we are liberal about mark order on one letter (questionable penmanship is not our concern) but preserve order across letters, which is meaningful reading order. See the tanach.us Psalms 124:4 note.

Masoretically-blessed oddities (not charities)

The features below would make a naïve checker blink — two accents crowding one letter or one word, or the same divider written twice in a row — but none of them is a quirk of LC, BHS, or WLC to be forgiven. They are official masoretic tradition, attested in the standard manuscripts. The checker accepts them; its only real decision is one of representation — whether to keep both accents as a sequence, fuse a pair into one token, carry a single accent, or collapse a repeated divider to one — and, as the telisha gedola exhibit shows, that decision is (for the multi-accent cases) a choice among readings that all parse cleanly.

telisha gedola + geresh/gershayim (five words)

Five WLC words carry both a “telg” (telisha gedola) and a “gerstar” (a geresh or gershayim). (In 2K17:13, the geresh results from our charitable interpretation of a geresh muqdam.) This double accent is not a quirk of WLC, BHS, or the LC: it is attested in the standard manuscripts. In three of the words, the two accents sit together on the first letter of the word (G5:29, Ts2:15, and 2K17:13). In the other two words, the telg sits on the first letter but the gerstar sits on a later letter (L10:4 and Ee48:10). Presumably this is because the stress is not initial in those two words, and the naqdan (the pointing-scribe) wanted to preserve the prepositive and impositive placement of telg and gerstar respectively.

In the three same-letter words, the LC has the gerstar first and the telg second. (See UXLC notes G5:29 and Ts2:15, and the UXLC change record for 2K17:13, 2020.09.22-2.) So, the same-letter words have their accent swapped compared to the order in the cross-letter words.

The checker reads a telg and a gerstar on a single letter in the order they are written in the manuscript — gerstar-first in the three same-letter words. That order is not the checker's own invention; it is inherited from the source data. In WLC 4.22's original Michigan-Claremont (M-C) encoding, the two accents stand in manuscript order: gerstar-first in the three same-letter words, telg-first in the two cross-letter words. My conversion of that encoding to Unicode preserves that order, so the “word” column of the table below shows all five words exactly as the manuscript orders them — gerstar-first in the three same-letter words — and the checker now reads them in that same order, rather than floating the prepositive telg to the front of its own reading. The checker can do this because it allows a telg and a gerstar to appear in either order; i.e., either order is considered grammatical. The telg-then-gerstar order appears normally, i.e. across separate words, about 175 times in Tanakh, while the gerstar-then-telg order appears only about 18 times.

In the two cross-letter words the telg leads instead — but there, too, that is simply the manuscript order. The telisha gedola is prepositive and is written at the front of its word wherever it is chanted, ahead of the later letter that carries the gerstar, so the telg-first order is forced by prepositivity rather than chosen by the checker. Either way, same-letter or cross, the checker preserves the order it finds.

Each verse continues to parse cleanly if either accent is dropped from these five telg + gerstar words. I mention this to show that the checker deems these verses grammatical no matter which of the various reasonable chanting interpretations are given to the two accents on these words:

Performing only one of two accents the manuscript has is itself an accepted choice in the reading tradition — compare the two accentuations of the Decalogue (Exodus 20 and Deuteronomy 5) and Genesis 35:22. The table below shows, for each word, the double accent form and the two single-accent “thought experiments.” The same-letter words are shown in manuscript accent order (gerstar-first).

versewordtelggerstarsame letter?
G5:29זֶ֞֠הז֠הז֞הyes
Ts2:15זֹ֞֠אתז֠אתז֞אתyes
2K17:13שֻׁ֜֠בוּש֠בוש֜בוyes
L10:4קִ֠רְב֞וּק֠רבוקרב֞וno
Ee48:10וּ֠לְאֵ֜לֶּהו֠לאלהולא֜להno

The tables further below show two examples of the checker's actual parse tree (one same-letter case, one cross-letter).

Tsefaniah 2:15

UXLC | MAM

0slqc
1atncslqc
2zaqqcatnczaqqcslqc
3pashczaqqpzaqqcatncrevczaqqczaqqcslqc
4gerppashcpashpzaqqptippatnplgmprevppashpzaqqppashpzaqqptippslqp
5telgppashp
ger2telgmah pashmun zaqqpashzaqqtipmun atnlgmmun revpashzaqqyetmun zaqqtipmer slq

Levit 10:4

UXLC | MAM

0slqc
1atncslqc
2zaqqcatnczaqqcslqc
3revpzaqqctippatnprevpzaqqctippslqp
4pashpzaqqppashczaqqp
5telgppashc
6gerppashp
mun revpashmun zaqqmer tipmun atnmun revtelgger2mah pashmun zaqqtipslq

A note on the trees above: in both the same-letter word (here Zephaniah 2:15, whose tree shows the gerstar before the telg) and the cross-letter word (Leviticus 10:4, telg before gerstar), the two accents are kept distinct — each read in the order it is written in the manuscript — never merged into a single unit, whether or not they share a letter. That is the contrast with Ezekiel 20:31 below, whose two accents do merge.

For MAM's own documentation notes on these five words — what each manuscript and edition reads, and in which order the two accents are taken — see the deep-dive translation.

Mahapakh + qadma (Ezekiel 20:31)

UXLC | MAM

In Ezekiel 20:31, נִטְמְאִ֤֨ים carries both a mahapakh and a qadma on its alef. It is the only word in Tanakh with two conjunctive accents on one letter. The checker accepts it outright: the scanner fuses the pair into one mahapakh!qadma token, which the grammar parses as an ordinary accent. As with the telg words, both accents survive; only the representation differs. The telg and its gerstar are two disjunctives, which the grammar admits as a two-accent sequence, so the checker keeps them as a sequence; the mahapakh and qadma are two conjunctives sharing one letter, which the scanner instead fuses into a single token. Either way both accents survive — the difference is sequence versus fused token. The two conjunctives do have a reading order — qadma before mahapakh, as the MAM note below stresses — but the fused token does not try to encode it.

That this double accent is intentional in the LC is supported by the word’s masorah qetannah note, ל̇ בטע̇ (“[this is the] one [word in all Tanakh that appears] with [these] accents [arranged like this]”):

Image source: 286A column 3 line 21

MAM has this double accent, and has a documentation note citing support for it from three standard manuscripts (Aleppo, Leningrad, Cairo) and their masorot. MAM also cites Yeivin 28.1 p. 232. MAM spells out why this double accent is puzzling yet standard:

זאת התיבה היחידה בכל המקרא שיש בה שני טעמים מחברים בהברה אחת. הקדמא קודמת למהפך בקריאה, כמו בעוד שש מקומות במקרא (שבהם הקדמא במקום הראוי לגעיה והמהפך בהברת הטעם), כגון: ויקרא כה,מו; במדבר [כ,א].
This is the only word in all of Tanakh with two conjunctive accents on one letter; the qadma precedes the mahapakh in chanting, as in six other places where [(on a single word)] a qadma occupies a syllable fit for a ga‘ya and a mahapakh occupies the stressed syllable, for example: Lev. 25:46; Num. 20:1.

The note names only those two as examples — and writes “Num. 21:1” where it means Num. 20:1. The other four verses alluded to above are as follows, making six total: Ezek. 43:11; 2 Chron. 35:25; Dan. 3:2; Ezra 7:24. See the full note on the MAM-with-doc Ezekiel page. Because the manuscripts agree, this double accent is whitelisted rather than treated an error.

The instructive contrast is Lev. 25:20, the only other prose word with two accents on one letter (a mahapakh and a tipeḥa). There the editions do not agree — MAM keeps only the tipeḥa and WLC tags the word anomalous — so it may well be an error in the LC, and the checker flags it (as a lexical error). Same surface shape, opposite verdict, decided by whether the sources agree. Its full treatment is on the Goerwitz page. The poetic merkha!azla of Psalms 56:10 is the same story on the poetic side: the double accent in the LC has no support from other manuscripts, so it, like Lev. 25:20, is flagged as an error.

How forced is the fusion? Unlike the telg readings above — where every alternative parses cleanly and the choice is one of faithfulness — here the grammar all but dictates it. The table below runs the verse through the real checker under each of the five ways to present the pair: the fused token, dropping either accent, and keeping both as a sequence in either order. Only two parse: the fused mahapakh!qadma token and the qadma-then-mahapakh sequence — and those coincide, because the same pashta phrase rule that accepts the fused cluster also accepts the cross-letter qadma mahapakh pair, in exactly the reading order the MAM note describes (qadma before mahapakh).

Why the other three readings fail is worth spelling out, because it is not merely that “a word needs an accent.” The pashta here is served by three conjunctives: a “telq” (telisha qetanna) on the preceding word אַתֶּם, then the qadma and mahapakh sharing this word’s alef. Once a telq heads the chain, the grammar admits only telq qadma mahapakh pashta (or its one-token analogue telq mahapakh!qadma pashta) — the qadma is obligatory, because the masoretic rule (Yeivin §246) is that a telq is always followed by qadma. So keep the mahapakh, drop the qadma leaves a telq directly before a bare mahapakh — a sequence the tradition never produces and the grammar has no rule for — and the parse falls into error recovery (verdict ERROR, not even a clean non-parse). Dropping the mahapakh instead leaves a qadma with nothing between it and the pashta, also unruled (qadma must be followed by mahapakh or merkha); and the mahapakh-then-qadma order is simply backwards. A bare mahapakh pashta is a perfectly legal one-servus pashta elsewhere, so dropping the qadma would parse if this pashta stood alone — but it does not, because the telq makes the chain three deep.

readingverdict
fuse into one mahapakh!qadma token (what the checker does)clean
keep the mahapakh, drop the qadmaERROR
keep the qadma, drop the mahapakhERROR
keep both, as a qadma then mahapakh sequenceclean
keep both, as a mahapakh then qadma sequenceERROR

The checker’s parse tree for the verse, with the fused token shown as mahapakh!qadma:

0slqc
1atncslqc
2zaqqcatnczaqqcslqc
3pashczaqqptipcatnprevpzaqqctippslqp
4pazppashctevptipppashpzaqqp
5gerppashp
mun paztelq qom gertelq mah!qom pashzaqqtevmer tipmun atnrevpashmun zaqqtipslq

Double tsinnor (Psalms 17:14)

UXLC | MAM

Psalms 17:14 carries two צנור (tsinnor) marks in a row — on the adjacent words בַּחַיִּים and the qere וּצְפוּנְךָ — the only place in the Three Books where two tsinnor occur consecutively. A disjunctive divider written twice in a row looks like an error, but it is no quirk of LC, BHS, or WLC: Sassoon 1053 clearly shares the doubled mark, and Breuer (following Wickes) calls it the sole Scriptural example of two consecutive tsinnor — the poetic counterpart of the doubled zarqa that Yeivin notes for the prose system. It is masoretic tradition, not a leniency to be forgiven.

As with the telg and Ezekiel 20:31 exhibits, the checker accepts the verse and the only question is one of representation — but here, instead of keeping both accents (as a sequence or a fused token), it drops one. The two marks are the same divider written twice, and the underlying LALR(1) grammar cannot parse that repetition directly. Rather than extend the grammar for this single verse, the checker just deletes one of the two tsinnor from the copy it feeds the parser, and parses what remains. This is a hack for our own convenience, not a more faithful reading: collapsing the repeat throws away a mark the LC really does write. It is tolerable only because the doubled mark is legitimate — so this is not a verse we would want to flag in any case — and because the deletion touches only the parser's input: the recorded accents are left untouched, so the WLC-vs-MAM comparison still sees both marks. (That the workaround is a pre-parse deletion rather than a grammar rule is also why there is no alternate-reading tree to show here.)

The full account — the Leningrad and Sassoon 1053 images, MAM's documentation notes, and Breuer's structural analysis (the doubled tsinnor as a substitution for the pair of big-revia subdividers the oleh-weyored rule would otherwise place) — is on the Psalms 17:14 double-tsinnor deep dive.