Showing posts with label lexicography. Show all posts
Showing posts with label lexicography. Show all posts

2011-03-16

Dictionary Updates: Kriaras, Vol. XVII; Trapp, Fasc. VII

New volumes of Kriaras' and Trapp's dictionaries of Greek are out. Kriaras covers Vernacular Early Modern Greek, and Trapp covers (mostly learnèd) Late Mediaeval Greek, with some overlap. For background on these dictionaries—and on the coverages of the dictionaries of Greek in general—see my earlier post on Dictionary coverage of Greek.

Trapp's Dictionary, Fascicle 7 of 8, runs from right after προσπελαγίζω to the end of sigma. Currently it is available only in electronic form, from the Austrian Academy of Sciences Press; a print edition is expected in a month or so. As the linked blurb notes, Trapp has not ruled vernacular texts out of scope (and is more inclusive of non-literary texts than Kriaras); it is addressing the delay in the completion of Kriaras by using already published word lists of Early Modern Greek, to give indicative coverage. Trapp is also including words from later papyri; while its brief is ostensively late mediaeval, it does look backwards as well, given the gaps in lexicography.

2011-03-31
Kriaras' Dictionary, Vol. 17 of an expected (20?) covers πνεύσις through προβίβασις. The volume is already available in online bookstores—I've just ordered it from Patakis. The announcement linked from the Greek Language Gateway (Πύλη για την ελληνική γλώσσα) notes that the volume will be made available online by them as well, as part of Emmanuel Kriaras' Collected Works. Vol 17 isn't up, but the others are, as scans (in the kind of page-by-page interface that makes me relieved to have paper copies :-) ). The abridged dictionary, covering the first 14 volumes, is also online in a searchable interface, through the Electronic Node (Ηλεκτρονικός Κόμβος) site—both from the Centre for the Greek Language.

2011-01-11

στήτη, a post-Homeric ghost word

I posted in November about Leo Allatius, who coined a new word in the Greek literary corpus through a misreading of Pindar—or rather, perpetuating a mediaeval misreading of Pindar. But with the transmission of Classical literature as haphazard as it was, Allatius was not the only writer to have come up with such creative misreadings.

In the sixth verse of the Iliad, Homer introduces the feud between Achilles and Agamemnon over Briseis:
ἐξ οὗ δὴ τὰ πρῶτα διαστήτην ἐρίσαντε
Ἀτρεΐδης τε ἄναξ ἀνδρῶν καὶ δῖος Ἀχιλλεύς.
from the time when first they parted in strife,
Atreus' son, king of men, and brilliant Achilles

διαστήτην ἐρίσαντε literally means "the two of them stood apart contending". But διαστήτην is a strange verb for those who post-date Homer: it is too archaic to be understood readily.

The verb is archaic enough to lack an augment—the prefix that obligatorily indicates past tense in Classical Greek: where Homer used δια-στήτην /dia-stɛ́ːtɛːn/, later Greek would expect δι-ε-στήτην /di-e-stɛ́ːtɛːn/. Later on still—by the end of Classical Greek—the dual number of the verb fell out of use; Greek by then would expect not the dual διεστήτην, but the plural διέστησαν.

So διαστήτην did not look familiar to speakers to later Greek as a verb. What it did look like, though, was a feminine accusative noun, since -ην is the first declension ending for that case. Since spaces were not normally marked between words in Ancient writing, it would be easy for ΔΙΑΣΤΗΤΗΝΕΡΙΣΑΝΤΕ to be read as διὰ στήτην ἐρίσαντε, "contending for an X".

Reading διαστήτην as διὰ στήτην means that you have unearthed a brand new noun in Homer, and now you need to come up with a meaning for it. It's not the first time readers of Homer were faced with such a challenge, and the way out was provided, as it often was, by context. Achilles and Agamemnon were contending for Briseis; since Briseis was a woman, and στήτην seems to be a feminine noun, it follows that Homer has the noun στήτη, meaning "woman".

Some Homerically clueless poet later on, in the same mindset as Allatius, followed along with this misconstrual of Homer: he used στήτη as a noun in his verse, to mean "woman". In fact he went a step further than Allatius: Allatius cited πεδαφρόνων just as he read it in his Pindar; but this poet was pretending to write in Doric, so he switched dialects in the inflection, and came up with the genitive στήτας, not στήτης.

To make things even worse, another poet, within the next century, used this ghost word στήτας again, in a poem clearly derived from the first poet's conceit. ("Of the writer nothing is known; he was obviously acquainted with the [first poem]".) Both poets were pedants, more concerned with crafting poesie concrete than using words anyone had heard of. Clearly Homeric learning had fallen unpardonably far.

The poets in question are Theocritus (in his poem shaped like a Pipe) and Dosiadas (in his poem shaped like an Altar). In the 3rd century BC. You can see their handiwork at theoi.com, reproducing the 1912 Loeb edition and translation.

In the 3rd century BC, of course, Homeric scholarship was just getting started: Aristophanes of Byzantium might have worked in Alexandria at the same time as Theocritus. And we can be smug about misreadings like διὰ στήτην now, but five centuries after the Iliad was written, scholars had to start somewhere trying to make sense of Homer's antiquated Greek. Those scholars' attempt to make sense of Homer led them to invent Western grammar. So we should cut Theocritus and Dosiadas some slack.

Here are the phantoms of στήτη in action:
LSJ, στήτα:
στήτα, ἁ, pseudo-Doric, = γυνή [woman], Theoc.Syrinx 14, Dosiad.Ara 1. (The form arose from a false reading of Il.1.6, διὰ στήτην ἐρίσαντε having quarrelled about a woman, cf. Eust.21.43, Sch.D.T. p.11 H.)
Dosiadas, Altar, 1
Εἱμάρσενός με στήτας
πόσις, μέροψ δίσαβος,
τεῦξ’
I am the work of the husband of a mannish-mantled quean, of a twice-young mortal

Note that in the 1912 Loeb, writing to a more literate audience, J M Edmonds could afford to translate obscure Greek into obscure English: quean, "1. a disreputable woman; specifically: prostitute; 2. chiefly Scottish: woman; especially: one that is young or unmarried".

Theocritus, Syrinx, 13-20:
ψυχὰν ᾇ, βροτοβάμων,
στήτας οἶστρε Σαέττας,
κλωποπάτωρ, ἀπάτωρ,
with which heartily well pleased, thou clay-treading gadfly of the Lydian quean [i.e. Omphale], at once thief-begotten and none-begotted
Scholia in Theocritum, Syrinx 14:
Without the scholia to Theocritus, we'd be even more lost than Theocritus was before Homer:

στήτας οἶστρε Σαέττας: τουτέστιν ὁ οἶστρον ἐμβαλὼν τῇ Λυδῇ γυναικί. φασὶ γάρ, ὅτι ἡ Ὀμφάλη ἡ Λυδὴ οἶστρον εἶχε περὶ τὸν Πᾶνα πολύν. τὸ δὲ στήτη ἡ γυνή, Σαέττης δὲ τῆς Λυδῆς.
That is, the gadfly poking the Lydian woman. For he is saying that Omphale, the Lydian, had a great gadfly (= sexual excitement) about Pan. And στήτη means "woman", while Σαέττη means "Lydian woman".

I think this is our only source form knowing what a Saetta was. It's not our only source for knowing what a στήτη is:

Scholia on Homer, Scholia Recentiora [more recent scholia] by Theodore Meletiniotes (codex Genevensis gr. 44), I 6:
[διαστήτην] διὰ τὴν στήτην, διὰ τὴν γυναῖκα.
διαστήτην: for the στήτην, for the woman.
Hesychius
στήτα· γυνή
στήτα: woman

Though Hesychius does Contain Multitudes:

Hesychius:
στήτην· ἔστησαν, δυϊκῶς
στήτην: "they stood", in the dual

Some scholars, at least, had worked out what had gone wrong—fully and astutely (Melampous or Diomedes' grammatical commentary), or almost but not quite (Eustathius of Thessalonica):

Eustathius of Thessalonica, Commentary on the Iliad, Vol. 1 p. 35 (van der Valk):
Ἰστέον δὲ [...] ὅτι περιέργως τινὲς ἐπιβαλόντες Θεοκρίτου στήτην τὴν γυναῖκα εἰπόντος γράφουσιν ἐνταῦθα «διὰ στήτην ἐρίσαντο», ἵνα λέγῃ ὁ ποιητής, ὡς διὰ γυναῖκα ἤρισαν. ὁ δὲ τούτοις προσέχων εἴη ἂν φιλόκαινος.
Note that oddly enough, some authorities, imposing Theocritus' στήτη "woman" here, read this as διὰ στήτην ἐρίσαντο, so that the Poet ends up saying that they "contended for a woman". To pay attention to such readings would be an infatuation with novelty.
Commentary on Dionysius Thrax's Art of Grammar, by Melampous or Diomedes, p. 11 (Hilgard, Grammatici Graeci vol. 1.3)
Ἕως ἐνταῦθά ἐστιν ὁ ὅρος τῆς γραμματικῆς. Εἴπωμεν οὖν αὐτόν· «γνῶσις τῶν παρὰ <τοῖς> τὰ ἔμμετρα καὶ ἄμετρα γράψασιν ὡς ἐπὶ τὸ πλεῖστον εὑρισκομένων». Διὰ τί δὲ εἶπεν «ὡς ἐπὶ τὸ πλεῖστον»; Ἐπειδή τινες λέξεις ἅπαξ που ἢ δὶς εἰρημέναι εἰσίν, ἃς οὐ πᾶσα ἀνάγκη εἰδέναι τὸν γραμματικόν, οἷον οἱ γρῖφοι. Τί δέ εἰσιν οἱ γρῖφοι; Τὰ ζητήματα τὰ δεινά· [...]
ἢ ὡς ἐν τῷ βωμῷ τοῦ Δοσιάδου ἡ γυνὴ εἴρηται στήτη, ἐπειδή τινες τὸ παρ’ Ὁμήρῳ διαστήτην ἐρίσαντε οὕτως ἐξηγήσαντο, διά τινα γυναῖκα. Σύριγξ δὲ καὶ βωμὸς ποιήματά τινά ἐστιν ἐμμέτρῳ τῷ σχήματι καὶ τῇ διατυπώσει τὴν ἐπωνυμίαν ἔχοντα. Τὰ οὖν τοιαῦτα ζητήματα εἰ μὲν ἐπίσταται ὁ γραμματικός, ἐπαινετέος ἐστίν, εἰ δὲ μή γε, οὐκ ἔστι μέμψεως ἄξιος.
This much is the definition of grammar. Let us add: "the knowledge of most things written in verse and prose". Why "most"? Because there are some words which have been used just once or twice, which it is not essential for the grammarian to know, such as riddles. And what are riddles? Difficult questions [...]
Or, as in Dosiadas' Altar, where στήτη is used for "woman", because some people interpreted Homer's διαστήτην ἐρίσαντε as referring to a certain woman. And the Pipe and the Altar are poems in verse form and named for their appearance. Now if a grammarian knows about such matters, he is to be praised; but if not, he does not merit condemnation.

...Read more

2010-11-13

Ghost words revived in Allatius

The canon (patchwork though it is) of Greek lexica that I described a while back has a fair representation of German scholarship: Lust Eynikel & Hauspie, Bauer Danker Ardnt Gingrich, Trapp.

The oddity is that German scholarship wasn't represented there for the Classical period. Yes, LSJ is a major work, and DGE is more comprehensive still (if only I would live to see it completed). But it's odd that the Germans didn't corner the market in Ancient Greek lexicography in the 19th century, when they ruled Classical philology.

As it turns out, Wilhelm Pape's Handwörterbuch der griechischen Sprache (revised by Max Sengebusch in 1880) is competitive with LSJ, with 99,000 headwords against LSJ's 116,000. Given that LSJ, published in 1940, covers an additional 60 years worth of papyri and inscriptions, Pape's dictionary has nothing to be ashamed of.

I've recently had occasional to compare it against the other lexica of Greek: I keep trying to fill in gaps in lexical coverage hither and thither—for reasons that should be obvious to the regular readership I've been neglecting. Of those 99k, 545 lemmata are absent from the other lexica, but still turn up in the TLG corpus. That's a lot of lemmata, given that dictionaries have been making a point of filling in each other's gaps (Trapp in particular).

There are a few reasons for Pape's 545 not to have been recorded anywhere else. Pape has not picked up the allergy to Late Antiquity that LSJ did. This allergy has meant that pagan late antiquity is the least well covered period in Greek lexicography. As more inscriptions and papyri turned up, LSJ displaced late entries with the newly found earlier entries. With these late entries attributed to no author more specific than "Eccl." or "Byz.", LSJ wasn't desperate to hold on to them, to begin with.

Pape has the same dismissive "Byz." attribution for its late entries; but it has not undergone the same kind of cull—to the benefit of the Greek Anthology and Oribasius.

And, to my surprise, of the scholiasts. The scholiasts are something of a discomfort to lexicographers. Scholiasts provide their own definitions of Classical words, so Classical lexicographers care about the words the scholiasts use. But scholiasts explain Classical words to post-Classical audiences, by using decidedly post-Classical vocabulary. Which means that documenting the scholiasts puts Classicists in the Mediaeval Greek business. (Not always wisely, as seen with LSJ's mishandling of στοίχημα "wager".)

There is a smaller group of lemmata in Pape which the other lexica overlook for different reasons—and which by rights should not be turning up in a Classical corpus at all. The text of the Classical canon is described by lexicographers following standard editions; but the manuscripts they work from had a lot more variability than the lexicographers now need to account for. Our modern edition of Pindar has established that Pythian Ode 8.74 reads
πολλοῖς σοφὸς δοκεῖ πεδ’ ἀφρόνων

to many he seems wise among fools

So we are not interested that in some mediaeval manuscripts, the last two words πεδ’ ἀφρόνων "with the non-prudent" were run together as πεδαφρόνων "of the after-prudent"—or, as Pape has it:
πεδά-φρων, ον formerly appeared in Pind. P. 8.74, and was explained as Aeolic for μετάφρων, someone who is wise later, after the deed; Böckh writes it as πέδ' ἀφρόνων.

Once Böckh settled that πεδάφρων did not exist in Pindar, πεδάφρων ceases to be of interest for Classicists.

But the people who read those mediaeval manuscripts thought πεδάφρων existed. The scholia are commentaries on mediaeval versions of the Classical texts, and reflect the mediaeval understandings of those texts. If the scholia describe these misreadings, they perpetuate them, even after we have cleaned up the source text (to our best judgement). So our Thucydides 8.91.3 now reads
ἦν δέ τι καὶ τοιοῦτον ἀπὸ τῶν τὴν κατηγορίαν ἐχόντων, καὶ οὐ πάνυ διαβολὴ μόνον τοῦ λόγου

This was no mere slander, there being really some such plan entertained by the accused.

Somewhere along the line διαβολὴ μόνον "mere slander" was metanalysed as διαβόλιμον ὄν "being liable to slander". The scholiast accordingly tries in the margin to make sense of the word they saw in the main text. Our main text is emended, but the margin still counts as text in the Greek corpus:
διαβόλιμον ὄν: ἤτοι καὶ οὐκ ἔχοντος τοῦ λόγου διαβολὴν (?), οὐδαμῆ εἶχεν ὁ λόγος διαβόλως (?)
being liable to slander: namely, not having slander in speech (?); the speech was in no way slanderous (?)

(You can thank Karl Hude, who edited the scholia, for the incredulous question marks.) Thomas Magister also recorded the word in his "Selection of Attic nouns and verbs"—although he was reluctant to go so far as to commend it:
Διαβόλιμον Θουκυδίδης λέγει τὸ διαβεβλημένον· καὶ οὐ πάνυ διαβόλιμον ὂν ἀπὸ τῶν Μεγάρων τὴν Σαλαμῖνα παραπλεῖν. σὺ δὲ διαβεβλημένον λέγε.
liable to slander is what Thucydides calls something slandered. "Without it being liable to slander that they should sail by from Megara to Salamis". But you should say "slandered".

You should also not trust Thomas' reading of Thucydides; the sailing from Megara to Salamis is from Thucydides 8.94.1, a few pages on, and the mission talked about as διαβόλιμον was heading to Eetonia near Piraeus.

The mediaeval manuscripts were not just read by the Byzantine scholiasts and lexicographers, though. The manuscripts—and the scholia and lexica explaining them—were how the Byzantines learned the Classical language, on which they modelled their own literary language. That means that διαβόλιμον is of interest in the study of later Greek. Sure, it's a ghost word, like dord, and it was never part of any spoken form of Greek.

But if the text of Pindar that Byzantine writers read said πεδάφρων, then writers would assume it to be real enough: no less real than μετάφρων, which they could have coined in their own, Atticist learnèd Greek. No less real, for that matter, than embiggen.

The main reason these ghost words were taken up was that the ghost words did make sense on their own: they could be analysed according to the rules of Greek compounding. If you know how Classical Greek works, you already know that διαβόλιμον really does mean "liable to calumny", because -ιμος is a real suffix (cf. ἀγώγιμος "carriable"); and once you know that πέδα is Aeolic for μετά, you know that πεδάφρων must somehow mean "after-mined" (cf. ἔκφρων "out of one's mind"). The misreading of διαβολὴ μόνον or πεδ’ ἀφρόνων is not alien to spoken language after all; it's the same metanalysis that gives us adder or newt—or helipad or pollute. So it may look like Byzantine Greek writers were asking for this kind of bogosity, by slavishly relying on flawed manuscripts for their models; but what they came up with was not that different to what spoken language has done elsewhere.

And so it is no surprise that πεδάφρων shows up in the TLG too, even if it is not in the TLG's Pindar. It turns up in someone who has read Pindar in those corrupted manuscripts, and is happy to show his reading off. But this is not a scholiast or lexicographer. This time, it's a reader who is producing literature of his own: Leo Allatius, in his poem Hellas, written in the 17th century. (A couple of centuries before Böckh.)

My Ancient Greek isn't as good as I might make it seem, and Allatius' run-on syntax doesn't make it easier; but this is how I think he uses πεδάφρων—in the same genitive πεδαφρόνων as in the pre-Böckh Pindar.
ΕΛΛΑΣ γάρ εἰμι, τέκνον, ΕΛΛΑΣ, ἧς κλέος
ἄσβεστον ἔργοις τιμίοις πεπραγμένον,
διελθὸν ἔπτη πρός τε γαῖαν, καὶ διὰ
πόντου πέρασσε νυκτός, ἠδ’ ἠοῦς λάχος,
ἠδ’ εἴ τι ποῦ ἐστι λαιά, κἀπιδέξια
συμφραδμόνεσσι τηλόθεν γ’ ᾠκισμένον,
πεδαφρόνων γὰρ οὐδὲ μικρόν μοι μέλει,
ὃ ΓΑΛΛΙΚΟΙΣ στήθεσφι τοῦδ’ ἄχρις χρόνου
ἄφθαρτον ἐμπεφυκὸς ἀγλαΐζεται. (vv. 196–204)
For I am Hellas, child, Hellas, whose inextinguishable
glory, carried out in honourable deeds,
is past, has crouched down to the ground, and has
passed through the sea of the night: this is the doom of the dawn;
but it is also what is there in a place, that is dwelled in
left and right by counsellors from afar,
(for I do not care the slightest for those wise after the fact):
a thing which is splendid in Gallic breasts
growing incorruptible to this day.


Pindar is not the only Classical author whose fractured words turn up in both Allatius and Pape. ναυβάτης "ship-walker" turns up in several Classical sources—Aeschylus, Sophocles, Euripides, Herodotus, Thucydides, Xenophon). The most obscure source is Lycophron:
καὶ τὰς Ἐρεμβῶν ναυβάταις ἠχθημένας προβλῆτας ἀκτάς (vv. 827–828)
and the jutting shores of the Erembi, abhorred by mariners

The TLG's edition of Lycophron dates from 1964; Pape in 1880 records ναυάτης as Lycophron's reading, and also has it as a variant reading in Euripides, IT 1380. Our Lycophron and Euripides have now been cleaned up to read ναυβάτης; Allatius' had not, so that his poem has the first appearance of ναυάτης in the TLG corpus:
Ἀργοῖς ἐρετμοῖς πόντος, εἰ καὶ μαίνεται,
κλύδωνί τ’ οἰχθείς, ψάμμον ἐκβράζει βυθοῖς,
βύκτας ἀέλλας ἡμερώσας ναυάτης,
οἰηΐοις τε νηὸς ἰθύνας δρόμον,
πρὸς ὅρμον ἔλσεν ἀσκηθὴς σόον σκάφος. (vv. 233–237)
Though the sea rages, and throws up sand from the depths,
departing with slow oars and through the billow,
the mariner calms the blustering whirlwind,
and straightens his course with the ship's rudder,
to shelter the vessel unscathed in its anchorage.

(If you go searching the TLG for ναυατ-, btw, ignore Ναυάται in Germanus I, Narratio de haeresibus et synodis ad Anthimum diaconum 48. His Nauatae predate Allatius by eight centuries, but they are Novatianists, a sect, as you would expect in an anti-heretical text.)

The misreading of ναυβάτης as ναυάτης is unsurprising: they sound the same in Modern Greek, [naˈvatis]. Unlike πεδάφρων or διαβόλιμον, ναυάτης doesn't mean anything on its own in Greek. It was taken up because it still looks like another real word of Greek, ναύτης "shipper = sailor", which has survived into the Modern language. Copyists simplified ναυβάτης into ναυάτης, and then figured that Euripides was just using ναύτης with an extra α in the middle. The Ancients did strange this like that, after all, didn't they?

(The Baroque did strange things too, just in a different way.)

Outside the lexicographers and scholiasts, there aren't many instances of such rejected readings turning up in literary Greek. ἠπίαμα "cure", as defined in Pape, is a metanalysis of Herodotus 3.130 ἤπια μετὰ τὰ ἰσχυρὰ προσάγων "supplying both mild and strong (medicine)"; the misquote turns up in Constantine Porphyrogenitus De virtutis et vitiis II p. 11, and Suda, delta 442; but that doesn't count as new usage. The run-in εὐ ναιόμενος "well-dwelling" of Iliad 14.255, which Pape allows as a single verb, turns up as a distinct verb not just in commentators, but also in Galen (Kühn 18b p. 763). But that's it. (I thought I saw an instance in Gregory of Nazianzen's aping of Homer, but I can't find it.)

That may mean that Allatius' Greek is more derivative than his forebears; I have my doubts, given how Byzantine literary culture worked. Since I did not track down words already defined in Lampe and Trapp, it's likelier that any earlier such repurposings had already been dealt with there.

… And so I return to blogging. Missed it.
...Read more

2010-06-15

A Turkish etymology for both α and σιχτίρ?

Language advisory

In the last obscenity-filled post on this blog, Pierre left a comment on α σιχτίρ "fuck off", which is derived from Turkish:
The Turkish is sıçdırmak ( ﺼﭽﺩﺭﻣﻕ ) with a chim, rather than a kha, and it gets "shit" right back into the context. Actually, it is a causative form and means "to make (someone? / yourself?) shit" and it appears to be imperative. My guess is that "ey sıçdır" is all Turkish and means "Go take a shit."

In fact, there really is a Turkish verb sikmek "to fuck" ("Turkic cognates include Azeri sikmək and Uzbek sikmoq"), and this discussion thread goes through the diffusion of siktir into the Balkans and Armenian.

What I did *not* know from the thread though, is that the interjection siktir in Turkish has a variant hassiktir. And that makes me look at the α in α σιχτίρ, and speculate whether it too is Turkish in origin—and not Ancient Greek, as is normally assumed.

α in α σιχτίρ or ά γαμήσου corresponds to get in get fucked or get lost. It's a hortative particle, and I think what is commonly assumed about its origin is wrong.

Modern Greek has four similar hortative particles.
  • άμε is a verb in origin; it's derived from Classical ἄγωμε "let's go", but in Modern Greek is used as "go!", as an exhortation. We know that because in Early Modern Greek, ἄγωμε was used to mean "go!" instead of "let's go": (ἄγωμε μὲ τὴν κάμηλον τὴν μακροσφονδυλάτην "go with the long-neckled camel!", Entertaining Tale of Quadrupeds 768)
  • άντε has been argued (tortuously) to be a verb in origin as well: ἄγε δή "go indeed!" The etymology makes no sense, and though the Triantaphyllides dictionary vacillates, and the other dictionaries are unrepentant, the obvious derivation is from the Turkish exclamation haydi. (I have 12 pages of an uncompleted paper arguing this; there are straightforward cognates throughout Turkic, and in Russian and Ukrainian via Tatar.) Notwithstanding its origin, άντε does act in some ways as a verb, including picking up a plural ending (άντεστε), and taking subjunctive complements (άντε να δεις "go to see, go and see")
  • The other two particles are άι ~ άει and α. They are assumed to be related; this is the entry from the Triantaphyllides dictionary on the pair:
    α2 & άι interj.: before exclamation (cf. άντε); depending on context and intontation expresses (a) indignation, annoyance, dismissal; "go, get": (with an imperative or να + subjunctive with a comparable meaning) α/άι πνίξου / παράτα μας / χάσου / να χαθείς "go drown! / go leave us alone! / get lost!" / || (with σε "to", article and noun) α/άι στην ευχή / στο διάβολο / στο καλό "to the blessing! / to hell! / to the good!". α/άι στη δουλειά σου "to your business": "mind your own business, continue what you are doing". || friendly reproach: ~ να χαθείς! "get lost!" (b) exhortation: Άι στο καλό, παιδί μου, και πρόσεχε "go to the good (farewell), my child, and be careful": go on, go to the good [= farewell]. || wondering about what will happen: α/άι να δούμε πώς θα τα βγάλουμε πέρα "let's see how we get through this". α/άι να δούμε τι θα γίνει "let's see what will happen". [α: truncation of άι· άι: άε < Ancient ἄγε "onwards!" (imperative of ἄγω "go") deleting intervocalic [ɣ] and diphthongised]

    (It's α2, to distinguish it from the interjection "ah!", α1.)

Now, unlike άμε and άντε, άι and α do not act like independent verbs. They are prefixed to imperatives, which means they are not verbs with a dependent verb: they are acting like interjections—or serial verbs. άμε can't precede an imperative; άντε can, because άμε is not an interjection and άντε is. (άι and α do precede the subjunctive as well, but the subjunctive is also used as a gentler command; so it's consistent with άι and α behaving as interjections.)
  • άμε/άντε/*άι/*α να δεις ποιος είναι "go to see who it is"
  • *άμε/άντε/άι/*α δες ποιος είναι "go, see who it is"

They also cannot appear on their own as verbs (or for that matter as interjections): they must always introduce something.
  • άμε/άντε/*άι/*α "go on!"

And άι/α has no flexibility with what prepositions it can take: it cannot take από "from", meaning "go past":
  • άμε/άντε/*άι/*α από την αγορά "go past the market"

The only prepositional phrase άι/α can take is σε "to" with a definite article:
  • άμε/άντε/άει/α στο διάολο "go to the devil" (to hell)

Surely that means άι/α is behaving like a verb here? Well, no. If it is a verb, why the constraint on having a definite article?
  • άμε/άντε/*άι/*α σε κανένα μπουντρούμι "go to a dungeon"

Now, it turns out that you can use σε phrases with a definite article on their own, as an oath, a blessing or curse, hoping that someone ends up there. You can't leave the article out if you do that: it won't be the same oath:
  • στο καλό!/στο διάολο! "to the good" (farewell), "to the devil"
  • *σε καλό!/*σε διάολο! "to good" (farewell), "to a devil"

These expressions of course imply "go!", but they have still settled into a template of requiring a definite article, and expressing wishes. άι/α is being prefixed to those established expressions: that does not mean it is acting as a real verb. You can't use άι/α before στο if an oath is *not* involved:
  • άμε/άντε/άι/α στο διάολο "go to the devil"
  • άμε/άντε/??άι/*α στο γιατρό "go to the doctor"

Which suggests that, whatever άντε was originally, it is now (also) a verb; and whatever άι used to be, it is now not behaving as a verb, but just an exclamation, a particle introducing verbs and oaths.

In fact, that's reason enough to suspect άι did not start out as a verb at all, but as an interjection—such as, say, I dunno, the Turkish interjection hay. The thing is though, there are several instances from the Historical Dictionary of Modern Greek (Modern dialect dictionary) of "missing link" particles in dialect, between the verb ἄγε and άι:
  • Leucas: άγε, Lesbos: άγι, Sikinos: έγε
  • Cyprus: άγι̮α
  • Thessaly, Cephallenia, Cyme, Leucas, Siphnos etc. άε
  • Maina: χάε

Still, multiple causation does happen, and the very similar exclamation hay! may have encouraged άγε > άι to be restricted to exclamation-like use.

But there's something interesting about the Triantaphyllidis definition. In all its examples but one, it uses άι and α interchangeably. All the examples I've given have been interchangeable as well, although I hesitated over ??άι/*α στο γιατρό "go to the doctor". The one time Triantaphyllidis does not use α but only άι, is in the following pair:
  • *άμε/*άντε/άι/α στο καλό "oh, to the good" (euphemism for "oh, to the devil" = "well I'll be! what the deuce!")
  • άμε/άντε/άι/*α στο καλό "go to the good" (= "farewell")

στο καλό is ambiguous between a literal blessing, and a euphemistic curse. άι can be used both the bless and to curse. α is not used to bless: only to curse. Curses such as can be found at slang.gr:
  • α σιχτίρ "get fucked!"
  • α γαμήσου "get fucked!"
  • α να χαθείς "get lost!"
  • α στο διάολο "go to hell!" > ασταδγιάλα ~ ασταδιάλα
  • α να σε γαμήσω "I'm gonna fuck you!" (not as foreplay, but as one man threatening another—hence the joke reply given there by Vrastaman, "I'd rather we just stay friends".)
    There's a further complication that the phrase is actually used as a deferred threat: "A peculiar expression said (usually twice) instead of 'I will fuck you' during stand-offs, but usually with a desire to avoid trouble with someone who is trouble anyway. Like 'count yourself lucky', but leaving everything open, especially if the other talks back." (Halikoutis) (Pritsapirdulas adds in comments: "We also say it (1) when something is broken and we can't fix it; (2) when we react to something startling us.") But a curse it still is.

  • α πάγαινε ~ α πάαινε ~ ρε α πάαινε "get lost", where Standard Greek has retained this dialectal imperative "be going" only in the context of this curse:
    ρε α πάαινε: an abbreviated form of the expression "pardon me sir/madam, could you possibly relocate yourselves a smidgeon in the opposite direction from me? I thank you in advance for your understanding." It is used to show beyond doubt that the speaker is not disposed to be serviceable towards his interlocutor, and that the discussion is probably coming to its definitive conclusion right about now:
    —Pardon me sir, could you please park a bit further on, so my car can fit in too?
    —Ρε α πάαινε, wanting to park, no less! As if you had a car back in your village, you FUCKING HILLBILLY! (acg)

So α is used consistently in contexts reminiscent of α σιχτίρ "get fucked!" I suspect now that Greeks heard both hassiktir and siktir, reanalysed the former as a sixtir, and related the a back to άι—but only in contexts to do with cursing, like the original hassiktir. This would have been helped along by the existence in Greek of α1, the interjection "ah!"
  • In the expression α γεια σου "ah your health!" ~ α μπράβο "ah bravo" = "that's more like it!", it's not immediately obvious which of the two α is involved; it could be either the pure exclamation α1 ("aha! that's more like it!"), or the hortative α2 (ά μπράβο = άντε μπράβο: "go on! that's more like it!") The restriction of α otherwise to curses makes me suspect the former, but I'm being schematic there.

...Read more

2010-03-26

RIP: Tassos Karanastassis

Tassos A. Karanastassis (Τάσος Καραναστάσης), lecturer at the University of Thessalonica seconded to the Centre for Byzantine Studies, passed away last week, entirely too young.

He finished up at the Centre for Byzantine Studies; but for much of his career, from 1980 to 2003, Tassos worked at the Dictionary of Mediaeval Greek Vernacular Literature. The dictionary was established by the now 103–year-old Emmanuel Kriaras in 1968; there is a good reason why the dictionary bears his name, and Kriaras continued to have the final word over all work up to 1997. But in his time there, Karanastassis had the day-to-day charge of the dictionary: he was its soul and its motor and its thrall.

His involvement with the dictionary ceased in 2003, earlier than he would have preferred; Karanastassis went back to his dissertation (which we hope to see published soon), and his research. The dictionary has gone on without him, and I wish it well; but it will not soon see someone with that degree of immersion, dedication, and easy familiarity with half a millenium's worth of words. A familiarity, I am told, that translated into 4,000 marginal notes to Trapp's overlapping Lexikon der byzantinischen Gräzität. I hope they too will eventually see the light of day.

I profited from his easy familiarity while George Baloglou and I were translating the Entertaining Tale of Quadrupeds: more than once during our inquiries he'd pause, look up, mutter "that can't be right", walk to the shelf, and fish out a novel sense of στασίδι from a six-hundred year old legal deed. ("Fishing spot", not "pew"—teaching me for the first time that Early Modern Greek was not as like to the language I speak as I'd assumed.) He had volumes of facts and contexts and connections filed away in his head, in a heterogeneous throng like the jars of pencil stubs and pottery shards he kept by his desk. And he always delighted to gather more: my unconscious lapses into Cretan phonology (δεκαρά), as much as mediaeval words for saucers (σαρσαρόλι).

That delight is what made him the dictionary's soul and its motor, as its thrall. He had objected to me saying as much about him on my web site, and he deferred to his forebear. (He was mortified when I guffawed at a paper I was reading on the premises in '96: "Shut up, the Professor is inside!") But Karanastassis has a large share of the responsibility for Early Modern Greek now having a dictionary. And his contribution has not been adequately acknowledged. His name does appear on the cover of the 2001 abridgment at least; but the whole dictionary is his, as much as it is anyone's. And that deserves to be said more.

George knew him longer than I did, and has put up his own reminiscences of him, more detailed than you'll find here. I'm glad to know Tassos stopped by this blog from time to time, even as his illness started to take its toll. I smile to read he had smirked to George about the same obscene citation from 1383 that I posted about last month. But by last month, Tassos was no longer able to stop by and smirk once more.

Others are better qualified than I to praise his literary scholarship—the encyclopaedic knowledge that gave him license to draw the long bow lines, and connect the unsurmised. The dozens of young scholars that served time at the dictionary know better than I how sound a mentor he was. The people of Kallikrateia, where his final resting place is, had more of a sense than I of what he was like outside the office and away from the pencil stubs.

For my part, I pause at an image, and at a put-down. Someone once groused at me, "Karanastassis thinks he's labouring sub specie aeternitatis." That wasn't intended as a compliment. But while I was gratified by Tassos' excitement when he solved a problem, or his good humour as he told a story entirely too long, the image that abides with me is Tassos labouring "under the aspect of eternity": pensive with his melancholy mustache and shock of grey hair, his blue eyes staring into the distance, sorting through words and facts and contexts, bringing them into deliberate, and unrushed, order.

He is eternity's now. Ελαφρύ το χώμα που τον σκεπάζει. Light be the earth that covers him.

Nick Nicholas and Tassos Karanastassis, 1996

NB: Tassos Karanastassis is not to be confused with his namesake Anastasios Karanastassis, another Greek lexicographer, who wrote the Academy of Athens' Dictionary of the Greek Dialects of Southern Italy.
...Read more

2010-02-15

New TLG words in DGE VII

As I posted last month, the new volume of DGE (Diccionario Griego–Español) has appeared, spanning ἐκπελλεύω–ἔξαυος. As with any lexicographic work of an older language, some philology and textual emendation has been involved; this paper by Eugenio Luján Martínez gives four such instances, in Epicurus, Aretaeus, Nicander, and Galen.

I have gone through this volume and the TLG texts dated from before i AD, to find words not given in the other dictionaries out there (LSJ, LSJ Supplement, Bauer, Lampe, Trapp). I'm posting them here for interest. Note that the list may be small, but that's because the major gaps are in papyri (which are not in the TLG)—not ancient literature, which is already (nominally) well-ploughed land. (There should be quite a few more words from TLG AD texts.) DGE's entries are still much more detailed than LSJ's, and its coverage of antiquity broader. I'm also omitting proper names, which I have been treating differently in lemmatisation.

I'm leaving out the glosses, because the DGE is a commercial product and all. Most of them you can guess the meaning of, if you know your way around Greek vocabulary...

  • ἐκπυρώδης, ες.
  • ἐλάφινος, η, ον. (But Trapp has the variant ἐλαφινός.)
  • ἐλέκεβρα, ας, ἡ. (ἠλεκέβρα, ἰλλεκέβρα, ἐκλεκέβρα).
  • ἐλέσσω.
  • ἑλωρεύς, έως, ὁ.
  • ἕμα, ματος, τό.
  • ἐμβρωσί, τό.
  • ἐμπιεστός, ή, όν.
  • ἐμπύρευσις, εως, ἡ.
  • ἔνδεινος, ον.
  • ἐνναγώνιον, ου, τό.
  • ἔνοψ, πος, ὁ.
  • ἐναντιοεργός, όν.
  • ἐνεδρευτός, ή, όν.
  • ἐνιζυγίς, ίδος, ἡ.
  • ἐντεροκοιλιακός, ή, όν.
  • ἐντραπής, ές.
  • ἐντυπάδιον, ου, τό.
  • ἐξάρακτος, ον.

...Read more

2010-02-13

Etymologies and attestation of μουνί

(See also μουνί vs. monín; μούτζα, μουνί and Tzetzes.)

OK, let's draw this talk of μουνίν to some sort of close. I'll present the first attestations of the word, as given in Trapp's and Kriaras' dictionary; and then I'll reproduce Moutsos' presentation of the various proposed etymologies, with a few of my comments.

The attestations are given with date of authorship, followed by date of earliest manuscript. The manuscript date matters because, as you'll have already seen from TAK's comments in previous threads, you can't trust mediaeval scribes not to interfere with the language of what they're copying.
John Tzetzes, Theogony (12th century/ca. 1400)
οὐκ αἰσχύνεσαι, αὐθέντριά μου, νὰ γαμῇ τὸ μουνίν σου παπᾶς;
Aren't you ashamed, my lady, to have a priest fuck your cunt?

The reading is preserved only in one manuscript, ca. 1400 ("turn of 14th century"). There is a possibility that the manuscript scribe has introduced this, pejorating whatever Tzetzes originally had written. The other manuscript preserving the passage, from the 15th century, clearly saw something worth censoring in the original, to have left the second half of the verse out. But given that it also censored the previous verse, mentioning a priest as lover, we can't be sure what it censored was a translation like "fuck your cunt", or something closer to the Proto-Ossetian's "have a love affair".

Entertaining Tale of Quadrupeds 467 (ca. 1364/15th century)
διὰ νὰ σηκώνῃς τὴν οὐράν, νὰ δείχνῃς τὸ μουνίν σου
(Sheep to Goat): You're here to lift your tail and show your cunt off!

The translation is due to one George Baloglou and one Nick Nicholas.

Miklosich & Müller, Acta et Diplomata Graeca Medii Aevi, Vol. II p. 53: Church synod condemnation of Constantine Cabasilas (24 August, 1383)
τέταρτον· ὅτι ἐβάπτιζε ποτὲ παιδίον, ἵστατο δὲ ἐκεῖσε καὶ γυνή· ἔχρισεν οὖν τὸ βρέφος τῷ ἀγίῳ μύρῳ, εἷτα λέγει πρὸς τὴν γυναῖκα· φέρε μοι τὸ μουνίν σου ἐνταῦθα, ἵνα χρίσω αὐτὸ, καὶ οὐ συγκάπτῃ
Fourthly: that he was christening a child once, and there was a woman standing there too; so he anointed the infant with holy myrrh, and then said to the woman, "bring me your cunt here, for me to anoint it; it won't swallow it up."

I don't quite understand what exactly συγκάπτῃ means here, and I'm not sure I want to. Yes, he was excommunicated. Legal processes, such as this, are usually boring, but at times can give invaluable linguistic evidence. Even though in this case the grammar has clearly been antiquated, the vocabulary has not.

Mass of the Beardless Man (ca. 1500/1515~1519) (Β 165, Α 499)
γραίας πορδὴ μαστίχα σου, γαδάρας μουνὶ πουγγί σου
an old woman's fart is your mastic, a donkey's cunt is your purse

The Mass is a relentlessly scatological parody; as the previous quote shows, not all churchmen were saintly, then any more than now. It was printed in 1553, but the two manuscript versions are if anything even filthier—as is shown here.

Mass of the Beardless Man (ca. 1500/1515~1519) (Α 375)
Μαγαρίζομέν σε, κὺρ Φασούλη σπανέ, καὶ ὑβρίζω τὴν ὡραίαν πατσάδαν σου ὡς ἀντίτυπον γαδάρας <τὸ> μουνίν.
We pollute thee, Sir Bean Beardless, and I curse thy pretty beard as a copy of a donkey's cunt.

Glossae Graecobarbarae (end of 15th century/1614) [cited in Meursius]
τῆς γυναικὸς τὸ αἰδοῖον, ὅπερ καλοῦσι μουνὴν.
The woman's pudendum, which they call μουνί.

The Glossae Graecobarbarae have survived only in citations by the lexicographers Meursius and DuCange; they've been claimed to originate in Cyprus, at the end of the 15th century. (Beaudouin, Mondry 1884. Dialecte chypriote. pp. 109-131 cites the glosses and compares them to Modern Cypriot.)

Stefano de Sabio, Corona Pretiosa (1527) [cited in Meursius]
μουνὴ. Cunnus. αἰδοῖον γυναικὸς.
μουνί. Cunt. "(Ancient Greek) woman's pudendum."

The Corona Pretiosa was published by Stefano de Sabio in 1527, with a reprint in 1543; it's a glossary that translates Modern Greek into Latin, Ancient Greek and (apparently) Italian. I'd like to register my astonishment that this is the first time I've heard of it. I'd also like to register my astonishment that nowadays I can *expect* a 1527 book to be digitised and online. It isn't, but all the other lexica are.

Johannes Meursius: Glossarium Graeco-barbarum (1614)
Μουνή. Membrum muliebre. [cites definitions from Corona Pretiosa and Glossae Graecobarbarae]
Μουνί: Female member.

Little gripe to Kriaras' dictionary: if they're going to cite words as being cited in Meursius, DuCange, Vlachos and Somavera, they really should also have mentioned De Sabio and the Glossae: they push the date back a lot.

Alessio da Somavera (Alexis de Sommevoire), Tesoro della lingua greca-volgare ed italiana (1709)
Μοῦνα, ἡ. μαϊμοῦ μὲ τῆν οὐράν, ἡ. (ζῶον) Mona, gatto fariano. (animale) // Μουνάρα, ἡ. Natura granda di donna. // Μουνί. βλ. Σάρκα. // ἡ Σάρκα. Le parti vergognose, honestamente parlando.
Μοῦνα, ἡ. Monkey with a tail (animal). // Μουνάρα, ἡ. Large feminine organ. // Μουνί. See Σάρκα "flesh". // Σάρκα. The embarrassing parts, to speak bluntly.

Somavera is more hesitant than previous lexicographers, but he does note the augmentative μουνάρα. He also shows that the Venetian monna "monkey", which I mentioned confused matters in Venetian, had also entered Greek at the time.


Now to the etymologies. None of them are straightforward phonologically: there is no obvious Ancient word starting in /mon/ or /mun/, which could account for it.
A friend of mine said he always assumed it was derived from μόνος "only, unique" (as in mono- in English), because there's only one of them. That's in contrast to testicles, presumably, but it doesn't exactly distinguish vaginas from penises though.

No, I'm not going further with that proposal.

Faced with this difficulty, Hatzidakis arrived at the ingenious (too ingenious) parallel of /evnuxos/ "eunuch" > /munuxos/. Each of the steps posited for that transition has precedent in Greek:
  • eunúkʰos
  • evˈnuxos, through regular phonetic change
  • *ˈvnuxos, through aphaeresis
  • ˈmnuxos, through assimilation
  • muˈnuxos, through epenthesis

This allowed people to look for etymologies of μουνίν in something like *βνίν. I admit to some residual scepticism; as I said, the epenthetic /u/ in /munuxos/ could be copying the latter /u/, which wouldn't apply to /vnin/; and neither /mn/ nor /vn/ is always broken up in Greek: /mnimori/ "memorial stone", /keravnos/ "thunder". So if we could find a less awkward etymology for μουνίν, we'd use it.

Like, say, Venetian mona, as I had at first leapt at. But as I've argued, the evidence from Italiot Greek is that the Venetian word probably does have a Greek origin after all, so it doesn't help get rid of the problem.

So, let's see who Moutsous reports has had a go. I've already mentioned the later etymologies:
DuCange (1688): βουνή, from βουνός "hill, mound". As in mons Veneris
Moutsos cites Psichari and Rohlfs as rejecting it, and there's no good reason for /vun/ to go to /mun/.
Koraes (1835): μύλλον "lip"; cf. μυλλός "cake shaped like a vagina" (Athenaeus 14.647a), and μυλλάς "prostitute".
μυλλός and μυλλάς are derived from μύλλω "to fuck"; μύλλον apparently is unrelated, and there's no obvious reason for /myl/ to go to /mun/, either.
Hatzidakis (1892): εὐνή > *εὐνίον "bed" > *βνίν.
It seems a bit stretched, although I did point out the parallel in Modern Greek with carriola "cradle with wheels" > καριόλα "bed" > "whore". Hatzidakis knew it was stretched too, and didn't want to rule out mona
Filintas (1934): μνοῦς "down" > *μνίον
Moutsos dismisses this; the proposal "drew no attention as being entirely hypothetical". I'm not as sure: the form is semantically possible, and phononologically less indirect: we need only posit *mnin and not *vnin. The word did stick around long enough to show up in the Graeco-Latin glossaries as a gloss of pluma "feather", both as μνοῦς and as the vernacular diminutive μνούδιον.


Moutsos' proposal is that this is a nominalised infinitive of βινεῖν "to fuck". We have several such fossil infintives in Modern Greek: φαγεῖν "to eat" > φαγί > φαΐ "food", πιεῖν "to drink" > πιεί "drink" (dialectal); φιλεῖν "to love" > φιλί "kiss". Moutsos adds γαμήσειν "to get married > to fuck" > γαμήσι "fucking", parallel to λύσειν "to untie" > λύσι "untying" (dialectal); that I'm not as convinced of.
  • On a scale of more to less phonological plausibility—intermediate steps postulated: we have (1) μνοῦς > *μνίον (2 steps); (2) βινεῖν (3 steps); (3) εὐνή > *εὐνίον (4 steps).
  • On a scale of more to less semantic plausibility, we have (1) βινεῖν; (2) μνοῦς > *μνίον; (3) εὐνή > *εὐνίον.
  • On a scale of morphological plausibility, we have (1) βινεῖν; (2) μνοῦς > *μνίον; (3) εὐνή > *εὐνίον. βινεῖν uses a nominalised infinitive, which is attested as a process, but rare. The dimunitives μνίον and εὐνίον are both unattested, and -ίον did stop being a productive suffix sometime in Early Middle Greek. At least μνούδιον shows the word stuck around in the vernacular for a while (the Graeco-Latin glossaries admit colloquial words); I see no evidence that εὐνή made it to the Koine.

On balance, all three have problems, but "bed" has the most problems; and I guess "fuck" has the least (though not by as much as Moutsos thinks).

The other evidence that Moutsos gives is:
  • βινέω seems to have survived into Proto-Pontic: βιντώ "to be in a rut" (> βινητιῶ), βίντος "gadfly".
  • Circumstantial influence of a nominalised τὸ βινεῖν in Alexis, cited in Plutarch:
    τὰς ἡδονὰς δεῖ συλλέγειν τὸν σώφρονα.
    τρεῖς δ’ εἰσὶν αἵ γε τὴν δύναμιν κεκτημέναι
    τὴν ὡς ἀληθῶς συντελοῦσαν τῷ βίῳ,
    τὸ φαγεῖν τὸ πιεῖν τὸ τῆς Ἀφροδίτης τυγχάνειν·
    τὰ δ’ ἄλλα προσθήκας ἅπαντα χρὴ καλεῖν
    The wise man knows what of all things is best,
    Whilst choosing pleasure he slights all the rest.
    He thinks life’s joys complete in these three sorts,
    To drink and eat, and follow wanton sports;
    And what besides seems to pretend to pleasure,
    If it betide him, counts it over measure (Alexis fr. 271 Kock; Plutarch: Moralia 21e)

    Bless those old-school translators for using verse. Damn clever: "to be lucky with Venus" (Ἀφροδίτης τυγχάνειν) as a euphemism for βινεῖν, which happens to rhyme with φαγεῖν and πιεῖν "to eat and drink"—two infinitives that happen to have survived as nouns in Modern Greek. I don't think this is overwhelming evidence that τὸ βινεῖν was a commonplace colloquial expression, let alone an expression that turned into "cunt"; but it's cute anyway.
  • μουνίν and associated compounds are common in Early Modern Greek. True, but they're mostly in the Mass of the Beardless Man (μουνιοτζακάτος, καβουριομουνομέτωπος, σκατόμουνος), which is reasonably late, and doesn't prove anything about etymology.

Moutsos also derives Italiot munno (and Erice Sicilian monnu) from *μοῦνος, explained as an augmentative of μουνί attested in Modern Greek. The augmentative I know is the one Somavera recorded, μουνάρα; but slang.gr confirms the existence of μούνος, adding the improvised proverb κάλλιο μούνος και στο χέρι παρά κώλος και καρτέρι, "A bush in the hand is worth two arses in the bush". I think.) Now, it's true that switching genders can act as an augmentative, although that's because diminutives old and new are neuter, so this can be viewed as a back-formation. But a masculine could also be just an archaism.

To explain: Ancient masculine ποῦς, ποδός "foot" survives in Modern Greek through the diminutive neuter πόδιον > πόδι. The masculine πόδας is also found, and is how ποῦς, ποδός would have developed on its own (switching third to first declension). πόδας is normally interpreted as an augmentative: if people remember that a masculine was bigger than a neuter, and the neuter is now the normal term, then the masculine must be an augmentative of the neuter.

But πόδας is (I think!) the normal Cretan term for "foot", which suggests its an independent survival, unaffected by πόδιον. And that's a problem with munno. μουνίν looks, at first glance, like a diminutive of *μοῦνος. We know of no Ancient noun like μοῦνος. (We do have the Ionic μοῦνος "only, unique", corresponding to μόνος in the rest of Greek; but we've already rejected that track.) It's because we don't have an Ancient *μοῦνος that we've ended up looking at *vnin forms.

But Bova is evidence that there was a *μοῦνος form at some stage after all. And Moutsos' proposal has no room for a *μοῦνος. (Neither does Hatzidakis'; Filintas' does only with accent shift.) Bova could be doing the same gender-switch as Modern Greek, forming a local augmentative. But it really does look more like an original unattested Ancient form, of which μουνίν is the derived form.

But until someone comes up with an Ancient form that can explain *μοῦνος, Moutsos' βινεῖν is the best proposal on the table. Not overwhelmingly good or unproblematic; but that's the thing with etymology. And scholarship. Sometimes, we only have weak hypotheses. As the great Greek humourist (and amateur student of mythology) Nikos Tsiforos once put it, "Scholars argue when they don't know what's going on. When they do know what's going on, they just say '1 + 1 = 2', and they're done."
...Read more

2010-01-19

DGE Vol VII

Volume VII of the Diccionario Griego-Español (ἐκπελλεύω–ἔξαυος), intended to be the most comprehensive dictionary of Ancient and Early Middle Greek, has been published in 2009, and is available for purchase. (I've just ordered it.)

I found the new volume by googling; you'd be none the wiser about that from the DGE's own web page, which hasn't been updated since the second edition of Volume I. I've already commented on the melancholy slow pace of the DGE (begun in 1980) elsewhere, and that redoing Vol I was a really really inefficient use of time. But the appearance of VII is to be welcomed—the more so since VI appeared in 2003. I've just harvested from DGE five lemmata missing in Lampe for Gregory of Nyssa; the DGE continues to fill in gaps the others don't. (Gregory of Nazianzen wasn't as lucky though.)

2009-10-07

The Motley Word

I continue the random miscellanea postings with a website I did not know about, and stumbled on because of a posting I will write next week. The Motley Word (Παρδαλή Λέξη) is a crowdsourced dictionary for Greek dialects, like Urban Dictionary and its Greek counterpart, slang.gr

The Motley Word has all the poor quality you'd expect of a crowdsourced project without critical mass of participants; they don't even provide for correction yet. So it's not going to supplant the Academy's dialect dictionary, the Historical Dictionary Of Modern Greek, any time soon. Sure.

But that's harder to say when the Historical Dictionary hasn't budged past delta in twenty years. (I see there's at least funding to digitise their holdings.) And the glossaries of mainstream dialects, as opposed to the more distinct variants like Pontic or Tsakonian, have been pretty motley themselves.

More meta-importantly, the Motley Word shows that speakers of the dialects still care, and it's still a resource. A resource I'm going to make use of soon—here's a hint, but I haven't started writing yet, so shhhhh...

So: excellent work! (And thank God noone's done Tsakonian :-)

2009-08-12

Old Man Hare

[EDIT: followup post]

As I already mentioned in the past, the occasional Early Modern Greek word ends up in LSJ, because it has been used in a scholion to explain an Ancient word, and LSJ figured they'll take all the help they can get.

Such a word is λαγόγηρως. Literally, it's "Old Man Hare". Actually, literally, it's "Hare Old Man", but that just wouldn't work in English. As recorded in LSJ, it's used in the scholia to Lucian Dream 24 to gloss μυγαλῆ "field-mouse". You won't find it in the 1906 Rabe edition of the Scholia to Lucian, which the TLG has: it's "ap Bast. Ep. Crit. p. 169". This is an instance of Classicists' infuriating habit of using abbreviations without explaining them anywhere (and I checked). After some googling, I worked out it means that the gloss is mentioned in Friedrich Bast's 1805 Lettre critique de F.-J. Bast à M. J.-F. Boissonade, sur Antoninus Liberalis, Parthenius et Aristénète. (So Rabe's edition does not contain every single piece of Byzantine commentary ever authored on Lucian? Grrr.)

The word λαγόγηρως is also used in Suda, the mixmatched 10th century encyclopaedia, to gloss μύξος—although that doesn't help us much, because the only thing we know about a μύξος is that Suda says it's a λαγόγηρως.

So it's likely a Modern Greek word, and given the scholion to Lucian, it's likely a field-mouse, or some other rodent of that ilk. Old Man Hare shows up in other dictionaries too, but it does not show up in the big contemporary dictionaries of Greek; so other lexicographers are on their own. Trapp shrugs and says it's just "an animal". (He does at least record the more modern-looking variant λαγόγερος.) Kriaras can afford to go further, particularly given where it has been compiled (I'll explain in a sec): it reports that a λαγόγερος is "a kind of rat, a μυγαλῆ", and the passage it cites in response is an Early Modern falconry manual, in which the Old Man Hares are seized by birds of prey.

So those aren't Human Old Men being seized, and while they could be hares, there's no reason to think the Lucian scholion is wrong: it's some sort of rodent.

It's also a μύξος, whatever on earth that is, and as I was perusing recent additions to Suda On Line, I noted the newly translated entry on μύξος, expressing some puzzlement about glossing an unknown word with another unknown word.

At this stage, I had not checked LSJ, and I had not checked Trapp (which would have told me nothing anyway), and I certainly had not checked Kriaras. Instead I noticed that the two unknown words were not the same flavour of unknown. Suda didn't just say a μύξος is an Old Man Hare: it said a μύξος is an Old Man Hare παρ’ ἡμῖν. That παρ’ ἡμῖν means "with us"; and in Suda's way of structuring definitions, it means "in our language". As in, our vernacular, not Ancient Greek.

So without looking at Kriaras, I realised this word was at least Early Modern Greek, and quite likely Modern Modern Greek. I popped across to the online Triantafyllidis dictionary, and didn't find Old Man Hare there: so it's not a word that's made it to the Contemporary Standard. But figuring that there are always surprises to be had on the Interwebs, I googled λαγόγερος just in case.

I found Old Man Hare in a Greek digital photography forum. Like me, the photographer was an urbanite who wouldn't know a field-mouse from a dormouse (which was the Suda translator's first surmise). I mean... I don't know: *are* they the same thing? But the chap took the pic, recorded the place where and when the photo was taken—midday, near Edessa, in Greek Macedonia; and added what the locals call the beastie. Ladies and Gentlemen, courtesy of poster "Junior", meet Old Man Hare:

He exists, and he certainly looks like what Lucian's scholiast had in mind—and what the falconry writer established was an appropriate afternoon snack for an eagle.

And of course, some Greek dialectologist somewhere has recorded the fact that in Edessa (and probably elsewhere) this beastie is called an Old Man Hare. Because the Modern Greek dialect dictionary is still stuck at delta, it's not straightforward to find out who—although at least a draft of the remaining letters is now prepared. But Kriaras' dictionary staff have a fair collection of dialect glossaries on site, so they would have had the wherewithal to figure it out. And even if they didn't, Edessa is just a 94 km drive away from downtown Salonica. You'll probably run into Old Man Hare before you get to the waterfalls. (That's why the Other Language's name for Edessa is Vodena, "waters".)

Three other hits of note for Old Man Hare online. One was the Suda On Line entry, cached when it wasn't yet translated. One was from another Greek forum, this time ecological, recording Old Man Hare as one of the animals you might be surprised to find in the vicinity of Thessalonica. The Byzantine pronunciation /laɣoɣiros/ seems to survive for our rodent friend, because one posting later, courtesy of poster Kostas Karpadakis, Old Man Hare shows up again, this time spelled as λαγόγυρος, "Hare Roundabout":

I still can't tell you whether it's dormouse or field-mouse. Or hamsteroid. I showed the three links to Nikos Sarantakos (he whose Magnificent Blog I keep extolling), and he said that even before he got to the forum mention, he'd worked out this must have been a χαμστεροειδές. That kind of nonce macaronic coinage—American stem, archaic suffix—is pretty damned funny if you're steeped in the angst of Greek language history.

*I* can't tell you, but Karpadakis did not take the photo himself: he linked to zoology site in Novi Sad Uni, and the photo is labelled s-citellus02.jpg. A citellus is none of the above: as the site itself says, this is a Spermophilus citellus, which in English is the European ground squirrel, aka European Souslik.

And when I google his "historically wrong" spelling λαγόγυρος, I get not three or four hits, but 321. Confirming it as the Spermophilus citellus or Citellus citellus. Some more links: citellus #1, citellus #2, citellus #3. And a table of beastie names at Nature Names for Tourists: Local names for distinctive European mountain wildlife:
EnglishPolishSlovakSloveneRomanianBulgarianGreekscientific
European ground squirrel or sousliksuseł moręgowany syseľ pasienkovýtekunicapopândăulлалугер λαγόγυρος; σπερμόφιλος Spermophilus citellus; also Citellus citellus

So the Bulgarians got their word for the beastie from the Greeks.

In fact, that /laluɡer/ in the Bulgarian may suggest the modern spelling λαγόγυρος, which in Byzantine Greek would have been /laˈɣoɣyros/, is more accurate, and the Old Man bit was a written correction... Nah, the /u/ is in the wrong spot. Some Bulgarian phonological process, I guess: the Greek Macedonian pronunciation would be /laˈɣoɣirus/, which doesn't explain лалугер. So I'll stick with the assumption for now that this was originally Old Man Hare reanalysed as Hare Roundabout, until I hear something to the contrary. And that the accepted modern spelling is λαγόγυρος, although that form isn't in Triantafyllidis' dictionary either.

The third link for λαγόγερος, Sarantakos dismissed as "an incredible concoction", and I'm not disagreeing. The article was by an Italian classicist, and it was dealing with the Suda entry on μύξος as a whole, which goes into bizarre beliefs about donkey urine. The identity of Old Man Hare comes up just before the conclusion.

I've never studied Italian, but what with Esperanto, Latin, French, and a fair exposure to Classical Music in my youth, I can sort of read it. It's helped working in a French & Italian department, and not being as embarrassed about speaking in Super Mario Bros. Italian as I am about speaking in Pepe Le Pew French. I could even make the minimal effort of dealing with the mojibake of the page, such as by, I dunno, switching my browser encoding to Latin-1.

But once I'd worked out the article's claim, I had no motivation to proceed further. A dictionary I had not checked was Sophocles', and the author accepted and elaborated on Sophocles' surmise that an Old Man Hare was a kind of fish. Because λαγώς was also a word for sea-hare.

Wee tim'rous beastie, you're not a sea-slug of the Aplysiomorpha clade perchance, are you?

No, I didn't think so.

That is ungracious and horrid of me. I had the benefit of late Aughties Interwebs, this guy... well, this guy was writing in 2006, so he did as well. Hm. I had access to Modern Greek (even if it was in digital photography forums, detouring via Novi Sad Uni's Zoology department), this guy likely didn't. Still, the guy was not writing after the compilation of LSJ, and even if Suda is obscure about Old Man Hare, the scholiast to Lucian is not. Sophocles at least offered a definition for λαγόγηρως, but Sophocles *is* dated, and you can do better than that in general. The falconry manual may well have been inaccessible; but Suda has a whopping big "in our vernacular" in there, he could at least have asked some Greek contacts.

The blunder here is violating the Common Sense precept of etymology: If you're looking for Latin etymologies, you start your search on the Tiber. I tracked down the origins of that saying in my personal blog just before—taking a break from obsessing about constructions of Franco-Canadian identity; and it sounds a hell of a lot better in the original German. I'm always a little miffed when Classicists get as far as Byzantine Greek, and don't notice the live language spoken on the other side. (Witness LSJ's definition of στοίχημα: it's "wager", not "deposit", and they got "deposit" by a casual reading of Eustathius' scholion.)

But now at least, we have photographic refutation. In Edessa, you search for etymologies on the river Voda.

A sentence I smirk at, I confess, given where you search for the etymology of the river Voda. Now unsurprisingly renamed Edesseos.
...Read more

2009-07-21

Lerna VIIc: Variants

The various counts of lemmata that I've been putting out for the last while have made little mention of the difficulty in deciding whether two forms belong to variants of the same lemma, or distinct lemmata. The judgement call is difficult enough within a homogeneous language, with slight variations in derivational morphology. It's even worse with a large linguistic span like we've been dealing with, with lots of dialectal variation, phonological change across time, and spelling mutability.

So confronted with two similar nominative singulars in the vocabulary, or two 1st person present indicatives, you need to decide whether you'll count them as the same lexeme or not. And how liberal you are in your counting will decide how many lemmata you count, and how many you dismiss as variants.

To illustrate with English: you won't count color and colour as distinct lemmata, nor publicise and publicize. You won't count recieve as a distinct lemma from receive: misspellings happen, and you still need to count those misspellings as something. Though imbed is not a misspelling of embed, but a different derivation, you'd still want to conflate them too as the one headword.

OTOH, some people draw a distinction between racist and racialist. You may not, but if enough people do, you have to treat them as distinct. (The fact that the OED has decided they're now the same thing does not mean the entire language community has.) That's a judgement call in itself: people are uncomfortable with morphological variation as much as with any other kind, and people swore to the OED that there's a meaning distinction between gray and grey too. So there's no right answers. But there are arbitrary decisions.

Dictionaries normally make those arbitrary decisions for you, because they cross-reference variants to main headwords, and you can rely on their judgement. But dictionaries' judgement can be as arbitrary as any other's: the distinctions can be fine—particularly with variant suffixes, as we'll see. And some dictionaries conflate variants more effectively than others. Kriaras has a lot of phonetic variability, and hides a lot more spelling variability, because of its normalised modern spelling. But it does a reasonable job of indicating which forms are just variants.

By contrast, LSJ's more discursive entries include derived lemmata and similar lemmata, as well as simple variants, in the one entry. So it's harder for automated processing of dictionaries, such as the TLG lemmatiser does, to trust that two words cited in the one entry are in fact the same lexeme. Because the automated processing errs on the side of caution, it will distinguish variants as lemmata more than it should, which inflates the lemma count. The TLG lemmatiser currently seens two or more lemmata in some 12,000 LSJ entries; eyeballing, I'd say maybe a fifth of those could arguably be conflated. There are also a number of lemmata in later dictionaries (Trapp and Kriaras most notably) which could also arguably be conflated with lemmata in LSJ—but haven't been yet, because I haven't been through all 70,000-odd of their lemmata manually. (I'm not bold enough to guess how many.)

So there is a margin of overcounting lemmata; OTOH, there are many more instances where one could argue lemmata are undercounted, because the lemmatiser conflates variants. *I* won't argue that, because I've been responsible for a lot of the conflating. But this has been driven by a particular take on the vocabulary of Greek: if you're searching for words through a search engine, you're likelier to care about meaning than inflection, and you'll want the search to retrieve any word instance that looks enough like your word to match. You won't want to skip matches just because the spelling is slightly off. So for that purpose, less lemmata, meaning more search hits per lemma, is a good thing.

This has meant that, where there was doubt about whether a form is a distinct lemma or not, e.g. in LSJ cross-references, or variation between LSJ and Trapp, I've usually conflated those forms that I've been through manually. That's not everyone's purpose. If you're trying to inflate the count of distinct words of Greek, it's definitely not your purpose. Even if that isn't what you're trying to do, there will be disagreements on how much to conflate.

Part of those disagreements are tied to the dead tree: paper dictionaries have been understandably more reluctant to conflate variants that are alphabetically distant from each other. But there have been times that the choice to conflate hasn't been clear; and times when I've decided not to conflate. (The TLG search engine does display cross-referencing to the user, so that decision is not fatal.)

So what sorts of things may or may not be conflated as the same lemma in Greek? This laundry list will cover at least some of it:
Dialectal phonological differences
For some of the ancient dialects some of the time, these differences are so predictable, they're not even mentioned in LSJ entries. The big example is the different treatment of Proto-Greek */aː/ in Doric and Aeolic (stays α), in Ionic (goes to η), and in Attic (η except after ε, ι, ρ). The second example is Aeolic being accented as far back in the word as humanly possible, something which in the rest of Greek regularly happens for just verbs and compounds. There's more, and together they all mean that you shouldn't count Attic βοηθέω, Ionic βωθέω, and Doric βοαθοέω as different verbs for "help". Nor for that matter Aeolic βαθόημι: the different inflection is normal for Aeolic, and α is a plausible fate for /oaː/. (Just like the Ionic ω is.)
Classical spelling variations
If the odd inscription spells βοηθέω as βοιηθέω, that still isn't reason enough to call it a different lemma. Some spelling variation is endemic to the Classical language, because the pronunciation of those words was in flux—as occurs at any stage of any language. Usually, dictionaries will sweep this spelling variation under the carpet too, in generic cross-references like "κρεω-: see κρεο-". Once LSJ has worked out that λιπ- was the original pronunciation of compounds meaning "lacking in", every compound that can be spelled with λειπ- is listed under λιπ- .
... With the inconvenient exception of λειπογνώμων "without an inspector (or tariff, or distinguishing mark)", for which no ancient authority gives a λιπ- spelling, because the pronunciation was already changing. LSJ does dismiss the mediaeval spellings of the word as λιπογνώμων; but really, who could blame the mediaevals?
Late spelling variations
λιπογνώμων is very, *very* far from the only instance when mediaeval or even Hellenistic writers could not cope with the increasingly historical spelling of Greek, and spelled words more creatively than LSJ allows. LSJ is an historical dictionary, so its business is to go back to the original phonology of the word where possible, and dismiss everything outside its ambit. But if the manuscripts and papyri spell the word differently, that does not mean it's a different word. I've spent several entertaining months helping the lemmatiser cope with λλ as an alternate spelling of λ, or ι as an alternate spelling of ει, or σζ as an... interesting spelling of σ.
Hypotheticals
In making sense of the derivation of words, grammarians would often suggest what they thought the real underlying form of a word was. For example, in discussing ἀζηχής "continuous", they would suggest the form is underlyingly ἀδιηχής "unseparated" or ἀδιεχθής "unhostile". I have on occasion conflated these hypothetical derivations with the word they're explaining, when there was not anything else to be done with them.
Hypercorrections
With written Greek increasingly remote from the spoken language, Byzantine writers did increasingly oddball things to make their words sound high-falutin'. Theodore Metochites' Homer-Through-The-Looking-Glass is one egregious instance, the result of too much Classical learning ("if Homer said νοῦσος for νόσος, then I'll say σουφία for σοφία"). Theodore Studites' is another, and the result of not enough Classical learning: I'll never quite get over him working out that the imperfect of περισσεύω is περι-έσσευε. Throughout the period, there's a persistent tendency to stress words on the "wrong" syllable. I assume the authorities don't discuss it because it was beneath their notice; and I assume the Byzantines did it Because They Could.
This means that Byzantines have made up a bunch of variants for lemmata which did not exist, quite artificially. However artificial they are, they turn up in the corpus, so they need to be accounted for; but they shouldn't be counted as new words. They're old words in clown outfits.
Language change
But Classical words don't only turn up in the corpus in either chlamys or clown outfits. They also turn up in denim. Hm, I don't know if that analogy is going to work...
Variation in words is also going to result from natural language change. (In fact, differentiating natural and artificial language change is harder than it looks.) As we saw in earlier episodes, the *anr- stem of Proto-Greek, "man", turns up as ἀνήρ, ἀνδρός in Classical Greek, as ἄνδρας, ἄνδρα in Byzantine Greek, and as ἄντρας, ἄντρα in Modern Greek. The third declension of the original has been done away with, and the Ancient /ndr/ cluster is now spelled differently, leaving the eye-pronunciation /nðr/ to learned forms.
For the purposes that the TLG lemmatiser is put to—searching words in a diachronic corpus—that's not enough reason to call them different lemmata. ἄντρας is the natural development of ἀνήρ, it means the same thing, so they count as the same thing. But morphologically ἄντρας has already moved on from ἀνήρ, and counting them as the same is a diachronic artifice. It's only because we're trying to cover three thousand years in the one corpus, that we're conflating words three thousand years apart.
Variations in compounding
Putting stem A and stem B together forms you a new compound AB. That compound should for the most part count as the same word, regardless of slight differences in how the compound is put together. So ὑδατοφόρος "water-bearing" does not show up in the dictionaries, but ὑδροφόρος "water-bearing" does; they should count as variants of the same lemma, since hydro and hydato are allomorphs of the same noun, /húdɔːr/ gen. /húdatos/ "water". This, already, is a conflation lexicographers will not be equally eager to embrace.
Similarly, Classical Greek has δαφνηφόρος "carrying laurels" (as a ceremonial act), using -η- as a combining vowel; Late Greek no longer used -η- as a combining vowel, so the word turns up there as the more regular δαφνοφόρος. I'm disinclined to call these distinct lemmata. Others may not be.
Variations in inflection
Here things really do start getting murky: if the inflection of two forms is slightly different, though their stem is the same, should they count as the same lemma? I have allowed them to some times, especially if there is a distinction between an earlier and a later, more transparent or commonplace form. So when Theodore Studites, with his shaky command of classical morphology, uses εὐέλπης, εὐέλπες instead of εὔελπις, εὔελπι, or when he forms the aorist passive participle of εὐκρινέω as εὐκριθέν, implying a citation form εὐκρίνω, I smile benignly, and allow that he's gotten the Classical form wrong—not that he's come up with a brand new lemma. When the variants are contemporary, I'm more reluctant to make that conflation. So I have let ἀθλεύω and ἀθλέω remain distinct verbs. That's not to say I've never done such a conflation—especially when LSJ has said "= sq."; but I've had less of a motivation to.


So I have come up with a count of lemmata that conflates some variants and deflates others. In conflating variants, I am coming up with a lower count of lemmata than others might. Which means there is a question mark over the count of 173,000 lemmata I've claimed.

Well, good. Like I keep saying, there's a question mark over any count of words of any sort. But given that the lemmata contain multitudes, it's worth uncovering how many of those multitudes there are. So I'm going to count up the variants.

The first warning about these counts is that this is still not how you're going to beat English in counts of words. The OED counts some 250,000 lemmata (and its coverage of Old and Middle English is *not* exhaustive); with variants, that gets up to 615,000. There will be a lot more than 173,000 variants to count in the Greek corpus; but there won't be 615,000.

The second is to explain what I'm counting.
  • I allow any variation in the lexical rot, recorded in the lexical database, to count as a distinct variant. So ἀνήρ, ἄνδρας, and ἄντρας count as three variants, and Doric ἀγέννατος and Attic ἀγέννητος as two.
  • I do *not* count dialectal or diachronic change in the same inflection paradigm (including tense stems) as different. So I don't count both Ionic χώρ-η and Attic χώρ-α, or Doric δωρ-ίσδω and Attic δωρ-ίζω, as different.
  • I don't count dialectal variants in preverbs, because they are not part of the root; so ξυμμαχέω is not counted separately from συμμαχέω.
  • I don't count uppercase and lowercase variants of the same lemma. (But I do count them when they are distinct lemmata.)
  • I don't count active and passive variants of the same root verb as distinct.
  • And I don't count the adhoc respellings that the lemmatiser can do on the fly, to recognised deviations particularly in diplomatic editions.

So. I had 214,381 lemmata in the corpus. Without proper names and Milesian numbers, that came down to 172,646. How many variants does that translate to?
  • Variants: 362,947
  • Without numbers: 352,895
  • Without names and numbers: 286,652.

(I'd guessed around 350,000 variants a couple of postings ago. That's pretty good. It would be even better, if my guess hadn't excluded proper names...)

This amounts to 1.7 variants per lemma. I'll admit to some surprise that the OED ratio of variants to lemmata is more like 2.5: Greek historical spelling should allow for comparable confusion. My suspicion is it does, and I'm discounting adhoc misspellings which the OED doesn't.

Names are slightly more variable than normal words: 66,243 variants for 37,306 names, which is a ratio of 1.8:1. Foreign names in particular get mangled in several ways, including creative hellenisations: that's why there are 75 different variants of "Muhammad", and 43 variants of "Lombard". (Those examples aren't fair, since "Muhammad" includes the Turkish "Mehmed", and "Lombard" also includes the earlier "Longibard"—so again, I'm conflating variants more than some might.)

That count can be whittled down further of course.
  • If we discount all variants noted as hypothetical—which were made up by grammarians, and were not used in the actual language, we come down to 274,650.
  • If we ignore variation in accentuation, which is mostly a Byzantine hypercorrection, we're down to 265,233.
  • If we ignore the uncertainty between ει and ι, which bedevilled the Koine, we're down to 260,362.
  • Ignore double consonants: 253,891.
  • Ignore the distinction between η and α, as a brute-force levelling of Doric and Attic, and we're down to 247,294.
  • Ignore the distinction between smooth and rough breathing (which occasionally tripped scribes up): 245,672.

Even with these common causes of variation excluded, that's still some 73,000 added variants (41%) that I have not counted as distinct lemmata. That means that one could argue some of them should be counted as separate—although I have trouble seeing how a consistent criterion could be devised, especially over such a large timespan.

So the lemma count may be an underestimate, because of different judgements on what counts as distinct; but at its most inflated, the lemma count will no more than double. In reality, I think the debatable instances are closer to 20% than to 70%. So no, not even this way are we getting to 5,000,000 lemmata.
...Read more

2009-07-10

Lerna VId: A correction of lemma counts

Last post had its share of egg on my face, showing systematic overcounts of word forms in the corpora. This post is another healthy serving of omelette, correcting the lemma counts given in Lerna VIa. The overall story is:
  • There are less distinct word forms in the PHI #7 corpus than I thought
  • There are less scribal alternate forms left in PHI #7: if an editor thought they knew better than the scribe, the scribe's form is left out of consideration
  • There is less dialectal and orthographic wiggle-room allowed to PHI #7
  • So as a result of all this, the count of lemmata distinctive to PHI #7 has crashed: ignoring proper names, 3,800 lemmata that the lemmatiser thought it saw in PHI #7 are no longer there.
  • The count has still crashed, even though I've added a fair few lemmata to deal with PHI #7—the most frequent names, the overlaps with Trapp's dictionary, a few stragglers from DGE—as well as some dialectal grammar and some more respelling rules. I've picked up around 800 non-names and 1200 proper names; so I'm down by 1800 lemmata from before, rather than 3800.
  • I could have kept going to add more names than that, but it's been two weeks already, for gorsakes.
  • OTOH, because I've added extra names in particular, recognition of the TLG has slightly improved. So there are a few more lemmata for just the TLG-based corpora. (*Very* few.)
  • I also did some debugging of orthographic variation in lemmata, which resulted in some conflation of variants.
  • So if you ignore proper names, the TLG lemma count... actually ended up losing a few lemmata. (Again, *very* few: a couple of hundred lemmata each way.)


So.

LemmataExcluding Greek NumeralsExcluding Proper Names
TLG + PHI #7216,234 214,381211,794 209,952175,791 172,646
TLG (viii–XVI)201,680 201,823197,448 197,591162,219 162,009
LSJ (viii-VI)159,636156,720124,215
Mostly Pagan (viii–IV)99,426 99,48598,593 98,65276,145 76,067
Strictly Ancient (viii–iv)66,437 66,39066,078 66,03155,003 54,898


I also had a tally including also-rans analyses:

Lemmata
TLG + PHI #7220,560 218,727
TLG (viii–XVI)206,161 206,470
LSJ (viii-VI)166,387
Mostly Pagan (viii–IV)107,257 107,512
Strictly Ancient (viii–iv)73,427 73,532


In all of this, I've not been paying the PHI #7 corpus that much attention, though I did make a point of slipping it into the LSJ corpus. (The LSJ coverage of inscriptions and papyri are in fact why I called up PHI #7 in the first place.) I knew there would be extra lemmata there, and this lemma count is the PHI #7 disc's chance to shine. PHI #7 has added 6.5% more word instances to the TLG's, but 16% more word forms, and 6% more lemmata! That's phenomenal!

... What on Earth am I talking about? Remember Zipf's Law: the cumulative number of word forms that turn up is inversely proportional to the instance count for each word form. It's a Long Tail. If you add 6% more word instances, by the time you're already at 95 million instances, you should be getting... well, I can't do the maths, but you should be getting at most hundreds of new lemmata, not (as the table above shows) 12,000, of which only a couple of thousand are proper names. The 10,000 more lemmata of ordinary vocabulary shows you that the inscriptions and papyri—the Greek of daily life and of far flung dialects—has a very different vocabulary from the Greek of literature.

Of course, that you get 16% more word forms in PHI #7 means there's a lot of different inflections in the corpus that lie outside the TLG's ambit, because of all the non-literary dialects represented in the inscriptions. It also means a lot of misspellings that didn't belong in the TLG, as well.

In VIa, I went into an extended riff extrapolating how many more lemmata of Greek could turn up. Let me attempt that again, this time with more detail on proper names—but *not* including proper names in the final estimate.

The reason proper names don't belong in a final tally is worth restating, because not enough people are laughing at the notion. When we want to know how many words of English there are (which we shouldn't, but I've already been through that), we don't add the New York State White Pages to the Oxford English Dictionary, and we don't start screen-scraping geonames.org. We recognise that proper names are a different kind of thing from normal words (although the boundaries are fuzzy); and we also recognise that it's problematic to say a name belongs to one language and not another.

Does Κόρινθος count as a Greek name, even though it has the prehellenic telltale -νθ-? Well sure it does. Does Ομπάμα count as a Greek name? Or Σαίξπηρ for Shakespeare? Surely not. But what about the older declinable transliteration Σακεσπήριος? Doesn't that at least look Greek? What about Αὐρήλιος? But then again, what about Ἰσαάκ? Is Αμπντουλάχ not a Greek name? But does it become a Greek name when it was hellenised, as the Byzantines did, as Ἀβδελλᾶς? And is counting these names as part of the vocabulary of Greek a meaningful thing to do?

Well, better not to count proper names in the final tally at all; but let me add the counts I do know of, just in case someone is curious.
  • Right now, the TLG lemmatiser knows about almost 42,000 proper names. That includes most names of the Strictly Classical canon; a fair few names from later literature (including lots of Byzantine surnames), the names in Smith's Dictionary of Greek and Roman geography , and the thousand-odd names I was shovelling in over the past fortnight, to deal with the inscriptions and papyri.
  • Pape-Benseler went into its second edition in 1863, which increased it by a third. It covers geographical, personal, and mythological names in Ancient literature, and has some coverage of later stages. It has good coverage of such inscriptions as were known at the time, and is starting to notice papyri—though remember, this is thirty years before the discovery of Oxyrhynchus. And the dictionary is reasonably good about conflating variants.

    Benseler does not say how many names he has in total, but he does say that Alpha under his revision went from 3820 names to 6120. Extrapolating based on LSJ, that should mean 38,000 names overall. There are clearly lemmata in Pape-Benseler that aren't in the TLG lemmatiser: I add 500 names because of dealing with PHI #7, and that was only dealing with names occurring 10 times or more in PHI #7. How much more am I missing? No idea. But I'd be surprised if it was more than 10,000.
  • In the following, I need a sense of how many of these names are personal, and how many are geographical. The Heidelberg word lists for papyri are a bit more reluctant to conflate variants than I prefer, but at least they list personal and geographical names separately: 8838 personal, 2637 geographical. Good enough for me, I'll say personal numbers :: place names are 4:1.
  • 1863 is a long time ago in epigraphy, and the Lexicon of Greek Proper Names has been running for the past three decades to record the torrent of names found on inscriptions. It avoids mythological names (which are covered well enough in literature and Pape-Benseler), and it also does not do geographical names. It's ongoing, but its online search knows of 35,000 distinct names of people (whereas Pape-Benseler has 38,000 names of people, places, and gods). Now, the TLG lemmatiser recognises 17,600 distinct names, personal and geographical, in the ancient inscriptions on the PHI #7 disc. Guessing that 14,000 of those are personal names (4:1 ratio), that means it's missing at least 21,000 personal names.
  • The Leuven projects recognise 16,000 personal names in the papyri (with 7,000 extra variants), using the Duke Documentary Papyri corpus. The TLG lemmatiser recognises 9,600 distinct names, personal and geographical, in the same corpus on PHI #7. Guessing that 7,700 of those names are personal, it's missing at least another 8,000 names.
  • Some of the Leuven names will overlap with LGPN; but the Egyptian names won't. Let's say that all up, we're owed at least another 27,000 personal names. And using that 4:1 ratio again, another 5,000 place names. Heidelberg counts 9,000 personal names to Leuven's 16,000, and Heidelberg counts 2,600 geographical names; extrapolating up, that's consistent with 5,000.
  • That's not even scratching the surface of Byzantine and Modern names (let alone Σαίξπηρ or Ομπάμα, or the Thessalonica and Environs phone book). But so far, we can guess 42+27+5=74,000 names.
  • Flipping things around, there are 72,000 unrecognised capitalised words in PHI #7. That does not mean 72,000 missing names: lots of these will be misspellings of known names that the lemmatiser isn't dealing with, or different inflections of the same name. And those names are in the scope of LGPN and Leuven. I'd say the personal names are already accounted for in the 27,000 (say) personal names of the two initiatives.
  • There are a further 42,000 unrecognised capitalised words in TLG. Most of these won't be in LGPN and Leuven—though some will be in Pape-Benseler. Most of these by far are from post-Classical texts, and they include ancient gazeteers. (Ptolemy's Geography alone accounts for close to 3,000 unrecognised names.) How many of these are legitimate novel proper names? Again, no idea, but by this stage we're getting into one-offs, because all proper name word forms occurring more than 7 times in the TLG have been added to the database. I'll guess 30,000. There'll be some overlap with Leuven and LGPN, but not a lot, because many of these names are Byzantine.
  • As mentioned, the TLG is maybe 70%, maybe 75% complete for Byzantine literature, and only starting to go into Early Modern literature. It does have a lot of Byzantine surnames through church deeds (which account for 5,000 unrecognised capitalised words); so it'll have a reasonable cross-section. I haven't gone through the Byzantine proposopographies though (285-641, 642-1265, Palaeologan), to work out how many surnames they've unearthed in sum.
  • And I have not spent quality time with the Attica or Thessalonica or Nicosia phonebooks.
  • So at least 70,000 proper names to go, adding up to something like 110,000 proper names, and that count only goes up to the Fall of Constantinople.

Anyone who wants to start boasting of the 110,000 proper names of Two And A Half Thousand Years of Greek needs to be smacked upside the head with all three volumes of the Dictionary of American Family Names, and have the printout of all 8,000,000 places on geonames.org dropped on their foot. Because all of those count as proper names of One Year of English, by the same criterion.

(The Blogger Writing These Lines enjoyed contributing to the Dictionary of American Proper Names, even before he realised its value as a tool of percussive persuasion.)

So. Banishing proper names, we're left with 173,000 lemmata, as guesstimated. How much is left to go again? As it turns out, I'm doing the same guesstimates as before—but they make more sense without including proper names:
  • I keep my guesstimate of 20,000 lemmata more from Trapp (including texts not yet added to the TLG and volumes not yet published), and 10,000 lemmata more from Kriaras (ditto). That's 203,000.
  • There are words in LSJ that are not represented in this corpus. The biggest gap is the mediaeval Latin-Greek glossaries, with 1,000 missing lemmata; but there are several other oddities. The latest I've encountered, under ἐλεφαντουργική "of or pertaining to ivory-working": the 1161 AD commentary to the astrologer Paul of Alexandria, writing in 378—and last published in 1588. (The irony here is, the same adjective turns up in the rather more mainstream Heliodorus, a century beforehand.) But again, once the PHI #7 texts are in, and with the changes in text editions between the original LSJ and the TLG—not to mention the rejected scribal forms—I don't think there's more than 3,000 lemmata to add. That takes us to 206,000.
  • I'm inclined to revise my extrapolation for DGE downwards. Volume I updated may have 3500 lemmata not in LSJ, but it's competing not only with Bauer, Lampe, and Trapp, but also with the LSJ Supplement—which on its own adds 10,000 lemmata to LSJ, and which also has made a point of covering more inscriptions and papyri. I haven't taken the time to do any counting with DGE. It's a long plane trip tomorrow to Montreal—so maybe I will.

    But there's no way Volume I has 3,500 lemmata not also in LSJ/Bauer/Lampe/Trapp/LSJSupp. DGE looks like taking 20 volumes if and when it finishes. (I wasn't planning on living until 2100 AD to find out.) If there's just 500 novel lemmata in Volume I, that means 10,000 novel lemmata all up; if 1000, then 20,000, as I proposed last time. I'm feeling jaundiced, but I'll still give them 20,000. That takes us to 226,000 lemmata, up to the fall of Candia.


Ούφ. On those figures, English still wins, :-) though not by much. The level of precision I've given is of course illusory, and in a following post I will tackle what is a more sensible question: how much vocabulary do you need to recognise n% of a text. But these counts should at least be indicative.
...Read more