Showing posts with label word separation. Show all posts
Showing posts with label word separation. Show all posts

Apr 17, 2015

Writing as encoding

Unicorn seal from Harappa, c. 2200 BC. Image: Harappa.com.

This project explores the history of medieval information technology by modelling writing as code. But what, exactly, does writing encode? Answers to this question are complex, significant, and richly productive.
what does writing encode?

The easy answer is that writing encodes information. The earliest forms of writing seem to have been used for accounting. However, there is some debate among historians of writing over whether visual semiotic systems that encode non-linguistic information should be considered ‘writing’. This is perhaps a semantic quibble, but it underlines a very important historical development: at some point, every writing system commonly used today was adapted for linguistic information. Since the most commonly and frequently used method that humans use to create, store, transmit, and process information is verbal – that is, humanly usable information is mostly linguistic – this tight linkage of writing and language became overwhelmingly powerful. It is easy to think of writing only as a representation of language, and of spoken language in particular.

Mar 25, 2015

Reading twitchily, then and now

A physician reading. Image: US National Library of Medicine.
[Right: Historiated initial from the Articella. Bethesda, United States National Library of Medicine, MS of the Articella, fol. 19v. Oxford, 13th century.]

After expressing all kinds of scepticism in two previous posts (here and here) about Paul Saenger’s arguments concerning word separation, I think it only fair to lay out an aspect of reading with spaces that, according to more recent studies on the neuropsychology of reading, Saenger got (mostly) right. It has been empirically demonstrated that word separation by space does speed up reading for skilled adult readers. Furthermore, it does so in a range of writing systems, including writing systems that use scriptio continua, that is, those that do not conventionally separate words by space, such as Chinese or Thai. In other words, even when skilled readers of such scripts are confronted with unconventionally word-separated text, they process the text more quickly.

Feb 20, 2015

Reading with spaces in Anglo-Saxon England

Boisil teaching Cuthbert. Image: British Library.
In the Life of St Cuthbert composed by Bede c 721, there is an episode in which Cuthbert asks his mentor, the saintly Boisil, to recommend a book that can be read in one week. Boisil suggests the Gospel of John, and provides a copy consisting of seven quires (codex habens quaterniones septem), which the two of them read together, one quire a day, until Boisil, as he has predicted, dies at the end of the seven days.

[Image: London, British Library MS Yates Thompson 26, fol. 21r, detail. Durham, late 12th century. Miniature from Bede's Vita Sancti Cuthberti.]

Whether or not this story can be entirely accepted as historical fact is perhaps doubtful; it occurs in a text that is concerned more with promoting the saintliness of Cuthbert than with what we might consider historical accuracy, and I suspect that Bede wants us to see a parallel between the seven days of reading, after which Boisil goes to his eternal rest, and the biblical seven days of creation, at the end of which God rested. As well, we might wonder at Boisil’s recommendation of the Gospel of John as a book it would take a week to read; in a modern English translation this text is about the length of a longish short story, and a skilled modern reader could easily read through it in an hour or two. Each quire (quaternion) in Boisil’s manuscript would be the equivalent of 16 pages, but Boisil expected Cuthbert to take a day to read each quire. What is even more striking is that Bede considers a week’s time to be a quick read of the Gospel of John; he implies that ordinarily it would have taken longer. Medieval people must have been slow readers.

Well, yes, by evidence such as this, medieval people were slow readers by our standards. But to discover why, we have to consider the evidence more carefully.

Feb 11, 2015

Spaces and silence

 London, British Library MS Harley 1775 (Harley Gospels), fol. 373v. John 1.14 per cola et commata, in scriptio continua.

In 1997, Paul Saenger pulled off a feat I admire: he published a lengthy, detailed, and learned book entirely about the little spaces between words. Space Between Words is still the place to go for a painstaking account of the history of word separation in the medieval West, from the first texts separated by space in late 7th-century Ireland to the eventual adoption of canonical word separation for all Latin and vernacular texts in western Europe by the end of the Middle Ages. Although early Greek and Roman texts separated words by interpuncts (little dots), later Latin texts were written in scriptio continua (or scriptura continua): a continuous string of characters without spaces to mark word boundaries. Beginning in the late 7th century, Irish scribes introduced spaces at irregular intervals to create what Saenger calls ‘aerated’ text , and, by the 11th century, scribes in northern Europe were separating Latin text ‘canonically’ – that is, the way we do now in standard written English, with spaces between words. This history can be verified by checking the manuscripts that Saenger cites as examples, if anyone has the gumption to follow up all the items in Saenger’s impressive nineteen-page list. I, for one, am quite happy to take Saenger’s word for it.

London, British Library MS Additional 89000 (St Cuthbert Gospel), fol. 1v. John 1.14 with canonical word separation.

More controversial, however, is the argument that Saenger builds on top of this history, an argument encapsulated in the subtitle of the book: The Origins of Silent Reading. Briefly, Saenger claims that word separation was, in the Middle Ages, ‘the crucial element in the change to silent reading’. Silent reading, in turn, facilitated ‘reference reading’ – the technique of scanning texts quickly to find specific items of information – and contributed to the shift from the idea of reading as a public activity, as it was in the ancient world, to the idea of reading as a private activity, as it is in the modern world. I am, of course, greatly oversimplifying Saenger’s ideas here, but if you want all the details, you should read the book.

Saenger’s thesis may be attractive, not least the argument that a seemingly innocuous and subtle encoding practice – introducing a bit of white space between words in written texts – has had such far-reaching technological and social effects. But is there a straightforward causal relationship between word separation and silent reading? I’m not so sure.

Oct 21, 2013

Word by Word


Franks Casket, front panel. Image: British Museum.

Corpus Glossary, CCCC MS 144, fol. 58v (detail)

Vespasian Psalter, British Library MS Cotton Vespasian A.1, fol. 24r (detail). Image: British Library.

Cambridge University Library MS Kk. 5.16, fol. 128v (detail)

These four documents all include English text in some form, and all date from the eighth century. The first, the front panel of the Franks Casket, features an Old English riddle in runes. The second, the Corpus Glossary, provides meanings of Latin terms in Latin and sometimes in Old English. The third, the Vespasian Psalter, gives an interlinear Old English gloss on the Latin text of the Psalms. And the fourth, a copy of Bede’s Historia ecclesiastica, includes the Old English text of Cædmon’s Hymn as an annotation to the Latin text. All these documents provide graphic evidence of the way early English writers thought of their language as being divisible into word units.

Jul 25, 2013

Reading in two languages

Yin Liu 

The oldest surviving English translation of any part of the Bible can be found in this manuscript:



The translation consists of the smaller words written above the larger main text. It is not a Bible translation in the sense that we customarily think of – that is, a text that can be read in its own right. Rather, this translation takes the form of an interlinear gloss, intended to help an English speaker read the Latin text.

The image is from a page of the Vespasian Psalter, British Library MS Cotton Vespasian A.1, a copy of the Psalms in Latin. The manuscript was made in the second quarter of the 8th century, probably in Canterbury. The gloss was added just over a century later, in the middle of the 9th century. The image shows the start of the psalm Caeli enarrant, Psalm 18 (Psalm 19 in most modern translations). The script of the main text, in Latin, is in English uncials; the gloss is in an insular pointed minuscule, in a Mercian dialect of Old English. The English gloss provides a word-for-word translation of the Latin: thus in omnem terram is glossed in alle eorðan (‘into all the earth’). This way of presenting a text in two languages probably seems straightforward and self-evident to us now. Think how often students annotate their textbooks in just the same way when learning a new language or reading a text in a language in which they are not fluent; and in modern linguistics, interlinear glosses, laid out in much the same way, are a regular feature to assist readers in understanding examples of speech or text in many different languages.
This bit of parchment shows an example of medieval text being encoded, structured, and presented as data

Nevertheless, that bilingual interlinear glosses are among our earliest surviving examples of English text (and continue to be used through the Middle Ages and into the modern period) should give us pause. Before a gloss is added, the text on the manuscript page can be read as a relatively simple linear transcription of speech. But once the interlinear gloss appears, the reader is challenged to regard the text on the page as a much more complex structure, existing in two dimensions rather than one. No longer is there a single sequence of linguistic units to follow, but two parallel sequences, linked by a one-to-one correspondence between individual elements. Furthermore, the two sequences are not equal in value; the Latin sequence is privileged not only visually (it is written in a larger and more prominent script) but also by dictating the order of elements on which the English gloss depends, even when normal Old English word order would be much different from Latin.

We may notice also that word-division in the two sequences does not always correspond. For example, the Latin text frequently joins the conjunction et (‘and’) to what we would consider the next word, with no space between: etopera, etnox, etipse. The English gloss separates out the conjunction and or ond (abbreviated with a symbol that looks like ‘7’) so that it is recognised as an individual linguistic unit: 7 werc, 7 neht, 7 he. This may seem trivial, except that word-separation by use of white space had only just been developed as an encoding convention by scribes such as these in the British Isles, and it had deep and far-ranging repercussions for reading practices throughout Europe and into the present day (Saenger 1997). Among its effects was a shift in the meaning of the word ‘word’. In this text, Latin verbum and English word mean ‘utterance, something said’. But separated script visually fragmented the stream of language into discrete units, which could then be processed and presented in new, non-linear ways.

This bit of parchment shows an example of medieval text being encoded, structured, and presented as data: tokenised and then arranged so that relationships between the tokens are visually apparent. Medieval English readers, grappling with a text in a foreign language, implemented reading aids such as interlinear glosses that allowed people to receive not only auditory but also visual linguistic information, and so created ways of understanding that depended ever more heavily on technologies of writing and of the book.

Addendum

Some useful qualifications to these remarks, and a much better linguistic analysis, can be found in Alderik H. Blom, Glossing the Psalms: The Emergence of the Written Vernaculars in Western Europe from the Seventh to the Twelfth Centuries (De Gruyter, 2017), 161-173.