Saturday, 12 April 2014

Handling Transcriptions


Making transcriptions of records is not as common amongst genealogists as you might expect, but why is that? What do we need in order to create useful transcriptions? If we’re part of the minority who do make them then where should we attach them?

Because of the availability of online data sources, and the ease with which digital copies can be created (owner permitting of course), many people believe they do not need full transcriptions of records. They might claim that since they can visit an online image, or they have a digital scan in their own data collection, then they can read it perfectly well without having it typed out. Whether it’s a baptism entry, a newspaper report, or a census page, many genealogists therefore find they have a growing collection of equivalent JPEG files sitting on the periphery of their data.

What I mean by this is that such a file can be pointed to, or referenced, by other data, but it cannot reference anything itself[1] or be textually searched. This means the information is not truly integrated into your data. The arguments for adding mark-up to a transcription in order to achieve this are almost exactly the same ones that I made for using mark-up in authored narrative at Semantic Tagging of Historical Data. This allows, for instance, references to people, places, events, dates, etc., in that transcription to be connected to the relevant entities in your data.

A transcription requires more though. It also requires a way of indicating transcription anomalies — parts that deviate from the normal flow — such as marginalia, footnotes, interlinear/intralinear notes, struck-out text, and uncertain characters or words. Both the uncertain characters and the uncertain words may require annotation to provide suggestions and possibilities, both of which must be honoured during searches. A transcription also requires an indication of any original emphasis, such as italics or underlining. NB: the original use of italics, underlining, footnotes, etc., in something being transcribed is different to their deliberate use in a written report, and so must use a distinct form of mark-up.

Traditional editorial notations for transcriptions are not well-suited to digital text as they do not facilitate efficient and accurate searching. TEI has comprehensive sets of mark-up for handling transcription issues but falls short when applied to genealogical data, and probably historical data in general. It is certain that some specialised mark-up is required, but how you visualise a transcription on-screen is a separate consideration. The same mark-up could alternatively show multi-coloured and hyper-linked text, or the plain editorial notation. That sort of flexibility only comes from using a computerised annotation rather than human annotation.

The fact that both transcription and authored narrative may co-exist in the same written report led to STEMMA® unifying them in its own mark-up. Those distinct usages — for transcriptions and for generating new narrative (e.g. essays, reports, inference, etc.) — have some similar and markedly different characteristics as follows:
  • Transcription (including transcribed extracts) — requires support for textual anomalies (uncertain characters, marginalia, footnotes, interlinear/intralinear notes), audio anomalies (noises, gestures, pauses), indications of alternative spellings/pronunciation/meanings, indications of different contributors, different styles or emphasis, and semantic mark-up for references to persons, places, groups, animals, events, and dates. The latter semantic mark-up also needs to clearly distinguish objective information (e.g. that a reference is to a person) from subjective information (e.g. a conclusion as to whom that person is).
  • Narrative work — requires support for layout and presentation. Descriptive mark-up captures the content and structure in a way that provides visualisation software with the ultimate control over its rendering  It needs to be able to generate references to known persons, places, and dates that result in a similar mark-up to that for transcriptions. The difference here is that a textual reference is being generated from the ID of a Person entity, say, as opposed to marking an existing textual reference and possibly linking it to a Person with a given ID. Also needs to be capable of generating reference-note citations and general discursive notes.
Actually, transcription isn’t just an action associated with a manuscript or typescript document; it could be associated with speech too. In those circumstances then it must reflect speech levels and emotional emphasis, but I haven’t even thought about that field yet.

As you can imagine, in order to generate a quality transcription, and to incorporate semantic links and annotation, a very good software tool is needed. It would be something like a specialised word-processor tool, but most of us are left using general-purpose word-processor tools that have none of the required facilities. This will be a secondary reason why so few transcriptions are made.

So where do I attach transcriptions in my own data? In order to explain, I first need to convey something of the structure of my data.

STEMMA Entity Linkage

This simplified view of the rich connections in the STEMMA tapestry doesn’t show its places, or groups, or lineage links between people, or hierarchical/protracted events. That would be too complex! What it does show is a network of multi-person events and the relationship of sources to those events. Notice that the sources are attached to the events, and not to the people. As already explained in Evidence and where to Stick It, the vast majority of our evidence – if not all of it – relates to events; things that happened in a particular place at a particular time. In other words, our entire view of history rests on discrete and disjointed pockets of evidence describing a finite set of events. Everything else is inference and interpolation creating as smooth a picture as we can.

So what is the general form of these underlying source entities in the data? Our real-life sources may be remote, such as a document in an archive or a book in a library, or local, such as a family letter or a photograph. In both cases, we may have a digital scan of the items. STEMMA[2] has two important concepts that it employs for sources:

  • Resource – This describes some item in your local data collection, including not just files on your disk, but also physical artefacts or ephemera.
  • Citation – Despite the name, this is merely a link to some source of information. A traditional printed citation may be generated from it, but this software entity also incorporates collections, repositories, and even attribution; possibly chaining them together.

Either or both of these may apply, therefore. A full transcription would be associated with the Resource entity that would describe any physical or digital edition of the associated material. In the case where you may have transcribed a document in an archive, or even from one of the online content providers, the transcription should still be placed in a Resource entity rather than a Citation entity, even though the latter is possible.

Genealogist Janice Sellers, in her blog-post at Transcription Mentioned on Television, explains how transcriptions of documents are valuable for sharing the details with family and friends. She recounts how she tried to convince a well-known British TV program to advise their guests to make transcriptions of their historical documents and heirlooms.

STEMMA’s mark-up is primarily about semantics. Shallow semantics would mark an item as, say, a person reference but without forming a conclusion about who the person was. Deep semantics involve cross-linking references to persons, places, groups, events, and dates, to the relevant entities in your data. I have previously tried to convey this using the worked example of an old family letter at Structured Narrative.

Genealogist Sue Adams has taken the concept of semantic mark-up in transcriptions to a deeper level on her Family Folklore Blog. Her worked examples clearly demonstrate the temporal nature of historical semantics. Anyone with a passing interest in the Semantic Web and RDF is encouraged to read about “temporal RDF” and consider why it doesn’t yet exist. You may find a lot of theoretical work that considers things like temporal graphs but very few real examples like hers. In an ideal world, the developers of such technology would be working closely with the people who need to utilise it.




[1] I’m ignoring the issue of meta-data held within an image until a future post. The issue here is one of the text in an image making discrete references to its subjects rather than anything to do with image cataloguing.
[2] STEMMA V2.2 — which includes important refinements here — has just been defined but, at the time of writing, I am still preparing to painstakingly update the Web site. The landing page will indicate when this is complete.

Thursday, 3 April 2014

What to Share, and How – Part II



In the first part of this blog spot, What to Share, and How, I suggested that if collaborative Web sites were designed to accommodate creative works, rather than mere trees, then it would better accommodate family history, and it would encourage more sharing by ensuring accreditation and integrity. I now want to suggest how this might work in practice. I also want to conclude with a potential sting-in-the-tail for those, like me, who believe that simple trees cannot be copyrighted.

So what do I mean by a creative work here? Many of us have had to write up some form of narrative, whether for a client or for publication in an article, book, or blog. Such work is no different from an original work of research or fiction in that it is automatically copyright by virtue of the Berne Convention. Hence, this would be a prime component of the improved sharing.


STEMMA could take this further since it has the ability to package an integrated set of data that includes both narrative and transcriptions, and the entities that they reference such as people, places, events, and groups. The whole bundle could be indexed by the people (which includes their lineage), or a timeline, or their locality.


All of these components are cross-linked, thus making it an integrated bundle. The lineage section connects the people in the normal way, according to their biological lineage, but there may be multiple, disjoint trees. In other words, the bundle may represent distinct sets of people.

If such a bundle were uploaded to a collaborative site then none of it need be undone. Instead, each of the lineage sections would be anchored to a corresponding person entity in a lineage-based framework. The overall framework would be constructed based on the lineage of all the uploaded contributions, thus making it dynamic.


In this ideal world, therefore, there would be no need to edit or copy other people’s contributions. Multiple contributions could be associated with a single person entity (and their family) in the overarching framework. The accreditation and integrity of individual works would be preserved, and citations (or attribution) used when necessary. This is a collaborative model far-removed from what we have now, and I’ve glossed over issues of voting up/down contributions that may disagree, but let me know if you like the concept.

I want to conclude this post, though, with something a little unexpected. In the earlier piece of this two-part blog, I explained that mere collections of facts available in the public domain cannot be copyrighted. This is true, but does that description include the family trees that we currently see online? Dick Eastman recently blogged on this subject at Genealogical Privacy, and he explains the “legal and practical” fallacy of many researcher’s views that they own the data they’ve collected, and that publishing it allows others to freely steal it. The premise for this is that the data is “freely available to everyone in the public domain”. Although he does add the caveat that this is the case in the US, and so recognises that someone in the UK, say, may have had to pay for the information they’re publishing, there are a couple of other issues. Not all data may have been in the public domain, although this is usually associated more with data related to recent generations. Also, as we all know, the published details may not be clearly visible in the public-domain data; meaning that some effort may have gone into determining someone’s true lineage.

The reason I am picking on these points, and in doing so questioning my earlier statements, is that legal precedents may exist in similar, but non-genealogical contexts. One that I am aware of is a case of copyright that was tried in England in 1868, and due to its unusual nature is still referenced in many academic books on copyright law. The reason I am aware of this is that one of my ancestors was on the receiving end, and the judgement went against him, breaking him in the process.

The case is that of Morris v. Ashbee. William Ashbee, in order to create a new trade directory of London, took an existing trade directory, compiled by John Morris, and gave the alphabetical list of names to his canvassers to check. Although Ashbee didn’t pass-off the earlier directory as his own, the judgement went against him because Morris had incurred the labour and expense of getting the information and of making the compilation. Ashbee had therefore benefitted from Morris’s work and that was considered an infringement of copyright. Ashbee later went bankrupt and died not long after.

Although this case was never envisaged in the context of genealogy, the general concept of benefitting from someone else’s labour and expense being an infringement of copyright must have potential implications for people who advance their genealogy by “stealing” data from others — even though it may have been publicly available.

Friday, 28 March 2014

What to Share, and How



The subject of public family trees frequently does the rounds. In this post, I want to examine what people would like to share — software permitting — and the functional requirements of that sharing.

At the fore of the current posts on this subject is probably James Tanner’s Genealogy’s Star blog. Although there are too many posts there for me to cite separately, they consider the differences between having public trees versus private trees, using online trees versus local genealogy programs[1], and unified online trees versus user-owned ones. I’ll use the following table in order to put these terms into perspective.


Public
Private
Online
Unified

User-owned
Local


In other words, a tree maintained using a local genealogy program is private to you, although you could give a copy to someone else. When such a tree is hosted online then there is a choice. If it is part of a unified tree then it is necessarily public, but if held as a separate user-owned tree then it could be either public or private.

There are both advantages and disadvantages to maintaining local trees, and a comprehensive list was recently posted by Renee Zamora on Renee's Genealogy Blog, but why do people want to create online trees? One reason is simply for use as “cousin bait”, and attracting distant relatives with a view to controlled sharing. This is the closest description of my own situation as my definitive data is held on a local machine. It is a position that will become increasingly difficult to maintain as online data is judged more and more by the sources it elects to cite. Some people do not want to share their data publicly for more delicate reasons, and a good case was presented by Kerry Scott at Why Don’t People Post Public Family Trees?

My own reasons for not sharing more data online are deeper than the implication above that I simply want to use a local program. My STEMMA® R&D project is one factor since I am developing software to support that custom data representation. Possibly more important, though, is the fact that my data is far from being a simple family tree. It is a representation of general micro-history that incorporates family history, family trees, pedigrees, timelines, narrative, etc. I cannot, therefore, share everything since there is no standard for this type of data, and no sites (at the time of writing) that are properly structured for this type of data.

James Tanner’s post Family Trees: Unified vs. User Owned caused me to think more about what I would like to share, and how, so I will try to expand on my brief response to him. I detest our industry’s preoccupation with “family trees”, and the way that it leads newcomers into believing that is the be-all and end-all of genealogy. I don’t know of a single experienced genealogist who only wants to collect names, dates, and places associated biological lineage in order to create a tree. They’re all interested in family history, and all aspects of that history. Although I’m out on a limb by declaring an interest in general micro-history. including the history of places, groups, and non-relatives, this is merely a superset of family history.

A very significant issue with any type of historical work is that it is a creative work. It involves research, thoughtful analysis, and some skill in writing it up accurately and interestingly. This is more than just an assembly of facts that anyone could find in the public domain. Even when a public tree is given source citations, it would be little more than an assembly of such facts. If it were possible to share our data as creative works then our requirements would suddenly align with those of authors of other online works, whether fiction or non-fiction. It struck me how close those requirements are to the issues people currently raise as obstacles to sharing their genealogical data publicly. For instance:

  • Attribution – Ensuring that their authorship is acknowledged. Allowing their work to be cited by the work of others as opposed to having it plagiarised.
  • Integrity – Allowing other researchers to see their work, but not to edit it. Their work could be connected to a central tree for indexing purposes but should not be assimilated entirely into the tree in order to preserve its structure or narrative form.
  • Drafts – Allowing revisions of their work, and possibly the addition of tentative items that they don't want to expose until they're more confident of them.
  • Longevity – Ensuring that their work will persist after they are no longer able to contribute.
  • Privacy – Allowing certain information to be disclosed at some point in the future (e.g. some respectable point after their death).

Obviously I cannot speak for everyone out there, but if this were possible now then I would gladly share all of my research. However, I would clarify that a tree-based site that simply accommodated rich-text notes is not what I’m thinking of. It would have to fully accommodate a structured representation of historical data that includes all of the items I mentioned above, including narrative, and yet could be indexed by a tree, a pedigree chart, or a timeline, etc. This is certainly possible and is one of the goals driving STEMMA development.

I can’t quite work out the dynamics behind the industry advertising and the tools that we’re provided with. As I said, the concept of a family tree is endemic, but whether the advertising influences tool development, or vice versa, is hard to determine. As a software developer myself, I sometimes wonder whether developers see our tools more as a technical challenge than something that has to satisfy the requirements dictated by real genealogy. Collaborative Web sites, where we build a single picture of something, are a good example. Ignoring those sites that are wiki-based collaborations, everything I have seen is related to “unified family trees” rather than anything involving events, places, and narrative. The fact that even these existing sites are problematic supports my view that they are considered to be challenging. Although I demonstrated that other forms of collaboration are possible at Collaboration Without Tears, I also feel that it should be possible to upload “rich” (see above) user-owned data contributions to hang off a unified lineage-based framework. This step would be more significant than it may sound but I’ll defer any detailed presentation until another post — if there’s any interest, of course.




[1] It would be restrictive to term these ‘genealogical database programs’ since a local program does not necessarily have a database. As explained in DoGenealogists Really Need a Database?, a memory-resident database might be constructed on-the-fly from permanent and definitive data.

Wednesday, 19 March 2014

Related Entities

No, not that sort of relation this week. This brief post espouses some thoughts on the esoteric subject of “related” places and groups. This should be of interest to software people but maybe not to everyone else.

To give my readers the chance to decide, let me explain the basic problem. Places are acknowledged to be fundamentally hierarchical — meaning that each place, whether a house, street, town, county, etc., is part of some bigger place — but every so often something breaks that neat picture. A place may be split into smaller ones, or small places joined together to form a bigger one, or a place torn down and replaced with an entirely new one at the same location[1].

This is a thorny issue that I don’t believe anyone has succeeded in handling gracefully. When the name of a place has changed at some point, or its boundaries have moved, then that is a simpler problem that I have already addressed. I have even addressed the more complicated situation where the parent place (i.e. the bigger one to which it belongs) has changed. In all these cases, the entity may be considered to be the same one, albeit with some new or changed properties.

The anomalous relationships I want to discuss may be categorised broadly as:

  • One-to-Many. This involves a splitting of an entity. An example for a place would be where a country has been split into several smaller ones, or a family farm has been divided between a number of siblings.
  • Many-to-One. This involves a joining of smaller entities. An example for places might be where someone has purchased adjoining land, or contiguous houses, and made them into one.
  • One-to-One. This involves a connection between two different entities that isn’t modelled by a normal hierarchy. It could be used as a catchall since the potential circumstances are more varied. For instance, the two entities may co-exist but still be related (see below), or there may be gap between the demise of one and the creation of the other. An example of the latter might be the redevelopment of an area of housing involving the digging-up and the laying-down of totally different roads. Two houses, of different dates, might then be related by their physical location.

In Revisiting the Family Group, I made the point that this issue must be solved in a consistent fashion for both places and groups. The first two categories have an obvious correspondence since groups may merge or splinter. The aforementioned post even provides a military example where two regiments were merged to form a new one. Luckily, both the Place and Group entities of STEMMA contain Creation/Demise elements that indicate when it came into being and/or ceased to exist. This turns out to be the ideal position to document the one-to-many and many-to-one transformations since the two forms would not co-exist. For convenient management, I elected to collect the links into a single place, thus providing a SplitTo and a JoinFrom sub-element.


In the one-to-one category, there’s a Group situation that has no equivalent for a Place: where one entity has inspired another. This is not the same as a splinter, which would otherwise divide the group membership, and an example I’m particularly familiar with is the creation of FHISO. This group was formed by a small set of members from BetterGEDCOM but the former group was unchanged.

So is this a complete solution? The answer has to be ‘no’ but I’m floating these ideas to get some constructive feedback, and also to indicate what failings the approach has. My goal in this exploration is to find a balance between structure and narrative; using the latter to differentiate the finer points of some generic connection or transformation event. One area it may not address is the Conurbation: a collection of neighbouring cities or towns that have a name independent of their respective parent places.

There is a need to accommodate these because they appear in such records as census returns. The example I will use for the purposes of illustration is The Potteries: an area of North Staffordshire, England, which encompassed the towns of Tunstall, Burslem, Hanley, Stoke, Fenton, and Longton. These towns, and several villages, were later given the single name of Stoke-on-Trent, a polycentric town that eventually became a city, although there was a settlement of the same name before that. An example occurrence in the records may be found for the place-of-birth of Joseph Davies in 1861[2].

I’ve picked this case since it appears in my own data. As a singular case, there may be a way of handling it as an older version of the Stoke-on-Trent town/city, but in the general case of UK Conurbations there may be instances where the natural parent place (e.g. a county) may be different for the constituent towns and villages. Short of having multiple parent entities (of different types), I cannot see a better scheme than relying on a one-to-one association at the moment.

** Post updated on 26 Dec 2015 to align with the changes in STEMMA V4.0 **



[1] STEMMA® makes a precise distinction between the terms ‘place’ and ‘location’. A definition and discussion may be found at STEMMA Places.
[2] "1861 England, Wales & Scotland Census",  database, FindMyPast (www.findmypast.org.uk : accessed 18 Mar 2014), household of Joseph Davies (age 30); citing RG 9/2292, folio 47, page 3; The National Archives of the UK (TNA).