Showing posts with label Artefacts. Show all posts
Showing posts with label Artefacts. Show all posts

Friday, 4 January 2019

Organising More Resources


Time for a rant since I keep on seeing the same old suggestions for organising digital resources that fail at "naming 101".  Organising by either person name, location, or date are not only restrictive but impractical and unhelpful in many circumstances; and trying to put these terms into file or folder names is a road to nowhere.



So, Tony, what is your problem here? Well, there are several issues:

  • There's a huge difference between the physical organisation of material (digital or otherwise) and the indexing according to your mode(s) of access.
  • Coding in a file or folder name forces you to make a limiting choice. The classic example is a group photograph that has several people of different surname and family connections.
  • Access isn't always by just one category of index-term: you may want, for instance, all resources related to Proctors, in the city of Nottingham, during the 1950s; or, all Jessons who were present at a particular event.

I've said before that physically organising by the nature or provenance of the material is not only the archival way, but it has advantages for maintenance, changes of ownership, and even making inferences (e.g. identifying a person in a photograph). Anyone who splits up an inherited photographic collection should have their fingers taped together. In contrast, indexing helps you access your material according to various categories, such as surname, location, etc.

In Organising Digital Resources, I described the difference between these two concepts, and how indexing for your mode(s) of access can be done according to multiple inclusive categories; you're not forced to choose just one. Unfortunately, although we have real archivists in our community, the tendency is still to miss the analogy between their professional organisational schemes and the digital world, and so oversimplify such things to coding surnames, etc., into file and folder names. Professionals in the digital world do not do this, and their schemes would also be the ones used by archives to implement their own schemes.

Perhaps the best arguments against the way things should be done are: (1) that your software of choice may be rather limited, or (2) that browsing the resources in the absence of any specialised software leaves them hard to understand.

In a much older article, Organising Photographs, I mentioned the use of meta-data, and how this could be used to add important information (visible only to software) to images, or to add index-terms to all file types (including images) in order to aid in their access from a simple Windows search. This is not unlike writing information (or source labels) on the back of physical photographs, except that software could use the digital equivalent to help access them. A major goal of that article was to show that digital resources could be indexed using very simple software technology available on all our computers, in contrast to using some highly specialised genealogical software. Well, some people would still not like the invisible nature of that meta-data, and it's still poorly supported by standards, and hence by different computer operating systems.

So let's explore the analogies with physical artefacts. If you had a photograph album then you would probably have written details underneath each picture. If you didn't have an album — just a biscuit tin on the top of your wardrobe — then you may have written details on the back of each picture. However, I have some WWI photographs of soldiers that were sent to their families as postcards. That means you cannot write on either side without damage to the precious original, and in that case you might just have a separate piece of paper, or better still an envelope, with the salient details written on it.

One alternative for digital  resources might be to use a simple non-specialised bit of software such as Excel, which uses multidimensional hierarchical indexes — effectively a scaled down version of an OLAP database. It's proprietary, yes, and it's opaque, yes, but it can be kept alongside the digital resources.

Because I arrange my own digital material hierarchically, akin to a micro-archive (see Hierarchical Sources), then I also use text files to describe the material at each level (e.g. for fonds, series, or even items), and this presents a very simple alternative that is a closer analogy to "separate paper" idea: buddy files. For instance, in a collection of photographs, I would have a single text file of a fixed name (e.g. Description.txt) with all the details of what the collection is, where it came from, when, and who had it before that. Alongside each photograph (i.e. at the item level) I often have a buddy file of the same name with not just a plain-text description of the where, when, and who, but tags that can be searched on.

For example, suppose I have an image of family photograph called Picture1.jpg then I might also have a Picture1.txt text file with tags as follows (descriptive text not shown here).

Figure 1 – Picture1.txt and Picture1.jpg

This has effectively three categories of hierarchical index-terms, and a search through the buddy files for #Proctor, #Nottingham, #1950s would throw up the name Picture1 whose associated image could be viewed unaided by any specialised software.

Here's another example for comparison:

Figure 2 – Picture2.txt and Picture2.jpg


How you name the item-level files is partly irrelevant, as long as they're unique at each level of organisation. For photographs, you could invent your own scheme of codes, similar to an archive, or use something more meaningful — it doesn't matter because anyone browsing the images directly would also have the buddy-file details on hand. For images of material obtained from an archive, though, and this would include any census images for England and Wales, I would strongly recommend using the assigned archival codes in the image names.

So, this is probably the simplest scheme possible, and it doesn't rely on hidden meta-data, or databases, or specialised genealogical software. Such software could still index your resources, as I've already explained, but this scheme provides your digital resources with plain-text notes and index-terms of their own — ones that would follow them if ever they were transferred elsewhere or duplicated for someone.

The organisation of your physical artefacts should follow the precedent set by archives, so why not do digital resources in a similar fashion. For instance, my extended family has a large stone chess board, originally seated on a wooden table, dated 1859 with the name of my ancestor carved into it, and with the initials of an in-law who was a stone mason. I may have images of it but the artefact has a fundamental existence of its own, and I would need to index this in my software. This may be unusual, but many of us have letters, certificates, ephemera, other original documents, and photographs. Older photographs were obviously printed and so scans are derivative, but modern photographs are "born digital" and so it's the printed forms that are derivative. Either way, keeping paper-based copies is always wise. Believe it or not, there are people who recommend scanning old photographs so that the paper copies can be thrown out — no taped fingers for them; I recommend those nice white jackets that fasten at the back. :-)


So wouldn't it be a little messy to select one of these textual buddy files from the search results, find the corresponding image file(s), and then open it? Well, no, not at all. It's extremely easy for some programmer to create a tiny program to do this for you, much like the code attached to the aforementioned 'Organising Photographs' article. Put simply, you could right-click on the buddy file you want, and select 'Open with <ProgramName>', and that program would find the image file for you and automatically open it in place of the text file. To be more bullet-proof, it would be best to use a special file type rather than *.txt (e.g *.meta), in which case the program could be registered as the one to always use for that file type, and you would merely have to double-click on the *.meta buddy file. There must be a commercial opportunity here.


A software tool was developed to demonstrate this basic principle under Windows, and it may be downloaded for free from Dropbox folder. No installation is necessary, and the folder includes both a PDF user guide and the original source code. A description of what it does may be found at Tool for Annotating Image Files.

Thursday, 15 May 2014

Ashes to Ashes, Artefacts to Bits



What do you do with artefacts in your family history collection? Are they connected to your software entities (e.g. in some database) and/or any digital images of the items?


An artefact is “an object made by a human being, typically one of cultural or historical interest”.[1] The word is occasionally confused with ephemera, which are:

“Things that exist or are used or enjoyed for only a short time”, and “Collectable items that were originally expected to have only short-term usefulness or popularity”.[2]

For the purposes of this genealogical blog-post, I will consider ephemera to be a subset of artefacts and will not mention them specifically.

Now that we’ve established the context, we should be able to see that this discussion is about physical items, including original photographs, actual letters, original documents, medals and other awards, paintings, clothing, jewellery, furniture, personal possessions, and family heirlooms. I’m particularly interested in this subject as I have several examples myself, including photographs (which many people may have, possibly stuffed in a biscuit tin), a police award, medals, military documents, an army uniform, and personal letters.

A recent post in a LinkedIn group caught my attention as it was about ‘Archiving your digital artifacts’. This confused me slightly since I would not have used that term to describe digital files. It was basically about their long-term storage (locally or in the cloud) and the preservation of their integrity and fidelity. It is true that there are preservation issues for digital data but this post conveniently skipped over the issue of real artefacts. More interestingly, though, a couple of responses suggested that people may be more interested in the digital editions because “…creating the copies is how we are able to preserve some physical artifacts. Some artifacts just don't last that long”. This belief contrasts sharply with the attitude in archives and museums where preservation of the originals is a prime objective.

So, should we be interested in preserving our own originals, or should we just endeavour to keep images of them? Preservation is not always easy, and very few of us are experts in that field. It is usually considered to be something of interest to those aforementioned institutions rather than to genealogists and family historians. Let’s look at probably the most common case: original photographs. If you’re sharing them with family and friends then digital copies, and other types of reproduction, are worthwhile and easily generated. If you want a copy for frequent consultation then a digital copy may also prevent excessive access to a delicate original. If you turn an original photograph over, though, then you may find invaluable annotation or notes in someone’s original hand. For instance, I have a picture here of a solider in 1915, the back of which is actually a postcard sent by him, from France, to his wife in Nottingham, England. Yes, an early form of “selfie”!

How, too, would a mere digital image convey the full shape, quality, texture, etc., of a wedding dress, or of an army uniform? Some of my relatives have a chess tabletop, carved in some type of stone — originally seated on a wooden table that’s now long gone — and dated “MDCCCLIX [1859] January I”. This was a present to my ancestor, Henry Procter, who is named at the top of it, and who was married a couple of months earlier. The initials at the bottom suggest that it was from a member of his wife’s family who I happen to know was a stone mason. The issue here is that although its preservation is easier, the digital images that I personally have of it do not do it justice. A project to recreate a supporting table is planned which would allow it to be put on show again, and possibly to enable it being used as a games table again.

Certainly, one feeling that underpins this fixation on digital copies is that it’s the information that’s being preserved. A scan of a photograph, or of a letter, preserves the information therein, suggesting that there’s no fundamental difference between the original and a good copy. There is some truth in this — otherwise we wouldn’t be content with those census scans that we all have — but anyone who considers their artefacts to be treasured memorabilia would disagree.

In Handling Transcriptions I explained that STEMMA® uses a Resource entity to describe both digital files and physical items, including any images of the physical items.


Some interesting fallout from this occurs when sharing such data. There may be many copies of your data — if you’re so inclined — but only one copy of the artefacts. When sharing your data, you will most likely be sharing just the digital contributions, and that means that any association between artefacts and images thereof must be broken.

If we’re famous then we may decide to bequeath our collection to a local archive, and presumably they would help with the issues of organisation and preservation. For the vast majority of us, though, we must learn how to become micro-archives.

This is a seriously neglected issue in private collections (i.e. outside of archives, museums, and libraries). I don’t have any answers for how to preserve our varieties of artefact, or the best practices for cataloguing them, but we should be learning from those institutions that do this. What I do know is that we’re not encouraged to record their presence in our family history collections, or even given the basic tools to accommodate them. The relentless march of commercial software is steadily turning genealogy into a digital-only world!



[1] Oxford Dictionaries Online (http://www.oxforddictionaries.com/definition/english/artefact : accessed 15 May 2014), s.v. “artefact; ‘artifact’ is the US spelling.
[2] Oxford Dictionaries Online (http://www.oxforddictionaries.com/definition/english/ephemera : accessed 15 May 2014), s.v. “ephemera.

Saturday, 12 April 2014

Handling Transcriptions


Making transcriptions of records is not as common amongst genealogists as you might expect, but why is that? What do we need in order to create useful transcriptions? If we’re part of the minority who do make them then where should we attach them?

Because of the availability of online data sources, and the ease with which digital copies can be created (owner permitting of course), many people believe they do not need full transcriptions of records. They might claim that since they can visit an online image, or they have a digital scan in their own data collection, then they can read it perfectly well without having it typed out. Whether it’s a baptism entry, a newspaper report, or a census page, many genealogists therefore find they have a growing collection of equivalent JPEG files sitting on the periphery of their data.

What I mean by this is that such a file can be pointed to, or referenced, by other data, but it cannot reference anything itself[1] or be textually searched. This means the information is not truly integrated into your data. The arguments for adding mark-up to a transcription in order to achieve this are almost exactly the same ones that I made for using mark-up in authored narrative at Semantic Tagging of Historical Data. This allows, for instance, references to people, places, events, dates, etc., in that transcription to be connected to the relevant entities in your data.

A transcription requires more though. It also requires a way of indicating transcription anomalies — parts that deviate from the normal flow — such as marginalia, footnotes, interlinear/intralinear notes, struck-out text, and uncertain characters or words. Both the uncertain characters and the uncertain words may require annotation to provide suggestions and possibilities, both of which must be honoured during searches. A transcription also requires an indication of any original emphasis, such as italics or underlining. NB: the original use of italics, underlining, footnotes, etc., in something being transcribed is different to their deliberate use in a written report, and so must use a distinct form of mark-up.

Traditional editorial notations for transcriptions are not well-suited to digital text as they do not facilitate efficient and accurate searching. TEI has comprehensive sets of mark-up for handling transcription issues but falls short when applied to genealogical data, and probably historical data in general. It is certain that some specialised mark-up is required, but how you visualise a transcription on-screen is a separate consideration. The same mark-up could alternatively show multi-coloured and hyper-linked text, or the plain editorial notation. That sort of flexibility only comes from using a computerised annotation rather than human annotation.

The fact that both transcription and authored narrative may co-exist in the same written report led to STEMMA® unifying them in its own mark-up. Those distinct usages — for transcriptions and for generating new narrative (e.g. essays, reports, inference, etc.) — have some similar and markedly different characteristics as follows:
  • Transcription (including transcribed extracts) — requires support for textual anomalies (uncertain characters, marginalia, footnotes, interlinear/intralinear notes), audio anomalies (noises, gestures, pauses), indications of alternative spellings/pronunciation/meanings, indications of different contributors, different styles or emphasis, and semantic mark-up for references to persons, places, groups, animals, events, and dates. The latter semantic mark-up also needs to clearly distinguish objective information (e.g. that a reference is to a person) from subjective information (e.g. a conclusion as to whom that person is).
  • Narrative work — requires support for layout and presentation. Descriptive mark-up captures the content and structure in a way that provides visualisation software with the ultimate control over its rendering  It needs to be able to generate references to known persons, places, and dates that result in a similar mark-up to that for transcriptions. The difference here is that a textual reference is being generated from the ID of a Person entity, say, as opposed to marking an existing textual reference and possibly linking it to a Person with a given ID. Also needs to be capable of generating reference-note citations and general discursive notes.
Actually, transcription isn’t just an action associated with a manuscript or typescript document; it could be associated with speech too. In those circumstances then it must reflect speech levels and emotional emphasis, but I haven’t even thought about that field yet.

As you can imagine, in order to generate a quality transcription, and to incorporate semantic links and annotation, a very good software tool is needed. It would be something like a specialised word-processor tool, but most of us are left using general-purpose word-processor tools that have none of the required facilities. This will be a secondary reason why so few transcriptions are made.

So where do I attach transcriptions in my own data? In order to explain, I first need to convey something of the structure of my data.

STEMMA Entity Linkage

This simplified view of the rich connections in the STEMMA tapestry doesn’t show its places, or groups, or lineage links between people, or hierarchical/protracted events. That would be too complex! What it does show is a network of multi-person events and the relationship of sources to those events. Notice that the sources are attached to the events, and not to the people. As already explained in Evidence and where to Stick It, the vast majority of our evidence – if not all of it – relates to events; things that happened in a particular place at a particular time. In other words, our entire view of history rests on discrete and disjointed pockets of evidence describing a finite set of events. Everything else is inference and interpolation creating as smooth a picture as we can.

So what is the general form of these underlying source entities in the data? Our real-life sources may be remote, such as a document in an archive or a book in a library, or local, such as a family letter or a photograph. In both cases, we may have a digital scan of the items. STEMMA[2] has two important concepts that it employs for sources:

  • Resource – This describes some item in your local data collection, including not just files on your disk, but also physical artefacts or ephemera.
  • Citation – Despite the name, this is merely a link to some source of information. A traditional printed citation may be generated from it, but this software entity also incorporates collections, repositories, and even attribution; possibly chaining them together.

Either or both of these may apply, therefore. A full transcription would be associated with the Resource entity that would describe any physical or digital edition of the associated material. In the case where you may have transcribed a document in an archive, or even from one of the online content providers, the transcription should still be placed in a Resource entity rather than a Citation entity, even though the latter is possible.

Genealogist Janice Sellers, in her blog-post at Transcription Mentioned on Television, explains how transcriptions of documents are valuable for sharing the details with family and friends. She recounts how she tried to convince a well-known British TV program to advise their guests to make transcriptions of their historical documents and heirlooms.

STEMMA’s mark-up is primarily about semantics. Shallow semantics would mark an item as, say, a person reference but without forming a conclusion about who the person was. Deep semantics involve cross-linking references to persons, places, groups, events, and dates, to the relevant entities in your data. I have previously tried to convey this using the worked example of an old family letter at Structured Narrative.

Genealogist Sue Adams has taken the concept of semantic mark-up in transcriptions to a deeper level on her Family Folklore Blog. Her worked examples clearly demonstrate the temporal nature of historical semantics. Anyone with a passing interest in the Semantic Web and RDF is encouraged to read about “temporal RDF” and consider why it doesn’t yet exist. You may find a lot of theoretical work that considers things like temporal graphs but very few real examples like hers. In an ideal world, the developers of such technology would be working closely with the people who need to utilise it.




[1] I’m ignoring the issue of meta-data held within an image until a future post. The issue here is one of the text in an image making discrete references to its subjects rather than anything to do with image cataloguing.
[2] STEMMA V2.2 — which includes important refinements here — has just been defined but, at the time of writing, I am still preparing to painstakingly update the Web site. The landing page will indicate when this is complete.