Saturday, 18 December 2021

Genealogical Journalism

 

In these days of massive transcription projects for historical newspapers, genealogists have realised just how much of a goldmine they are for micro-history, in addition to national or global history, but what does the future hold for them?

NB: This article is written from personal knowledge of British and Irish newspapers, and presumes that there are similar issues elsewhere.


 

I have found the material in old newspapers invaluable for reconstructing ancestral history, and not just for reconstructing families. For instance, articles such as A Sad Career have received very positive feedback from people involved with these projects, and from companies hosting newspaper databases. This style of research is akin to journalism, but is more accurately a form of meta-journalism that puts together an historical account from multiple journalistic articles of the time.

One of the biggest facilitators of this, for me, has been the British Newspaper Archive (BNA), which has transcribed an enormous number of national, regional, and specialist papers. They weren't the first, though, and my earlier research relied on the Gale database of '19th Century British Library Newspapers’, which used to offer free online access (i.e. connectivity from home) through libraries. The Gale database is now almost impossible to find because it has been displaced by the BNA database. Libraries found it cheaper to include BNA access, and yet the agreement was not the same: access requiring you to be physically present in the library.

Another difference is the available search operators. Finding specific material needs search operators that allow you to precisely specify your subject matter, and eliminate stuff you're not interested in. This is crucially important if your subject of interest just happens to contain words that appear in a different but more-prevalent context. To this day, the Gale search operators were better, but the range of newspapers is far better in the BNA. Unfortunately, the more relaxed search criteria of engines such as Google have set a trend that means you are more likely to be deluged with the wrong stuff. In essence, you're invited to enter a bunch of words and some software algorithm then decides which search results would be more appropriate.

But what do we mean when we say micro-history? Well, papers from the 19th century reported much more local news than modern papers do, and regional papers would often cover every little misdemeanour, transgression, accident, celebration, or other event. There was a very high chance of finding references to an ancestor, their family, or even something happening on their street; the events did not have to be of national importance. This provides a very rich substrate in which to find events of their everyday lives, and local events that may have affected their lives. For instance, in Jesson Lesson, a case is made for the movements of a family having been forced by industrial strife and change in the work of framework knitters.

In Where is Nottingham Castle?, an account of riots that destroyed the Castle was wholly constructed from newspaper accounts, leading to a follow-up article revealing the life of one poor sole who demonstrated such bravery on the gallows in the aftermath.

This degree of detail persisted into the early 20th century, but it can be seen to be diminishing after WWII. This is where the future becomes uncertain for newspaper archives because the BNA has only transcribed to about 1950, and I am trying to get a statement about why this is so, or what the future intention is. I am sure I recollect a comment that copyright is an issue, but I have no record of this that I can cite. Assuming that the BNA will transcribe more recent editions, and so document the more-modern events that directly affected our own lives, then the papers will not be the same

Many papers have since been consolidated such that the remaining local papers are usually no more than free papers containing less news than advertising. The real newspapers have become bigger and so much more focused on national and global events. They are not going to send reporters into the inner-city streets, or out into the provinces, unless there is a fairly big or important story to capture.

But the nature of the papers has changed in addition to their ambit. Many national papers now have their own digital archive, which could make the work of non-proprietary archives easier — assuming they want to extend their range, and that some agreement could be reached in sharing the data. A more subtle change has been the shift from objective news to opinion: editorials are bigger and each paper has a significant letters page for public opinion. This is interesting because future researchers will no longer be building opinions of their own solely from newspaper accounts (whether subjective or not), but will have to consider opinions of the time. These will clearly be very subjective, and so requiring a more critical eye, but it can be argued that they will give a multi-dimensional perspective.

So, if our existing newspaper archives later cover years beyond 1950 then we are not going to see the same level of detail since there will be fewer papers, and a gradual loss of micro-history. We can expect to see coverage of the race riots through 1950s Britain, and the social, political, and cultural changes that we now think of as the "sixties", but not the fact that your ancestor was fined 6 shillings for making a noise after leaving the pub in a state of intoxication, or someone on your ancestor's street was in court for not emptying their "night soil" correctly.

If not in the papers then where might future genealogists look for micro-history? A strong candidate would be social media, although far from being objective. Unfortunately, much content of this ilk is considered ephemeral and transient, and may never be preserved. Even comments on special-interest groups are unlikely to get saved (remember Google+?) and so the future looks grim for our inconsequential trivia, but this is likely to be a future blog topic.

 

[Edit 22 Dec 2021: After three exchanges with BNA, trying to get a statement on the coverage of further dates (beyond 1950), I gave up. They sent me stock replies about covering more titles, and also suggested contacting the British Library for editions they don't yet cover, but seemed to miss the fact that this was for an article actually about newspaper archives.

However, looking at their knowledge base, it seems that copyright is supposedly a problem: "We currently have more newspapers from the 1800s than the 1900s because we need to get permission from the copyright holders to publish more modern content. We are constantly working on this, so over time, you will see more appear." I can understand that each batch of new dates has to be agreed with the copyright holder first, but this does not really explain why there are virtually no editions of any paper after 1950.

On the subject of copyright, there was an interesting knowledge-base article about making transcriptions: "If you plan on creating your own transcripts from the newspapers, then please be advised that unsigned newspaper text goes out of copyright 70 calendar years after the year of publication, and signed newspaper text goes out of copyright 70 calendar years after the death of the author(s)", which would coincidentally be applicable to unsigned text from the early 1950s.]

 

[Edit 23 Dec 2021: Aha, someone finally understood. Here's their full reply: "I've done some digging and spoken to the team to try and get a bit more information so I hope the below helps: Copyright around newspaper content is complex. There is not just the ownership of the content that the publisher holds to consider, but the nature of the material within the newspaper itself like adverts, photographs, illustrations and other content. When it comes to more modern newspapers, such as those from the 1960s and 70s, we work with current copyright holders and publishers of those newspapers and they are the ones who decide what we can publish on The British Newspaper Archive as they are the ones who will hold copyright on the digitised newspapers that we add to our Archive. When it comes to adding more modern content from newspapers we have already digitised it just comes down to the nature of our agreement to publish the newspapers in our Archive with the current copyright holders AKA the publisher, which can differ from publisher to publisher." This still doesn't explain why it seems to be across the board (i.e. all titles), unless the consolidation I mentioned means there is only a very few copyright holders involved.]

Saturday, 13 November 2021

SVG Family-Tree Generator (v6.1)

The release of SVG-FTG V6.0, back in March, has been very successful. There was an article in the UK Family Tree magazine, and an online interview with me. Details of availability can be found on the summary page: SVG-FTG Summary.

It is now time to release V6.1. The major feature in this version is an "album" of all your images that can be browsed or searched from your web browser. This feature was actually deferred from V6.0 because of the enormous number of other features and improvements already in there.

Although an experienced software developer, I am also a genealogist and a user of this software. Hence, new features are usually added because I have a genuine need for them, and the software is made freely available both because it was never written for commercial gain and because I expect many other genealogists to have similar goals.

One of these goals relates to images. SVG-FTG V6.0 supported thumbnail images in the person-boxes, and arbitrary images in the biographical or historical notes for either persons or places; however, I needed something more than this. Expecting all images to be relevant to a specific person, or even to a family, would put SVG-FTG in my own firing line that criticises other software. I have images that relate to groups (e.g. weddings or parties) and some that relate only to places. I wanted to be able to catalogue them such that I could search for people and/or places, and view associated textual notes, or simply to browse them. I wanted this to be possible on my local computer or on a website. But most importantly, I didn't want my images modified, renamed, copied, or moved in any way from my originals — a requirement that all responsible genealogists will immediately understand.

Image Album

SVG-FTG 6.1 provides a tool to maintain distinct groups of images (termed "galleries") and add meta-data to the individual images. This can include a caption, a place description, associations to persons in the current tree, and general textual notes. The images can be of anything you want, including groups and document scans, and so are not constrained to be of specific persons.

At a point of your choosing, all the images in all your galleries can be indexed for presentation and searching using a Web browser; this indexed set is termed an "album".

NB: None of your images are moved or modified in anyway. Separate storage records all this meta-data, but it is indexed so that your images can be viewed using a typical filmstrip tool, and searched from an interactive application added to your family tree.

There are two separate pieces of software that benefits from this meta-data:

  • The first is a small tool for displaying the properties of image files in a Windows directory. It is called MetaProxy, and was described previously at MetaProxy V3.0. It is a free tool that requires no installation, and it links images and their meta-data properties so that they are automatically displayed together. MetaProxy is not an essential tool, but it provides a very simple way of capitalising on all the information you would have added for each image, effectively making your Windows directories a bit more sophisticated than simple groups of named image files.
  • The second bit of software is a new SVG-FTG interactive application, called 'Image Album'. The application talks to an album filmstrip — displayed in another tab — so that searches can be performed and results displayed in the filmstrip. For instance, I could request that all the images showing my old school be presented, or just the pictures of my school with me in them, or pictures of my school with either/both my brother and me in them. Selecting one of the images from filmstrip shows any textual details you may have added for it.

When the interactive application is added to your tree, you will usually get an extra 'Album' button in the bottom-left of each person-box. This button supports the following operations:

  • Click: Selects (or deselects) the current person so that they can be referenced by a subsequent search.
  • Ctrl+Click: Clears all the highlights for currently selected persons.
  • Shift+Click: Launches the search box.
  • Alt+Click: Tells the album filmstrip tab to present all images, from all galleries (i.e. ignoring any search criteria).

The search box allows you to include any/all/none of the selected persons, a specific place description, and a gallery name. For instance:

 


When you press the 'Search' button then the meta-data is searched in the current browser tab. If there are no matches then you will see a message, otherwise the album filmstrip is asked to present the matching images. The album filmstrip is automatically launched in a separate tab, and if you are using viewpoints then they will all communicate with the same filmstrip tab.

The image filmstrip shows how many images there are in total and which of those is currently shown in the main panel. There are buttons to the left and right allowing you to cycle through the current images. Alternatively, you can click on a specific image in the strip of images in the lower panel.

 


In response to a search then the list of images is reduced to just those matching the search criteria. Clicking on the main image will show (or hide) a panel of text showing the notes that you created for the image.


The full list of images can be restored by clicking on the album title, or from your tree by using the search box or Alt+Click operations.

Help Text

All of the forms that have a menu bar will now show an extra option: Help. Selecting this will display a simple page of help relevant to that particular form.

GEDCOM V7.0

After a very long hiatus, GEDCOM 7 was released by Family Search this year. The goal has been to create a more precise and unambiguous specification, and to provide a more solid foundation for the representation of family trees.

SVG-FTG V6.1 imports both GEDCOM 7 and 5.5(.1), and offers a choice of exporting as GEDCOM 7 or 5.5.1. This will not be of much use until other companies support the new GEDCOM specification.

Image Scaling

In response to a question from one of the SVG-FTG users, I've had to look at image scaling. In V6.0, when expanding a thumbnail image into another browser tab, it was scaled to fit the window size, but if the image was a small low-resolution one then this could result in a very blurry presentation.

The low level function that handles this (su_showFile) has now been modified so that an image will be displayed original size if it will fit, otherwise it is scaled down to fit while retaining its aspect ratio.

Non-Standard GEDCOM Dates

By far, the biggest issue with importing GEDCOM files has been the near-ubiquitous inclusion of non-standard dates by the software that generates them, i.e. date formats that clearly do not conform to the GEDCOM specification. It is too early to tell whether GEDCOM 7 will help, or how databases of stored dates — especially those held by online genealogy software — will be made compliant.

SVG-FTG V6.1 now accepts many of the most common date variations in order to make the import a smoother operation, and lessening the need to manually "cleanse" the data before loading.

Thursday, 30 September 2021

A Citation Generator

 

[NB: For the latest information, see Decision Tree Generator.]

This title will have hopefully grabbed the attention of many genealogists. The article does describe a citation generator, and it will be demonstrated below — directly in this blog — but the true subject matter is considerably more general: a decision tree generator.


Some time ago, I was asked about the possibility of a Web-based application to guide a user around resources, and to direct them to where they want to go. The idea was for it to ask relevant questions and then present them with the information they need, together with hyperlinks for them to select. The example I was initially given was to locate sources for vital events in different regions of the world. This is a fairly straightforward and useful idea, but the general idea has many more possibilities.

Unfortunately, the reality of such systems can be complex. Even a sequence of just twenty questions (as in the game) could result in a million possible answers, and nearly a million distinct questions. Trying to design and maintain such a system requires specialised software.

Although relevant software products already exist, I offered to write one for free because (a) I had already developed graphical tools for SVG-FTG that could be re-used, (b) I could see many different applications for this if designed properly, and (c) I wanted something that was simple but customisable. Decision trees can be used to guide an end-user around resources, find something in a catalogue, identify something, implement a hierarchical help system, or construct something (e.g. a citation string). Ideally, the tool used for design and maintenance should be blind to the final usage, and be equally applicable to all these cases and others.

The resulting software (D-Tree) is a free tool that allows you to graphically design a question-answer decision tree that can incorporate yes-no questions, multiple-choice menus, input fields, and more. Following the SVG-FTG precedent, at any point during the design, the user can ask to generate an HTML implementation of the tree that could be tested in their default browser, and later deployed on the Internet. Since there are many items of software that may be described as a decision tree then this description is important; D-Tree is a design tool that helps you create and maintain decision trees, and generate HTML implementations of them when required.

In a nutshell, the questions and input fields gather a set of information from the end-user during the traversal of the tree. At an endpoint to the traversal, that information can be presented to the user, or used to give an overall answer, direction, or advice.

Where is D-Tree?

Users of SVG-FTG will see many similarities in the appearance, functionality, and installation of D-Tree. The distribution kit, PDF documentation, and some samples (including the citation generator demonstrated here) are described on the summary page: D-Tree Summary.

The PDF documentation includes a user guide, a guide to the internals for people who want to add extra node types and customise things, and an installation guide.

What does D-Tree look like?

D-Tree presents a canvas upon which flowchart symbols can be placed and moved around. Such symbols are not as common as they used to be, but there is a standard and D-Tree tries to stay compatible with it. Lines can be drawn between these symbols to represent the possible flow of control.


Many of the functions provided in D-Tree will be similar, if not identical, to the ones in SVG-FTG. For instance:

  • Manual arrangement of nodes or automatic layout.
  • Finding nodes by caption or by key.
  • Multiple selection of nodes for moving or deleting.
  • Copy and paste of nodes.
  • Zoom levels for viewing more or less detail in the tree.
  • Array of settings to modify the generated tree.

As with SVG-FTG, the designer (D-Tree) is Windows-based but the generated HTML should work in all modern browsers.

Citation Generator

One of the applications that I wanted to use D-Tree for was to design a Q&A approach to generating citations. Yes, there are several citation generators around, but all the ones I am aware of consist of a large number of templates into which your citation parameters are inserted. Selecting the appropriate template for a given situation can be a genuine hassle because there are so many possible nuances to the basic array of source types. Even with a hierarchical index, it can be hard to translate your knowledge of the citation scenario into a subset of applicable templates, whether using a flat index or otherwise.

Evidence Explained[1] contains a huge index of its quick-reference guides, but anyone who has actually read the book will know that the chapters give solid advice on how to handle the nuances; the quick-reference pages are merely a guide to the most common situations in each source category.

Below is a simple generator for reference note citations to a published book. This is not a product itself and so don't get hung up about missing nuances or the colour choice. It is a demonstration of the possibilities, but also a starting point for people to develop their own. In other words, you can take this free design tool (D-Tree) and write your own, better, citation generator. Or, you can take this illustration as a starting point and add to it since the corresponding tree definition file is one of the samples in the installation kit.

The questions take you through some of the nuances specific to this one source type. At each point in the traversal of the tree, the 'Help' button will display relevant help text, and this is all part of the design available in that sample file.

The HTML implementation of the decision tree uses CSS to style it. The standard CSS file defines an area in the centre of your screen for the tree traversal, which is good for local testing but may not be what you want when integrating it into your website. The standard CSS also use a blue theme for the main tree and a green theme for the hierarchical help, as demonstrated by the example above. But you can easily nominate a CSS file of your own to change colours, fonts, frame sizes, question and menu layouts, input fields, etc., and ensure, for instance, that the result works equally well on a hand-held device.

So what if you also use a template-based system? Well, it could actually dovetail with these trees, if its design is open and flexible enough. Rather than the terminal nodes (i.e. endpoints) of the traversal using the gathered information to construct a citation string, it could be passed on to another system to generate one.



[1] Elizabeth Shown Mills, Evidence Explained: Citing History Sources from Artifacts to Cyberspace (Baltimore, Maryland: Genealogical Pub. Co., 2009).

Sunday, 8 August 2021

Stop Adding Incorrect Details for My Family

 

This is a woefully common reaction in online genealogy. Whether someone is overwriting your details in a shared tree, or publishing incorrect details about your family in their own tree, it can be frustrating when you see incorrect details published and then blindly copied by others.

But wait a moment, why does your opinion carry more weight than theirs? If you politely tell someone that their details are wrong, and they either ignore you or tell you to go forth and multiply, why would you expect them to change anything? If you provide them with a reference to a census page or parish record (or other) and they say ‘Oh, I already have a John Smith in my tree, so thanks but no thanks’ then you may begin to see the problem.

Yes, there are a few edge-cases, such as someone telling you your own father was Peter rather than Paul, or insisting that you were born in Kathmandu rather than the East End of London, but we are looking at the more general case here.

Far too often, a record by itself does not sufficiently identify a person, or their relationships to other people. Even a group of corroborating records may be insufficient to connect them. But useful evidence is not necessarily the same as a direct answer; to show that John Smith1 cannot be right, yet without confirming that John Smith2 is correct, is still a useful product of your research. In a surprisingly large number of cases, there will be no direct evidence, and the naive expectation fuelled by advertising that you can build your tree from online records, becomes a sad disappointment. To get past such brick-walls, you need to get into inferential genealogy, and this then implies that you need to write up the how and why of your research for it to be accepted. In fact this is a general tenet that applies to all claims, with more details being required where there was more inference. If you need to convince another researcher that their information is incorrect then you need to provide such an explanation.

This is a lesson yet to be learned by the hosts of online records and online trees. Their insistence on the simplistic research model of ‘search here and build it there’, compounded by ‘if you cannot find it, copy it’, has already done untold damage to our collective research.

A few years ago, I drew an analogy between research in academic fields and in online trees at Research in Online Trees, pointing out that in all academic fields (including professional and academic genealogy), it is not a direct A-to-B process. There are intermediate steps where researchers may build upon, or even challenge, past work by other researchers. But for that to happen, dedicated researchers need a way to write up and publish their research for others to find.

Note that attaching  a write up to a particular person or family in your tree is not good enough: even if the text is searchable via search engines such as Google (which it rarely is because it is stuck in a database) then it may not be specific to that person or family. It could involve multiple generations in more than one family. As with those other academic fields, a comprehensive write up will also need basic formatting, and especially a mechanism for footnotes and citations.

Many researchers who fall into this category, and who are not solely interested in academic journals, will resort to blogs. These can be highly effective ways of making such arguments, but they are ignored by the hosts of online trees. Yes, they can be cited (sort of), but you are forced to find them using a search engine. If your surnames also happen to be common terms then this is not going to be very productive. Unsurprisingly, I have explained how these hosts can easily fill that gap, allowing their users to search registered blogs along with their own online records, and without diverting traffic away from those blogs or risking any type of copyright: Blogs As Genealogical Sources. Reactions — few as they were — implied the idea was too complicated, or that it would cost too much money to implement, or that it would involve payments to the blog authors, or simply that no one else does it. All of these notions were ill-founded, but their lack of both vision and responsibility will leave us wanting in terms of a more robust model for sharing and collaboration in general.

Monday, 2 August 2021

Is Pinterest a Valid Source?

 

At the end of 2019, I made a case for online trees being a valid source, although with some caveats. I recently thought about a similar case for Pinterest, but the situation is not the same there so I wanted to dig over the main issues.


 Like many people, I started a Pinterest account when it first appeared, and then got disinterested when I realised that it was all smoke and mirrors (or images thereof), that my feed could not be tailored to deliver what I really wanted to see, and when I got deluged by unwanted advertising. In fact, I have just closed my account as it had no value for me.

Pinterest has been criticised for many reasons, including some content being pornographic or obscene, being overtly political, hosting commercial scams, spreading misinformation (especially medical), or focusing on people's eating disorders or weight problems. But what was it supposed to be?

According to Wikipedia, Pinterest is an "American image sharing and social media service designed to enable saving and discovery of information ... and ideas", but the reality is much more mundane. It is now basically an image sharing site with no obvious purpose. You see images were supposed to be just a taster that encouraged people to pin them, and click on them to get information; an image on its own — with no caption, link, or accompanying information — is a dead-end.

Let me pick a specific case, one that initially encouraged me to look at Pinterest: images of old places. I love to see historical pictures of my home town, but on Pinterest they invariably contain no details, or any caption. If I wanted to search for an image of a particular place then I cannot — the search bar simply finds boards of that name shared by other users. If I happen upon a rare or interesting image there then I might, if I'm lucky, recognise the place, but what about the date, or the photographer, or the story behind the picture?

If I was doing this as part of a research project then I have a deeper problem: provenance. Where did the image come from? Who took the original, and is it in copyright? Pinterest does have a mechanism for a copyright holder to get material taken down, but this is fighting against the tide because it already makes it so easy for people to share anything and everything that they might find, online or otherwise. At the very least, it should have implemented a mechanism identifying the initial point of entry of an image onto Pinterest, i.e. who first loaded it, and where from.

The situation is more complicated than this, though, and doomed to failure in the hands of people who treat it like stamp collecting. I have several images in my blog posts that I have taken pains to get permission to display from the copyright holders, and I had shared those same posts via Pinterest using those images. And yet I have these images in isolation on other people's Pinterest boards. That is, pinned images, divorced from my blog, without the associated information, and without any provenance or attribution. The images had been appropriated to sit in someone's "gallery" of images that they like, but that serves no purpose beyond the private pleasure of such hoarders.

So, if Google turns up some image during your research that resides on Pinterest, what do you do? Would there be any point in citing it at all, in the way you might for other social media? Google does provide a search-by-image mechanism through which you might be able to identify a non-Pinterest copy — ideally being older and with more details — but then a Google search could equally have found that, so what purpose does Pinterest serve? As a means of sharing, it is naively structured and simply exacerbates sharing issues already present on the Internet. But as a source of information that is worth reading and citing then it is a non-starter.