Showing posts with label Research. Show all posts
Showing posts with label Research. Show all posts

Thursday, 7 April 2022

What a Mess

This will be my last blog post for the foreseeable future, and probably forever. This is not a matter of free time, or of advancing years, or even of competing tasks, but of a complete disillusionment with modern genealogy.

I will continue to research in my spare time, but this will be to produce something of standing for my family and extended family to read; I have lost faith in the public world of genealogy. But let me explain in more detail because this has been on the horizon for a while, and yet previously planned retreats have all fallen down for different reasons.

I am use to academic research, and how academic research works in other fields. The purpose of research in those other fields is to find answers — truths — and to produce a valuable collective body of work through collaboration. Virtually by definition, it is not a commercial goal.

I have previously pondered over the nature of genealogy (What is Genealogy?), and considered its difference from family history, but there is a more systemic difference that touches on collaboration, software, and commercial forces. Although genealogy has a well-respected academic side, it generally considers the internet, and digital resources in general, as only good for derivative sources such as images and transcriptions, and not for publications. This field has high standards and produces quality work in traditional publications such as books and journals, but the internet is considered inappropriate for publication due to its transient and ephemeral nature.

If we look to commercial genealogy then we see two quite different worlds: that of the generation of derivative sources that people can search, and that of online trees. Other than improving the search tools and options, I have no real criticisms of the many digitisation and transcription projects, but for online trees then I have many. In fact I have written so many articles on this subject that I won't even begin to enumerate links to them. Irrespective of whether we are considering "unified trees" or "user-owned trees", there are fundamental issues with their structure and the process by which they are generated.

In terms of structure, a tree is appropriate for representing biological lineage, but dreadful for representing history — can you imagine a family tree attempting to detail, say, the events of WWII? But non-biological lineage, such as fostering and adoption, or even weaker associations between people, break this visualisation and can result in a cat's cradle of complexity and confusion. A tree is also limiting in terms of proof arguments (particularly if they reference multiple individuals, families or generations), citations that refer to actual claims (as opposed to simple hyperlinks saying where you got your information), and linking to external resources (images or document scans) that are not specific to single individuals.

But worse than this is the process by which we are expected to construct such trees. We are all probably aware now of the variable quality of trees — although I still find it vexing when I see 'trees are not a valid source' (it depends on the claim) — and that trees can persist online long after someone may have dabbled for a few months using a free trial or a subscription birthday present. There is no responsibility taken by the respective companies for the accuracy of what their subscribers publish, and they appear to be disinterested in why academia looks down on these published works. It is impractical for these companies to fact-check stuff, and so I am not suggesting that is the solution, but they do not acknowledge, publicly, that the simple paradigm of building trees directly from their raw digitised records is naive (despite their advertising). There are many difficult cases of family reconstruction that require effort — possibly an enormous amount of effort — to get around missing information, ambiguous information, or even deliberately obfuscated information, and so make a case for what really happened in the past.

But two experienced researchers might reach different conclusions, both of which appear to fit available information, and so how should that be dealt with? Well, the red mist and edit wars commonly associated with "unified trees" are not the answer. If left to software people then they might suggest transactional get and commit operations, analogous to those in software source-control systems. If you don't know what these are then it's probably best not to ask; they're complicated, generally with horrible user interfaces, and even get software people into trouble.

Well, why don't these companies look at how collaborative research works everywhere else? I can't believe that they're ignorant of it, and so I can only assume that they fear it would be too complicated for their subscribers, or that it would cost them money, or even that it's just a huge step into the unknown and they don't want to kill their cash cow.

Collaborative research elsewhere is not a linear one-step 'raw-data leading to final conclusions'; it's stepwise, and involving prior work by other researchers. Researchers can then look at the work of others and build from it (or refute it). This means real written work, with real citations, is a starting point as claims have to be justified, not just by pointing to data that appears to confirm them, but by explaining why, and why not something else.

OK, so not everyone will be able or willing to produce such written work, but there are people who do, and regularly do so: bloggers. I have already made a case that online genealogy companies could take advantage of this in a way that requires minimal investment, would not run into copyright or attribution problems, and would increase traffic to the respective blogs — surely, a win-win (Blogs as Genealogical Sources). Briefly summarised, the author of a blog article would give permission to the genealogical company to list the corresponding URL in one of their databases, and would provide meta-data to ensure that it showed up in the results of appropriate searches. The genealogy company would store such information in a database of so-called authored works (i.e. the URL, name of author, article title, and meta-data), but would not copy the body of the works. When these works showed up in a genealogical search, the end-user would click on one of them in order to be directed to the original blog article.

Yes, there would be some smaller issues such as the rating these works, or citing them, and so on, but it's academic as there has been no subsequent engagement — Zero, with a capital Z — by any of the companies, including the ones I approached directly.

Modern genealogists rely on the search functions within these online companies, and possibly on Google (although woe betide we have to research a surname such as 'covid'), but they would be less likely to find relevant printed books or journal articles. This sort of scheme could even be extended to cover non-internet sources, but there is yet another possibility, one that flies in the face of the view that research has to be written up in paper-based journals.

People who have researched in other fields may be aware of sites such as arXiv.org (the 'X' is actually representing the Greek letter chi, and so the site name is pronounced as "archive"). These contain online articles, submitted online, and viewed online. They are much more accessible and searchable than the old paper-based journals, and it is entirely possible that this could be done for genealogical research, but it would take the initiative away from any forward-looking genealogy company. Does that matter to genealogists? Probably not as there are many searchable resources that do not fall under their control. Would it contribute to the accuracy and a truly collaborative approach in modern genealogy?

I wish I could be optimistic here, but I'm not!

Tuesday, 25 January 2022

Crystal Balls

 

About the time of my previous post on Genealogical Software, I came across a forward-looking article by Dick Eastman that mentioned push-button genealogy: What I See in My Crystal Ball: The Future of Genealogy Research. After picking myself up off the floor and composing myself, I just had to comment on the fallacy of this.


In all fairness, these mentions were really talking about automated record matching, and the article acknowledges that the process is not quite there yet, especially not "for all ancestors". However, in this vision of the future, we would eventually be able "to push a button and see a filled-in pedigree chart within seconds". So what's wrong with this picture?

Well, it assumes that records are clear, unambiguous, and cover everyone. Of the two-dozen or so research articles that I have published, more than 75% involve difficult cases of establishing identities and relationships where those facts were either disguised, deliberately hidden, or simply not documented in the records. If someone changed their name, or tried to conceal the parentage of a birth, then the "reliable records" are not going to help, no matter how much the developers "fine-tune the algorithms".

At best, this push-button genealogy will generate a valuable set of hints, and we all know how unreliable such hints can be. But on the whole, wouldn't the list be mostly accurate, once the technology has matured more; wouldn't it only be a minority of cases that needed the heavy-lifting? Well, no, once you have an error then it propagates downwards through subsequent generations, and all would have to be unpicked if that error ever got detected. As they say: one bad apple can spoil the bunch.

Difficult identity cases need a wider perspective (basically the FAN principle), and a consideration of both context and non-written sources. If you ignore such things as ephemera or oral history, instead expecting direct answers to be found in the so-called reliable records, then you will fail at some point.

This all reminds me of the case that got me into genealogy, a case recounted on a very early episode of Mondays with Myrt. This was a promise to my mother that I would find her estranged five siblings. They were all taken away from the family home back in 1947, and then sent to various foster homes or officially adopted. My mother was the oldest at the time (9 years), and it had been so long that she could only just recall their names and ages.

My first plan was to hire a professional but they told me that it simply wasn't possible. You see, all the relevant records were destroyed within 25 years of the event, and so there were no records to go on. I didn't accept the impossibility of this endeavour because it was important to me, and to my mother. My mother's natural parents (who I found had divorced shortly after the event) had passed on well before I started, and so my research had to identify extended family (aunts, uncles, cousins) and actually talk to real people, much as a detective would do. I did consult records but they were not obvious ones, and only selected on the basis of cases made from such testimony. Most of those siblings had changed their name (two had even changed their given name as well as their surname), but the effort was successful: they were reunited on my mother's 70th birthday.

Maybe you can now understand why I take such developer-based claims with a large pinch of salt. Genealogical research needs real people, and its evolution is not helped by wild and misdirected claims about what software can do for them.

Saturday, 18 December 2021

Genealogical Journalism

 

In these days of massive transcription projects for historical newspapers, genealogists have realised just how much of a goldmine they are for micro-history, in addition to national or global history, but what does the future hold for them?

NB: This article is written from personal knowledge of British and Irish newspapers, and presumes that there are similar issues elsewhere.


 

I have found the material in old newspapers invaluable for reconstructing ancestral history, and not just for reconstructing families. For instance, articles such as A Sad Career have received very positive feedback from people involved with these projects, and from companies hosting newspaper databases. This style of research is akin to journalism, but is more accurately a form of meta-journalism that puts together an historical account from multiple journalistic articles of the time.

One of the biggest facilitators of this, for me, has been the British Newspaper Archive (BNA), which has transcribed an enormous number of national, regional, and specialist papers. They weren't the first, though, and my earlier research relied on the Gale database of '19th Century British Library Newspapers’, which used to offer free online access (i.e. connectivity from home) through libraries. The Gale database is now almost impossible to find because it has been displaced by the BNA database. Libraries found it cheaper to include BNA access, and yet the agreement was not the same: access requiring you to be physically present in the library.

Another difference is the available search operators. Finding specific material needs search operators that allow you to precisely specify your subject matter, and eliminate stuff you're not interested in. This is crucially important if your subject of interest just happens to contain words that appear in a different but more-prevalent context. To this day, the Gale search operators were better, but the range of newspapers is far better in the BNA. Unfortunately, the more relaxed search criteria of engines such as Google have set a trend that means you are more likely to be deluged with the wrong stuff. In essence, you're invited to enter a bunch of words and some software algorithm then decides which search results would be more appropriate.

But what do we mean when we say micro-history? Well, papers from the 19th century reported much more local news than modern papers do, and regional papers would often cover every little misdemeanour, transgression, accident, celebration, or other event. There was a very high chance of finding references to an ancestor, their family, or even something happening on their street; the events did not have to be of national importance. This provides a very rich substrate in which to find events of their everyday lives, and local events that may have affected their lives. For instance, in Jesson Lesson, a case is made for the movements of a family having been forced by industrial strife and change in the work of framework knitters.

In Where is Nottingham Castle?, an account of riots that destroyed the Castle was wholly constructed from newspaper accounts, leading to a follow-up article revealing the life of one poor sole who demonstrated such bravery on the gallows in the aftermath.

This degree of detail persisted into the early 20th century, but it can be seen to be diminishing after WWII. This is where the future becomes uncertain for newspaper archives because the BNA has only transcribed to about 1950, and I am trying to get a statement about why this is so, or what the future intention is. I am sure I recollect a comment that copyright is an issue, but I have no record of this that I can cite. Assuming that the BNA will transcribe more recent editions, and so document the more-modern events that directly affected our own lives, then the papers will not be the same

Many papers have since been consolidated such that the remaining local papers are usually no more than free papers containing less news than advertising. The real newspapers have become bigger and so much more focused on national and global events. They are not going to send reporters into the inner-city streets, or out into the provinces, unless there is a fairly big or important story to capture.

But the nature of the papers has changed in addition to their ambit. Many national papers now have their own digital archive, which could make the work of non-proprietary archives easier — assuming they want to extend their range, and that some agreement could be reached in sharing the data. A more subtle change has been the shift from objective news to opinion: editorials are bigger and each paper has a significant letters page for public opinion. This is interesting because future researchers will no longer be building opinions of their own solely from newspaper accounts (whether subjective or not), but will have to consider opinions of the time. These will clearly be very subjective, and so requiring a more critical eye, but it can be argued that they will give a multi-dimensional perspective.

So, if our existing newspaper archives later cover years beyond 1950 then we are not going to see the same level of detail since there will be fewer papers, and a gradual loss of micro-history. We can expect to see coverage of the race riots through 1950s Britain, and the social, political, and cultural changes that we now think of as the "sixties", but not the fact that your ancestor was fined 6 shillings for making a noise after leaving the pub in a state of intoxication, or someone on your ancestor's street was in court for not emptying their "night soil" correctly.

If not in the papers then where might future genealogists look for micro-history? A strong candidate would be social media, although far from being objective. Unfortunately, much content of this ilk is considered ephemeral and transient, and may never be preserved. Even comments on special-interest groups are unlikely to get saved (remember Google+?) and so the future looks grim for our inconsequential trivia, but this is likely to be a future blog topic.

 

[Edit 22 Dec 2021: After three exchanges with BNA, trying to get a statement on the coverage of further dates (beyond 1950), I gave up. They sent me stock replies about covering more titles, and also suggested contacting the British Library for editions they don't yet cover, but seemed to miss the fact that this was for an article actually about newspaper archives.

However, looking at their knowledge base, it seems that copyright is supposedly a problem: "We currently have more newspapers from the 1800s than the 1900s because we need to get permission from the copyright holders to publish more modern content. We are constantly working on this, so over time, you will see more appear." I can understand that each batch of new dates has to be agreed with the copyright holder first, but this does not really explain why there are virtually no editions of any paper after 1950.

On the subject of copyright, there was an interesting knowledge-base article about making transcriptions: "If you plan on creating your own transcripts from the newspapers, then please be advised that unsigned newspaper text goes out of copyright 70 calendar years after the year of publication, and signed newspaper text goes out of copyright 70 calendar years after the death of the author(s)", which would coincidentally be applicable to unsigned text from the early 1950s.]

 

[Edit 23 Dec 2021: Aha, someone finally understood. Here's their full reply: "I've done some digging and spoken to the team to try and get a bit more information so I hope the below helps: Copyright around newspaper content is complex. There is not just the ownership of the content that the publisher holds to consider, but the nature of the material within the newspaper itself like adverts, photographs, illustrations and other content. When it comes to more modern newspapers, such as those from the 1960s and 70s, we work with current copyright holders and publishers of those newspapers and they are the ones who decide what we can publish on The British Newspaper Archive as they are the ones who will hold copyright on the digitised newspapers that we add to our Archive. When it comes to adding more modern content from newspapers we have already digitised it just comes down to the nature of our agreement to publish the newspapers in our Archive with the current copyright holders AKA the publisher, which can differ from publisher to publisher." This still doesn't explain why it seems to be across the board (i.e. all titles), unless the consolidation I mentioned means there is only a very few copyright holders involved.]

Saturday, 5 June 2021

When a Place Unlocks a Bit of History

This small post recounts a heart-warming case where knowledge of a place unlocked a bit of family history, and painted events in the lives of my parents and their friends that no document could ever yield.

I have long held that a place is more than just a name, or even just a string of jurisdictional names; it has a history and properties all of its own. I have written many posts that have tried to get this message across — often focusing on the legacy of the GEDCOM model — and have even devised a research data model capable of recording the history, documents, images, name changes, and connections of a place in a fashion analogous to that of persons.

My parents, Stan and Ann, were young when they married: my mother was only 17, and she had me by the age of 18. My father had just completed his two years of National Service when he took a job with Nottingham knitwear manufacturers Harry Prew-Smith Ltd, and it was here that he and my mother first met. They were married towards the end of 1955, but he left that job after less than a year to try and find a better one. He tried working on the railway, first in a shunting yard and followed by a spell in the parcel yards (about seven months in all), after which he tried selling vacuum cleaners and then brushes (about six months in all). He then took a job with Stanton Iron Works but had to give it up after about ten months because of knee problems. He often told me how the heat and heavy manual work there was extremely tough, and how he would come home soaked in sweat, but with not an ounce of fat on him — he weighed less than ten stone when he left. But he then blagged his way into a cooking job, based almost entirely on what he had learned during his National Service. This was certainly not the end of his job-changing but it will suffice to set the scene for this article.

I still remember the St. Ann's district where I lived until the age of five, and where my grandparents continued to live after that. It consisted of Victorian back-to-back terraced houses with little if any amenities: no hot water on tap, no central heating, and no inside toilet. The district had a reputation, but it also had a community — something sadly lost when the slum-clearance programs of the 1970s not simply demolished the houses, and dispersed the residents, but also tore up the streets to lay down different ones. This left people clinging to their memories of that community and any old photographs because there was virtually nothing left that they could point to and identify with.

Facebook groups can be a great comfort to such people, and I was a member of several groups that shared photographs and memories. It was here, in 2018, that I connected with David Marriott, a man who remembered my parents, and my father's family. He even admitted to a romantic "fling" with my father's younger sister, Gill, and that she had a best friend, Christine Broadley, whom she went everywhere with; the pair even dressed the same. He recalled that my father and his younger brother, Ray, were older than him, but my cousin (also a David) was about the same age. There was a youth club — the Sycamore Youth Club, run by an Eric Wheatley — on Lavender Street, and my father and cousin used its basement as a gymnasium since it had the weights and other equipment down there.

He then shared a photograph of himself taken in about 1957; he knew it was in my "mother's living room", but not really where or by whom.

 


Figure 1, David Marriott, around 1957, displayed with his kind permission.

I took one look at the awful wallpaper behind him and thought 'I know that pattern. I've seen it before somewhere'. So I went through my own photograph collection looking for something similar, and came up with the following:

 

Figure 2, Ann Proctor, around 1957.

A closer match you could not find: the wallpaper is the same, the seat and cellar door (to the right) are the same, and the lighting is the same. The picture of my mother was taken by my father with a homemade pinhole camera (real ones were very expensive back then), at 9 Manning Grove, Nottingham, and developed by himself. David's picture was not only taken at exactly the same place but probably on the same evening, too.

My grandparents lived at 15 Manning Grove, and 9 Manning Grove was the house of Keith Walker, which my parents rented for about a year. My mother and Keith Walker's wife used to make the tea at the gymnasium.

That was enough to jog David's memory. "Goodness," he said, "I recognised her straight away". He recalled that my father was working night shifts at Stanton Iron Works, and so she was on her own, except for yours truly who was only about one year old. Because she was a bit lonely and scared in someone else's old house, David, Gill, and a few more friends used to go round to keep her company, eat chips (that's "fries" to my American friends) and drink all of her tea. I passed these recollections onto Gill, my cousin David, and my mother (my father having sadly died a few months earlier) and they all wanted to know how the others were doing. My mother said they used to have a great laugh on those late evenings, and it really helped a young mother cope.

That's real history, painted in vivid colours!