Showing posts with label Blogging. Show all posts
Showing posts with label Blogging. Show all posts

Thursday, 7 April 2022

What a Mess

This will be my last blog post for the foreseeable future, and probably forever. This is not a matter of free time, or of advancing years, or even of competing tasks, but of a complete disillusionment with modern genealogy.

I will continue to research in my spare time, but this will be to produce something of standing for my family and extended family to read; I have lost faith in the public world of genealogy. But let me explain in more detail because this has been on the horizon for a while, and yet previously planned retreats have all fallen down for different reasons.

I am use to academic research, and how academic research works in other fields. The purpose of research in those other fields is to find answers — truths — and to produce a valuable collective body of work through collaboration. Virtually by definition, it is not a commercial goal.

I have previously pondered over the nature of genealogy (What is Genealogy?), and considered its difference from family history, but there is a more systemic difference that touches on collaboration, software, and commercial forces. Although genealogy has a well-respected academic side, it generally considers the internet, and digital resources in general, as only good for derivative sources such as images and transcriptions, and not for publications. This field has high standards and produces quality work in traditional publications such as books and journals, but the internet is considered inappropriate for publication due to its transient and ephemeral nature.

If we look to commercial genealogy then we see two quite different worlds: that of the generation of derivative sources that people can search, and that of online trees. Other than improving the search tools and options, I have no real criticisms of the many digitisation and transcription projects, but for online trees then I have many. In fact I have written so many articles on this subject that I won't even begin to enumerate links to them. Irrespective of whether we are considering "unified trees" or "user-owned trees", there are fundamental issues with their structure and the process by which they are generated.

In terms of structure, a tree is appropriate for representing biological lineage, but dreadful for representing history — can you imagine a family tree attempting to detail, say, the events of WWII? But non-biological lineage, such as fostering and adoption, or even weaker associations between people, break this visualisation and can result in a cat's cradle of complexity and confusion. A tree is also limiting in terms of proof arguments (particularly if they reference multiple individuals, families or generations), citations that refer to actual claims (as opposed to simple hyperlinks saying where you got your information), and linking to external resources (images or document scans) that are not specific to single individuals.

But worse than this is the process by which we are expected to construct such trees. We are all probably aware now of the variable quality of trees — although I still find it vexing when I see 'trees are not a valid source' (it depends on the claim) — and that trees can persist online long after someone may have dabbled for a few months using a free trial or a subscription birthday present. There is no responsibility taken by the respective companies for the accuracy of what their subscribers publish, and they appear to be disinterested in why academia looks down on these published works. It is impractical for these companies to fact-check stuff, and so I am not suggesting that is the solution, but they do not acknowledge, publicly, that the simple paradigm of building trees directly from their raw digitised records is naive (despite their advertising). There are many difficult cases of family reconstruction that require effort — possibly an enormous amount of effort — to get around missing information, ambiguous information, or even deliberately obfuscated information, and so make a case for what really happened in the past.

But two experienced researchers might reach different conclusions, both of which appear to fit available information, and so how should that be dealt with? Well, the red mist and edit wars commonly associated with "unified trees" are not the answer. If left to software people then they might suggest transactional get and commit operations, analogous to those in software source-control systems. If you don't know what these are then it's probably best not to ask; they're complicated, generally with horrible user interfaces, and even get software people into trouble.

Well, why don't these companies look at how collaborative research works everywhere else? I can't believe that they're ignorant of it, and so I can only assume that they fear it would be too complicated for their subscribers, or that it would cost them money, or even that it's just a huge step into the unknown and they don't want to kill their cash cow.

Collaborative research elsewhere is not a linear one-step 'raw-data leading to final conclusions'; it's stepwise, and involving prior work by other researchers. Researchers can then look at the work of others and build from it (or refute it). This means real written work, with real citations, is a starting point as claims have to be justified, not just by pointing to data that appears to confirm them, but by explaining why, and why not something else.

OK, so not everyone will be able or willing to produce such written work, but there are people who do, and regularly do so: bloggers. I have already made a case that online genealogy companies could take advantage of this in a way that requires minimal investment, would not run into copyright or attribution problems, and would increase traffic to the respective blogs — surely, a win-win (Blogs as Genealogical Sources). Briefly summarised, the author of a blog article would give permission to the genealogical company to list the corresponding URL in one of their databases, and would provide meta-data to ensure that it showed up in the results of appropriate searches. The genealogy company would store such information in a database of so-called authored works (i.e. the URL, name of author, article title, and meta-data), but would not copy the body of the works. When these works showed up in a genealogical search, the end-user would click on one of them in order to be directed to the original blog article.

Yes, there would be some smaller issues such as the rating these works, or citing them, and so on, but it's academic as there has been no subsequent engagement — Zero, with a capital Z — by any of the companies, including the ones I approached directly.

Modern genealogists rely on the search functions within these online companies, and possibly on Google (although woe betide we have to research a surname such as 'covid'), but they would be less likely to find relevant printed books or journal articles. This sort of scheme could even be extended to cover non-internet sources, but there is yet another possibility, one that flies in the face of the view that research has to be written up in paper-based journals.

People who have researched in other fields may be aware of sites such as arXiv.org (the 'X' is actually representing the Greek letter chi, and so the site name is pronounced as "archive"). These contain online articles, submitted online, and viewed online. They are much more accessible and searchable than the old paper-based journals, and it is entirely possible that this could be done for genealogical research, but it would take the initiative away from any forward-looking genealogy company. Does that matter to genealogists? Probably not as there are many searchable resources that do not fall under their control. Would it contribute to the accuracy and a truly collaborative approach in modern genealogy?

I wish I could be optimistic here, but I'm not!

Friday, 26 March 2021

SVG Family-Tree Generator (v6.0)

 

The word is catching on about this free tool (SVG-FTG, for short), whether for publishing family trees on your website or blog, or simply for providing an interactive visualisation for the benefit of you and your family. As a result of this, some considerable effort has been made to produce a much-enhanced V6.0.

 


In fact, during its time of steady incremental growth, this is by far the biggest set of changes as they are greater than all the previous changes combined.

So what's new in V6.0? Well, there are three main areas of change:

  • Applications and Services: Packaged interactive applications and services that can be selected from a simple menu. No more coding.
  • Viewpoints: The ability to view and maintain different parts of a large tree separately.
  • GEDCOM Export: SVG-FTG already supported import from GEDCOM files, but it now supports export to GEDCOM files, too.

What is SVG-FTG?

SVG-FTG generates Scalable Vector Graphics (SVG), in combination with HTML, CSS and Javascript, to display interactive family trees in your website or blog. Unlike normal images, SVG is a format that does not go all fuzzy when you zoom in. There are several packaged applications that can be run from your tree, and an open framework to develop your own. It supports complete control over layout, thumbnail images, hover text, HTML biographical or historical notes, scrolling/zooming of individual trees, GEDCOM, timeline reports, and linked trees. The designer is Windows-based but the output is neutral and runs in all modern browsers. Note that the output is non-proprietary, royalty-free and needs nothing to be installed first; it is therefore ideal for sharing with friends and family.

Where is SVG-FTG?

There have been several previous posts about SVG-FTG: Interactive Trees in Blogs Using SVG, More on SVG Family Trees, and SVG Family-Tree Generator (v5.0).

Details of availability can now be found on the summary page: SVG-FTG Summary.

Presentation

The general presentation quality has been improved again, particularly under high magnification (i.e. when zooming in to a high degree). For example, there is now no visible overlap where lines join. The colours have been standardised to remove previous differences between SVG-only and mixed HTML/SVG modes.

A number of presentational enhancements are placed under the control of the end-user:

  • It is possible to nominate a background image upon which your tree will be drawn, including its opacity and whether it is repeated across the available height and width. The image at the head of this article shows an example.
  • Improved choice between scrolling and scaling of large trees. In other words, whether you want to see everything at once and zoom in to see the detail, or to see a part of the tree at normal magnification and pan around to see the rest.
  • Person-boxes can now be opaque or translucent, without changing the visible colour. This will be a consideration if you have a background image, or if you are using "fanned" lines rather than the normal horizontal/vertical ones.
  • You can nominate stock images, according to sex (e.g. head-and-shoulders silhouettes), to display in the absence of thumbnail images for persons.
  • You can add custom CSS class names to person-boxes, family-circles, lines, and notes panels, either for application purposes or to change their presentation.
  • Default size of person-box buttons has been raised from 10x10 to 12x12 pixels, but this may be changed via the settings form.
  • There are now fields in the settings form to change the size of person-boxes and their separation.
  • SVG-FTG never worked properly before with Internet Explorer (IE) 11 because it is such a non-standard browser; but this has now changed.

The following image shows a partial tree demonstrating the possibility of bigger buttons (in both default and icon modes) and the use of stock images for person-boxes having no thumbnail image. It also demonstrates the use of opaque person-boxes in conjunction with "fanned" lines.

 


Applications and Services

One of the main goals of SVG-FTG was to produce interactive trees — not just static images. That means being able to utilise a tree as the user interface (UI) to different applications and services, or in other words to make a tree do things for you. Previous versions already offered some examples in the form of 'Timeline Reports' and pop-up 'Information Panels', but they sometimes required editing of the code.

This is arguably one of the biggest changes to SVG-FTG as those previous applications, and several new ones, have now been packaged up. This means that their definitions and configurations have been placed in a separate registration file, and the end-user just selects the required ones from a simple menu; there is no longer any requirement to see or change code as it's now generated for you, based on your selections.

If you require applications to be configured differently (e.g. change the mouse-click operations, change the button allocations, or even to add new buttons) then it can be done with a small adjunct to the standard registration file; you don't have to edit the distributed standard one.

Additional applications (i.e. in addition to the existing Information Panels and Time Reports) include:

Expand Notes

If you are working in SVG-only mode, or you have lengthy biographical notes, possibly with multiple images, then the existing Information Panels may be insufficient. This application allows them to be displayed in a separate browser tab instead. For instance:

Notes for family of Henry Proctor and Elizabeth Turton

 

Married 2 Oct 1858 at Nottingham St Nicholas.

Find Persons

With a large tree, it can be hard to find specific persons. This application provides a floating (i.e. movable) search box that allows you to filter a list of the available person captions until you can see the one(s) that you want. You can then select from the list which will cause them to become highlighted.


As you type a partial caption name into the search box, the list of possibilities is filtered. You can select multiple captions, and the borders of the selected ones are highlighted as shown above. As well as highlighting a relevant person-box, a selection also ensures that it is visible by automatically scrolling the tree as necessary.

The 'Clear' button removes those highlights, and the 'Hide' button collapses the search box until you need it again.

Ancestor Links

This application is similar to the standard RootKey feature provided by SVG-FTG, except that it is dynamic. Clicking on the button (a tree icon by default) for any person will show their maternal and paternal ancestral paths, all the way to the top of your tree. For instance:


Alternatively, performing a shift+click operation on the button will just highlight the relevant person-boxes of the ancestors, as in the Find Persons application, above. An alt+click operation will show both modes together. A control+click operation on the button clears the current highlighting.

The application collaborates with the older RootKey feature to provide a "start-up person" in your tree. Such a person is highlighted and scrolled into view.

Compendia

This application uses pairs of titles and URL data links held in the "program data" tab of persons (and families). Clicking on the button (a document icon by default) will present a list of any such references in a floating dialog, and clicking on any of the entries will show their contents in a new browser tab.

These references may be to articles, blog-posts, documents (e.g. PDF), images, or anything with the URL link.

 



A shift+click on that button will highlight all persons or families that have at least one data link in common with the clicked one, implying that they share references in the same article or appear in the same image. An alt+click on the button will highlight all persons (and families) having any such references available. A control+click on the button will clear any current highlights on all elements.

Linked Trees

This application allows you to link together separate trees such that clicking on a person-box button will take you to some associated person in a target tree, or present you with a menu if there are several to choose between. The application is general-purpose, but it is also employed by the Tree Viewpoints, described below.

Application Development

The packaging framework for interactive applications and services is open, meaning that it is documented and it can be added to, allowing custom applications to be shared with other users. If you are a developer, or a "power user", then you can write your own applications and register them for selection by end-users using the same mechanism provided for registering and configuring the predefined ones.

There are many new features for helping such development, including several libraries of code for accessing 'navigation data' (data about persons or families, and their relationships), 'notes data' (i.e. biographical notes, images, links), and 'program data' (application-defined data). Services, as distinct from applications, include reusable UI elements, such as a message-box dialog for reporting errors or asking questions, and a menu dialog for making a choice from several alternatives.

Tree Viewpoints

A disadvantage of presenting all your persons and families at the same time is that (a) it rapidly becomes unreadable, (b) the lines become messy because they usually have to cross over each other, and (c) it consumes more memory. Most products avoid these issues by keeping all the persons and families in a database, and only showing you a selected group at once, which inevitably means that you no longer have any control over their layout.

A viewpoint is a view onto a sub-tree of your loaded persons and families. What this means is that although the full tree will still show every person and family that you have defined, a viewpoint can be used to focus on just a few of them. You can have many viewpoints defined, and persons or families may appear in multiple viewpoints if required.

When a viewpoint is loaded into the Tree Designer, things operate virtually the same as when a full tree is loaded, except that there is less clutter and complexity. A typical usage of viewpoints is to break down a large tree by surname or by family groups.

It is very easy to flip between your viewpoints in the Tree Designer, or to find a person among multiple viewpoints. A Viewpoint Manager is provided to help with creating, deleting, and modifying viewpoints, but also to manage the allocation of persons and family groups to at least one viewpoint. A typical goal is to view your final family tree as a set of sub-trees, each in a separate browser page, and all connected by hyperlinks. In order to achieve this, the Viewpoint Manager helps you manage the dividing-up of the full tree, and ensuring that no person has "fallen through the cracks".

When the HTML (or SVG) code is generated for your browser, by the 'View' button, then it utilises the Linked Trees application. This general-purpose application allows users to link their trees according to different person roles, but the viewpoint feature also uses it to link the viewpoints according to persons in common between them.

The following illustration of clicking on a button in a person-box and selecting an option to go to another viewpoint in a second tab is based on a sample distributed with SVG-FTG.


An important feature of the Linked Trees application is that it names the tabs, which allows it to maintain a finite set of them. If a target tree is already loaded then control is transferred there rather than loading it yet again. Having parallel access to these "tabbed trees" is very powerful given that they can each run their own applications.

GEDCOM

Previous versions already allowed the importing of GEDCOM data, but it is now possible to export your current tree in GEDCOM format. It should be noted that the tree definition format, as used by SVG-FTG, and GEDCOM serve quite distinct purposes, and so totally lossless round-trips (e.g. when importing a GEDCOM file and re-exporting it) is not guaranteed.

As well as exporting all the person and family details (including their notes) that you have loaded in SVG-FTG, this operation also saves your place-key mappings, your tree layout (i.e. person-box coordinates), and any local settings on your persons and families. It saves a copy of the remote (URL) version of any image references that employ place-keys (i.e. the mechanism used to access both local and remote resources in parallel). It does not save any details from the 'Program data' tabs or any of your viewpoint definitions.

The existing GEDCOM import has also been enhanced in the area of personal names, particularly where a person has multiple names, or they have been stored in itemised format rather than as a single string. The import operation also acknowledges any adopted or fostered status of children in a family.

Settings

The standard settings form has been split into one for core settings (as used by everyone) and one for advanced settings (as used more by developers or "power users"). The first of these is very similar to the form used in previous versions, except for a few new options and control over any stock images to be displayed in the absence of thumbnail images. The advanced-settings form contains options relevant to the size and position of person-boxes, button configuration, code generation, and nominating a background image.

Media Enhancements

Users may have experienced problems with record terminators when moving files between Microsoft and Apple systems, or even when loading GEDCOM files exported by certain systems. SVG-FTG now endeavours to identify which terminators are being used, and then honour appropriately. This even includes the non-standard CR-CR-LF that often results when a file generated on one system is processed on a different one.

Through a limitation of a code library being used by SVG-FTG, it was not possible to specify PNG images in the Tree Designer, although all image formats are accepted in the output files by the browser. This has now been fixed by adding custom code to specifically support PNG. The change should be seamless so that you can select either PNG, JPG, or the other formats, in the same way.

Sundry

The edit-person form can now capture multiple personal names, as distinct from the long and short person-box captions. This is particularly important if you intend to interface to other systems, such as through GEDCOM export.

The editing of HTML notes for persons or families has been enhanced so that it automatically finds the corresponding start and end tags, and offers forms for editing selected elements. Those forms will preserve any attributes that it does not yet support.

When creating either a person or a family, a default key name is automatically generated. This can, of course, be modified if you don't like it, but it is designed to help make the process quicker.

Previous versions of SVG-FTG offered three sex options: male, female, and unknown/unspecified. This is now supplemented by a further indeterminate/other option.

The 'Find Person' menu option in Tree Designer now scrolls a selected person-box into view if it is not currently visible. It will also find a person across any viewpoints that you may have defined. The option can now be accessed using a Ctrl+F shortcut in addition to the normal menu option.

SVG-FTG relies on a number of external resources for the output HTML/SVG to work, and this includes JavaScript files, CSS files, and icons. Although these were held in a folder on a neocities.org site, all references now use a custom domain in order to obscure and decouple that connection (i.e. https://parallaxviewpoint.com/SVGcode).

Thursday, 6 December 2018

Research in Online Trees


My previous post, The Future of Online Trees, prompted a flurry of reaction, most of which was positive; however, it did suggest, implicitly, that many people find it hard to think beyond their entrenched views, and that my explanation may have assumed too much.

This follow-up collects together some of the explanatory comments that I'd since posted around the Web, and tries to make a coherent  argument for what genealogical research should entail.


The previous article made several negative comments about existing online trees, including:

  • That assembling their associated conclusions directly from raw digitalised information is not always easy, and gets very hard as you go further back in time (e.g. before census returns and civil registration).
  • That it's hard to tell naive trees from properly researched ones, and that, no, a bunch of citations are not a useful indicator.
  • That naive trees usually persist long after a creator may have  abandoned them, and could steer new researchers down the wrong path.
  • That proof arguments (i.e. reasoned explanation), as opposed to simple proof statements (i.e. citations), are almost never provided.
The primary issue behind my article was that many genealogical conclusions require their research work to be written up in order for them to be assessed, and for that research to be cited either by trees or other research work. Proof statements alone are only applicable when the sources offer direct answers and do not conflict with each other, but identity problems and family reconstruction can require lengthy arguments that examine multiple sources. The results will often address groups of correlated people rather than just some specific person, and so it's not realistic to expect that work to be tucked away in a single person entity (on a tree) or in a single person page.

An associated issue is that the contributors to online trees — and probably the users of genealogy software in general — routinely talk about individual "claims", and the supporting sources for those claims, as though they're all independent of each other. This is fallacy! The idea that a specific claim can be justified in isolation, and linked directly to one or more sources that give a direct answer, is a huge oversimplification of the research process, and yet this is a mindset that is hard to argue with.

One of the positive things I suggested in my previous article (possibly the only one) was that there are researchers who do publish their work online (e.g. in blogs), and that online trees could reference their research as "authored works": a recognised source category that supplements those of "original" and "derivative". There is no issue at all with representing this in GEDCOM files — the data format most often used to transfer data between two places — nor any significant issue with online providers recognising such work as a specific source category.

Many traditional genealogists write-up their work in academic journals, but this is more about kudos than about helping a  community of genealogists; few of us will be subscribers to these journals. This is a shame because they cannot distance themselves from online genealogy, nor ignore the associated problems, because we're all tarnished with the same brush. If we describe our work as "genealogy" then it will be linked automatically to the prevailing impression of its most common form: online trees.

It may be hard to see what I'm getting at if you haven't participated in research in other fields. All the fields that I am aware of, such as in science and medicine, rely on published works. This could be in journals or online, but by far their biggest difference from what is currently considered genealogical research is that newer works cite older works. The consensus is then built up through layers of research, each of which may support or refute previous work. There's a saying about standing on the shoulders of giants, and it makes perfect sense: someone could have spent a lifetime solving one particular deep mystery, and so to expect someone else (beginner or otherwise) to find the same answer directly from raw online information is unrealistic. I cannot think of any other area of research that works as genealogy currently does, and where conclusions are either copied blindly from those of someone else or constructed independently from raw information. This is a little like a surgeon creating an independent textbook themselves by simply dissecting the evidence — a cadaver in this case. Knowledge and progress come from sharing research, and by building on the research of others. It's step-wise, progressive, and takes time. And without seeing any written research then you cannot tell whether someone made their conclusions in 30 minutes or 30 years.

So what size of work are we talking about? Is it just a single paragraph? Well, it could be, or it could be a couple of thousand words, as with several blog articles that I've encountered. I have two unpublished works of 5000 words, myself, that I want to contribute to the community, and for posterity, but also a work-in-progress that is already at 10,000 words — such is the complexity.

A note on the use of wikis as a medium for collaborative research is necessary because they were mentioned by a few people. It is true that wikis can be, and are, used for such research purposes, but they have significant weaknesses. They are often limited in the richness of their presentation — usually amounting to more of a protracted discussion, as the old BetterGEDCOM wiki demonstrated — but genealogical research requires support for rich formatting, images, tables, and citations. Not all blogs offer this, but there are usually ways of achieving it (see Summarised Blogger Tips for instance). Wikis have little, if any, editorial control, and no attribution support beyond their confines. Also, that they constitute a confined medium — forcing people to contribute outside any personal medium or prior work — would put too many people off. By contrast, blogs are not confined, they may be linked or associated with other work by the author, and their articles have immediate attribution. Wikipedia was also mentioned as an example of successful collaboration, but it has strict rules that prevent original work or theories being presented. It relies on secondary sources, and so implicitly collects information that is already in the mainstream. This certainly doesn't prevent edit wars but it does place it apart from collaborative wiki-based genealogy.

So, my suggestion is to separate attributable research work from tree-based conclusions, and to cite such work for the harder cases rather than just some raw information. This suggestion is not rocket science so why aren't we doing it?


Saturday, 1 December 2018

The Future of Online Trees


Many people have written about the ills of online family trees, including me, but that won't stop me writing about them again. Yes, they're full of errors — more than some people want to admit — but is there an answer? Is there a future for them?

I first want to summarise some of those problems, and then suggest that the current situation is not sustainable. Much damage has already been done to the collective knowledge of our family histories, but also to the reputation of genealogy itself. If online trees are wrong then those errors will propagate through simplified collaboration, they will tarnish the interpretation of DNA data, and they will stymie any attempt to use AI technology to suggest further connections.


One basic problem is that users are encouraged to create their trees directly from raw digitised resources, including transcripts, held by the online provider. Although it is possible to include data from external sources, there is some debate about whether all providers acknowledge them when grading their trees. So what's the problem here? Is there any other way to create a tree?

Well, yes and no! A tree very rarely captures any of the logic employed when its relationships were determined, or of the histories of the individuals in the tree — usually prerequisite knowledge for solving the hard identity and relationship problems. Yes, they may include citations or electronic bookmarks to show where information came from, but that isn't the same as explaining why the information is correct or relevant. Also, they are incapable of supporting complex conclusions that have correlated information from multiple sources.

This sets the stage for errors that are unchecked, and when coupled with the ability to easily copy data from one tree to another then it means that the errors will proliferate. Not only that, the source of an error, whether known or not, will likely be permanent, even if the author abandoned their tree.

I came upon such a case just the day before writing this article. I don't want to pick on this specific person but it typified a very real problem with online trees. I had contacted this person because their tree appeared to offer solutions to a deep problem that I was collaborating on with another person. On looking at the tree, I could see that the conclusions were very wrong (e.g. conflating two given names to create a name found nowhere in the evidence, simply to tie-off an apparent anomalous reference), but there were no sources given. On approaching the owner they admitted that they had received a gift of membership to this provider, during the period of which they copied data from a cousin's tree, and that they were no longer a fully-paid member. So even though they no longer dabbled with genealogy, their naive tree would outlive their membership and be used by other beginners, ad infinitum. Ideally, the industry should be generating something of historical value, but all we have is a cauldron of ills and woes.

Such cases are compounded by the fact that no provider I know of offers a mechanism to flag a tree, or part thereof, as tentative, under-construction, or otherwise work-in-progress. This would be useful for experienced researchers as well as beginners, and could go some way to stopping the proliferation of erroneous conclusions.

A more complete solution would be to make room for research narrative — not just pieces of plain text stuffed into the records for specific persons or families — and I have proposed this several times before, e.g. Feeding the Trees. The problem here is that there isn't much inclination amongst providers to do this, possibly because they don't want to be first, or rich-text (formatted with images, tables, citations, etc.) is somehow harder to handle, or simply that they don't see the need when everyone has a word processor.

All is not lost, though, because there are genealogists out there who write-up their research and make it publicly accessible, usually via a blog. You must be able to see where I'm taking this, now. If there was a way to reference such research in an online tree then it would relieve the more-casual user of the burden of the heavy lifting. In effect, rather than reference raw digitised records directly, and with no logic explained, they could reference work contributed by others — work that explained how conclusions were reached as well as citing the relevant sources.

Well, I have proposed this before, too: Blogs as Genealogical Sources suggested that published research (especially in blogs) could be indexed as external sources by providers. Also, that the indexing of an article URL would be done with the author's permission, and with no bulk text being copied or indexed by the provider — actions that would violate copyright and divert traffic away from that author's article. This win-win scenario wouldn't require a large investment by the provider, and so it was strange that only one showed any interest and asked me for more details. Such work can already be found by researchers, but only via general-purpose search engines such as Google. How much more convenient if genealogy subscribers saw them as another source group and were allowed to perform genealogical searches that included them.

I had heard rumours of a past RootsTech announcement by Findmypast of an initiative called "verified reference trees" but I could find no official mention of it. It sounded good but would probably be impractical — who could produce them? How many would you be likely to generate? How much of a given pedigree could one contributor realistically cover?

What I'm saying here is that online trees need a rethink if they are to have a future. Simply putting more and more digitised resources at their disposal is not bad, but there are fundamental problems that need to be addressed, ones that will ultimately come back to bite those providers who are ignoring the future.

Will we still be working on the same trees in 20 years time, and in the same way? Some pundits have suggested that our family histories will have all been resolved by then through collaborative unified trees. ROTFLMAO! There are already more inaccurate conclusions online than accurate ones, and it is not easy to tell which are which — citations alone do not cut it! Some other pundits have suggested DNA testing will solve everything. ROTF again! Even limiting yourself to biological lineage, and ignoring much family history, there are cases for which DNA can never help. For instance, read my conclusion at Jesson Lesson.

There are two very broad categories of user that providers need to consider: those who are experienced and likely to do a lot of the research leg-work, sharing their research details with others; and those who want to play with online tools to create a simple tree of their own. Although the latter category may be able to contribute valuable knowledge and lore from recent generations, most would struggle to reconstruct older generations from raw information alone for the simple fact that it's hard.

So whose work should they use? Well, trees don't need to be officially "verified". Works of written research can be read by any user who can then judge for themselves on the depth of the research and of the soundness of the conclusions. I'm not suggesting that these users can't use raw information, only that published research will give them more rungs on the ladder; bigger pieces of the jigsaw to assemble.

Another RootsTech announcement that I recall was that of employing AI to find new connections in the data. This is worth a paragraph of its own before leaving this subject because AI is akin to black magic for many people. Applying AI to raw information is a recipe for disaster because, as any researcher knows, it can never be taken at face value. AI would simply be correlating information from many record groups and attempting naive identity resolution and family reconstruction. It would know nothing of real-life situations and the reasons that records may have lied or that families had moved, or of real human relationships and the reasons they may have been forged or broken. The naivety of such algorithms would ultimately manifest as the appearance of fictitious people, or of duplicates for the same person (merging of inappropriate people would also be a risk but probably to a lesser extent). There will always be the need for human researchers, and any application of AI must have their additional inputs.

So, has anyone seriously tried to envisage what genealogists will be doing in a couple of decades time? Will the same providers still be offering online trees for users to build from raw information, and will beginners be researching the same people as we are doing now? Will all of today's trees still be online then, and will the error rate in those trees be reduced? If so then how? Product Managers beware! Your industry's future is at stake unless some visionaries step forward. Genealogy is big business at the moment, but it cannot remain as it currently is. Companies may be throwing their hat at DNA testing, but what happens when that fad passes its peak and it is taken for granted? A company's valuation is dependent on a predicted future for its market, so buyer beware also.

[see follow-up at Research in Online Trees]