Wednesday, 22 January 2014

Using Feedburner with Blogger



Once you have your Blogger account working, you may have heard about a free tool provided by Google called Feedburner, but what is it? Even if you already use it, you may not be quite sure what it does. As well as trying to explain this as simply as I can, I also want to highlight some potential log-jams that you may encounter.


Once you have started publishing to your blog, you will want people to find and read it. Relying on them finding it by accident through a Web search is not really going to work so you will be sharing each post into news streams such as those in Google+ (Circles and Communities) and Facebook. However, it’s easy to miss stuff in these streams, and if someone wants to follow your every word then they will need a way for them to subscribe and be notified when something new has been published.

RSS (Really Simple Syndication) and Atom are both mechanisms for the syndication of Web feeds. In other words, they allow changes on one site to be syndicated to one-or-more other sites. In the context of a blog, this simply means that they allow users to see when new material has been published. When users subscribe to a Web feed, they may see these changes through a feed reader (such as Google Reader, Bloglines, or NetVibes) or by email. A feed reader is sometimes called an aggregator program as it allows you to bring together content from multiple feeds (e.g. from different blogs) so that you can view them all in one place.

A feed reader just needs the URL of the target blog (e.g. http://parallax-viewpoint.blogspot.com) for it to subscribe. It’s not unlike an email program in that it periodically checks your subscribed feeds and shows the number of new posts. You can then decide to selectively download and read them. This is all very well if you’re familiar with those tools, or you don’t mind learning new tools. However, many people would prefer to just receive a simple email containing the latest blog post when it appears.

So where does Feedburner fit into this? Well, Blogger can publish updates via four different feed URLs, such as:

Atom feeds:-
http://parallax-viewpoint.blogspot.com/feeds/posts/default
http://parallax-viewpoint.blogspot.com/atom.xml
RSS feeds:-
http://parallax-viewpoint.blogspot.com/feeds/posts/default?alt=rss
http://parallax-viewpoint.blogspot.com/rss.xml

It’s possible for other sites and tools to make use of any of these, and you would have no ability to see the total number of your subscribers. Those subscribers may not see identically-rendered copies of your post either. What Feedburner does is redirect all these URLs to a new data feed of its own, such as:

http://feeds.feedburner.com/blogspot/xxxxxx

The xxxxxx is a string of characters generated for you. This redirection means your subscribers will all feed from the same place, and Feedburner can then generate subscriber statistics for you.

In order to set this up, you must first have a Feedburner account. Go to Feedburner and sign in with your Google Account. Put your Blog URL (e.g. http://parallax-viewpoint.blogspot.com/) into the 'Burn a Feed Right This Instant' and click ‘Next’. Specify a feed title, and take note of the ‘Feed Address’ that it creates for you as you’ll need to tell Blogger about it. The Feedburner account should now show up in your Google dashboard (https://www.google.com/settings/dashboard) along with your Blogger account, etc.

Then go to your Blogger dashboard. Select Settings→Other from the left panel, and go to the Site Feed section. In the ‘Post Feed Redirect URL’ field, enter the URL of your Feedburner feed (e.g. http://feeds.feedburner.com/blogspot/xxxxxx). Set the ‘Allow Blog Feed’ field to “Full”.

OK, you now have a Feedburner feed. Now let’s make it easy for email subscribers. Go back to your Feedburner account, select the Publicize tab, and then the ‘Email Subscription’ entry on the left. Click ‘Activate’. You can now customise some settings under the ‘Email Subscription’ section such as the title/body used for confirmation emails, the title used for notification emails, and the time-of-day when notifications should be sent.

Back in the Blogger dashboard, select Layout on the left. Pick a panel (usually the right panel) and ‘Add a Gadget’. Choose the ‘Follow By Email’ gadget, provide a label for the email address field, and specify the Feedburner feed URL, as shown above. This will provide a very simple field on each blog page into which a user can enter an email address. Those users will be sent a confirmation email which they must respond to in order to receive email notifications from your blog.

So far, so good; I hope. A quick search of the Internet, though, shows lots of people have problems getting email subscription working, so what’s the problem? Well, Feedburner keeps a copy of the last few blog posts so that it can compare them with the latest information from your original blog feed, and so only notify people of updates. However, it has a space limit of 512KB. Note that this is the size of the HTML rather than of your original text, and it does not include any images or attachments. It’s therefore a little difficult to gauge. The two main issues are exceeding this size limit, and having unrecognisable content in your feed.

Preparation of your blog post in Microsoft Word can be a cause in both of these issues, but this has already been covered in my previous post at: Using Microsoft Word with Blogger.

In the Feedburner dashboard, under the Troubleshootize tab, there are tools to validate the original (Blogger) feed and the Feedburner feed. If these show errors such as the following then there’s a simple explanation:

Undefined entry element: georss:featurename 3 occurrences
... 2" width="72" /><thr:total>0</thr:total><georss:featurename>Naples, FL,

Undefined entry element: georss:box 3 occurrences

nt>26.1420358 -81.7948103</georss:point><georss:box>25.913972299999998

The problem here is that you’ve set a Location in the ‘Post Settings’ down the right-hand side of your blog-post.  This generates geocoding data for you but — at the time of writing ― it’s not a in a format expected by Feedburner, and so it throws-up as a result. If you simply unset that Location property on your post then this error should go away.

Another reason for exceeding the size limitation is if you generate a lot of lengthy blogs, or even a few very lengthy ones. By default, Blogger provides details of the last 25 posts to Feedburner and the sum total of this may be too great. This can be restricted with a parameter on the end of your original feed address. Click the ‘Edit feed details…’ button at the top of the page, and add a max-result parameter to the end of the ‘Original Feed’ address, such as:

http://parallax-viewpoint.blogspot.com/feeds/posts/default?max-results=3

I specified a very low value for myself since, although I only generate about one post per week, they tend to be quite large.

Sunday, 19 January 2014

You’re Probably Right



When anyone mentions statistics or probability being applied to genealogical research then there’s usually a sharp reaction. There happen to be some valid questions that would benefit from thoughtful discussion but unfortunately the many knee-jerk reactions tend to be for all the wrong reasons.


It’s hard to find a single reason why this topic gets such an adverse reaction since the arguments made against it are rarely put together very carefully. I have seen some reactions based purely on the fear that any application of numbers means that assessments will be estimated to an inappropriate level of precision, such as 12.8732%. That’s just ludicrous, of course!

In this post, I won’t actually be making a case for the use of statistics since I am still experimenting with an implementation of this myself and it isn’t straightforward. What I will try to do is identify what is and is-not open to debate, and ideally to add some degree of clarity. Although I have a mathematical background, this only briefly touched on statistics. It is a specialist field, and many folks will have a skewed picture of it, whether they’re mathematically inclined or not. It is also a technical field and so a few symbols and numbers are inevitable but I will try and balance things with real-life illustrations.

Statistics is generally about the collection and analysis of data. Despite what politicians might have us believe, statistics proves nothing, and this is important for the purposes of this article. Statistical analysis can demonstrate a correlation between two sets of data but it cannot indicate whether either is a consequence of the other, or whether they both depend on something else. The classic example is data that shows a correlation between the sales of sunglasses and ice-cream — it doesn’t imply that the wearing of sunglasses is necessary for the eating of ice-cream.

Mathematical statistics is about the mathematical treatment of probability, but there is more than one interpretation of probability. The standard interpretation, called frequentist probability, uses it as a measure of the frequency or chance of something happening. Taking the roll of a die as a simple example, we can calculate the number of ways that it can fall and so attribute a probability to each face (1/6, or roughly 16.7%). Alternatively, we could look at past performance of the die and use that to determine the probabilities; a method that works better in the case where a die is weighted. When dealing with the individual events (e.g. each roll of the die), they may be independent of one another, or dependent on previous events. A real-life demonstration of independent events would be the roulette wheel. If the ball had fallen on red 20 times then we’d all instinctively bet on black next, even though the red/black probability is unchanged. Conversely, if you’d selected 20 red cards from a deck of playing cards then the probability of a black being next has increased.

The other major interpretation of probability is called Bayesian probability after the Rev. Thomas Bayes (1701–1761), a mathematician and theologian who first provided a theorem to expresses how a subjective degree of belief should change to account for new evidence. His work was later developed further by the famous French mathematician and astronomer Pierre-Simon, marquis de Laplace (1749–1827). It is this view of probability, rather than anything to do with frequency or chance, which is relevant to inferential disciplines such as genealogy. Essentially, a Bayesian probability represents a state of knowledge about something, such as a degree of confidence. This is where it gets philosophically interesting because some people (the objectivists) consider it to be a natural extension of traditional Boolean logic to handle concepts that cannot be represented by pairs of values with such exactitude as true/false, definite/impossible, or 1/0. Other people (the subjectivists) consider it to be simply an attempt to quantify personal belief in something.

In actuarial fields, such as insurance, a person is categorised according to their demographics, and the previous record of those demographics is used to attribute a numerical risk factor (and an associated insurance premium) to that person. This is therefore a frequentist application. Consider now a bookmaker who is giving odds on a horse race. You might think he’s simply basing his numbers on the past performance of the horses but you’d be wrong. A good bookmaker watches the horses in the paddock area, and sees how they look, move and behave. He may also talk to trainers. His odds are based on experience and knowledge of his field and so this is more of a Bayesian application.

Accepted genealogy certainly accommodates qualitative assessments such as primary/secondary information, original/derivative sources, impartial/subjective viewpoint, etc. When we consider the likelihood of a given scenario then we might use terms such as possible, very likely, or extremely improbable, and Elizabeth Shown Mills offers a recommended list of such terms[1]. Although there is no standard list, we all accept that our preferred terms are ordered, with each being between the likelihoods of the adjacent terms. These lists are not linear; meaning that the relative likelihoods are not evenly spaced. They actually form a non-linear[2] scale since we have more terms the closer we get to the delimiting ‘impossible’ and ‘definite’. In effect, our assessments asymptotically approach these idealistic terms, but never actually get there.

As part of my work on STEMMA®, I experimented with putting a numerical ‘Surety’ value against items of evidence when used to support/refute a conjecture, and also on the likelihood of competing explanations of something. This turned out to be more cumbersome than I’d imagined, although a better user interface in the software could have helped. The STEMMA rationale for using percentages in the Surety attribute rather than simple integers was partly so that it allowed some basic arithmetic to assess reasoning. For instance, if A => B, and B => C, then the surety of C is surety(A) * surety(B). Another goal, though, was that of ‘collective assessment’. Given three alternatives, X, Y, & Z, simple integers might allow an assessment of X against Y, or X against Z, but not X against all the remaining alternatives (i.e. Y+Z) since they wouldn’t add up to 100%.

Although I didn’t know it, my concept of ‘collective assessment’ was getting vaguely close to something called conditional probabilities in Bayes’ work. A conditional probability is the probability of an event (A) given that some other event (B) is true. Mathematicians write this as P(A | B) but don’t get too worried about this; just treat it as a form of shorthand. Bayes’ theorem can be summarised as[3]:

           P(A)
P(A | B) = ―――  P(B | A)
           P(B)

It helps you to invert a conditional probability so that you can look at it the other way around. A classic example that’s often used to demonstrate this involves a hypothetical criminal case. Suppose an accused man is considered one-chance-in-a-hundred to be guilty of a murder (i.e. 1%). This is known as the prior probability and we’ll refer to it as P(G), i.e. the probability that he’s Guilty. Then some new Evidence (E) comes along; say a bloodied murder weapon found in his house, or some DNA evidence. We might say that the probability of finding that evidence if he was guilty (i.e. P(E | G) is 95%, but the probability of finding it if he was NOT guilty (i.e. P(E | ¬ G)[4] is just 10%[5]. What we want is the new probability of him being guilty given that this evidence has now been found, i.e. P(G | E). This is known as the posterior probability (yeah, yeah, no jokes please!). The calculation itself is not too difficult, although the result is not at all obvious.

           P(E | G)            95%
P(G | E) = ――― P(G) = ――――――――――――― x 1% = 8.8%
             P(E)          (95% x 1%) + (10% x 99%)

This may just look like a bunch of numbers to many readers, but the mention of finding new evidence must be ringing bells for everyone. If you had estimated the likelihood of an explanation at such-and-such, but a new item of evidence came along, then you should be able to adjust that likelihood appropriately with this theorem.

So what about a genealogical example? Well, here’s a real one that I briefly toyed with myself. An ancestor called Susanna Kindle Richmond was born illegitimately in 1827. I estimated that there was a 15% chance that her middle name was the surname of the biological father. If we call this event K, for Kindle, then it means P(K) is 15%. This figure could be debated but it’s the difference between the prior and posterior versions of this probability that are more significant. In other words, even if this was a wild guess, it’s the change that any new evidence makes that I should take notice of. It turns out that the name ‘Kindle’ is quite a rare surname. FreeBMD counted less than 100 instances of Kindle/Kindel in the civil registrations of vital events for England and Wales. In the baptism records, I later found that there was a Kindle family living on the same street during the same year as Susanna’s baptism. Let’s call this event — of finding a Neighbour with the surname Kindle ― N. I estimated the chance of finding a neighbour with this surname if it was also the surname of her father at 1%, and the probability of finding one if it wasn’t the surname of her father at 0.01%. What I wanted was the new estimation of K, i.e. K | N. Well, following the method in the murder example:

           P(N | K)              1%
P(K | N) = ―――― P(K) = ――――――――――――― x 15% = 94.6%
            P(N)           (1% x 15%)+(0.01% x 85%)

This is a rather stark result from the low probabilities being used. I’m not claiming that this is a perfect example, or that my estimates are spot on, but it was designed to illustrate the following two points. Firstly, it demonstrates that the results from Bayes’ theorem can run counter to our intuition. Secondly, though, it demonstrates the difficulty in using the theorem correctly because this example is actually flawed.  The value of 1% for P(N | K) is fair enough as it represents the probability of finding a neighbour with the surname Kindle if her middle name was her father’s surname. However, the figure of 0.01% for P(N | ¬ K) was really representing the random chance of finding such a neighbour if her middle name wasn’t Kindle at all. What it should have represented was the probability of finding such a neighbour if her middle name was Kindle but it wasn’t the surname of her father. However, it failed to consider that the two families may simply have been close friends.

There is no room for debate on the mathematics of probability, including Bayesian probability and the Bayes’ theorem. The application of this mathematics is accepted in an enormous number of real-life fields, and genealogy is not fundamentally different to them. As part of my professional experience, I know that many companies use Bayesian forecasting to good effect in the analytical field know as business intelligence. The only controversial point presented here is the determination of those subjective assessments. All of the fields where Bayes’ theorem is applied involve people who are quantifying assessments that are based on experience and expertise. We already know that genealogists make qualitative assessments but would it be a natural step to put numerical equivalents on their ordered scales of terms. We wouldn’t argue that ‘definite’ means 100%, or that ‘impossible’ means 0%, but employing numbers in between is more controversial even though we may use a phrase like “50 : 50” in normal speech.

I believe there are two issues that would benefit from rational debate: where those estimations come from, and whether it would be practical for genealogists to specify them and make use of them through their software. Although businesses proactively use Bayesian forecasting, the only examples I’ve seen in fields such law and medicine have been ex post facto (after the event). For my part, I find it very easy to put approximate numbers against real-life perceived risks, and the likelihood of possible scenarios. I have no idea where these come from, and I can’t pretend that someone else would conjure the same values. Maybe it’s a simple familiarity with numbers, or maybe people are just wired differently – I really don’t know!

Even if this works for some of us, it is unlikely to work for all of us. By itself, though, this is not a reason for dismissing it out-of-hand, or lashing out at the mathematically-inspired amongst the community. A potential reaction such as ‘We happen to be qualified genealogists, and not bookmakers’ would say more about misplaced pride than considered analysis. Genealogists and bookmakers are both experts in their own fields. When they say they’re sure of something, they don’t mean absolutely, 100% sure, but to what extent are they sure?



[1] Elizabeth Shown Mills, Evidence Explained: Citing History Sources from Artifacts to Cyberspace (Baltimore, Maryland: Genealogical Pub. Co., 2009), p.19.
[2] If you’re thinking “logarithmic” then you would be wrong. The range is symmetrically asymptotic at both ends and so is hyperbolic.
[3] This simple form applies where each event has just two outcomes: a result happening or not happening. There is a more complicated form that applies where each event may have an arbitrary number of outcomes.
[4] I’m using the logical NOT sign (¬) here to indicate the inverse of an event’s outcome. The convention is to use a macron (bar over the letter) but that requires a specialist typeface.
[5] Yes, that’s right, 10% and 95% do not add up to 100%. The misunderstanding that they should plagues a number of examples that I’ve seen. The probability of finding the evidence if he was guilty, P(E | G), and the probability of finding the evidence if he was not guilty, P(E | ¬ G), are like “apples and oranges” because they cover different situations, and so they will not add up to 100%. However, the probability of not finding the evidence if he was guilty, P(¬ E | G), is the inverse of P(E | G) and so they would total 100%.

Monday, 13 January 2014

Using Microsoft Word with Blogger



The general consensus on this is ‘don’t do it’. However, it can easily be made to work, and it allows a much richer and better styled content to be posted.



In a slight break from my usual genealogy posts, I want to pass on some of my experiences with blogger.com in case they help someone else. Depending on the feedback from this one, I have further experiences that I may post.

The word blog was formerly weblog (i.e. a contraction of ‘Web log’) and was coined by Jorn Barger in December 1997[1]. The original content of a blog was commentary on Web links, or personal thoughts and essays. Contrary to recent reports[2], the blog is not dead; it has merely diversified. To understand how it has diversified, we need to look at the essential structure of a blog.

A blog is basically a serialised publication. Posts are made if-and-when the author(s) deem appropriate. The blog URL address will take you to the latest post, although any previous post can be revisited by using its specific URL address. Readers can subscribe in order to get a notification when a new post appears, and they can usually comment on posts to participate in some interaction with the author or other readers.

This basic structure has allowed the blog to be applied far beyond any concept of an online personal diary. Although many blogs are still concerned with news items in the author’s field of interest, the concept of an interactive serialised publication has found new uses such as advertising, special interest micro-publications (e.g. cooking recipes, or car restoration stages), blog fiction (i.e. serialised publication of narrative chapters), and technical presentations (for education, or for research and discussion).

The relevance of this bit of blog history and analysis is that some uses require more care and attention to their preparation than others. If your post is more than a few paragraphs, and you want to use a specific layout, or include endnotes or source lists, then the formatting tools provided by most blogs are too primitive. The longer you anticipate the relevance of your post to be, then the more effort you will want to invest on it. Yes, you can usually switch to editing raw HTML — as Blogger allows — but that’s outside of the skill-set of most authors. Also, why bother if you have access to a word-processor such as Microsoft Word.

So what is the issue with Microsoft Word? At first glance, it appears trivially easy to copy-and-paste from your Word document into the Blogger Compose window, and that was the route I used with my first few posts. Unfortunately, when I added support to notify subscribers via email, using a tool called feedburner, then it failed and no one was notified. Microsoft Word generates a lot of tags that are specific to Microsoft Office and this has two consequences: those Office-specific tags failed ‘validation’ by feedburner because it didn’t understand them, and the sheer volume of these tags (many of which are quite superfluous) regularly breaks a size limit within feedburner.

The consensus on Blogger forums is simply to “flatten” the post by removing all Word formatting (e.g. by pasting it into a simple text editor and copying it back), and then resurrect the formatting, as best you can, using Blogger’s own features. If you’ve used Word deliberately in order to craft a good presentation then this sort of help can be both frustrating and annoying. The suggestion may even be impossible because, as I’ve already indicated, Blogger’s features are more primitive as it’s not a professional formatting tool.

So what’s the answer? When your Word version is ready to be transferred, make sure you first save your definitive copy back to its native *.doc(x) file. Then, go back to the Save-As dialog and scroll down the ‘Save as type’ list to find the entry ‘Web Page, Filtered (*.htm, *.html)’. Save a temporary copy in this format, say on your desktop so that it doesn’t conflict with your master copy. Word will now be displaying an HTML version that has all the Office-specific tags filtered out, and you can safely copy-and-paste from what you see on your screen to the Blogger Compose window. When you’re done, you can delete the temporary copy from your desktop.

This copy-and-paste does not transfer your images but that's only a small issue. Your blog post will contain empty frames where your pictures should be, but these can be removed and your original pictures uploaded to Blogger and inserted at the correct position. The fidelity is generally very good, although it’s always wise to look at a ‘Preview’ before publishing a new post.



[1] Rebecca Blood, "Weblogs: A History and Perspective", Rebecca's Pocket, 7 Sep 2000 (http://www.rebeccablood.net/essays/weblog_history.html : accessed 12 Jan 2014).
[2] Jason Kottke, "The blog is dead, long live the blog", Nieman Journalism Lab, 19 Dec 2013 (http://www.niemanlab.org/2013/12/the-blog-is-dead/ : accessed 12 Jan 2014).