A quick note on testing research
"So when you use a test that has no bias whatsoever, i,e, a double blind randomized placebo controlled trial, homeopathy always fails. Always."
Uh. This is from a
comment on Andrew Collin's distressingly
misinformed blog post on Gillian Mc Keith. Throughout the comment thread produces comments which indicate a complete misunderstanding of science and alternative medicine (short story- individual doctors may be bad, but alternative therapists are almost always worse). I've mostly agreed with what people who disagree with him say, but I have to disagree with the commenter here.
Whenever we conduct a double blind randomised placebo controlled trial, we must be aware of certain things. First of all, while the scientists involved will do their best to randomise, they will not have done so perfectly. They probably recruited from a sample space that wasn't the entire country, and there will always be a non-compliance issue- firstly with people refusing to take the trial and some failing to follow the drug regime properly. The former is typically controlled more than the latter, which is much more difficult to measure. However, having a blinded placebo should theoretically protect us against most of those effects, but its true to say that all studies are not a truly random sample of the population.
Now, even if we ignore those issues (which, to be fair, we often can), then every single study we do is balanced to have a false positive rate. That is, when we calculate whats called our test statistic, we obtain the probability that we would have seen that result by mere chance. Usually, if there was a 5% or less chance of that occuring, we say that its likely that there is some difference between placebo and the active treatment. This means that if I was to conduct a perfect study 20 times on a homeopathic remedy, even if it was ineffective I would expect to think it was effective in exactly one of those studies. Be aware that sample size shouldn't affect this, because I have explicitly designed my study to HAVE a 5% false positive rate. Of course the fact that I needed to run 20 studies to get a positive is rather telling, but it does contradict the quote above.
But heres whats worse, and this is something that the excellent Mr Goldacre often forgets. When I run my experiments I make certain assumptions. In particular I usually assume that the mean of my outcome is normally distributed (effectively the probability of seeing a particular value is shaped like a bell curve, so I'll see less and less on either side of the mean). Now this is always not true. Unless we set up our example very carefully its unlikely that we'll actually get a perfect normal distribution (height is a stereotypical example of a normally distributed variable, but a normal distribution assumes one can have negative values, and one cannot have negative height). This means the theoretical probabilities that we've calculated are not going to be exactly true in practice. That said, theres a theorem, called the central limit theorem which says the more observations we get, the closer our means get to being normally distributed. This means that the bigger our survey gets, the happier we can be with our assumptions.
Yet this comes with an even bigger disclaimer. For most drugs its not actually hard to show a difference. The standard statistical test assumes that the outcome is identical for placebo and a drug. If the treatment and the placebo are physically different in some way this will almost never be the case (indeed I suspect that it is almost always true, which is a probabilistic statement we don't need to worry about too much here), so we need to worry about clinically significant differences- thats the standards most drugs need to meet (actually, they usually need to beat their competing drug).
So whats my point here? My point is that while double blinded placebo controlled trials are a good standard for the industry to have, theres an argument to be had that they can be misinterpreted, and caution must always be used when applying their results. In particular, pointing at any one study with an impressive result is a terribly bad idea. We can become confident when several such, independently conducted studies, get the same results. The fact that these trials can be misinterpreted is why homeopaths can mislead by pointing to studies where they have succeeeded- when actually looking over the body of research the trend is the reverse.
Labels: rant, statistics
Another rant on luck (a pointless shout against the void)
It is a subject of some amusement among my friends to chide me on dice beliefs. Many gamers become superstitious about dice, and attempt to convince me that they are unlucky or lucky, and that they have seen strange things.
So, here is an admission. Luck does, of course, exist. Random results are not always ``fair'', in a uniform way. There will be games where most rolls go against you (actually we'd expect that to happen about 50% of the time with two players), and perhaps even a run where this will happen. So yes, one can have bad luck in a particular instance or game. But there is no such thing as someone being intrinsically lucky or unlucky. If someone has had a terrible run of luck, in life and dice, then I would probably call them unlucky, but going forward I wouldn't expect them to be particularly unlucky in the future.
This is the gambler's fallacy, the belief that after a coin has flipped heads 10 times it is now extremely likely that it will flip tails (if anything, from a Bayesian point of view its rather more likely that it will flip heads!).
For those of you who are convinced that you are just unlucky, and can give me countless anecdotes of it, you just aren't. Focusing on dice here, here are the possibilities:
1)You keep using the same dice which are actually biased
2)Your opponents are cheating, either by using biased dice, or tricks
3)You are rolling the dice in such a way that they roll lower
4)You only remember the good times and not the bad
5)You have an inherent universal property called "luck" which inflicts you and causes you to magic dice to bad results.
4, is of course, the most probable situation. If we were to apply occam's razor, we would go with it. Luckily, we're all scientists here, so we can do some experimentation!
1 and 3 can be eliminated together. Get your dice and roll them repeatedly, making sure to record each result, trying to roll as you usually do. Then count up the frequency of the result, and apply something called the
chi squared test, where you sum up the squared difference between the number you expect and the number you acheived, and test it against a probability distribution. If you make your sample large enough, and your significance level low enough, you can be pretty certain of seeing whether your dice are biased or not.
If they are not, then 1 and 3 are eliminated. Hoorary! Supposing they are, and we want to fix the problem, we can conduct further experimentation- simply try rolling with some different dice, ideally produced by a different company, and see if your results are different. If they are, its your dice. If not, its your rolling. Both problems have simple solutions. For 1, buy new dice, for 3, buy a dice cup to roll your dice in- that should prevent you from having an unhealthy influence (or, alternatively, use your skills to win at games of dice!)
Number 2 can be tested in much the same way. You will need to meticulously record every single roll your opponent makes, and do so without alerting them that you are doing it. This isn't an easy target, but is vital to ensure that your test is fair and unbiased. If you discover that they are cheating... well the solution isn't easy there, but thats a whole 'nother post.
I've actually cheated a little here. I've yet to eliminate 5. Its a little harder to get rid of. But it can be done. What we need to is make sure that if we are getting bad results, that we eliminate any possible other cause. So we need to roll a range of dice, and we need to roll them using a range of different methods, on a range of different surfaces. We can also start dealing ourselves hands of cards, as we are apparently looking for a universal luck reduction principle here. If all of these lead to a bias, then perhaps the universe is truly against us....
Labels: rant, statistics
G-spot
So xkcd wrote an
amusing comic about
this article. Some researchers determined that the G-spot might not exist. Their proof? They asked some women whether they thought they had one. They expected identical twins to say yes, but this didn't happen....
Sigh. So they asked some people whether they had an elusive sexual feature, and if they said no, assumed they were necessarily correct? As responses within the article indicate, many may simply have not had a sexual experience which allowed them to reach such a spot. 56% of women surveyed said they had one, and there was no genetic link. The abstract does suggest that it could be that the other women simply haven't found it, but then postulates that its really because it does not exist.
Sadly I do not have access to the article in question- at least online. I have no idea if my library carries the journal or not. Apparently the women in question filled in a thorough questionnaire. If I am correct that those women who said they did not have one may not have found it in some cases there should be a difference in the sexual behaviours of the two groups. Sadly without access to the survey I can proceed no further, but the claims this paper make are interesting. Universities do love a good press relief after all, and its much less impressive to say "study shows some women cannot find their g-spot, but is unable to show whether this means it is there or not".
Labels: rant, statistics
Experimental design
I enjoy making long blog posts about subjects which no-one will understand. So for your(lack of) entertainment, heres a mini-post on experimental design, the field of my research:
Most scientists, at some point, will need to perform experiments. Typically the process they are looking at will not be devoid of error. That is, when you measure someones height, you are unlikely to a-get the same answer each time and b-get the correct answer (most people would round up to 1cm at least, which imposes inaccuracies in your measurements). Such errors mean that even if you are certain as to what causes the process, your predictions would be out by some amount.
Such errors can usually be given a probability distribution, and we want to do our best to minimise these, so our predictions can be as accurate as possible. If we have a model to describe a process, we might know the form, but not the parameters of said model. Lets suppose we have a model for the temperature of a meal. We say that
temperature=a+b*time cooked for. We know age, and we can measure height, but for us to make predictions we need the values of a and b. Well we can run experiments, cooking our items for certain amounts of time, and then measuring the temperature. If there was no error, there would be no need to run more than two experiments, as we could perfectly estimate these parameters. However, we exist in the real world, and experimental error is a fact of existence. So lets suppose our food can be baked for between 10 and 200 minutes. If we had 10 cakes to test this with, a naeive approach might be to take these evenly spaced across time. In fact if you use statistical theory you can do much better, and take 5 observations at 10 minutes, and 5 at 200. Why? Well we only have two things to estimate, a and b, so we only need to take a minimum of two observations as we know. The thing we need to estimate here is the amount of error in our estimations, so repeating our observations at these point helps to minimise it.
This result is counter-intuitive, and important, because it extends. If we had two things we could vary (time cooked and weight of cake), an instinct for many experimenters would be to vary only one thing at a time- so look at changing time, while holding weight fixed, then holding time and varying weight. The best thing to do is to vary both at once, because this allows you to see if weight and time are interacting in any way.
Now this is an extremely simplistic look at the subject, with lots of the subtlties glazed over, but some important things to note are:
Experimental design has been demonstrated to be effective in multiple situations. Given a set of goals the experimenter wants to acheieve, statisticans can almost always find a design that will do better than the experimenter currently use.
Huge amounts of scientists are completely ignorant of this field.
Its an interesting field, with many problems left to solve, and one that most people, including many mathematicians, are entirely ignorant of.
Labels: phd, statistics
The importance of conditional probability (why screening is not always a good idea)
Jade Goody died from cervical cancer. This has caused some people to call for the screening age for cervical cancer to be lowered to the age of 20. A laudable goal, perhaps? Indeed, if we have tests, why don't we screen for every type of cancer?
Sadly, unless a test is very good, there will be many false positives- a simple example: Suppose 1 in a 1000 people gets cervical cancer, and lets suppose our test is 99% accurate. That is, it will miss the disease one percent of the time, and falsely claim the disease is present when it is not 1% of the time. Bayes theorem allows us to calculate the probability that someone who is declared positive for the disease actually has it. Rather than subject you to the equation, I will explain how it actually works.
Suppose I scan 1000 people. Of those, 1 of them will (on average), have the disease, so we correctly identify this with a 0.99 percent probability. There are 999 people remaining, and our test is 99% accurate, so thats 9.99 people diagnosed falsely with the disease.
So thats 0.99 who actually have the disease, and 9.99 who do not! And this example is actually generous: Generally speaking tests are MUCH worse than this, and I'm not sure the disease is even that prevalent.
Now one can increase these probabilities greatly by repeating the test, providing that we accept that our patient has the disease only if both are positive. Still, this is a lot of cost, and worry for the patient who has endured this. This is why screening tests are generally saved for those at risk to the disease. Bear in mind that we only have finite resource, and if the NHS spends a lot on screenings tests, while some people who wouldn't be picked up won't be, who knows who will suffer thanks to the massive waste in resources in checking these hundreds who do not actually have the disease.
This result is not immediately obvious, when looking at probabilities, but it is vital, and sadly not known by the majority of people
Labels: maths, rant, statistics