Skip to main content

Posts

Showing posts with the label Science

Don't Know? Find Out!

In What We Know We Don't Know , Hillel Wayne crisply summarises a handful of research findings about software development, describes how the research is carried out and reviewed and how he explores it, and contrasts those evidence-based results with the pronouncements of charismatic thought leaders. He also notes how and why this kind of research is hard in the software world. I won't pull much from the talk because I want to encourage you to watch it. Go on, it's reasonably short, it's comprehensible for me at 1.25x, and you can skip the section on Domain-Driven Design (the talk was at DDD Europe) if that's not your bag. Let me just give the same example that he opens with: research shows that most code reviews focus more on the first file presented to reviewers rather than the most important file in the eye of the developer. What we should learn: flag the starting and other critical files to receive more productive reviews. You never even thought about that possi...

Carrot or Stig?

  The science first ... Stigmergy is a term for individuals collaborating or coordinating through the effect of their actions on the environment. It comes from the study of animal populations which demonstrate sophisticated emergent behaviours without any obvious control structure or even direct communication. If that sounds too abstract, think about wasp, ant, or termite colonies and the level of organisation suggested by their complex nest-building, foraging, feeding, and breeding patterns. You've probably heard about scent signals. Stigmergy observes that an individual will mark the path to a food source and that this influences the behaviour of other members of the community. They in turn also mark the path, reinforcing the signal and bringing still others to it.  If there are alternative routes to the food then the one which conveys most benefit, perhaps by requiring less energy to reach, will tend to get reinforced more. It can be instru...

Looking Forward to Risk Analysis

When I asked Twitter this: Anyone know of a course on risk analysis that could work as in-house training for software testers?   It's a topic we're interested in but the stuff I'm turning up is more to do with corporate analysis, register building , mitigation at the business level #testing #risk Paul Hankin  suggested a book, Superforecasting, The Art of Science and Prediction , by Philip Tetlock and Dan Gardner. As the title suggests the book is about prognosis rather than peril, but the needs of prediction and risk analysis overlap in interesting ways: understanding possible outcomes, identifying factors that contribute to those outcomes, and weighting the factors and their interactions. Tetlock and Gardner study forecasting and forecasters. They use a metric called the Brier score to assess the accuracy of forecasts, and, over time, and with repeated forecasts, it becomes clear that some people tend to make better forecasts than others. The Brier score...

Boxing Clever

Meticulous research, tireless reiteration of core concepts, and passion for the topic. You didn't ask, but if you had done that'd be what I'd say about the writing of Matthew Syed based on You Are Awesome — reviewed here a few months back — and now also Black Box Thinking . The basic thesis of the latter is captured nicely in a blog post of his from last year : Black Box Thinking can be summarised in one, deceptively simple sentence: learning from mistakes. This is the methodology of science, which has changed the world precisely because it is constantly updating its theories in the light of their failures. In a complex world, failure is inevitable. The question is: do we learn, or do we conceal and self-justify? Who wouldn't want to learn from their mistakes, you might ask? Lots of us, it turns out. The aviation industry tends to come out well in Syed's analysis. Accidents, mishaps, and near-misses are reviewed for ways in which future flights might be le...

Quite the Reverse

Cohen and Medley, in Stop Working & Start Thinking , say: Simple tests are not experiments ... A chef will bake a cake at different temperatures and find the one that gives the best results ... [but] we should only include [this test] in classical "science" if the "normal" situation is included as a control ... Every careful observation of a puzzling or new phenomenon should be matched to similar observations of well-understood or classical material. They go on to introduce some useful terminology: variables are the things that you will aim to alter in the experiment; all other factors that could vary, but which you will aim not to vary, are parameters . And they then describe three types of experiment concerned with investigating the possibility of a causal relationship between a variable, A, and an outcome, X. Deficit : run one experiment with A and one without A. Monitor the presence of X in both cases. If X is seen with A but not without A then pe...

A Different Class

In Stop Working & Start Thinking (which I also mentioned the other day )  Jack Cohen and Graham Medley want scientists to consider what science is and how they do it, as well as just getting on with it. To help explain this, they partition scientific answer-seeking like so: observation measurement investigation experiment And that's interesting in and of itself.  But the authors have been round the block and so recognise that this categorisation is not absolute, and that sometimes it might not be clear where a particular activity sits, and that some activities probably sit in multiple categories at different times and even at the same time. In science — and thinking, so this applies to you too, testers — generalisations are useful because they help us to frame hypotheses at relevant granularities. We’re all made up of atoms but a description of social deprivation in inner cities at an atomic level would unhelpfully obscure, for example, that higher-level conc...

Ignorance, Recognised

Stop Working & Start Thinking  is intended to help postgraduate students make profitable use of an essential piece of scientific equipment: their mind. I'm only a short way in, and finding it a bit dense at times, but there's already a few passages I'm loving. Here's one (page 15): Science asks questions, and it has a small variety of ways to look for answers. They are observation , measurement , investigation and experiment . Different kinds of problem need different approaches for their solution and one of the ways the experienced scientist knows which to use is that she or he has got it wrong many times in the past! This cannot be said too often or emphasised too much. Ignorance , recognised, is the most valuable starting place; all scientists should have many stories about where they were sure, and wrong; where they were ignorant but did not know it. Image: Goodreads Edit: I based two more posts on this book later: A Different Class and Quite the Rever...

Quality is Value-Value to Somebody

A couple of years ago, in It's a Mandate , I described mandated science : science carried out with the express purpose of informing public policy. There can be tension in such science because it is being directed to find results which can speak to specific policy questions while still being expected to conform to the norms of scientific investigation. Further, such work is often misinterpreted, or its results reframed, to serve the needs of the policy maker. Last night I was watching a lecture by Harry Collins  in which he talks about the relationship between science and democracy and policy. The slide below shows how the values of science and democracy overlap (evidence-based, preference for reproducible results, clarity and others) but how science's results are wrapped by democracy's interests and perspectives and politics to create policies. I spent some time thinking about how these views can serve as analogies for testing as a service to stakeholders. But Co...

Bug-Free Software? Go For It!

This post is a prettied-up version of the notes for my talk at the second Cambridge Exploratory Workshop on Testing last weekend. The topic for the workshop was When Testing Went Wrong .  Cold fusion is a type of nuclear reaction that, if it were possible, would provide a cheap, clean and safe form of energy. In 1989 two scientists, Fleischmann and Pons, made worldwide headlines when they claimed to have generated cold fusion in a test tube in their lab. Unfortunately, subsequent attempts to replicate their results failed, other scientists started to publicly doubt the experimental methodology, and the evidence presented was eventually debunked. Cold fusion is a bit of a scientific joke. Which means that if you are a researcher in that field - now also called Low Energy Nuclear Reactions - you are likely to have a credibility problem before you even start. And, further, that kind of credibility issue will put many off from even starting. Which is a shame because the potent...

What Do I Know?

My kids begged to go to the Funky Fun House this half-term. I've got nothing against these soft play barns particularly - I've been to stacks of them - but, for me, there's generally little that's funky about echoing industrial spaces crammed with primary-coloured foam, covered in crumbs and reeking of decades worth of half-eaten fish and chips. To be fair, though, this one is in a (ware) house and the girls do have a lot of fun . In fact, I used to have fun too when they wanted me to play on the thing with them. These days they just see me as the shoe and coat monitor and provider of snacks. (Oh, and somone to take the mickey out of in front of their friends.) And that's how it went down this time too, except that I was engrossed in a book called Are We All Scientific Experts Now? by Harry Collins . A book I read in its entirety at my sticky table, that blocked out the noise of the toddlers in the padded prison enclosure that I'd ended up sitting next t...

Testing Utility

Testing can take a lot of inspiration from the sciences and the scientific method and I've blogged about some concepts that I think cross over in the past. Here's a few examples: equipoise metascience mandated science The science around policy - and the policy around science - is particularly interesting because it mirrors in useful respects the relationship between a tester and a stakeholder. In  What makes an academic paper useful for health policy?  Christopher Witty looks at ways that scientists can better serve policy makers and  much of what he's saying is also relevant to testers who want to do their best to: put the most valuable information they can  into the hands of the stakeholders who are asking for it at a time where it's useful at a cost which is acceptable in a manner which is easily consumable and at the right level with caveats and methodology clear and biases minimised. Which is all testers, I hope. Image:  https://f...

Feyn Arts

The other day, I said I was reading Surely You Must Be Joking, Mr Feynman! by Richard Feynman and was captivated by it. I've finished it now, and I've pulled out a handful of quotes. I love this on bad (or as he puts it, cargo cult ) science and how strongly it relates to the way I want to perform and report testing: But there is one feature I notice that is generally missing in cargo cult science ... It's a kind of scientific integrity, a principle of scientific thought that corresponds to a kind of utter honesty - a kind of leaning over backwards. For example, if you're doing an experiment, you should report everything that you think might make it invalid - not only what you think is right about it: other causes that could possibly explain your results; and things you thought of that you've eliminated by some other experiment, and how they worked - to make sure the other fellow can tell they have been eliminated. Details that could throw d...

Testing Testing

Metascience, according to this article in Nature , is "the science of science ... It has its roots in the philosophy of science and the study of scientific methods" with a primary focus being the study of the reproducibility of experimental findings. The article points out "the decline effect , an idea ... that the size of an effect decreases over repeated replications," acknowledges experimenter expectancy effects and the power of double-blinding and expects that "self-examination can only strengthen the scientific process for all." And when we try something new in our testing, we subject that thing to testing, don't we? Don't we? Image:  https://flic.kr/p/q1ic

Equipoise

I took the National Institutes of Health 's course Protecting Human Research Participants  this week. It's aimed at people setting up experiments involving human subjects and covers areas such as risk (including identification, minimisation, compared to benefits - personal and communal), recruitment (including coercion, balance), rights of the participants (including consent, welfare, vulnerable groups) and statutes (including international research, differences between definitions in different US bodies). In a section on the design of a clinical trial (where interventions - such as drugs - are being compared for efficacy in treating some illness in humans), the term equipoise was introduced. The definition given is: Substantial scientific uncertainty about which treatments will benefit subjects most, or a lack of consensus in the field that one intervention is superior to another. and the course notes say: A state of equipoise is required for conducting research that...

It's a Mandate

If you stick around in testing for a while you'll doubtless encounter and perhaps get embroiled in discussions about whether testing is an art or a science. I won't rehearse the arguments - see The Appliance of Art  for a brief review - but I will mention a term I chanced across recently (although it seems to have been around since at least the 1980s) that looks interesting and relevant: mandated science . Salter, Levy, and Leiss   write : The term "mandated science" refers to the science that is used for the purposes of making public policy. Science, here, includes the studies commissioned by government officials and regulators to aid in their decision making. This scientific work is designed and carried out solely for the purpose of supporting particular regulatory decisions. According to them when a body commissions this kind of scientific work, it's clear that "the research should meet the test of good scientific work" but there's a tensio...

A Gradual Decline into Disorder

I like to listen to podcasts on my walk to work and I try to interleave testing stuff with general science and technology. The other day a chap from Cambridge University was talking about entropy and, more particularly, the idea that the natural state of things is to drift towards disorder. Entropy : "Historically, the concept of entropy evolved in order to explain why some processes (permitted by conservation laws) occur spontaneously while their time reversals (also permitted by conservation laws) do not; systems tend to progress in the direction of increasing entropy."  In an ironic reversal of its naturally confused state, my brain spontaneously organised a couple of previously distinct notions (entropy in the natural world and the state of a code base) and started churning out ideas: Is the development of a piece of code over time an analogue of entropy in the universe? Could we say that as more commits are made, the codebase becomes more fragmented and any orig...

The Appliance of Art

We're recruiting Senior Testers at the moment and one of the recent candidates said several times during interview that testing is a science. This is not a new topic in testing and his stance on it is not universally shared. Try  The Software Entomologist  or  IM Testy  or  Randy Rice  or consider the title of Myers' classic,  The Art of Software Testing . For what it's worth, I probably agree that testing is a science. Primarily,  the basic methodology of testers is also the basic methodology of scientists: formulate a model of the subject and then experiment to verify whether the model is a good fit with reality. It's indisputable that some testers have a nose for finding bugs, a gut feeling about risky areas, something in their bones that keeps them worrying away at a seemingly minor issue until they expose the major flaw, an intuition about where or how or when to poke the application under test in just the right way to cause it to ...