The thing I like about it is the variety of the talks. It’s not organized around a single programming language or technology, so it’s easy to go to talks outside your usual sphere of interest (it’s quite hard not to). I went to talks about MLOps (keynote by Luke Marsden), the Language Server Protocol (by Krzysztof Cieślak), a reactive JS compiler called Svelte (by Peter Allen), accessibility (by Svetlana Kouznetsova), Cloud Native ML (by Ant Kennedy), and building autonomous Mars rovers (by Mark Woods). I gave a talk on Single Cell data and algorithms (thanks to Steve Loughran for suggesting I submit it, as well as live tweeting it!).
Friday, 8 November 2019
Bristech 2019
The thing I like about it is the variety of the talks. It’s not organized around a single programming language or technology, so it’s easy to go to talks outside your usual sphere of interest (it’s quite hard not to). I went to talks about MLOps (keynote by Luke Marsden), the Language Server Protocol (by Krzysztof Cieślak), a reactive JS compiler called Svelte (by Peter Allen), accessibility (by Svetlana Kouznetsova), Cloud Native ML (by Ant Kennedy), and building autonomous Mars rovers (by Mark Woods). I gave a talk on Single Cell data and algorithms (thanks to Steve Loughran for suggesting I submit it, as well as live tweeting it!).
Saturday, 2 November 2019
How I manage my diabetes
In this post I describe the technology I’m using to manage T1D. Every PWD (person with diabetes) is different, so what works for me won’t necessarily work for others. Also, some of these things don’t even work for me all the time, such is the unpredictable nature of diabetes. So there are definitely improvements I could make. I’ll mention some of them at the end of this piece.
| Some of my diabetes kit |
Quite simply, the Libre is a superb piece of technology. It gives an amazing amount of insight into blood glucose levels - you can see what happens after you eat a particular food, or the effect of exercise, and even what happened to your levels during the night. I’ve been using one for over a year, and without it I really think I would be a lot more stressed about BG levels, and probably overcompensating for lows and oblivious to highs.
![]() |
| Nightscout |
![]() |
| Weekly BG summaries with dboard |
![]() |
| Carb counting a recipe with Ingreedy |
- Can I replace the Libre reader with the phone app?
- Do I need dboard given that both Libre and Nightscout have good analytics?
- Can I log insulin doses automatically (e.g. with CLIPSULIN)?
- Do I need to log meals and carbs?
- Is there a way of recording hypos? (Preferably automatically since when you are having a hypo you don’t think about logging stuff.)
- How can we make it easier to move someone’s entire T1D data between systems?
Friday, 27 April 2018
Pastures new-ish
On the personal side, my family and I lived in San Francisco during the early formative years of Cloudera, a time we will always treasure for the lifelong friendships we made.
On the professional side, it is no exaggeration to say that working at Cloudera has been the highlight of my career. I already knew that Hadoop was pretty special when I joined (I may have been biased as I was writing a book on it), but I had no idea how it would transform the industry and how it would be used in every sector you could imagine.
To all of you I have worked with over the last decade—at Apache, Cloudera and elsewhere, on many projects—I consider myself to be incredibly fortunate to have had the opportunity to work with you. Thank you.
So what’s next for me?
Jim Waldo, who worked on distributed systems at Sun, once said that he alternated six month periods between the lab and the outside world: in the lab he and his team built systems software, and in the outside world he saw how people used the system he was building. Doing so gave him valuable feedback on the system design, even though it was time away from being able to build the system.
In some ways this is another way of framing the explore/exploit tradeoff, where you decide between exploring new technological ground—building a new system—and exploiting that system to solve particular problems you are interested in, which is why you built the system in the first place. (Of course, this framing is oversimplified, since there are many people working on both parts simultaneously. It’s a useful way of thinking about things as an individual actor though.)
For the past few years I have been working on a few open source biology and healthcare projects (like GATK, Hail, and OHDSI). I think that the problems in biology are big enough and messy enough that new systems will need to be built. We can’t stop exploring the technological ground since the sheer amount of data will overwhelm even the best of today’s cutting-edge technology. (I like to cite the paper Big Data: Astronomical or Genomical? here for some concrete numbers.)
Having said that, there is still a lot of mileage left in our current crop of tools—which include Spark, TensorFlow, Jupyter, and the cloud. And this is what I am going to do: continue the work to apply tools like these to more bio projects, only now working as a freelancer. I plan to write more about what I’m up to on this blog, so please follow along.
Sunday, 15 April 2018
Type 1 Diabetes
Wednesday, 22 June 2016
Be Part of Something Bigger - Vote #Remain
Far from quelling the debate within the Tory party, the lead up to the referendum has had the opposite effect. The debate over the last couple of months has been increasingly toxic, with both sides making outlandish claims. Parts of the Leave campaign have been xenophobic and racist, in an attempt to scare people to leaving the EU - this is the true Project Fear. And then last week the appalling murder of the Labour MP Jo Cox brought about some reflection on how we’ve moved away from a more respectful, kinder politics. In the words of Stephen Kinnock, "When insecurity, fear and anger are used to light a fuse, an explosion is inevitable.”
But there is a referendum tomorrow, so we have a duty to vote. The vote is about a host of issues, and on all of them I believe we are better staying as a member of the EU. In my mind it boils down to being a part of something bigger than yourself. This is true on a personal level - being part of a company, a team, a club, or an organisation allows you to achieve more than if you go it alone. I’ve seen this in my professional life where loose-knit groups of programmers build open source software that an individual could never dream of. Of course, where there are many personalities pulling in different directions you get conflict, things get messy, compromises are needed, and you don’t always get your own way. But on some decisions you do have influence, and you do get to shape the results.
Being a part of the EU is about the UK being a part of something bigger, and being able to influence policy on issues that affect the UK. The world is a messy chaotic place, and there are many deeply-ingrained, complex problems that require complex policy interventions. Climate change, migration, tax havens, peace - to name a few - all of these need a coordinated approach that cross national borders. Leaving would squander our influence in attacking these problems, while doing nothing to solve them - for us or for the the rest of the world.
One of the more worrying themes of the Leave campaign is not to trust the experts. This allows them to conveniently dismiss the overwhelming opinion amongst economists that Brexit would mean the UK is worse off outside the EU. It’s like climate change denial, and running a country with that kind of gut-feeling policy making is terrifying.
Britain has been at its best when it has been an outward-looking nation, one that works with others and trades with others. That’s why I am going to vote to Remain in the EU.
Sunday, 5 July 2015
The Earth Moon Game
Before you read on, you might like to have a go yourself. If you don't have a tennis ball and basketball to hand, you can play with this online version I wrote.
My kids and I had a stall at our school fair this Friday where we played this as a game:
The Earth's diameter is 12,742 km and the Moon's is 3,475 km, so the Earth's diameter is about 3.7 larger. (We measured the basketball's diameter to be 23.5 cm, and the tennis ball to be 6.5 cm, so the ratio is about 3.6, which is pretty close!)
The Moon is (on average) 384,400 km from the Earth (the Lunar distance, measured from the centres of the two bodies), which is 111 times the Moon's diameter. Scaling this to the tennis ball, we get a distance of 111 × 6.5 cm = 7.2 metres.
Here's a picture showing the results at the end of the fair:
The basketball representing the Earth is in the bottom of the picture, and the tennis ball just visible at the top is the Moon, 7.2 metres away. To the right of the green tape are white flags that are the players' guesses for where the Moon would be.
It's striking that all the guesses were too low. This seems to be a mixture of two things. Firstly, people really do think that the Moon is closer than it actually is. Secondly, people tend to copy other people, so they would place their flags close to where the others were. (We told everyone that the Moon didn't have to be restricted to the green tape - that just happened to be how long it was.)
We saw a few interesting tactics though. One girl put one flag so it was the closest to the Earth compared to all the other flags, then another so it was the furthest out. She seemed to think that everyone else had either over- or underestimated the distance - which of course they had! (She didn't win though, as someone put their flag even further out later on.) Someone else put five flags over a range of about 25cm where she thought the Moon would be.
The most successful approach seemed to be for the player to stand where the Earth is, and have someone walk away holding the tennis ball until it subtends the same angle as the Moon does in the sky (or your mind's eye). This is easier said than done, however. The player in fourth place (who was about five years old) used this technique.
Here's the data plotted graphically, with each flag shown as a line. The blue line represents Earth, and the orange line the Moon.
Interestingly, the guesses did not benefit from the Wisdom of Crowds effect, where the average tends to be a good predictor of the actual answer:
The opening anecdote [of the book of the same name by James Surowiecki] relates Francis Galton's surprise that the crowd at a county fair accurately guessed the weight of an ox when their individual guesses were averagedFor the Earth Moon Game, however, the median distance was 2.6 metres, and the mean was 2.7 metres, which was 2.3 standard deviations (sd=1.96 metres) from the true distance, 7.2 metres.
Sunday, 19 April 2015
The Hay Dark Skies Festival, Reverend Thomas William Webb, and Jupiter
![]() |
| Young stargazers, Lottie and Millie |
The evening event was stargazing at Holy Trinity Church in Hardwicke, just outside Hay. Quite apart from the lack of light pollution, the location was a special one, since the vicar of the parish from 1856 until 1885 was Reverend Thomas William Webb, who in his spare time observed the night sky with telescopes and an observatory he had built himself.
![]() |
| Holy Trinity Church, Hardwicke |
In 1859, while at Hardwicke he wrote the classic book, Celestial Objects for the Common Telescope, the object of which was "to furnish the possessors of ordinary telescopes with plain directions for their use, and a list of objects for their advantageous employment".
The book remained in print well into the following century (and was recently republished by Cambridge University Press), and it's probably difficult to overemphasise the importance of this book in encouraging generation after generation of amateur stargazers.
In the words of Janet and Mark Robinson, who used to live in the vicarage and have edited a book about Webb,
Like Patrick Moore, he was an enthusiast who wanted to inspire as many people as possible to look through a telescope. Even at the choir party he "arranged the telescope and acted as showman and all in turn had a look at Saturn".Webb would no doubt have been pleased to see yesterday's gathering of enthusiastic amateurs (including the Robinsons) with an impressive range of telescopes, on a cold but very clear night. The highlight for us was seeing Jupiter and its four brightest moons (Io, Europa, Ganymede and Callisto) through a large reflecting telescope. We could even see the north and south belts, and the Great Red Spot (or Pink Splodge as Lottie named it).
![]() |
| Sunset. Venus is visible top centre |
Sunday, 8 March 2015
Tennis Ball Parabola
Millie filmed the video and edited it down to a shorter segment. I turned the resulting video frames into a series of JPEGs by running:
ffmpeg -i Tennis\ Ball.mp4 tennis-%03d.jpeg
Then I composed them into a single image using ImageMagick:
convert -compose lighten tennis-014.jpeg tennis-015.jpeg \
-composite tennis-016.jpeg \
-composite tennis-017.jpeg \
...
Millie then used Desmos (an online graphing editor) to superimpose a parabola on the image.
Update: Dima Spivak suggested I use the picture to estimate g, the acceleration due to gravity.
- My head measures 0.22 m (chin to crown), and is 49 pixels on the picture.
- The vertical distance, d, from the highest ball to the ball above Lottie's hands is 204 pixels, or 0.916 m.
- The time, t, it took to travel this distance was between 12 and 13 frames (it's hard to say more precisely than this from the picture), which at 29.97 frames per second is between 0.4 and 0.434 seconds.
Friday, 16 January 2015
Hadoop for Science
Open Data
Amazon S3 seems to be emerging as the de facto solution for sharing large datasets. In particular, AWS curates a variety of public data sets that can be accessed for free (from within AWS; there are egress charges otherwise). To take one example from genomics, the 1000 Genomes project hosts a 200TB dataset on S3.Hadoop has long supported S3 as a filesystem, but recently there has been a lot of work to make it more robust and scalable. It’s natural to process S3-resident data in the cloud, and here there are many options for Hadoop. The recently released Cloudera Director, for example, makes it possible to run all the components of CDH in the cloud.
Notebooks
By "notebooks" I mean web-based, computational scientific notebooks, exemplified by the IPython Notebook. Notebooks have been around in the scientific community for a long time (they were added to IPython in 2011), but increasingly they seem to be reaching the larger data scientist and developer community. Notebooks combine prose and computation, which is great for exposition and interactivity. They are also easy to share, which helps foster collaboration and reproducibility of research.It’s possible to run IPython against PySpark (notebooks are inherently interactive, so working with Spark is the natural Hadoop lead in), but it requires a bit of manual set up. Hopefully that will get easier—ideally Hadoop distributions like CDH will come with packages to run an appropriately-configured IPython notebook server.
Distributed Data Frames
IPython supports many different languages and libraries. (Despite its name IPython is not restricted to Python; in fact, it is being refactored into more modular pieces as a part of the Jupyter project.) Most notebook users are data scientists, and the central abstraction that they work with is the data frame. Both R and pandas, for example, use data frames, although both systems were designed to work on a single machine.The challenge is to make systems like R and pandas work with distributed data. Many of the solutions to date have addressed this problem by adding MapReduce user libraries. However, this is unsatisfactory for several reasons, but primarily because the user has to explicitly think about the distributed case and can’t use the existing libraries on distributed data. Instead, what’s needed is a deeper integration so that the same R and pandas libraries work on local and distributed data.
There are several projects and teams working on distributed data frames, including Sparkling Pandas (which has the best name), Adatao’s distributed data frame, and Blaze. All are at an early stage, but as they mature the experience of working with distributed data frames from R or Python will become practically seamless. Of course, Spark already provides machine learning libraries for Scala, Java, and Python, which is a different approach to getting existing libraries like R or Pandas running on Hadoop. Having multiple competing solutions is broadly a good thing, and something that we see a lot of in open source ecosystems.
Combining the Pieces
Imagine if you could share a large dataset and the notebooks containing your work in a form that makes it easy for anyone to run them—it’s a sort of holy grail for researchers.To see what this might look like, have a look at the talk by Andy Petrella and Xavier Tordoir on Lightning fast genomics, where they used a Spark Notebook and the ADAM genomics processing engine to run a clustering algorithm over a part of the 1000 Genomes dataset. It combines all the topics above—open data, cloud computing, notebooks, and distributed data frames—into one.
There’s still work to be done to expand the tooling and to make the whole experience smoother, nevertheless this demo shows that it's possible for scientists to analyse large amounts of data, on demand and in a way that is repeatable, using powerful high-level machine learning libraries. I'm optimistic that tools like this will become commonplace in the not-to-distant future.
Sunday, 11 January 2015
Marmalade
Sunday, 13 October 2013
Five years at Cloudera
Before I joined I had been working as an independent Apache Hadoop consultant for a year (probably the first Hadoop consultant anywhere), and was halfway through writing a book on Hadoop. The interview process had involved speaking to all four founders, and I remember when I came off the phone after the last call it was late in the UK but I couldn't sleep because the vision they had described was exactly what I wanted to see for Hadoop: a company that wanted to make Hadoop accessible to everyone, by making it easier to use and run, while maintaining a strong commitment to open source. The last point sealed the deal for me, and really at that point there was no way I could not join, and five years on I can say without exaggeration that it was the best decision of my professional life.
When I started I was living in Wales, which meant that on my first day I didn't see any of my new colleagues! That was remedied a few weeks later on when I visited California (and ApacheCon in New Orleans) in early November 2008. Initially the others were working out of a single room in AdMob's offices in San Mateo, but it wasn't long before we moved to a smart brick-lined office in Burlingame. I was around for the moving in day, which involved more flatpack assembly skills than programming.
From the very beginning we worked on making Hadoop easier to use, run, and support, and better integrated with other systems, so that it could enjoy broader adoption. That was borne out in the early projects at Cloudera which included creating training material, creating packages for Red Hat and Debian (CDH, and later Bigtop), writing tools for data ingest (Flume and Sqoop), creating a rich web UI for Hadoop users (Hue), as well as making contributions to the core project. I was mainly involved in the latter, which I did at the same time as completing the book in time for the Hadoop Summit 2009, which would never have been possible without the time and space my teammates gave me.
Over the first year I would visit every three months or so, and naturally each time the team would have grown. I always enjoyed meeting the new people who had joined since my last visit, but I realized that at such a formative time in a company's life, when the culture was being laid down that being closer to the team would make it easier for me to stay involved. The opportunity to move to California came up, and on the last day of October 2009 I arrived in San Francisco with my wife, Eliane, and two girls.
As anyone who has moved to a new country knows, there's a lot of things to sort out—somewhere to live, a school for the girls, reams of paperwork—and during this time the folks at Cloudera were incredibly helpful and supportive. When we moved into our new apartment (which Eliane had found a mere two weeks after we arrived) half of the engineering team turned up to help with Ikea flatpack assembly.
At the end of our three year sojourn in the US, we left having made many friends, sad to leave, but happy knowing we'd be living closer to our family again. Cloudera was an order of magnitude larger than when I had arrived, and was now an international company with offices in several countries across the world.
Over the last five years I've been lucky enough to have been given the freedom to work on many parts of the Hadoop stack, in different parts of the Hadoop community, and with different teams at Cloudera. In the course of doing so, I've worked with the most talented and intelligent group of people in my life. It's hard work, and challenging, but also a lot of fun and incredibly enriching. I have every reason to expect it to continue. Thanks Cloudera!
Update on October 14: reworded to state that ApacheCon 2008 was held in New Orleans, not California. Thanks to Isabel Drost-Fromm for pointing out the error.
Saturday, 30 March 2013
Making a Kitchen Table
Sunday, 3 February 2013
Have you put the chickens to bed?
The problem with the alert is that it is set to go off at sunset, which is all that IFTTT allows, and that's a bit too early as it's not dark enough for the chickens to be in their house. So we wait a bit, then we forget.
So I decided to write an Android app to send an alert a fixed amount of time (say 45 minutes) after sunset, so that when we received it, it would be dark, the chickens would be in their house, and we could close the door there and then.
This is the result:
Eliane is currently beta testing it, so we'll see how well it works. (Obviously the long term goal is an automatic sensor to open and close the chicken house door, but we're not there yet.)
Writing Android Apps
What's Next?
Source is on GitHub.
Monday, 31 December 2012
How far away is the sea?
You can try it out at http://how-far-away-is-the-sea.appspot.com/. It works well on phones too, so you can use it when you are out and about.
How does it work?
I used the dataset of land polygons from Natural Earth, which as the name suggests covers the whole world. The scale is 1:10 million, so inevitably there is some inaccuracy near the coast, particularly where it's wiggly.The app uses your current location (or a location you selected by clicking on the map) and computes the closest point in the set of land polygons. This calculation is performed using the JTS Topology Suite, a library for 2D spatial work, and it runs as a Java webapp hosted on Google App Engine.
Originally I used Geotools to perform the geospatial calculations, but unfortunately it doesn't run on GAE, so I wrote an offline tool to convert the Natural Earth shapefiles to a JTS binary format. JTS works fine on GAE, but it lacks a distance calculator. Luckily spatial4j has the requisite distance functions, and it too works on GAE.
The webapp exposes a simple query endpoint, so a request for the following URL, for example:
http://how-far-away-is-the-sea.appspot.com/query?lat=51.856479&lng=-3.13551
will return a JSON document with the closest point on the coast, whether the (origin) location is on land or at sea, and the distance in metres to the coast:
{
"latitude":51.856479,
"longitude":-3.13551,
"coastLatitude":51.55853913000007,
"coastLongitude":-2.984038865999878,
"onLand":true,
"distanceToCoast":34734.59501052392
}
The page that the user sees is a simple static HTML page that uses the Google Maps API (v3) to render the map and the markers, and jQuery to query the Java webapp.
The complete source code is on Github at https://github.com/tomwhite/how-far-away-is-the-sea.
Further ideas
Some of the polygons are a poor approximation to the coastline, so it would be nice to get a higher-resolution dataset. There are likely many potential sources, such as this one for the UK.It would be interesting to use the dataset to answer the question: "which is the furthest point from the sea [in the UK/in X/in the world]?". I'd like to find time to do that sometime. Adding in spatial indexes might be helpful too.
If you liked this app then you might like...
Is it day or night?Sunday, 16 December 2012
IFTTT
If this [trigger] occurs then perform that [action].
There are lots of triggers and actions, provided by channels. For example, the Weather Channel provides a trigger which fires at sunset. And the Google Talk Channel provides an action to send a chat message. I combined the trigger and action into a recipe called "Did you put the chickens to bed?" which will remind me (and Eliane) to close the chicken shed in the evening.
Sunday, 9 December 2012
Apportionment
[I wrote this in July, but never got round to posting it.]
Last weekend I visited the U.S. Capitol in Washington, D.C., with my family, and I learned that the House of Representatives has 435 seats which are appointed so that each state has a number of seats that is proportional to its population. It sounded simple when the tour guide said it, but I wondered how are fractions handled fairly? Simply rounding off quotas doesn't work—firstly because some states could get no seats, which would be unfair, and secondly, how do you make sure that the rounding is both fair and assigns all 435 seats?
When I got home I read about the apportionment problem, as it is known, which has a long and interesting history. Wikipedia [1] is a good read, as usual; and [2] goes into the history and mathematics of different apportionment algorithms in depth, at least one of which suffers causes a paradox. Here I'm interested in looking at the algorithm that is used today to calculate apportionments for the House of Representatives, and why it is considered to be the fairest.
The Algorithm
The algorithm in use today for apportioning seats is due to Huntington and Hill and is known as the Huntington-Hill method, or the method of equal proportions. It's best understood as a dynamic process, which works as follows:
To start, each state is given one seat. (This ensures that states with relatively small populations, like Wyoming, get at least one seat.) Then, each remaining seat is allocated in turn to the state is allocated to the state with the highest priority, where the priority of a state of population \(P\) and \(n\) previously-allocated seats is defined as
\begin{align} \frac {P} {\sqrt{n(n+1)}}\label{pri} \end{align}
We'll see why the priority is defined as it is below, but for now notice that it is approximately \(P/n\), so the seat is given to the state that has the least number of representatives per person, roughly speaking.
Results for the 2010 Census
Running the algorithm for the state populations from the 2010 Census (using a program I wrote [5]) gives the following apportionment, which agrees with the U.S. Census Bureau [3]. (The quota column is the percentage of the population for each state.)
| State | Seats | Population | Quota | People per representative |
|---|---|---|---|---|
| Alabama | 7 | 4802982 | 6.76 | 686140 |
| Alaska | 1 | 721523 | 1.02 | 721523 |
| Arizona | 9 | 6412700 | 9.02 | 712522 |
| Arkansas | 4 | 2926229 | 4.12 | 731557 |
| California | 53 | 37341989 | 52.54 | 704565 |
| Colorado | 7 | 5044930 | 7.10 | 720704 |
| Connecticut | 5 | 3581628 | 5.04 | 716325 |
| Delaware | 1 | 900877 | 1.27 | 900877 |
| Florida | 27 | 18900773 | 26.59 | 700028 |
| Georgia | 14 | 9727566 | 13.69 | 694826 |
| Hawaii | 2 | 1366862 | 1.92 | 683431 |
| Idaho | 2 | 1573499 | 2.21 | 786749 |
| Illinois | 18 | 12864380 | 18.10 | 714687 |
| Indiana | 9 | 6501582 | 9.15 | 722398 |
| Iowa | 4 | 3053787 | 4.30 | 763446 |
| Kansas | 4 | 2863813 | 4.03 | 715953 |
| Kentucky | 6 | 4350606 | 6.12 | 725101 |
| Louisiana | 6 | 4553962 | 6.41 | 758993 |
| Maine | 2 | 1333074 | 1.88 | 666537 |
| Maryland | 8 | 5789929 | 8.15 | 723741 |
| Massachusetts | 9 | 6559644 | 9.23 | 728849 |
| Michigan | 14 | 9911626 | 13.94 | 707973 |
| Minnesota | 8 | 5314879 | 7.48 | 664359 |
| Mississippi | 4 | 2978240 | 4.19 | 744560 |
| Missouri | 8 | 6011478 | 8.46 | 751434 |
| Montana | 1 | 994416 | 1.40 | 994416 |
| Nebraska | 3 | 1831825 | 2.58 | 610608 |
| Nevada | 4 | 2709432 | 3.81 | 677358 |
| New Hampshire | 2 | 1321445 | 1.86 | 660722 |
| New Jersey | 12 | 8807501 | 12.39 | 733958 |
| New Mexico | 3 | 2067273 | 2.91 | 689091 |
| New York | 27 | 19421055 | 27.32 | 719298 |
| North Carolina | 13 | 9565781 | 13.46 | 735829 |
| North Dakota | 1 | 675905 | 0.95 | 675905 |
| Ohio | 16 | 11568495 | 16.28 | 723030 |
| Oklahoma | 5 | 3764882 | 5.30 | 752976 |
| Oregon | 5 | 3848606 | 5.41 | 769721 |
| Pennsylvania | 18 | 12734905 | 17.92 | 707494 |
| Rhode Island | 2 | 1055247 | 1.48 | 527623 |
| South Carolina | 7 | 4645975 | 6.54 | 663710 |
| South Dakota | 1 | 819761 | 1.15 | 819761 |
| Tennessee | 9 | 6375431 | 8.97 | 708381 |
| Texas | 36 | 25268418 | 35.55 | 701900 |
| Utah | 4 | 2770765 | 3.90 | 692691 |
| Vermont | 1 | 630337 | 0.89 | 630337 |
| Virginia | 11 | 8037736 | 11.31 | 730703 |
| Washington | 10 | 6753369 | 9.50 | 675336 |
| West Virginia | 3 | 1859815 | 2.62 | 619938 |
| Wisconsin | 8 | 5698230 | 8.02 | 712278 |
| Wyoming | 1 | 568300 | 0.80 | 568300 |
The Mathematics
The algorithm finally settled on by Congress was chosen because it was thought to be the fairest. There are different ways of defining what "fair" means, and so it cannot be settled mathematically. In this context "fair" is taken to mean "minimizes the relative difference in representatives per person between states".
To see how the algorithm meets this definition of fairness, let's see what happens when we examine any two states to see if transferring one seat between them would improve the apportionment. This is the argument published by E. V. Huntington in [4].
Suppose after the apportionment, state \(A\) has received \(x+1\) seats, and state \(B\) has received \(y\) seats. Furthermore, also suppose that \(A\) is over-represented because the number of people per representative is less than for \(B\):
\begin{align} \frac {A} {x+1} &\lt \frac {B} {y}\label{Aover} \end{align}
We can check this in the case of California and New York:
\begin{align} \frac {37,341,989} {53} &\lt \frac {19,421,055} {27}\nonumber \end{align}
Or
\begin{align} 704565.83 &\lt 719298.33 \nonumber \end{align}
Now let's see what happens if we try to transfer one seat from \(A\) to \(B\)—does that make things fairer?
In the round when \(A\) won its last seat (number \(x+1\)), we know that its priority (defined by (\ref{pri})) was higher than \(B\)'s. That is,
\begin{align} \frac {A^2} {x(x+1)} &\gt \frac {B^2} {y(y+1)}\label{priority} \end{align}
(Note that even if \(B\) hadn't won its last seat (number \(y\)) at that point, the inequality still holds, since the number of seats it had would be less than \(y\).)
Again we can check this in the case of California and New York:
\begin{align} \frac {37,341,989^2} {52 \times 53} &\gt \frac {19,421,055^2} {27 \times 28}\nonumber \end{align}
Which is true. (The numbers also tally with the U.S. Census Bureau [6], and my program to calculate apportionments [5], where the priority value for California's last seat is \(711,308\), which is \(37,341,989/\sqrt{52 \times 53}\).)
Dividing (\ref{priority}) by (\ref{Aover}) we get
\begin{align} \frac {A} {x} &\gt \frac {B} {y+1}\label{Bover} \end{align}
which we can interpret as saying that \(B\) would be over-represented if one seat were transferred to it from \(A\). For our example of California and New York, this becomes
\begin{align} 718115.17 &\gt 693609.11 \nonumber \end{align}
The question now is, which over-representation is the smallest? That is, which is fairer, and therefore, to be preferred?
Using (\ref{Aover}), we calculate the relative difference before the transfer as
\begin{align} \newcommand{\slfrac}[2]{\left.#1\middle/#2\right.} \slfrac{ \left( \frac {B} {y} - \frac {A} {x+1} \right) } {\frac {A} {x+1}} = \frac {B(x+1)} {Ay} - 1 \label{Adiff} \end{align}
And, using (\ref{Bover}), the relative difference after the transfer is
\begin{align} \newcommand{\slfrac}[2]{\left.#1\middle/#2\right.} \slfrac{ \left( \frac {A} {x} - \frac {B} {y+1} \right) } {\frac {B} {y+1}} = \frac {A(y+1)} {Bx} - 1 \label{Bdiff} \end{align}
To compare these relative differences, note that we can rewrite (\ref{priority}) as
\begin{align} \frac {A(y+1)} {Bx} &\gt \frac {B(x+1)} {Ay} \end{align}
Thus
\begin{align} \frac {A(y+1)} {Bx} - 1 &\gt \frac {B(x+1)} {Ay} - 1 \end{align}
and the relative difference is smaller before the seat transfer (using (\ref{Adiff}) and (\ref{Bdiff})). So the original apportionment is optimal. There was nothing special about the choice of \(A\) and \(B\), so we can conclude that the apportionment is optimal overall.
Again, this checks out for our example. The relative difference for 53 seats for California and 27 for New York is \(0.021\), versus \(0.035\) for 52 for California and 28 for New York.
References
[1] United States congressional apportionment, Wikipedia. ↩
[2] Apportionment: Introduction, American Mathematical Society. ↩
[3] "APPORTIONMENT POPULATION AND NUMBER OF REPRESENTATIVES, BY STATE: 2010 CENSUS", U.S. Census Bureau. ↩
[4] The Apportionment of Representatives in Congress, E. V. Huntington, Transactions of the American Mathematical Society, Vol. 30, No. 1. (Jan., 1928), pp. 85-110. ↩
[5] A program to calculate apportionments, Tom White, July 2012. ↩
[6] PRIORITY VALUES FOR 2010 CENSUS, U.S. Census Bureau. ↩
Monday, 3 December 2012
d3troit
- For more information on Detroit's history, particularly its buildings, I highly recommend Dan Austin's http://historicdetroit.org.
- The Guardian has an amazing photo gallery of some of Yves Marchand and Romain Meffre's pictures of Detroit in ruins.
- How to Bring Detroit Back From the Grave by Josh Harkinson, Mother Jones.
- The Maker Culture is Reinventing Detroit by Gina Clifford, Wired.
- We didn't get to see it, but a friend highly recommended the heidelberg project, a street art project.
- Eliane on our visit: A journey into the past: Detroit in words
How I wrote the visualization
Thursday, 10 May 2012
Volcanoes!
We've been on a bit of a volcano tour recently. First we visited Lassen Volcanic National Park in October (climbing the Cinder Cone was a highlight), and we stopped in on Mount St. Helens visitor center on our way to Seattle last month. Yesterday we ventured into the Yellowstone caldera (the bit that blew out in the last eruption).
Before reading the book I hadn't appreciated how recent our understanding of Yellowstone's geology is. It was only in the 1960s that scientists combined new empirical data about the ages of different rock formations in the park with the then emerging theory of plate tectonics. One of the scientists was Robert Christiansen of the U.S. Geological Survey, who, with Richard Blank, collected samples from all over Yellowstone and pieced together the puzzle of how Yellowstone formed. (He also wrote the definitive account of Yellowstone's geology in 2001.)
They realized that the series of calderas between Oregon and Wyoming were all eruptions caused by what is now known as the Yellowstone hotspot over the last 16 million years. The continental plate is moving south west, which makes the newer volcanoes appear in the north east.
This diagram from Wikipedia summarizes it nicely:
Saturday, 4 June 2011
What's new in Apache Whirr 0.5.0-incubating
fully endorsed by the ASF. Please read the full disclaimer.
In this release the Whirr development team have added many new features while still making the core more solid. This post covers some of the more important changes. The full list can be found in the release notes.
Improving the new user experience
Orchestrating multiple services on cloud instances is a challenge to make simple, and Whirr has sometimes been a little fiddly to get running. SSH settings, in particular, have been a common sticking point with new users. The new Whirr in 5 Minutes guide walks through the minimum number of commands you need to type to get a simple 3-node ZooKeeper cluster running in a few minutes. From there you can move on to the Quick Start Guide and the Configuration Guide.The sample configurations in the recipes directory in the distribution contain useful settings for running the services on a variety of cloud providers. Users are always encouraged to share their working configurations with the community.
New services
Elastic Search and Voldemort have been added to the roster of services that come with Whirr. This brings the total to six; adding to Apache Cassandra, Apache Hadoop, Apache HBase, and Apache ZooKeeper.API improvements
Whirr is still a young project so it is not surprising that its API is rapidly evolving. In WHIRR-245, the demarcation between the user API (for users who control Whirr clusters from Java) and the service API (for developers writing new Whirr services) was clarified. The user API can be found in the org.apache.whirr package; whereas the service API is in org.apache.whirr.service.You can find out more about writing Whirr services in this presentation (PDF).
The firewall API that service writers use to open ports for services was simplified and made more powerful in WHIRR-275.
Overriding scripts
This feature was actually introduced in Whirr 0.4.0-incubating, but it's useful enough to mention here. In older versions of Whirr, if you wanted to make a modification to the scripts that run on cloud instances - to tweak some settings, for instance - you would have to upload your modifications (as well as all the other scripts) to a publicly available web server (Amazon S3 was a common choice), then point Whirr at the new location. Not particularly difficult, but a big enough barrier to discourage users from trying it.The new approach is to push scripts to nodes from the launching machine, so you can just edit them locally before launch. Full instructions are covered in the FAQ.
Running scripts on nodes
In 0.5.0 the scripts that run on cloud instances have been broken up to be more fine-grained, so many services have individual start and stop scripts (WHIRR-266). Combined with the ability to run scripts on sets of nodes in the cluster (by ID or role), users now have more control of the cluster once it has launched (WHIRR-173). Try running whirr run-script at the command line to use this feature. There's a contrib script to run the Yahoo! Cloud Serving Benchmark (YCSB) against an HBase cluster, which takes advantage of the run-script command (WHIRR-287).Also useful is WHIRR-291, which allows you to launch "blank" nodes with no services running on them (in a "noop" role), and then, with
whirr run-script, run arbitrary scripts on them to bring them into the state you want.Custom service builds
Developers who work on services supported in Whirr will find the ability to push a custom build to a cluster very useful for testing (WHIRR-220). For example, if you are working on a ZooKeeper feature, you can build a ZooKeeper tarball with your new feature, then launch a cluster that uses this tarball by specifyingwhirr.zookeeper.tarball.url as a local file:// URL pointing to your tarball. Whirr will push the tarball to a temporary blob store container, then each node will download from there.I used a variation of this feature to try out a nightly Hadoop 0.22 build on a small Whirr cluster. In this case the tarball URL is not a local file, so Whirr doesn't copy the tarball to a blob store since it is already accessible from the cloud.
Service improvements
Whirr is only able to exist because of the powerful abstraction that jclouds provides for interacting with cloud providers. A great example of this power is the API that jclouds provides for discovering the hardware capabilities of an instance running on any provider. WHIRR-282 took advantage of the jclouds API to find the number of cores on a node to dynamically configure the number of slots in a Hadoop cluster. Previously, you had to set this manually for each cluster to take full advantage of larger image sizes.This is just the beginning - there is more work to use memory capabilities to set configuration (WHIRR-229), and to use hardware capabilities generally in services other than Hadoop.
Cluster state storage
In previous releases of Whirr, information about launched instances was stored in a file on the machine that launched the cluster (~/.whirr/<cluster-name>/instances). With WHIRR-288, it's now possible to store this information in a blob store instead (such as Amazon S3, although any jclouds-supported blob store can be used), which is useful if you want to control clusters from multiple machines.Bring Your Own Nodes
Or just BYON, for short. Many users have requested the ability to deploy to privately owned hardware - and jclouds added this feature in 1.0-beta-9. Whirr now has preliminary support for BYON clusters. In a nutshell, you write a YAML file enumerating the nodes to deploy to - their addresses, access credentials, etc. - then Whirr will start services on them. The nodes just need to have a base OS like Centos or Ubuntu installed. You can find an example BYON configuration in the recipes directory of the download.BYON is also useful for testing locally by using VMware or VirtualBox to host target nodes.
A hummingbird
Last, but not least, Whirr finally has a logo! Many thanks to Alison Wong, who designed it and donated it to the ASF.
Credits
I would like to thank everyone who helped with the 0.5.0-incubating release. We have a growing community, and we welcome feedback and help from new users and developers. If you'd like to get involved you can start by downloading the new release and joining us on the mailing lists.What's next?
It's difficult to make firm predictions about the contents of the next release since Whirr is an open source project with many open issues, but the general themes include:- Adding more services. In tandem, we want to make it easier to write new services by pushing common patterns into the core (e.g. WHIRR-326 is one example of this).
- Improving existing services. By making them more flexible, better configured, easier to manage.
- Adding more cloud providers. The latest release of jclouds supports 30 providers, and we need help testing more of them with Whirr.
- Implementing services using other configuration management tools, rather than bash scripting. Andrei Savu is working on using Puppet to write new services (WHIRR-255).
- Supporting elastic clusters, so new nodes can be added to running clusters (WHIRR-214).






















