Thursday, 30 January 2014

Some tips for charts in Fusion Tables info windows

I've recently been working with some mortgage lending data from banks in Great Britain to produce a new mortgage lending maps site. Once again, I've used Google Fusion Tables to map the data because it's relatively quick and easy - further information can be found here. What's not so easy is getting the info windows to do exactly what you want, particularly when you want to include charts of the type shown below. In this post, I explain a little more about how you can get the info window to display such a chart, what can go wrong when trying to do so and what the code underlying code looks like.


The chart shown above is what you'll see when you click on any polygon in my mortgage lending map website. It takes the data from the underlying Fusion Table and provides a unique info window and chart for each postcode sector - in the case of the above it is for the Cardiff postcode sector of CF24 4, which as you can see had more than £117 million of outstanding mortgage debt owed by 839 households at the end of June 2013. The default info windows created in Fusion Table maps simply contain a number of default data records from columns from the underlying table. With this data, I was keen to show a visual comparison of lending mix in each area, and I wanted to do this using a horizontal bar chart. I'd done info window charts before (e.g. in my Deprivation in Scotland website) but these were line graphs showing change over time.

There is some general help on putting charts in info windows from Google, and this is a very good place to start, and you can also find a lot of explanation for what the different bits of chart code mean in Google's Charts Gallery, or in the Chart Feature List, but to help anyone who might be trying something similar to what I've produced, I thought I would provide an annotated code example. I also want to provide some general troubleshooting advice. Here's what the code looks like inside the Fusion Tables interface:



And here is a Word document with comments added explaining what each bit of code does. You'll notice that it's a bit messy but it produces a very nice looking chart. One thing that I noticed about all this is that if you want the numerical data to appear in the info window with comma separators - as above, for the £117 million figure - then it appears to disappear from the chart. This is what happened to me anyway. My advice would be to keep the number format for your data as 'none' in the Fusion Table field options. My other advice would be to remember to save a good version of your code but most of all to experiment with different things and then share the results.

I hope someone finds this useful! I know I would have before I got into this. 

Tuesday, 26 November 2013

A World of Open Access

I've recently been appointed as one of the Editors-in-Chief (with Alex Singleton) of a new Taylor and Francis/Regional Studies Association open access journal. For the two years prior to the launch of Regional Studies, Regional Science I was also heavily involved in the research and development of the journal. As we're about to start publishing papers, I thought I'd blog on the topic of open access more generally and include some interesting data from the Directory of Open Access Journals, the most authoritative source of open access information on the web. Those with an interest in open access will of course know all about DOAJ and the fact that there are now nearly 10,000 open access journals in the world. 9,990 to be exact (as of 26 November 2013).

However, few people have looked closely at the data on open access; probably because most people are still in debate about the merits and pitfalls of open access itself. The simple fact is that open access publishing is having a major impact on academia and the biggest journal in the world (by volume of papers) is now PLoS ONE, an open access title, as documented by several commentators including Heather Morrison at the University of Ottawa. Now for some data and charts...

The United States has over 1,200 open access journals and seven countries account for more than 50% of all open access journals. Journals come from a total of 124 nations (click charts to enlarge):

by country - full size chart


Open access publishing is currently in a major phase of expansion, but it's not new - for example, 22 new titles were launched in 1985. The peak year to date has been 2011 with 1,099 titles being launched:

by year started - full size chart


The vast majority of open access titles do not charge a publication fee (66%) but a substantial minority do (26%), including many of the best known titles:

by publication fee - full size chart

Your knowledge of, and exposure to, open access will vary greatly by discipline. There are over 500 open access titles in Medicine and more than 160 in Political Science but only 3 in Geology (according to DOAJ):

by discipline - full size chart

The majority of open access titles publish only in English (5,538 or 55%), with the next closest language of publication being Spanish (621 or 6%):

by language - full size chart

The data these charts are based on can be downloaded directly in CSV format from the DOAJ FAQ page. Just scroll down and look for the section entitled "How can I get journal metadata from DOAJ?".

There's a lot of activity in the field of open access but it is highly unequal in terms of its geographic, disciplinary and linguistic distribution. In the subject fields of Regional Studies and Regional Science (the subject areas for our new title) the open access landscape is considerably less crowded - particularly in relation to titles supported by major international learned societies and multinational publishing houses. Given this situation, we expect that Regional Studies, Regional Science (already known more commonly as RSRS) will play an important role in helping improve access to knowledge in regional research across a wide range of disciplines, with a focus on geography, planning and economics.

Look out for our first articles in December, by Andrew Beer (Adelaide, Australia) and Terry Clower (North Texas, United States), Sarah Ayres (Bristol, UK), John Gibney (Birmingham, UK) and Markku Sotarauta (Tampere, Finland).


Tuesday, 10 September 2013

The Age of Buildings in the City of Chicago

Following my last post, on the geography of New York City, I've been exploring other building-level datasets to see what they can offer us in relation to telling us more about the fabric of the cities we live in. This time, I've focused on Chicago's 'Building Footprints' dataset. It's not nearly as detailed as New York's PLUTO data but it does include variables on (e.g.) number of floors and year built. As with the NYC data, it is not perfect but we can still make good use of it to understand the development of the city and its structure. I've mapped the city using number of floors as a proxy for height and shaded it by building age to produce the following overview (blues = older buildings, reds = newer).


Besides looking relatively interesting, the above graphic also reveals something about the phased construction of the City of Chicago and - possibly - something more about the data itself. As with the New York City PLUTO data, I produced a chart of the 'year built' column just to give me some idea of its distribution. It looks better than the New York City chart but I'm still not convinced it is 100% accurate (the year built data run from 1852 to 2010). 


Were there really nearly 15,000 buildings constructed in 2006 and only 78 in 2000? Possibly, but it would be good to know more about the accuracy of the data. In total there are 820,154 building footprints in the dataset and there are a range of different columns - which you can read more about in the metadata file. Once again, it's pretty cumbersome to work with in a normal desktop GIS setting but my machine can just about handle it. 


Thursday, 5 September 2013

The Geography of New York City

In June 2013, the city of New York released as open data one of the most detailed, fascinating and user-friendly datasets ever. The Property Land Use Tax lot Output (PLUTO) dataset is essentially a record of every parcel of land in the city, what is on it and who owns it - but this is only part of it. See the full PLUTO data dictionary for more on this. Wired said the mapping elite were 'drooling' over it and there have been a few impressive visualisations already but I was keen to look at the data in more detail and then map land use patterns and get to grips with the dataset more generally. So, as an initial experiment, I mapped all 11 land use categories for the whole city in 3D (PLUTO has a field for number of floors so the maps below are extruded on this basis). Click on an image to enlarge and then flick through the images to compare land uses.












I've also put these images in a PowerPoint file in case anyone finds it useful... These visualisations in many ways tell us what many New Yorkers already know but the PLUTO data (n.b. I've used the ready-made MapPLUTO shapefile) offers everyone for the first time the opportunity to explore this open data and examine the geography of New York City as a whole in much more detail. 

Some further information about the dataset. There are 857,879 rows in the complete dataset and the MapPLUTO version has 85 fields so if you want to work with it then you better have a good computer. When you go to the download page you'll notice that the PLUTO dataset is available as one csv file while the MapPLUTO data is split into the five boroughs of New York City. 

This is an amazing resource but it is not perfect - as the Department of City Planning recognise when they say 'PLUTO is being provided ... for informational purposes only'. The data are only as good as the sources, and sometimes when you look closely things seem a little strange. For example, here's what you get when you chart the YearBuilt column for all buildings constructed since 1800 (click to enlarge). It's hard to tell but I reckon that from about 1980 onwards the YearBuilt column is pretty accurate but before that is is something of a best estimate - though I'd be happy to be proven wrong on this!


I'll probably come back and explore this again soon but that's all for now...


Footnote: 0.4% of tax lots and 1.0% of land remains unclassified. I produced the 3D maps in ArcScene and then annotated them in GIMP. I've just done these to explore at a basic level the characteristics of the dataset and the geography of land use in New York City.

Wednesday, 28 August 2013

Natural Earth for GIS data

People often ask me where they can get GIS data to use for projects, analysis and general mapping. In the UK we now have OS OpenData, which is very nice and very detailed. There's also a new GIS portal from the Office for National Statistics - built using the Geoportal Server. Other GIS datasets are available, from organisations like Natural England, but in this post I thought I'd highlight the excellent - and totally free - Natural Earth site which is very well known in the geodata community but perhaps not more widely. It really is absolutely fantastic.


A little bit of technical information....

  • Natural Earth Vector comes in ESRI shapefile format, the de facto standard for vector geodata. Character encoding is Windows-1252.
  • Natural Earth Raster comes in TIFF format with a TFW world file.
  • All Natural Earth data use the Geographic coordinate system (projection), WGS84 datum +proj=longlat +ellps=WGS84 +datum=WGS84 +no_defs
Whether you're looking for data to make a general world map or a more detailed local area map you will find what you are looking for here. It's not as detailed as OS data for Great Britain but then many GIS users don't need that level of accuracy. Another screenshot below showing the various download options...


Finally, taking my inspiration from the worst website in the world, and Ken Field's blog - I've made an absolutely awful map using some of this data*. Can anyone do a worse one?


*I actually did this as part of my work in developing some new GIS modules. I tried to pack in as much bad practice as possible just as an extreme example of what not to do.





Friday, 16 August 2013

Mapping flows in ArcGIS

This short post is about the process of flow mapping in ArcGIS and not really about the end results - though the maps are quite interesting. I've done quite a bit of flow mapping in the past and am now getting ready to work on the next set of Census flow data in the UK (which should be out in November) so I've been experimenting with some tools. I've written about this in the past in papers in Computers, Environment and Urban Systems and also Environment and Planning B but those papers are a bit long-winded! Other people have produced beautiful flight path maps so I thought I'd experiment with the same data using the relatively new ArcGIS XY to Line tool in version 10.0 (it can be found in ArcToolbox - Data Management Tools - Features - XY to Line at the bottom of the list of tools). For more on other methods and previous iterations of this kind of thing take a look at the work of Nathan Yau, Michael Markieta or James Cheshire.


For anyone wanting to map flows in ArcGIS, Michael Markieta's tutorial is probably a good place to start but be prepared for things to go awry in ArcGIS... When I mapped the 59,000 or so flight paths in the map above using XY to Line and one single dbf file (or csv, etc. - it makes no difference) the resulting shapefile only contained 16,066 rows. This happened every time I tried it and a couple of times my shapefile had only 73 rows. Another time it had ~14,000. That's why Markieta recommends splitting the file up - although I just cut it up into chunks of 16,000. Interestingly, I ran into exactly the same problem with my CEUS paper a few years back using Glennon's flow data model tool - though the limit was about 32,000 before it cut off. 

Another very annoying feature of XY to Line (for me at least) is that when you choose the Great Circle option under 'Line Type' it takes much longer to compute and the resultant shapefile is enormous. The shapefile for flight paths in the above map is about 10MB whereas the great circle version was over 450MB for one 16,000 chunk alone. Not sure if anyone else has run into this but it doesn't seem like a very efficient way of doing things! [Edit - as @baeing has reminded me, it's because shapefiles don't support curves - though geodatabases do.]

Once I had my complete shapefile I moved to QGIS, added in a world layer from Natural Earth and then experimented a little with symbology. I also experimented with different styles and projections to produce some of the maps below. That's all for now - I just hope ESRI are able to improve upon the current version of XY to Line because when it does work it is really fast (on my machine at least) and straightforward.

Very similar to above, minus text

Short haul, different projection

Short haul, different projection, borders

Slightly different symbology

And, yes, I know that flights from Australia or New Zealand typically go over the Pacific rather than the long way round! I'm just showing the XY to Line outputs as they are in this post.

Thursday, 8 August 2013

Employee Growth in London, 2001 to 2012

The Office for National Statistics has released a new dataset on the number of employees across London's 983 MSOAs. The data are sourced from the Inter-Departmental Business Register (IDBR) and they reveal some interesting trends. Naturally, I had to do a 3D map of this, so take a look at the image below for the obvious growth points...

Massive absolute growth in the City of London and Canary Wharf - and some other central MSOAs in Camden, Southwark and Westminster, but also massive growth in employment in Uxbridge.

Click on the image to enlarge

The number of employees in the City of London increased by 36%, compared to 267% in Canary Wharf and 170% in an Uxbridge MSOA. By contrast, one part of Islington had 78,600 employees in 2001 but only 57,000 in 2012 - a drop of 27%. This area of Islington is immediately north of the City of London and includes Clerkenwell and Finsbury.

If you're interested in this kind of thing it's definitely worth looking at the original dataset.