back

by riordan·13y ago·view on hn ↗
I'm part of the team at NYPL Labs that's been working on these historical geospatial projects for the past few years (The Vectorizer is the work of our own @MGA), and that's EXACTLY what we've been working toward. For almost 4 years, we've had staff and volunteers going over scans of geo-rectified (stitching and stretching a raster image so it aligns with geospatial coordinates) historical insurance maps of NYC meticulously extracting the information on there. Namely we go after the outlines of buildings to capture the amazingly detailed datapoints these maps had about every building in the city all the way back to the first half of the 19th century.

These are nice to have for researchers, but the real purpose of collecting this is, just as you note, to unlock the hidden historical geospatial data in textual materials. Once we've got all those names of places, their addresses, their lat/lon coordinates, and their timeframes of existence, we can start to search through texts to find linkages. Old city directories (they're basically books of ghosts) start to show you who lived and worked where [1] (and in the process starts to get you more names you can associate with these places), address matches in historical newspapers start to show you what happened in these places, and the maps start to become this geospatial backbone to traverse across tons of different datasets.

The Vectorizer is so freaking cool for so many reasons, but mostly because it's going to let us actually get through these insurance atlases to collect this data before we all die (one of our favorites is the 1854 William Perris Atlas [2][3] but it took nearly 3 years to actually get through the 64,000+ buildings in Manhattan south of 42nd st) so we can start doing this kind of querying with it. The real geniuses behind all this, our Geospatial Librarian Matt Knutzen and the team at Topomancy, have been working on an experimental gazetteer [4] so that we'll finally have this as a public web service for people to hack on all these places as we collect and conflate them. Give us a few months...

In the meantime, sign up for the Open Historical Maps project listserv [5] that some of the OSM crew is working on (including the geniuses at Topomancy).

Also, this came out of a historical geospatial hack day [5] we threw a few months back, which you should check out if you want to play around with some of our data sources for this kind of work or for building something else out of historical NYC's geospatial footprint.

[1]: http://andrewxhill.github.io/cartodb-examples/scroll-story/b... [2]: http://maps.nypl.org/warper/layers/861 Tileserver, please forgive me for linking to you [3]: http://aaronland.info/nypl-perris/ YEAH SHAPEFILES! [4]: http://vimeopro.com/openstreetmapus/state-of-the-map-us-2013... Schuyler Earle's presentation on their version of historical gazetteer they're building for the Library of Congress at State of The Map US 2013 [5]: http://www.nypl.org/blog/2013/07/12/maphack-hacking-nycs-pas...

1 comments
This sounds awesome. I've spent a lot of time digging around in the records at 30 Chambers and it is really hard to comprehend how much historical stuff is lying around in decaying pages.

Even just something as simple as showing an OSM with all the historical election districts / assembly districts over time for each census/election as map layers would visually convey to someone looking at an address what would otherwise take a decent amount of time to look up.

The NYC Dept. of Records has all the tax lot photos from the 1940 and 1980 canvas, so you could even build up a historical "street view" for the 5 boroughs. I've always found it annoying that they keep this data locked up and charge a decent sum for each photo. It is something that a decent microfilm scanner could make quick work of, but I don't know if they have plans to liberate all of that image data as part of the open data efforts.