Google actually scanned a huge number of newspaper microfilm and microfiche back in 2010-11: https://news.google.com/newspapers
It's an amazing resource, but it's hidden away and researchers within Google have trouble accessing it even if they knew it existed. I spent months with lawyers to get access to the original files for research and at one point they told me they were going to delete it! Whaaaaa!?!
It was something like 6 PB, which seriously, to Google isn't much, but the team that "owned" the data wasn't using it and to them it was just an expense. Ugh. People don't care about history.
(Each * here representing some unknown third party getting access to your email address, phone number, blood type and sexual preferences.)
New York State funded an effort to scan, but not digitize historical newspapers, and while the microfilm is stored in the state archives, the online versions are hosted by an eccentric guy who digitizes the microfilm as a hobby. The guy puts everything online, but makes it difficult to work with in a variety of ways.
According to the site, it specifically avoids "major" periodicals:
"A Collection of Interesting, Important, and Controversial Perspectives Largely Excluded from the American Mainstream Media."
It's a century's worth of NYTimes front pages. An amazing long term dataset that is fun to flip through to answer infinite questions you might have about both how the last 100 years really went down and also about the evolution of presenting information to readers.
https://static01.nyt.com/images/<year>/<month>/<day>/nytfrontpage/scan.pdf
seems to go back about 10 yearsIf you do find more, please let me know.
I was hoping scientific American or wsj etc would have them but couldn’t find any.
The Washington Post was recently caught scrubbing an unflattering quote from a 2019 profile of Kamala Harris:
https://reason.com/2021/01/22/the-washington-post-memory-hol...
Personally I had not seen anything from him stating he believed the security holes were fixed on a broad scale, or on any scale. But his recent video in November 2020 suggests he no longer had any substantial concerns regarding voting machine security: https://www.youtube.com/watch?v=cMz_sTgoydQ&t=521s
For COVID, back then, it was "tech bros are afraid of the flu (and maybe also racist)" and then "go hug a Chinese person", and then "closing flights from Wuhan is racist (even though China itself was doing it domestically)". It was funny to see the escalating tsunami of wrongness coming at the media, but of course, they would never admit wrong-doing.
For Kamala Harris, well she was (is?) hugely unpopular. She has terrible political baggage, has been personally responsible for terrible systemic racism, and the news coverage reflected that quite accurately up to a very specific point. Tulsi torpedoed her early on in the debates by bringing all this up.
However, when Biden said his VP choice would be a woman of color, there was an immediately 180 in coverage and retroactive editing to make it look like she was an exceptional candidate who has always championed racial issues and is a regular everyday person just like all of us.
Is that material still available anywhere? It was really extensive IIRC.
https://news.google.com/newspapers
It's fun to browse but I'm sad that the project seemed to just sputter out. At one time I thought it would grow to make historical newspapers searchable with coverage comparable to the books searchable through Google Books.
Sad they haven't made more of it really. Thanks for the pointer!
Alexa APIs, incidentally, were one of the first AWS products along with ECS (e-commerce service) and SQS in 2004.
As a side note, I'm loving the 90s tabloid layouts.
Well said, agreed! I should add Wikipedia.
At a first glance, they seem reasonable.
- 43% Direct support to websites. Keeping the Wikimedia websites online is about more than just servers. It also includes ongoing engineering improvements, product development, design and research, and legal support.
- 32% Direct support to communities. The Wikimedia projects exist thanks to the communities that create and maintain them. We strengthen these communities through grants, projects, trainings, tools to augment contributor capacity, and support for the legal defense of editors.
- 32% Direct support to communities The Wikimedia projects exist thanks to the communities that create and maintain them. We strengthen these communities through grants, projects, trainings, tools to augment contributor capacity, and support for the legal defense of editors.
- 13% Administration and governance. We manage funds and resources responsibly to recruit and support skilled, passionate staff who advance our communities and values.
- 12% Fundraising. Wikimedia is sustained by donations. Millions of remarkable individuals and institutions ensure that we have the necessary resources to continue our global mission.
[1] https://wikimediafoundation.org/support/where-your-money-goe...
As for archive.org they offer basically no transparency as to which sites are excluded from their archiving (and as far as I know they will remove all the content of a site once the owner asks them to).
Most internet organizations do that (free, charity or commercial). Good people are expensive. You need sysadmins, programmers, some graphics artist, probably more than a few lawyers etc.
And I wouldn't want them to use the cheapest, possible people that they can find for those roles. And asking good people to work for free or cheap, is just as shitty.
And a lot of free software and opensource foundation spend most of the money organizing conferences, so that people can meet in person and have presentations and working groups etc.. But when Wikipedia does the same its somehow wrong ?
The Internet Archive has proved itself so many times over, and is so underfunded, that it deserves all the money it can possibly raise - which will allow them to solve more of the problems they face.
I like them, but they seriously need a change in management.
They opened themselves up to millions of dollars (if not billions of dollars) of liability for no good reason. (Covid doesn't give you a pass to give away someone else's property without permission.)
Similar to you, I'm not going to donate to an org that is so reckless with donations. If they would admit it was an error in judgement and come to a quick settlement I might become a donor again.
Even still, the damage has been done -- it'll take a long time for publishers to trust the Internet Archive again.