Google got it on the first attempt even using the wrong name.
Too much data with poor search algorithms can be worse than a phonebook.
The bots are then run via morph.io. List of morph.io scrapers: https://morph.io/scrapers
But it looks like the bots aren't published (checked github.com/openc and morph.io). That's a very strange approach to open data like you noted but I like their goal.
For example, doing a search for Apple, and even limiting it to California, brings up a ton of junk that needs to be filtered out visually by the user:
https://opencorporates.com/companies/us_ca?action=search_com...
And you have to have some knowledge of corporate conventions to get at what you think you want. Here's the entry for Facebook, Inc., registered in Menlo Park:
https://opencorporates.com/companies/us_ca/C2711108
However, it's just a "branch". What you really want is Facebook, Inc. as incorporated in Delaware:
https://opencorporates.com/companies/us_de/3835815
The listing of subsidiaries is nice...I don't know how complete it is, but it contains the companies I expect (Parse, Oculus, Whatsapp...no Instagram though?)
https://opencorporates.com/companies/us_de/3835815/statement...
If ever there was an indication that open data can be disruptive and revolutionary, here it is.
Re the number of results, with over 95 million companies from hundreds of official sources, there are a lot of companies with the same or similar names -- that's part of the power, and why OpenCorporates is so widely used by journalists, anti-corruption investigators, lawyers, law-enforcement etc. The more we loosen the search, however, the more results you will get -- so it's a balancing act, and the advanced search and the API is part of us trying to make it work for users, and we'd love feedback on both -- just email us at community @ opencorporates dot com
For large listed companies such as Apple or Intel, it depends on what question you are trying to answer. If it's just an overview of the corporation, then Wikipedia or something like Yahoo Finance is the best route. The former gives a narrative overview, using the collected expertise of hundreds of contributors; the latter includes highly proprietary data (which was also for the most part collected by offshore humans) to build an overall picture.
However, for official, provenanced data under an open licence about legal entities, OpenCorporates is by far the best option -- the Facebook example is a good one. Contrary to many people's expectations, Facebook is not a California Corporation. However, because they operate in California they have to register as a branch (aka Foreign Corporation), and if you land on that page (https://opencorporates.com/companies/us_ca/C2711108) you'll see that we actually link to the home corporation, in Delaware (https://opencorporates.com/companies/us_de/3835815), and there we list subsidiaries and other branches. Why don't we have Instagram? You can check from the source of the subsidiaries we list (e.g. https://opencorporates.com/statements/37307675 which links to the SEC Exhibit 21 filing at http://www.sec.gov/Archives/edgar/data/1326801/0001326801150... ), and you can see there that they don't list Instagram (possibly because either there's no separate company for it, or because it doesn't count as a material subsidiary).
So the question often becomes (and news stories are problematic for multiple reasons), where can you find the linkages that can be parsed into structured, provenanced data in a reliable way. We're focusing on two areas: looking for sources of public data that can be combined together to give insight to all, and building up a community of users to help us do that: http://impact.opencorporates.com/contribute/
Please do consider joining us in this important mission, or just by pinging us with suggestions of how we can do better.