back

by moultano·16y ago·view on hn ↗
I work in search quality at Google, and while certainly not everyone agrees that it's a problem, a lot of people do.

I could write a lot about this, but the central issue is that it is very very hard to make changes that sacrifice on-topic-ness for good-ness that don't make the results in general worse. We're working on it though, and I suspect we'll never stop.

I think a lot of the promise lies in as you said, identifying the tangentially related article, or as I like to frame it, bringing more queries into the head. We've launched a lot of changes that do exactly this. (But you are right, it is difficult, and fundamentally so. Language is hard.)

3 comments
Can you please please please somehow tell someone important to get rid of those phony review sites, which when you get there have zero reviews of the item in question?
Give me some examples if you can. I'll send them around.
Here's a phony one: http://www.mahalo.com/3g-iphone-reviews

That site has a few thousand crappy reviews scraped from all over the place. It comes up on google all the time and it's something that I automatically ignore.

Can't you stop crawling sites that dish up your search terms for you (without any content)? A good example is eudict.com - search for an obscure word and this is the first that pops up.

It then returns a page with your search queries and no information.

What use is a site that simply returns search queries?

Why aren't sites that scrape content blacklisted?

>Why aren't sites that scrape content blacklisted?

The problem is more difficult than you'd think. For instance, virtually every news organization "scrapes" the associated press, but we wouldn't want to throw out every news organization.

Content-free search result pages are things we do try to remove, even manually if it becomes a big enough problem.

virtually every news organization "scrapes" the associated press

If they're not adding real value, like analysis or graphics or commentary or whatnot, why would you want to keep them if they're all just duplicates?

I had a friend work at a startup to solve this problem exact: we read virtually identical articles about the same bit of news on all the news sites. The startup was working on highlighting only the unique bits of each article and recommend the one article that seems to have the most pieces of information. You would read the one and skim to the unique bits of the others, and you would have gotten all angles and facts much more quickly.

Shame they closed it up.

We do filter near-duplicates within the same set of results. You'll likely see only one copy of an AP story with a link at the bottom saying something like "Repeat this search with the omitted results included"
"Why aren't sites that scrape content blacklisted?"

What like Google News?

It's a lame joke but it shows you how fine the line is between 'scraping' and 'aggregating' content.

Google News doesn't show up in search results the way, e.g., Mahalo might. The only search results I've seen that incorporate Google News are built right into the main results page; I don't click through expecting content and get a Google News page instead.

In fact, I haven't seen any Google-owned scraping or aggregating page in a result that I've clicked through. They are big believers in the theory that you should look at exactly one search results page, not a page that takes you to a page that takes you to (...) the result you actually wanted.

I haven't seen any Google-owned scraping or aggregating page in a result that I've clicked through.

What about Google Health results? Try [Whooping Cough] or similar. Top 'result' is a Google health page whose main column is all content republished from Medline. Right column is essentially 'more results' from News and Scholar.

It's not quite as bad as other paste-together pages of text and more results, but they're creeping in that direction.

It is a problem. Things like MFA spam and Google's recent trend of being less useful as a Search Engine for Programmers (ignoring quotes, ignoring punctuation like ? and !) are enough to make me switch, if I find something better. This article has convinced me to try out DDG for a little while.
>ignoring punctuation like ? and !

Google has done this since launch, I don't think it's a recent trend. I agree though that searching for programming information isn't great. (It's even harder with math.)

I doubt you'll find MFA spam to be better on DDG than on Google, but please, if you see a query where they are beating us. Send it over. :) I can guarantee you that I'll get a lot of eyes looking at it.

DuckDuckGo is actually worse for this query (they seem to not be ranking as high as they used to in Google SERPS), but it brings up an issue that I've never seen satisfactorily answered anywhere.

http://www.google.com/search?q=soap+with+flash+as3 brings up the following link from bigresource.com: http://www.bigresource.com/FLASH-SOAP-with-flash-AS3-PvTRLrv...

The BigResournce content is scraped from Actionscript.org and surrounded by BigResource's ads, which is pretty standard scraping. Clicking the link to "view original forum thread" redirects to a framed page with more BigResource ads, and the original content in a frame. The frame is handled internally by the site, so I'm doubtful they're even showing an actual link to their scraped content.

I've seen this specific site debated in this thread: http://www.google.com/support/forum/p/Web+Search/thread?tid=... and I know several users including myself have reported this site as spam (I even went the extra mile and changed Chrome to append "-site:bigresource.com" to queries using the default engine).

My question is this: Is there something that BigResource is doing that exempts them from being classified as spam? As near as I can tell, they add no extra value to the content for the user, and push the legitimate results they scraped further down in results (because they have many, many pages).

Update: Google now ranks it as highly as DDG. Great job!
Kudos, I was about to submit a longstanding complaint but it looks like it has been fixed. For a long time Google was autocorrecting my searches for gearman as "gearman" -- not just in a "Did you mean?" but actually giving me search results about germans instead of gearman. I'm not sure when it happened, but results are sane and accurate again, even for things like "gearman workers" and "gearman jobs".
er, typo fail. They were correcting "gearman" to "german."
It's over two weeks, but FWIW here it goes:

I got fed up of searching for CPU instructions and compiler intrinsics in Google. You get mostly pointless forum discussions or MFA kind of results.

Same happens with many other technical searches. It looks like the secret pagerank is gamed from both inside and outside.

Note: DDG's search is Yahoo Boss/Bing.

Now that 3 comments in a row mention it, what's DDG?
Duck Duck Go: http://duckduckgo.com/ by Gabriel Weinberg: http://www.gabrielweinberg.com/blog/ Gabriel has an awesome blog and is active here on HN.
DuckDuckGo