Monday, March 14, 2011

Optimizing for Vertical Search

Vertical search engines focus on specific niches of web content, including images, videos, news, travel, and people. Such engines exist to provide value to their user base in ways that go beyond what traditional web search engines provide.

One area where vertical search engines can excel in comparison to their more general web search counterparts is in providing more relevant results in their specific category. They may accomplish this by any number of means, including making assumptions about user intent based on their vertical nature (an option that full web search engines do not normally have), specialized crawls, more human review, and the ability to leverage specialized databases of information (potentially including databases not available online).

There is a lot of opportunity in vertical search. SEO professionals need to seriously consider what potential benefits vertical search areas can provide to their websites. Of course, there are significant differences in how you optimize for vertical search engines.

The Opportunities in Vertical Search

Vertical search has been around for almost as long as the major search engines have been in existence. Some of the first vertical search engines were for image search, news group search, and news search, but many other vertical search properties have emerged since then, both from the major search engines and from third parties.

This article will focus on strategies for optimizing your website for the vertical search offerings from Google, Yahoo!, and Bing. We will also spend some time on YouTube, which in January 2009 became the second largest search engine on the Web. First we will look at the data for how vertical search volume compares to regular web search. The data in Table 1 comes from Hitwise, and shows the top 20 Google domains as of May 2006, which is one year before the advent of Universal Search.





In May 2006, image search comprised almost 10% of Google search volume. Pair this with the knowledge that a smaller number of people on the Web optimize their sites properly for image search (or other vertical search engines) and you can see how paying attention to vertical search can pay tremendous dividends.

Of course, just getting the traffic is not enough. You also need to be able to use that traffic. If someone is coming to your site just to steal your image, for example, this traffic is likely not of value to you. So, although a lot of traffic may be available, you should not ignore the importance of determining how to engage users with your site. For instance, you could serve up some custom content for visitors from an image search engine to highlight other areas on your site that might be of interest, or embed logos/references into your images so that they carry branding value as they get “stolen” and republished on and off the Web.

1. Universal Search and Blended Search


In May 2007, Google announced Universal Search, which integrated vertical search results into main web results.

Thinking of it another way, Google’s web results used to be a kind of vertical search engine itself, one focused specifically on web pages (and not images, videos, news, blogs, etc.). With the advent of Universal Search, Google changed the web page search engine into a search engine for any type of online content. Figure 1 shows some examples of Universal Search results, starting with a Google search on iphone.

Figure 1. Search results for “iphone”



Notice the video results (labeled “Video results for iphone”) and the news results (labeled “News results for iphone”). This is vertical search being incorporated right into traditional web search results. Figure 2 shows the results for a search on i have a dream.

Figure 2. Search results for “i have a dream”



Right there in the web search results, you can click on a video and watch the famous Martin Luther King, Jr., speech. You can see another example of an embedded video by searching on one small step for man as well.

The other search engines (Yahoo!, Microsoft, and Ask) moved very quickly to follow suit. As a result, the industry uses the generic term Blended Search for this notion of including vertical search data in web results.

2. The Opportunity Unleashed

As we noted at the beginning of this article, the opportunity in vertical search was significant before the advent of Universal Search and Blended Search. However, that opportunity was not fully realized because many (in fact, most) users were not even aware of the vertical search properties. With the expansion of Blended Search, the opportunities for vertical search have soared.

However, the actual search volume for http://images.google.com has dropped a bit, as shown in Table 2, which lists data from Hitwise for February 2009.
Table 2. Most popular Google properties, February 2009

This drop is most likely driven by the fact that image results get returned within regular web search, and savvy searchers are entering specific queries that append leading words such as photos, images, and pictures to their search phrases when that is what they want.

For site owners, this means new opportunities to gain visibility in the SERPs. By adding a blog, releasing online press releases to authoritative wire services, uploading video to sites such as YouTube, and adding a Google Local listing, businesses increase the chances of having search result listings that may directly or indirectly drive traffic to their sites.

It also means site owners must think beyond the boundaries of their own websites. Many of these vertical opportunities come from one-time engagements or small additional efforts to maximize the potential of activities that are already being performed.

Optimizing for Local Search

Search engines have sought to increase their advertiser base by moving aggressively into providing directory information. Applications such as Google Maps, Yahoo! Local, and Bing Maps have introduced disruptive technology to local directory information by mashing up maps with directory listings, reviews/ratings, satellite images, and 3D modeling—all tied together with keyword search relevancy. This area of search is still in a lot of flux as evolutionary changes continue to come hard and fast. However, these innovations have excited users, and the mapping interfaces are growing in popularity as a result.

Despite rapid innovation in search engine technology, the local information market is still extremely fractured. There is no single dominant provider of local business information on the Internet. According to industry me rics, online users are typically going to multiple sources to locate, research, and select local businesses. Traditional search engines, local search engines, online Yellow Pages, newspaper websites, online classifieds, industry-specific “vertical” directories, and review sites are all sources of information for people trying to find businesses in their area.



This fractured nature of online local marketing creates considerable challenges for organizations, whether they’re a small mom and pop business with only a single location or a large chain store with outlets across the country.

Yet, success in these efforts is critical. The opportunity for local search is huge. More than any other form of vertical search, local search results have come to dominate their place in web search.

The regular web search results are not even above the fold. This means that if you are not in the local search database, you are probably not getting any traffic from searches similar to this one.

Obviously, the trick is to rank for relevant terms, as most of you probably don’t offer rental cars in Minneapolis. But if your business has a local component to it, you need to play the local search game.

Saturday, March 12, 2011

Using Advanced Search Techniques

One of the basic tools of the trade for an SEO practitioner is the search engines themselves. They provide a rich array of commands that can be used to perform advanced research, diagnosis, and competitive analysis. Some of the more basic operators are:



[-keyword]

Excludes the keyword from the search results. For example, [loans -student] shows results for all types of loans except student loans.

[+keyword]

Allows for forcing the inclusion of a keyword. This is particularly useful for including stopwords (keywords that are normally stripped from a search query because they usually do not add value, such as the word the) in a query, or if your keyword is getting converted into multiple keywords through automatic stemming. For example, if you mean to search for the TV show The Office, you would want the word The to be part of the query. As another example, if you are looking for Patrick Powers, who was from Ireland, you would search for patrick +powers Ireland to avoid irrelevant results for Patrick Powers.

["key phrase"]

Shows search results for the exact phrase—for example, ["seo company"].

[keyword1 OR keyword2]

Shows results for at least one of the keywords—for example, [google OR Yahoo!]. These are the basics, but for those who want more information, what follows is an outline of the more advanced search operators available from the search engines.

Friday, March 11, 2011

What search engines cannot see

It is also worthwhile to review the types of content that search engines cannot “see” in the human sense. For instance, although search engines are able to detect that you are displaying an image, they have little idea what the image is a picture of, except for whatever information you provide them in the alt attribute, as discussed earlier.

They can, however, determine pixel color and, in many instances, determine whether images have pornographic content by how much flesh tone there is in a JPEG image. So, a search engine cannot tell whether an image is a picture of Bart Simpson, a boat, a house, or a tornado. In addition, search engines will not recognize any text rendered in the image. The search engines are experimenting with technologies to use optical character recognition (OCR) to extract text from images, but this technology is not yet in general use within search.

In addition, conventional SEO wisdom has always held that the search engines cannot read Flash files, but this is a little overstated. Search engines are beginning to extract information from Flash, as indicated by the Google announcement at http://googlewebmastercentral.blogspot.com/2008/06/improved-flash-indexing.html.

However, the bottom line is that it’s not easy for search engines to determine what is in Flash. One of the big issues is that even when search engines look inside Flash, they are still looking for textual content, but Flash is a pictorial medium and there is little incentive (other than the search engines) for a designer to implement text inside Flash. All the semantic clues that would be present in HTML text (such as heading tags, boldface text, etc.) are missing too, even when HTML is used in conjunction with Flash.

A third type of content that search engines cannot see is the pictorial aspects of anything contained in Flash, so this aspect of Flash behaves in the same way images do. For example, when text is converted into a vector-based outline (i.e., rendered graphically), the textual information that search engines can read is lost.

Audio and video files are also not easy for search engines to read. As with images, the data is not easy to parse. There are a few exceptions where the search engines can extract some limited data, such as ID3 tags within MP3 files, or enhanced podcasts in AAC format with textual “show notes,” images, and chapter markers embedded. Ultimately, though, a video of a soccer game cannot be distinguished from a video of a forest fire.


Figure 2-21

Search engines also cannot read any content contained within a program. The search engine really needs to find text that is readable by human eyes looking at the source code of a web page, as outlined earlier. It does not help if you can see it when the browser loads a web page—it has to be visible and readable in the source code for that page.

One example of a technology that can present significant human-readable content that the search engines cannot see is AJAX. AJAX is a JavaScript-based method for dynamically rendering content on a web page after retrieving the data from a database, without having to refresh the entire page. This is often used in tools where a visitor to a site can provide some input and the AJAX tool then retrieves and renders the correct content.

The problem arises because the content is retrieved by a script running on the client computer (the user’s machine) only after receiving some input from the user. This can result in many potentially different outputs. In addition, until that input is received the content is not present in the HTML of the page, so the search engines cannot see it.

Similar problems arise with other forms of JavaScript that don’t render the content in the HTML until a user action is taken.

As of HTML 5, a construct known as the embed tag () was created to allow the incorporation of plug-ins into an HTML page. Plug-ins are programs located on the user’s computer, not on the web server of your website. This tag is often used to incorporate movies or audio files into a web page. The tag tells the plug-in where it should look to find the datafile to use. Content included through plug-ins is not visible at all to search engines.

Frames and iframes are methods for incorporating the content from another web page into your web page. Iframes are more commonly used than frames to incorporate content from another website. You can execute an iframe quite simply with code that looks like this:



Frames are typically used to subdivide the content of a publisher’s website, but they can be used to bring in content from other websites, as was done in Figure 2-21 with http://accounting.careerbuilder.com on the Chicago Tribune website.

Figure 2-21 is an example of something that works well to pull in content (provided you have permission to do so) from another site and place it on your own. However, the search engines recognize an iframe or a frame used to pull in another site’s content for what it is, and therefore ignore the content inside the iframe or frame as it is content published by another publisher. In other words, they don’t consider content pulled in from another site as part of the unique content of your web page.

Evaluating Content on a Web Page

Search engines place a lot of weight on the content of each web page. After all, it is this content that defines what a page is about, and the search engines do a detailed analysis of each web page they find during their crawl to help make that determination.You can think of this as the search engine performing a detailed analysis of all the words and phrases that appear on a web page, and then building a map of that data for it to consider showing your page in the results when a user enters a related search query.

This map, often referred to as a semantic map, seeks to define the relationships between those concepts so that the search engine can better understand how to match the right web pages with user search queries.



If there is no semantic match of the content of a web page to the query, the page has a much lower possibility of showing up. Therefore, the words you put on the page, and the “theme” of that page, play a huge role in ranking.

FIGURE 2-14.

Figure 2-14 shows how a search engine will break up a page when it looks at it, using a page on the Stone Temple Consulting website. The navigational elements of a web page are likely similar across the many pages of a site. These navigational elements are not ignored, and they do play an important role, but they do not help a search engine determine what the unique content is on a page. To do that, the search engine gets very focused on the part of Figure 2-14 that is labeled “Real content”.

Determining the unique content on a page is an important part of what the search engine does. It is this understanding of the unique content on a page that the search engine uses to determine the types of search queries for which the web page might be relevant. Since site navigation is generally not unique to a single web page, it does not help the search engine with that task.

This does not mean navigation links are not important, because they most certainly are—however, they simply do not count when trying to determine the unique content of a web page because those navigation links are shared among many web pages.

One task the search engines face is judging the value of content. Although evaluating how the community responds to a piece of content using link analysis is part of the process, the search engines can also draw some conclusions based on what they see on the page.

For example, is the exact same content available on another website? Is the unique content the search engine can see two sentences long or 500+ words long? Does the content repeat the same keywords excessively? These are a few examples of things that the search engine can look at when trying to determine the value of a piece of content.

Retrieval and Rankings

The next step in this quest for knowledge occurs when the search engine returns a list of relevant pages on the Web in the order most likely to satisfy the user. This process requires the search engines to scour their corpus of billions of documents and do two things: first, return only the results that are related to the searcher’s query; and second, rank the results in order of perceived importance (taking into account the trust and authority associated with the site). It is both relevance and importance that the process of SEO is meant to influence.



Relevance is the degree to which the content of the documents returned in a search matches the user’s query intention and terms. The relevance of a document increases if the terms or phrase queried by the user occurs multiple times and shows up in the title of the work or in important headlines or subheads, or if links to the page come from relevant pages and use relevant anchor text.

You can think of relevance as the first step to being “in the game.” If you are not relevant to a query, the search engine does not consider you for inclusion in the search results for that query.

Importance or popularity refers to the relative importance, measured via citation (the act of one work referencing another, as often occurs in academic and business documents) of a given document that matches the user’s query. The popularity of a given document increases with every other document that references it. In the academic world, this concept is known as citation analysis.

You can think of importance as a way to determine which page, from a group of equally relevant pages, shows up first in the search results, which is second, and so forth. The relative authority of the site, and the trust the search engine has in it, are significant parts of this determination. Of course, the equation is a bit more complex than this and not all pages are equally relevant. Ultimately, it is the combination of relevance and importance that determines the ranking order.

Popularity and relevance aren’t determined manually (those trillions of man-hours would require Earth’s entire population as a workforce). Instead, the engines craft careful, mathematical equations—algorithms—to sort the wheat from the chaff and to then rank the wheat in order of quality. These algorithms often comprise hundreds of components. In the search marketing field, they are often referred to as ranking factors or algorithmic ranking criteria.

Thursday, March 10, 2011

Crawling and Indexing

Imagine the World Wide Web as a network of stops in a big city subway system. Each stop is its own unique document (usually a web page, but sometimes a PDF, JPEG, or other file). The search engines need a way to “crawl” the entire city and find all the stops along the way, so they use the best path available: the links between web pages, an example of which is shown in Figure 2-11.


Figure 2-11


In our representation in Figure 2-11, stops such as Embankment, Piccadilly Circus, and Moorgate serve as pages, while the lines connecting them represent the links from those pages to other pages on the Web. Once Google (at the bottom) reaches Embankment, it sees the links pointing to Charing Cross, Westminster, and Temple and can access any of those “pages.”

The link structure of the Web serves to bind together all of the pages that were made public as a result of someone linking to them. Through links, search engines’ automated robots, called crawlers or spiders (hence the illustrations in Figure 2-11), can reach the many billions of interconnected documents.

Once the engines find these pages, their next job is to parse the code from them and store selected pieces of the pages in massive arrays of hard drives, to be recalled when needed in a query. To accomplish the monumental task of holding billions of pages that can be accessed in a fraction of a second, the search engines have constructed massive data centers to deal with all this data.

One key concept in building a search engine is deciding where to begin a crawl of the Web. Although you could theoretically start from many different places on the Web, you would ideally begin your crawl with a trusted set of websites. You can think of a factor in evaluating the trust in your website as the click distance between your website and the most trusted.