A digital library and/or search engine much be able to handle many different kinds of multimedia, huge numbers of simultaneous requests, and attempts by some parties to manipulate automated search engine processes.
In order to index effectively, search engines must use algorithms that can:
-Determine what content should be indexed
-Divvy up tasks between servers
-Analyze content
-Collect new URLs
-Ignore dupicate pages
-Add new URLS are added to a cue
-Ignore spam
-Avoid overloading servers which host the pages they are crawling
-Save page content for indexing indexing.
-Process simplistic or ambiguous queries like "the onion" and weed out content effectively
-multiple factors should be included in this determination
Suggested focii for improving search engine quality (Henzinger): spam, content quality, webmaster deviation from web conventions, duplicate hosts, and vaguely structured data.
Some websites use text, links, or cloaking mechanisms to improve their ranking in search result sets. This often includes adding content (key words, links, false content) in an attempt to fool search engine page ranking algorithms. Spam may be structured in the form of deceptive text (white text on a white background with text that is invisible to the user but not the webcrawler), link spam (or link farms which collect links pointing to every other page in that site), and providing entirely different content for the user and the webcrawler. Work can be done towards assessing the quality of content included in a given webpage. Some sites advertise false content (like celebrity names) in an attempt to redirect users to their site.
Most search engine users do not travel beyond the first page of search results. Since many website generate income from their traffic, there is great motivation to increases a page's ranking so that it appears on the first page of a result set. Search engine developers must constantly tweak their indexing and page ranking algorithms in order to stay one step ahead of those who wish to manipulate these processes in order to increase traffic to their own page.
David Hawking , Web Search Engines: Part 1 and Part 2 IEEE Computer, June 2006.
M. Henzinger et al. challenges in Web Search Engines. ACM SIGIR 2002.
Showing posts with label search engines. Show all posts
Showing posts with label search engines. Show all posts
Friday, October 16, 2009
Subscribe to:
Posts (Atom)