Calculus Infosoft
Web crawling

A Web crawler is a computer program that browses the World Wide Web in a methodical, automated manner. Other terms for Web crawlers are ants, automatic indexers, bots, and worms or Web spider, Web robot, or-especially in the FOAF community-Web scutter.

This process is called Web crawling or spidering. Many sites, in particular search engines, use spidering as a means of providing up-to-date data. Web crawlers are mainly used to create a copy of all the visited pages for later processing by a search engine that will index the downloaded pages to provide fast searches. Crawlers can also be used for automating maintenance tasks on a Web site, such as checking links or validating HTML code. Also, crawlers can be used to gather specific types of information from Web pages, such as harvesting e-mail addresses (usually for spam).

A Web crawler is one type of bot, or software agent. In general, it starts with a list of URLs to visit, called the seeds. As the crawler visits these URLs, it identifies all the hyperlinks in the page and adds them to the list of URLs to visit, called the crawl frontier. URLs from the frontier are recursively visited according to a set of policies.

Crawling policies

  • Selection policy
  • Restricting followed links
  • Path-ascending crawling
  • Focused crawling
  • Crawling the Deep Web
  • Re-visit policy
  • Politeness policy
  • Parallelization policy

Web crawler architectures

A crawler must not only have a good crawling strategy, as noted in the previous sections, but it should also have a highly optimized architecture.

Web crawlers are a central part of search engines, and details on their algorithms and architecture are kept as business secrets. When crawler designs are published, there is often an important lack of detail that prevents others from reproducing the work. There are also emerging concerns about "search engine spamming", which prevent major search engines from publishing their ranking algorithms.

OUR PRODUCTS
Email Portal Email Portal
A Live SupportA Live Support
Online ClassifiedsOnline Classifieds
Job Search portalJob Search portal
Online Community siteOnline Community site
Payment OptionsPayment Options
Software DevelopmentSoftware Development
+ view more
more
.............................................................................................................................................................................................................................................................

HOME   |   ABOUT US   |   PORTFOLIO   |   TEMPLATES   |   SERVICES   |   WEB HOSTING   |   PRODUCTS   |   OFFSHORE OUTSOURCING   |   ENQUIRY   |   CAREER