Contents
Sl no Topic Page no.
1 Introduction 1
2 Naming 1
3 Size 2-3
4 Accessing 3
5 Deep Resources 3-4
6 Distribution of Deep Web Sites 4-5
7 Crawling of the Deep Web 5-6
8 Conclusion 6
1
1 Introduction
The deep web (also called Deepnet, the invisible Web, DarkNet ,dark Web or the hidden Web)
represents all of the data, information, content and services that are hidden from standard
search engine crawlers and are therefore not easily found or accessed. The types of content
that exists in this region includes practically everything imaginable from product reviews to
videos to research papers and everything in between. In general, this hidden information is
more current and is more content rich than what is generally available on the surface.
There are several types of sites that can be categorized as it relates to the Deep Web,
most of these sites have made concsious decisions to make certain content hidden or invisible
from the search engines. The reasons for this include everything from proprietary information on
corporate intranets to subscription based systems that distribute content for a fee. The types of
sites include the following:
Visible: These are traditional websites (Web 1.0) that allow all of their content to be
easily found, accessed and indexed by the search engines,
Private: A site in this category basically has a single page with no accessible links on it
until a user is signed in with the appropriate credentials, these are typically banks,
brokers and other proprietary systems with confidential content,
Opaque: A site that has a home page with no links or obvious content, typically this
represents nothing more than a marker indicating where subsurface content is stored,
Proprietary: These sites are basically pay-per-view in that they require a subscription
and payments to access the content they contain,
Invisible: These are sites that contain no content that can be indexed by the search
engines, everything is encrypted and stored in formats that cannot be read by third party
services.
2 Naming
Bergman, in a seminal, early paper on the deep Web published in the Journal of Electronic
Publishing, mentioned that Jill Ellsworth used the term invisible Web in 1994 to refer to websites
that are not registered with any search engine.
Another early use of the term invisible Web was by Bruce Mount and Matthew B. Koll of
Personal Library Software, in a description of the @1 deep Web tool found in a December 1996
press release.
The first use of the specific term deep Web, now generally accepted, occurred in the
aforementioned 2001 Bergman study.
3 Size
There are widely differing opinions on this with the low estimates indicating about 75% and the
higher estimates approaching 99% of all internet based content is hidden, the reality is that it's
2
somewhere in between. Regardless of the exact amount it is clear that most of the content on
the Internet remains hidden from the search engines and therefore not easily found or accessed
by search engine users. Researchers and Internet guru's have known about this for several
years and have developed alternative methods for tapping into this repository.
Again, regardless of the actual volume most of the
information contained in the deep web is more valuable and
more current than information accessible on the surface.
Back in 2000 it was estimated that the internet contained
about 7,500 terabytes and around 550 billion individual
documents and files. In today's terms, an estimate of the
deep web content from UC Berkeley suggests about 91,000
terabytes, in contrast the surface web which contains only
Figure 1: figure depicting the
about 167 terabytes. To put this in some perspective the
size of Deep web vs Surface
Library of Congress, in 1997, was estimated to have around
web
3,000 terabytes of content.
Bright Planet, estimates that the dark web contains
500 times more content then the indexed surface layer. Considering that Google indexes about
8 billion pages, that number is incredible: 4,500,000,000,000 trillion pages of content! If you go
to the Bright Planet site it indicates that the deep web is actually thousands of times larger than
the surface, it's growing faster than I thought.
4 Accessing
To discover content on the Web, search engines use web crawlers that follow hyperlinks. This
technique is ideal for discovering resources on the surface Web but is often ineffective at finding
deep Web resources. For example, these crawlers do not attempt to find dynamic pages that
are the result of database queries due to the infinite number of queries that are possible. It has
been noted that this can be (partially) overcome by providing links to query results, but this
could unintentionally inflate the popularity for a member of the deep Web.
In 2005, Yahoo! made a small part of the deep Web searchable by releasing Yahoo!
Subscriptions. This search engine searches through a few subscription-only Web sites. Some
subscription websites display their full content to search engine robots so they will show up in
user searches, but then show users a login or subscription page when they click a link from the
search engine results page.
5 Deep resources
The iceberg is the perfect metaphor to describe the deep web because most of its mass is
below the surface. The deep web's informational mass is below the surface and not accessible
3
by reading static web pages. Deep Web resources may be classified into one or more of the
following categories:
Dynamic content: dynamic pages which are returned in response to a submitted query or
accessed only through a form, especially if open-domain input elements (such as text
fields) are used; such fields are hard to navigate without domain knowledge.
Unlinked content: pages which are not linked to by other pages, which may prevent Web
crawling programs from accessing the content. This content is referred to as pages
without backlinks (or inlinks).
Private Web: sites that require registration and login (password-protected resources).
Contextual Web: pages with content varying for different access contexts (e.g., ranges of
client IP addresses or previous navigation sequence).
Limited access content: sites that limit access to their pages in a technical way (e.g.,
using the Robots Exclusion Standard, CAPTCHAs, or no-cache Pragma HTTP headers
which prohibit search engines from browsing them and creating cached copies).
Scripted content: pages that are only accessible through links produced by JavaScript as
well as content dynamically downloaded from Web servers via Flash or Ajax solutions.
Non-HTML/text content: textual content encoded in multimedia (image or video) files or
specific file formats not handled by search engines.
6 Distribution of Deep Web Sites
To a large extent the content that exists just beyond the reach of the search engines includes
business intranets, academic research and content rich databases. The free web lives here as
well, this includes technologies like TOR, I2P and Freenet that allow content to be transferred
around without being traceable and with total anonymity of its users.
The term hidden may
be a misnomer here because
much of the deep web can be
found and accessed for free,
but one has to know how to do
it. In general, search engines
cannot access this content
unless the site chooses to
expose it to them which in
many cases they don't for
business reasons or can't for
legal reasons.
The idea is that around
95% of the deep web is
Figure 1: Distribution of Deep Web Sites by content type
contained in database tables
and other discrete formats, the
4
vast majority of this can be accessed for free. The accessible content that tends to be hidden
from view includes binary files like software applications, images, videos and other forms of
multimedia content. It also includes massive databases like the types listed below:
Governement databases which at the federal level include curated databases for almost
every agency and department within the government. An example of this can be seen at
the Medical Device database form published by the FDA.
Millions of topic specific databases exist like the plane crash database or the toxic
chemicals database. This content is generated and maintained by a whole range of
entities including industry and trade associations, all types of businesses, various
organizations and even individuals.
Academic databases and publications are one of the main constituents of the deep web,
having been the most prolific type of users on the net for the first 10 years means lots of
curated content. The content here is also the most discrete and therefore pure, as this
content makes its way to the surface it tends to be diluted for mass consumption.
The content that is completely inaccessible by search engines is made up of proprietory
databases that require some sort of membership, businsess intranets, and subscription
services. This content is intentionally made unavailable for reasons already mentioned like
potential legal issues, competitive advantages, propreitary information as well as others, most of
which are obvious. An example is Facebook which has to consider legal issues as well as
maintaining a competitive advantage to attract advertising dollars.
7 Crawling the deep Web
For much of the content that exists in the deep web there are a limited number of methods for
searching and finding it, in many cases there is only a single method of access exposed by the
content producer, typically via some web based query form. Using Amazon as an example with
their huge repository of activities, there are basically three methods to access that content:
Via the Amazon website directly,
Via Amazon partner sites, and
Via the Amazon web services.
The first two methods are fairly obvious but the web service based approach is gaining
popularity and becoming a ubiquitous model for interaction with Amazon services
There is another way to get to much of the deep web content through specialized search
engines that were built to harvest and index the content in these lower layers of the Internet.
There are literally hundreds of these search engines available to use but many are limited in
scope. The limitations are generally related to a certain type of content or a specific industry, so
while they reduce the number of places to search there is still no single source to go to like
Google is for the surface web. Here's a list of some of the popular deep web search engines:
Complete Planet: the deep web directory
The World Wide Web Virtual library
5
Infomine: Scholarly Internet resource collections
Intute( helps find deep web websites for research and study)
Dialog provides authoritative answers enhanced by ProQuest.
Google Scholar is an academic and research search engine
This is just a starter list but should provide a good jumping off point for exploring,
A set of trends are developing that should have a positive impact on our ability to
leverage the content buried in the deep web, these trends are the use of web services as a
method of syndicating both content and services, and the new application development
paradigm of mashups. Both of these trends conspire to generate a meaningful set of tools that
can be used to search out and find content that is currently only available to the elite users of
the Internet. A recent publication posted at SpringerLink entitled: Mashups over the Deep Web
discusses this new application paradigm and how it can be used to successfully traverse the
deep web.
8 Conclusion
The deep Web thus appears to be a critical source when it is imperative to find a "needle in a
haystack." The deep web represents a huge, almost undefinable repository of valuable content
and services that are generally considered more current and of higher quality than anything
found on the surface. But the lines between search engine content and the deep Web have
begun to blur, as search services start to provide access to part or all of once-restricted content.
An increasing amount of deep Web content is opening up to free search as publishers and
libraries make agreements with large search engines. In the future, deep Web content may be
defined less by opportunity for search than by access fees or other types of authentication.
Reference
[1] Deep Web(From Wikipedia, the free encyclopedia) [Online],1/11/2010,
Available: [Link]
[2] Dave Tribbett,The Deep, Dark Invisible Web [Online], 1/11/2010,
Available: [Link]
[3] Deep Web [Online], 1/11/2010, Available: [Link]