Showing posts with label Google. Show all posts
Showing posts with label Google. Show all posts

Thursday, July 7, 2016

Verbal Subject Analysis III: Webpage Databases (a.k.a. "Search Engines")

Human vs. Automatic Indexing

  • Both are related to the subject analysis of information resources.
  • Human indexing is used to describe the subject analysis of various periodical databases.
  • Automatic indexing is a term used for the subject analysis operations by the computer algorithms of various webpage databases (a.k.a. search engines).
    • Research from the 60s-80s were trying to get a computer to calculate what articles were about. The most frequent words, articles like a, an, the, etc., don't really tell you much about the article, neither do the least used words. The key is finding the sweet spot based  on what the author usually writes about.
Why Webpage Database?
  • It is always important to know the documentary unit of an information database. 
  • The adjective associated with database is always a cue to the documentary unit. 
  • Webpage databases are informational databases in which a webpage is the documentary unit. 
  • They are also known as search engines and discovered databases.
Analysis of Websites and their Structure
  • What are webpages? What are websites? Webpages:Websites as pages:books
  • Standards (or lack thereof) for the authoring of web sites and webpages
    • HTML and other markup languages
    • Editors
  • What are the implications of the lack of authoring standards for web-based information resources?
Location of Webpage Subject Metadata
  • In webpage headers: For individual webpages, subject metadata can be created by authors and included in HTML headers.
  • In separate metadata record databases:
    • Subject metadata can be created by intermediaries using Dublin Core schema
    • In search engines, subject metadata is inferred "automatically" by computer algorithm.
Search Engine Questions
  • For greater understanding we need to be able to answer:
    • Why do search engines produce different results the exact same query?
    • What is the principle for ranking the display of search engine records in response to a query?
The Term "Search Engine"
  • The term has become the common designation for webpage databases, However, in actuality, webpage databases have three parts:
    • Spidering/crawling software to collect webpages.
    • Indexing software to build the index of surrogate records.
    • Retrieval software to facilitate retrieval of surrogates.
Automatic Indexing in Context
  1. Obtain information resource - spidering/crawling
    • Steps for spidering/crawling:
      • Computers owned by search engine retrieve documents by clicking on all hyperlinks on each retrieved webpage
      • Determination is made whether a webpage needs to be indexed (because it is new) or reindexed (if it has already been indexed)
      • Determination is made whether reindexing is warranted
      • New webpages and those meeting criteria for reindexing are then placed in the indexing queue
  2. Describe information resource in surrogate record - read off webpages by indexing software
    • Left Side elements must be inferred by searcher:
      • Examine structure of retrieved records
      • Examine advanced search interface
      • Element sets are not standard, i.e., they will vary across search engines.
    • Right Side Content:
      • What is the source for the content?
      • Authority control?
  3. Subject analyze information resource in surrogate record - indexing software:
    • Verbal - inferred by computer algorithm
    • Classification - inferred by computer algorithm
    • Subject Indexing in Search Engines
      • The subject fields of webpage surrogate records include the words that describe what the webpage is about. 
      • Right side subject content is inferred through the application of proprietary algorithms.
      • Subject terms added to surrogate records are weighted:
        • Doc #1: SU = dogs (.99); breeding (.87);dachshund (.30)
        • Doc #2: cats (.92); dogs(.44); dachshund (.03)
        • The weights are computed by proprietary algorithm.
Retrieval from Search Engines
  • Unlike bibliographic databases, in which the ordering of retrieved surrogate records is reverse chronological, search engines use a relevance-based ranking.
  • The search engine component of a search engine takes the entered query and compares it to the terms to the index.
  • The documents that are retrieved first are those that contain a higher "relevance" score:
    • Doc #1: SU = dogs (.99); breeding (.87);dachshund (.30)
    • Doc #2: cats (.92); dogs(.44); dachshund (.03)
    • "dog" query would rank document #1 ahead of document #2
    • "breeding" query would rank document #1 ahead of document #2
    • "cats"query would rank document #2 ahead of document #1
How are Subject Weights Calculated?
  • Conventional methods (Dating from the 1950s) for automatically inferring what a document is about include the following three techniques:
    • Frequency of word occurrences
    • Location of words occurrences
    • Size of word occurrences
  • In the web era, however, these techniques did not scale well to meet the needs of databases containing billions of records:
    • Could facilitate retrieval of relevant documents, but could not distinguish between "good" and "bad" documents.
    • Were also subject to manipulation by authors desiring higher search engine retrieval (spamming)
Two responses to Early Indexing Failure
  • Yahoo! era (late 1990's)
    • Human indexing (website directories)
    • More discussion during lectures on classification. 
  • Google era (since 1999)
    • Additional criteria introduced to infer aboutness, e.g.,;
      • $ - paid submissions, such as Alta Vista
      • Quality - PageRank algorithm of Google
Google Approach to Authomatic Indexing
  • Issue addressed by Google concerns the quality problem: How to cause the "best" documents to rise to the top of a set of retrieved webpages.
  • Solution concerns identifying additional criteria to include int he subject weighting algorithm.
  • Google maintains additional metadata elements for each surrogate record in its index of webpages:
    • How many other webpages link to a given webpage
      • The more webpages (i.e. linkers) a dachshund webpage has poiting to it, the more quality it has.
      • This factors into the weight assigned to the "dachshund" descriptor inthe subject field of its surrogate record
    • Who are the linkers
      • Those linkers that have a higher quality rank are given more weight than those linkers with a lower quality rank.

Article Summary for Lecture # 10 - Northedge


and beyond:
information retrieval on the World Wide Web
Northedge defines a web directory as “a human compiled list of links to web pages, typically organized into a hierarchical structure of subject categories.” Back in 1994, a mere 3 years after Berners-Lee created the “Web”; there were less than 10,000 websites. This number inflated to almost 3.5 million in 1998, and in 2006, it was estimated to be at over 100 million. Imagine if those websites were books. Without anyone to organize and sort through all of them, it would take forever for us users to retrieve any kind of information, let alone navigate the sea of changes that authors and creators make on a daily basis to their sites. If a librarian is involved, the user can submit their queries to the librarian.
                In the case of the internet, search engines are the librarians. Several criteria measure the quality of the search engine, such as:
  • The size of the corpus – the more books the librarian can search, the better.
  • The speed of the answer – if we do not get our information quickly, we will find another search engine.
  • The availability of service – if it is not available when it is needed, the users are going to find another search engine.

·         The accuracy of results – if the information the user gets back is not what they are looking for, and then they will find another search engine that will return related results. However, if the three preceding criteria are not met, accurate data is not going to be important. (See this post).
Search engines require their users to submit their searches through a search box, which allows the user to choose whatever terms they like – unlike web directories, which constrain users to search using vocabulary chosen by the indexer. Since it might take a while for a search engine to sift through over 100 million constantly changing websites, it only makes sense to implement an indexing program (called a spider or robot). This program accesses web pages, analyses their contents and records the results in a database (referred to as an “index”), which enables fast access to sought information and bridges the gap between the search engine and the requested content.

                Today, one of the most used search engine is Google. Google’s software agent (indexer), called “Googlebot” continually locates billions of web pages, analyses the content, and save the result in the Google index. The algorithms it uses are a company secret, as they are what sets Google apart from its competitors (Bing, Yahoo, etc.). Googlebot breaks down webpages into words and examines their context within the page (position – is it in a header, sub header, body text, etc.) and sources are returned to user, based on the algorithms weighted scale, in order of assumed most relevant to least.

                In addition, while Google’s search box may seem to ask “What subject do you want information on?” in reality, it is asking “What word or combination of words will be most likely to appear on web pages that address the subject I am interested in, an least likely to appear on pages that are irrelevant to me?”. This may trip up users who are unfamiliar with how search engines work, and this may be the one negative Northedge presents about search engines – there is no one-to-one correspondence between words and meanings, and a single word may have multiple meanings (search for Java – the country – and only results about the computer programming language are returned). He also offers information on alternatives to search engines, which include META tags (the assignment of subject keywords by the web content creators), and folksonomies/tagging (creation of a taxonomy by the collective actions of users on the Web – see del.icio.us and flickr). These alternatives are somewhat controversial, because users and/or creators can deliberately assign misleading or inaccurate keywords to the content for financial gain or malicious reasons.

                Ultimately, Northedge offers insight into the possible future of web searches, computer-generated indexes, but the data contained in those indexes may be driven by data sets produced by human indexing techniques and human linguistic research. I agree with this assertation, because it seems as more technologies are developed and released, the search process becomes more streamlined and tailored to what the user REALLY wants from their search. This article is very informative, and if you want to know more about the inner-workings of search engines, this is a fascinating read. I definitely came away knowing more about what happens once I search for "cat videos" on Google. 
______________________________________
To read the whole article, see the citation below:

Northedge, R. (2007, April). Google and beyond: Information retrieval on the World Wide Web. The Indexer, 25(3), 192-195.

Friday, May 22, 2015

Information Tools and Next Generation Catalogs

The most widely-known information tool within the library is the library catalog. Traditional catalogs include books, CDs, DVDs, newspapers, magazines, microfilms, musical scores, etc. Next generation catalogs offer a social aspect within these catalogs. Patrons can tag, rate, and review materials, and is a wonderful attempt at merging older information tools with new information tools.

                   

Nowadays, if people have questions, they simply "Google" it. How can catalogs compete with the instant nature of Google? If someone has a question, Google can offer many resources within seconds, and traditional catalogs have tended to take time and expertise to use. However, with recent developments to catalogs mentioned above, usability is increasing, and libraries are creating search tools that mimic everyday tools used (tagging, bookmarking, reviewing, quick searches, etc.) It only makes sense for libraries to evolve and include these types of developments, as this is the future of information retrieval and dissemination.

Relevant readings:
  • Emanuel, Jenny. “Next Generation Catalogs: What Do They Do and Why Should We Care?”  Reference & User Services Quarterly 49.2 (2009): 117-120.
  • Yang, Sharon Q. and Kurt Wagner. “Evaluating and Comparing Discovery Tools: How Close Are We Towards Next Generation Catalog?”  Library Hi Tech 28.4 (2010): 690-709.