Showing posts with label retrieval. Show all posts
Showing posts with label retrieval. Show all posts

Thursday, July 7, 2016

Article Summary for Lecture # 10 - Northedge


and beyond:
information retrieval on the World Wide Web
Northedge defines a web directory as “a human compiled list of links to web pages, typically organized into a hierarchical structure of subject categories.” Back in 1994, a mere 3 years after Berners-Lee created the “Web”; there were less than 10,000 websites. This number inflated to almost 3.5 million in 1998, and in 2006, it was estimated to be at over 100 million. Imagine if those websites were books. Without anyone to organize and sort through all of them, it would take forever for us users to retrieve any kind of information, let alone navigate the sea of changes that authors and creators make on a daily basis to their sites. If a librarian is involved, the user can submit their queries to the librarian.
                In the case of the internet, search engines are the librarians. Several criteria measure the quality of the search engine, such as:
  • The size of the corpus – the more books the librarian can search, the better.
  • The speed of the answer – if we do not get our information quickly, we will find another search engine.
  • The availability of service – if it is not available when it is needed, the users are going to find another search engine.

·         The accuracy of results – if the information the user gets back is not what they are looking for, and then they will find another search engine that will return related results. However, if the three preceding criteria are not met, accurate data is not going to be important. (See this post).
Search engines require their users to submit their searches through a search box, which allows the user to choose whatever terms they like – unlike web directories, which constrain users to search using vocabulary chosen by the indexer. Since it might take a while for a search engine to sift through over 100 million constantly changing websites, it only makes sense to implement an indexing program (called a spider or robot). This program accesses web pages, analyses their contents and records the results in a database (referred to as an “index”), which enables fast access to sought information and bridges the gap between the search engine and the requested content.

                Today, one of the most used search engine is Google. Google’s software agent (indexer), called “Googlebot” continually locates billions of web pages, analyses the content, and save the result in the Google index. The algorithms it uses are a company secret, as they are what sets Google apart from its competitors (Bing, Yahoo, etc.). Googlebot breaks down webpages into words and examines their context within the page (position – is it in a header, sub header, body text, etc.) and sources are returned to user, based on the algorithms weighted scale, in order of assumed most relevant to least.

                In addition, while Google’s search box may seem to ask “What subject do you want information on?” in reality, it is asking “What word or combination of words will be most likely to appear on web pages that address the subject I am interested in, an least likely to appear on pages that are irrelevant to me?”. This may trip up users who are unfamiliar with how search engines work, and this may be the one negative Northedge presents about search engines – there is no one-to-one correspondence between words and meanings, and a single word may have multiple meanings (search for Java – the country – and only results about the computer programming language are returned). He also offers information on alternatives to search engines, which include META tags (the assignment of subject keywords by the web content creators), and folksonomies/tagging (creation of a taxonomy by the collective actions of users on the Web – see del.icio.us and flickr). These alternatives are somewhat controversial, because users and/or creators can deliberately assign misleading or inaccurate keywords to the content for financial gain or malicious reasons.

                Ultimately, Northedge offers insight into the possible future of web searches, computer-generated indexes, but the data contained in those indexes may be driven by data sets produced by human indexing techniques and human linguistic research. I agree with this assertation, because it seems as more technologies are developed and released, the search process becomes more streamlined and tailored to what the user REALLY wants from their search. This article is very informative, and if you want to know more about the inner-workings of search engines, this is a fascinating read. I definitely came away knowing more about what happens once I search for "cat videos" on Google. 
______________________________________
To read the whole article, see the citation below:

Northedge, R. (2007, April). Google and beyond: Information retrieval on the World Wide Web. The Indexer, 25(3), 192-195.

Thursday, June 2, 2016

Basic Retrieval Tools



What retrieval tool should you use?

  • For cited articles, consult a bibliography.
  • For library items, search a catalog.
  • For published articles, search a periodical database.
  • For archived items, use an archival finding aid.
  • For "published" web pages, use a search engine.

What is stored in retrieval tools?

Surrogate records - records that represent resources. If managed by a professional organizer they are structured surrogate records.

What are structured surrogate records?

Using a generic definition, structures are elements of individuals or things that exist across instances of those individuals or things (i.e. your skeleton, architectural elements of a house). In library terms, imagine the cards of a card catalog.
         
           Structures of a Card Catalog card:

    • Author
    • Title
    • Subject
What is the purpose?

To provide enough information for users to determine whether or not they want to actually pursue acquisition of the resource. The library wants to be economical with the user's time, regardless of resource type (book, artifact, or webpage).

What are some types of surrogates?
  • Appended lists (e.g. scholars who cite papers)
    • Citation is criterion - most common type of bibliography
      • cited scholarly articles
      • Cited books and other types of resources
    • Style guides serve as standardizing display mechanism over time:
      • Across authors
      • Across editors
      • Across journals
      • Across disciplines
  • Compiled lists (e.g. free-standing bibliographies)
    • Freestanding book-length works
      • Found in catalogs (e.g. A Bibliography of Austin Dobson)
      • Often the result of scholarly labor (i.e. used for tenure). including annotated bibliographies
    • Can be collocated based on various combined criteria:
      • Subject bibliographies
      • Author bibliographies
      • Bibliographies of a time period
      • Language, location, publisher, form bibliographies
  • Inventory lists (e.g. bookstores)
    • Surrogate list of resources held by bookstores. 
    • Intended to assist in book or other resource selection by patrons:
      • Mystical Unicorn for books (browse for "Sarah Eagle")
      • Amazon for books and other resources (keyword search for "Stephen King" and compare to advanced search capability to "limit" to an author search.
How are surrogates collocated?

Surrogates are collocated according to criterion (title, author, subject).

What is a library catalog?

A library catalog is a collection of surrogate records representing the collection of a single library (WorldCat). They have multiple access points that people can use to search for a resource like card catalogs or via electronic retrieval tools. Was created to help find what a user wants(author, title, name), to show what the library has (on an author, on a subject), or to assist in the choice of a book (bibliography, character).

What is an OPAC?

An OPAC is an Online Public Access Catalog. Interfaces for OPACs are either:

  • "Drop down box" (UA Library)
    • TRY IT! Keyword search for "dog" will retrieve different,limited results than a search for "dog" limited to the subject field.
  • "Tabbed" interfaces (UNT Library, WorldCat)

How do you search for Periodical Literature?

Indexes are a collection of surrogate records that represent the analyzed contents of resources like:

  • Journal articles
  • Conference proceedings/articles
  • Book Chapters
  • Websites
They have additional access points other than the traditional author, title, subject like issue, volume, year, etc.

TRY IT! Compare keyword search for "dog" vs. qualified search for "dog" as a subject in PubMed.

How do you search for specialized documents (images, artifacts,etc.)?

There are databases and indexes of digitized cultural heritage artifacts (e.g. Library of Congress, TIMEA). Some collections offer surrogate records that present long descriptions of archival collections (de Grummond Collection, Archives Portal Europe)


For Further Reading:


  • Creekmore, L. (2016). Training your eye to see structure. ASIST Bulleting 42(2):31-32.
  • NISO. (1998). Printed information on spines [click "Z39-41.pdf" link]. Bethesda: NISO Press.
  • Valente, C. (2009). Training successful paraprofessional copy catalogers. Library Resources and Technical Services 53:219-230.
  • Weinberg, A. (2016). In defense of the card catalog. Folger Shakespeare Library.