Showing posts with label review. Show all posts
Showing posts with label review. Show all posts

Thursday, July 7, 2016

Article Summary for Lecture # 10 - Northedge


and beyond:
information retrieval on the World Wide Web
Northedge defines a web directory as “a human compiled list of links to web pages, typically organized into a hierarchical structure of subject categories.” Back in 1994, a mere 3 years after Berners-Lee created the “Web”; there were less than 10,000 websites. This number inflated to almost 3.5 million in 1998, and in 2006, it was estimated to be at over 100 million. Imagine if those websites were books. Without anyone to organize and sort through all of them, it would take forever for us users to retrieve any kind of information, let alone navigate the sea of changes that authors and creators make on a daily basis to their sites. If a librarian is involved, the user can submit their queries to the librarian.
                In the case of the internet, search engines are the librarians. Several criteria measure the quality of the search engine, such as:
  • The size of the corpus – the more books the librarian can search, the better.
  • The speed of the answer – if we do not get our information quickly, we will find another search engine.
  • The availability of service – if it is not available when it is needed, the users are going to find another search engine.

·         The accuracy of results – if the information the user gets back is not what they are looking for, and then they will find another search engine that will return related results. However, if the three preceding criteria are not met, accurate data is not going to be important. (See this post).
Search engines require their users to submit their searches through a search box, which allows the user to choose whatever terms they like – unlike web directories, which constrain users to search using vocabulary chosen by the indexer. Since it might take a while for a search engine to sift through over 100 million constantly changing websites, it only makes sense to implement an indexing program (called a spider or robot). This program accesses web pages, analyses their contents and records the results in a database (referred to as an “index”), which enables fast access to sought information and bridges the gap between the search engine and the requested content.

                Today, one of the most used search engine is Google. Google’s software agent (indexer), called “Googlebot” continually locates billions of web pages, analyses the content, and save the result in the Google index. The algorithms it uses are a company secret, as they are what sets Google apart from its competitors (Bing, Yahoo, etc.). Googlebot breaks down webpages into words and examines their context within the page (position – is it in a header, sub header, body text, etc.) and sources are returned to user, based on the algorithms weighted scale, in order of assumed most relevant to least.

                In addition, while Google’s search box may seem to ask “What subject do you want information on?” in reality, it is asking “What word or combination of words will be most likely to appear on web pages that address the subject I am interested in, an least likely to appear on pages that are irrelevant to me?”. This may trip up users who are unfamiliar with how search engines work, and this may be the one negative Northedge presents about search engines – there is no one-to-one correspondence between words and meanings, and a single word may have multiple meanings (search for Java – the country – and only results about the computer programming language are returned). He also offers information on alternatives to search engines, which include META tags (the assignment of subject keywords by the web content creators), and folksonomies/tagging (creation of a taxonomy by the collective actions of users on the Web – see del.icio.us and flickr). These alternatives are somewhat controversial, because users and/or creators can deliberately assign misleading or inaccurate keywords to the content for financial gain or malicious reasons.

                Ultimately, Northedge offers insight into the possible future of web searches, computer-generated indexes, but the data contained in those indexes may be driven by data sets produced by human indexing techniques and human linguistic research. I agree with this assertation, because it seems as more technologies are developed and released, the search process becomes more streamlined and tailored to what the user REALLY wants from their search. This article is very informative, and if you want to know more about the inner-workings of search engines, this is a fascinating read. I definitely came away knowing more about what happens once I search for "cat videos" on Google. 
______________________________________
To read the whole article, see the citation below:

Northedge, R. (2007, April). Google and beyond: Information retrieval on the World Wide Web. The Indexer, 25(3), 192-195.

Tuesday, June 28, 2016

Article Summary for Lecture # 7 - Aitchison

The Thesaurus:
A Historical Viewpoint, 
with a Look to the Future


Four decades of use, experimentation, and development have allowed users, researchers, and catalogers to refine thesauri to be very effective search tools. Aitchison and Clarke discuss the history, making note of important, monumental events that help in the creation of what we know as the thesaurus. They draw on earlier printed histories of thesauri (specifically Gilchrist’s Thesaurus in Retrieval), and go on to define thesaurus as “a treasury or storehouse of knowledge, as a dictionary, encyclopedia and the like.” The primary purpose of thesauri is to match the vocabulary used by the indexer with the language of the searcher.

The first time the word “thesaurus” was used (in terms of information retrieval) was in 1957 by Peter Luhn of IBM, and had definitely evolved through the 1950s. One particular highlight on the timeline of the thesaurus is the Uniterm System, which used uncontrolled single words taken from the text of documents, which ultimately proved difficult since only single-word terms were available to deal with synonyms, homonyms, etc. Fortunately, it was superseded by vocabularies containing significant numbers of compound terms. In addition, during this time, the thesauri listed terms in alphabetical order, which was eventually carried into the standardization of format in 1967 when the Thesaurus of Engineering and Scientific Terms (TEST) was published.

This predominant feature alphabetically displays descriptors and non-descriptors - synonyms, broader, narrower, and related terms showing under each descriptor. A subject overview or systematic display was of secondary importance. The idea of a detailed classified arrangement was considered too complex. An example of this is the Descriptor Group Display, within which main groups are divided into subgroups, and within the subgroups, descriptors are further organized into clusters. For example:

14    DEMOGRAPHY. POPULATIONS
                14.01 POPULATION DYNAMICS
                14.01.01
                    CIVIL REGISTRATION
                    DEMOGRAPHIC STATISTICS
                    POPULATION DATA
                       USE: DEMOGRAPHIC STATISTICS
                    etc.

In order to section thesaurus information in this way, the classification scheme is an indispensable tool. When the editor works only with an alphabetical list, it is a sense of working blind, but if rigorous classification is developed, the compiler has a better chance of building accurate and meaningful relationships between the terms. In the early days, most thesauri were compiled manually, which was a massive and both a time and space investment (e.g. Thesaurofacet was held in more than 20 shoeboxes containing cards for 16,000 descriptors and 7,000 non-descriptors), and was greatly vulnerable to human errors or mid-process interruptions. This is where computer-aided compilation becomes handy. In the late 1970s, computer compilation was more common; however, there was no software to maintain a systematic display of the faceted thesaurus style.

During this time, access was usually limited to one workplace. There was either a large tome that stood by the bank of filing cards or optical coincidence viewer, and even computerized thesauri were limited in space. However, trained searchers became fluent with the process, and that alongside trained indexers they were able to fully harness the power of the thesauri to perform effective searches. Nowadays, pcs are everywhere, and each of them provide access to unlimited networks. In order to apply thesauri to information retrieval the authors feel the following challenges need to be addressed:
  • Access to information proceeds through any number of different portals, gateways, and search engines, many geared to particular audiences and subject areas. There is no universal thesaurus, but a multitude of different vocabularies for different applications.
  • In the publish one, re-utilize many times’ environment, it is hard to predict in which systems or networks a given document may eventually appear. Indexers must struggle to foresee all the needs that may arise for a given document.
  • With the data entry/indexing task distributed among a vast number of authors, webmasters, system administrators, etc., quality control cannot be enforced across organizational boundaries.
  • How can we train end-users to use a thesaurus properly? The experience of most information providers is that users do not want to cope with anything complicated, and the thesaurus is perceived as very complicated. Those beautifully presented systematic displays, carefully designed for selecting the right term(s) for each required concept, are often rejected as an unnecessary impediment and delay between the user and goal.

Confronting these challenges has recently led to two major trends in thesaurus developments:
  • Hunting for adaptations that will make a controlled vocabulary much quicker, easier, and more intuitive to use.
  • Drive to interoperability of systems, meaning to design vocabularies for easy integration into downstream applications such as content management systems, indexing/meta-tagging interface, search engines, and portals.

Current technology has users who are more than happy to browse through a simple classified directory, using point-and-click interaction with established headings instead of actively thinking of search terms. This is why some companies are working on developing taxonomies that will make things easier for the searcher, and perhaps even the indexer. There are even mentions of hiding vocabulary all together by implementing synonyms sets for selected terms that can be used to drive automatic expansion of free-text search queries.

Another topic current thesauri developers need to think about is interoperability. It makes things easier for users. Gone are the days of looking up in a printed thesaurus and then key selected terms into the indexing system. Now, copy-paste or clicking on them, a search system has to be capable of interacting with the thesaurus database.These newest concerns were reflected in the updated standards released in the Workshop on Electronic Thesauri held in 1999, which says, “The standard should provide for a broader group of controlled vocabularies than those that fit the standard definition of “thesaurus.” This includes, for example, ontologies, classifications, taxonomies and subject headings, in addition to standard thesauri. The primary concern is with shareability (interoperability), rather than with construction or display. Therefore, this new standard will probably not supersede Z39.19, but supplement it.

Overall, I think Aitchison and Clarke’s article is very thorough and offers a lot of insight into the world of thesauri. This is a great read for anyone interested in the development of thesauri or the organization of information. They have a ton of information and references to support their examples, and write in a way that is easy for even beginners to be able to pull information from the article and form new ideas and appreciation for the thesaurus in their life!
_______________________________________________________________________
For more information, check out the full article (citation below)!

Aitchison, J., & Clarke, S.D. (2004). The thesaurus: A historical viewpoint, with a look to the future. Cataloging & Classification Quarterly 37(3/4):5-21.

Thursday, June 9, 2016

Article Summary for Lecture # 4 - Carlyle

Understanding FRBR As a Conceptual Model
FRBR and the Bibliographic Universe

Allyson Carlyle, in her article, examines FRBR’s status as a model, and seeks to clarify what it is, what it is not, and what it attempts to do. In essence, FRBR is a conceptual model (or a model of a model – if one considers that a bibliographic record is a representation of a document). Carlyle begins with the deconstruction by reviewing the multiple definitions of “model”. A model can be:

1)      a representation of something (sometimes on a smaller scale)
2)      a schematic description of a system, theory, or phenomenon that accounts for its known or inferred properties and may be used for further study of its characteristics: a model of generative grammar; a model of an atom; an economic model
3)      a simplified description of a complex entity or process
4)      a preliminary work or construction that serves as a plan from which a final product is to be made: a clay model ready for casting.

           Being a conceptual model (sometimes referred to as an abstract model), it is entirely theoretical. This is a major strength, because it helps to facilitate understanding and allows manipulation of complex entities by making them less complex. Conceptual models can also model just about anything, and FRBR is a great way to make something that is abstract into something that is concrete. For example, trying to model love – an abstract concept. In order to make a model of love, we need to operationalize it. How many times two people kiss each other, how much time do they spend together, or whether or not they live together can be observable items that help model love.

        The FRBR Group 1 entities work and expression are abstract ideas that have a lot in common with love. In order to verify the existence of work and expression we need to know:

1)      What documents say about themselves and what others say about them
2)      What people say when they want to find a document – for example:
a.       Do you have Seamus Heaney’s translation of Beowulf? (expression)
b.      Do you have Stephen Hawking’s A Brief History of Time? (work)

Since FRBR is such a specific type of conceptual model, it can be referred to as an entity-relationship (ER) model. Carlyle describes ER modeling as a technique that specifies the structure of a conceptual model (i.e. the kinds of things that have to be in it and the properties those things may have). For FRBR, and an ER model, three things are allowed in it: entities (things – either physical or abstract), attributes (properties or characteristics of entities and relationships), and relationships (interactions among entities). 

       The developers of FRBR aimed to “produce a framework that would provide a clear, precisely stated, and commonly shared understanding of what it is that the bibliographic record aims to provide information about.” However, to clearly understand FRBR, it pays to look at other models.
  •  One Entity Model – the only entity recognized it “item” or ‘copy”
  • Two-Entity Model – recognizes editions as well as copies
  • Three-Entity Model – recognizes editions, copies, and “literary unit” (a.k.a. work)
  •  Four-Entity Model – recognizes editions, copies, works, and “text”

FRBR is new and different from these models in that it identifies and defines four entities, recognizes four entities simultaneously, and present a cataloging model using an ER modeling technique. However, according to Carlyle, one of the greatest challenges in implementing FRBR in a code of rules is determining which items will be assigned by catalogers to which set. In the implementation process, decisions about the boundaries of the abstract entities work and expression must be made. (e.g. will a movie version of an original textual work be considered an expression of that work, or will it be considered to be a new work with a derivative relationship to the original?

      Ultimately, Carlyle concludes that successful implementation of FRBR will help patrons successfully perform searches by presenting information about complex works in helpful and intelligent ways, and encourages the viewing of FRBR as a continuation/natural extension of cataloging models used over the centuries cataloging has been around. I would agree with her summarization and emphasis of embracing FRBR. As I have said in previous posts, the job of an information organizer to get quality information into the hands of users through the easiest means. FRBR takes complex entries, and helps break down the information (entities) in such a way retrieval is easy. If you are interested in FRBR, or think you should try and figure out what FRBR is about or clarify a muddy view of the model, this is the perfect article for you!
_________________________________________________________________________________
If you want to read more about this topic, be sure to check out the whole article (citation below)!
Carlyle, A. (2006). Understanding FRBR as a conceptual model: FRBR and the bibliographic universe. Library Resources & Technical Services 50:264-73

Tuesday, June 7, 2016

Article Summary for Lecture # 3 - Russel

Hidden Wisdom and Unseen Treasure: Revisiting Cataloging in Medieval Libraries


Beth Russell in her article “Hidden Wisdom and Unseen Treasure: Revisiting Cataloging in Medieval Libraries”, addresses the challenges of medieval catalogers and summarizes recent research and discovery in the field of medieval libraries. There are no two medieval libraries alike. Since there was no consensus between library professionals regarding the categorization and organization of texts, each library had its own practice for cataloging materials. However, the struggles of medieval catalogers are not too foreign from those of modern catalogers. Both seek to provide easy access to materials for patrons, and that is in itself the core of library cataloging.

When trying to accommodate users, medieval librarians did not have national and international standards upon which to begin building their system, they only had the needs of their patrons to consider. Most early examples of library cataloging take place in monasteries, where the most basic type of classification for books is use. For example, liturgical or service books were stored near the chapel since their function was for use in the chapel. In some situations, specifically Durham Cathedral’s catalog 1391-1395, an iron grille divided stored books, the inner portion being restricted use and the outer portion accessible to any patron, or monk in the case of Durham.

However, once collections started to diversify, there were two collections stored in different rooms with different keys (e.g. Sorbonne’s magna libraria and libraria parva), which was likely due to both secular and religious texts acquired at universities. In the beginnings, chained volumes (pictured above) were not included in early catalogs, as they did not need to be kept track of. In terms of what kinds of information were found on cataloged items, catalogers would often assign letters of the alphabet to the volumes or gave detailed descriptions of the shelf where they could be found. In addition, catalogers would physically describe the books, the number of volumes, their size, and the completeness of the set (if applicable). For multiple copies, catalogers would distinguish each copy by description. Trinity Hall, Cambridge, for example, listed a duplicate as “magna et pulchra” (big and pretty), which would help distinguish it from another copy of the same text that was small and ragged.

Later in medieval times, catalogers would record the opening words of the first leaf in the volume as an organizational tool. This shows that even in medieval times, catalogers realized that distinguishing among copies was necessary for duplicates. As organization evolved, there was an acceptance of alphabetical order for subjects, which began as early as the twelfth century. However, Russell points out that there is the issue of cataloger bias, and uses Glastonbury Abbey as an example. At Glastonbury, books with “interesting” subjects or whose author was not illustrious were cataloged by subject, whereas works by famous authors were filed under that author’s name (with no mention of subject). In this case, the cataloger’s opinion and knowledge defines whether the topic is interesting enough to have a subject heading, or if the author is well known enough to be filed by name.
As should modern libraries, when the use of books changed and when the number of books in the collections increased, so did the catalog and the way in which things were cataloged. Russell comments that "modern catalogers struggling to meet local needs in a cooperative electronic environment can look for inspiration and for examples of ingenuity and invention, to our early colleagues in centuries past, who dealt with similar problems in organizing the knowledge in their care." I believe it is extremely important to learn from the past to aid in future decisions – those who do not learn from history are doomed to repeat it. We, as librarians, catalogers, and organizers of information need to be inspired by the ingenuity of past ages, and understand that change isn't a bad thing, in fact, it might be the best thing for your collection. 

Overall, Russell's article was very entertaining and insightful. The organization of the information within this article felt chronological, though I felt like there could be some sort of better organization to the examples of cataloging in medieval times. There are so many examples thrown out, it definitely puts into perspective the non-conformative nature of the medieval catalog, which may be what Russell was going for. I would recommend this article to medieval historians, library and information science professionals, and library enthusiasts alike. I agree with her assertation that we modern information specialists should take note of historical practices and inventiveness and never feel unable to adapt to new usage and new organization of our collections.

______________________________________________________________________

If you want more examples, be sure to check out the whole article (citation below)!

Russell, B.M. (1998). Hidden wisdom and unseen treasure: Revisiting cataloging in medieval libraries. Cataloging & Classification Quarterly 26(3):21-30.