Showing posts with label indexing. Show all posts
Showing posts with label indexing. Show all posts

Tuesday, July 12, 2016

Article Summary for Lecture # 11 - Barite

The Notion of "Category":
Its Implications in Subject Analysis and in the Construction and Evaluation of 
Indexing Languages

       Mario Guido Barite, a professor and researcher at the School of Librarianship at the University of the Republic in Uruguay, attempts to tackle the notion of category (a basic intellectual tool for the analysis of the existence and changeableness of things) and proposes conceptual and methodological reexamination from a functional standpoint. Barite claims that "most classifiers or indexers assume the role of classificationist since the present state of indexing languages entails minor and major surgery be performed to adapt these languages to users' requirements." 
   
      Categories necessarily are the foundation of any organizational system of knowledge, however  category, characteristic, or class are sometimes used indistinctly. It is not possible to characterize categories in the Theory of Classification, as categories are extremely general abstract expressions. Categories are used as tools to discover certain regularities of the material world, but Barite suggest properties as a possible category to analyze the material world since categories are, in their basic nature, extremely simple notions. Within the Theory of Classification, categories are only relevant as instruments of analysis and organization of object, phenomena and knowledge, and classificationists are used in three precise activities Barite mentions:
  • design, planning, and structuring of indexing languages or systems of knowledge
  • modification or specification of classification tables
  • the evaluation and analysis of indexing languages and systems of concepts through a set of parameters capable of establishing the grade of reciprocal tension among related concepts and their relevance and validity. 
       Since it is not possible to isolate the notion of category from those of object and analyst there are several object attributes that Barite suggests condition its study:
  • Any object  is naturally dynamic and mutable - that being the case, in order for the analysis to be completed, the object must be captured at a certain time and abstraction from its reality is required at a given moment.
  • The object may be real or ideal - it may have existed as may be corroborated by its existence registers or maybe it only has an immaterial existence, not physical, due to its nature. These particular characteristics seem to obstruct the analysis since analysts are condemned to act by approximation. However, once conventions have been clearly established by conses, abstract objects are easily systematized after agreement has been reached regarding what a theorem is or certain chronological and factual conventions of the French Revolution - the difficulty of giving intellectual access to the concept diminishes.
  • Some objects have delimitation problems - attempts to produce a definition usually create discrepancies and shades of meaning among experts, so much that they may cause a certain aspect of the object to be placed within one category or the other. But we also have the difficulties posed by the concepts that do not attain conventional agreement. To exemplify, think of the difficulty of approving by consensus the basic statements towards the definition of the concept  labor flexibilization from the viewpoint of a sociologist with a Marxist orientation and another one of ultra-liberal ideas.
  • A large part of the objects belong to, or occur in a phase of the time-space continuum, or rather flow along a section of that continuum. - Due to their mutating and dynamic nature, some objects achieve various configurations and undergo a double influence: that of the processes occurring as a result of the action of internal agents, and that of the processes caused by external agents. This double influence is the determinant of each specific configuration, since any object is in a given time and in a given spatial situation, the synthesis of the impacts brought about by such agents. 
      He then goes on to decompose the notion of category to extract its most typical characteristics:
  • Every category is a sectorial one. 
  • Every category implies a specific level of analysis. 
  • Categories are levels of analysis external to the object.
  • Categories are mutually excluding.
  • Every category is highly generalizable. 
  • Every category may admit, with reference to an object, variable levels of subdivision.
  • Agreement has not been reached regarding a limited collection of categories.
Barite concludes that the proposition of greater attention on the definition of category, because it involves essential theoretical practical aspects for the reasonable command of the theory of concepts by specialists. I agree that it is important to dissect, correct, and fully understand terminology. If we can't fully understand the terminology within a an organization system, the system will not be efficient enough to get quality information into the hands of patrons with ease. The article is a bit philosophic, which made it somewhat of a struggle to read, but the heart of the article and Barite's ideas are spot on. I'd recommend this for organizers everywhere, as it really gets you to think about the elements of a system and the terminology used.

________________________________________________________________________
 For more information see the full article (citation below!)

 Barite, M. (2000). The notion of "category:" Its implications in subject analysis and in the  construction and evaluation of indexing languages. Knowledge Organization 27:4-10.

Thursday, July 7, 2016

Verbal Subject Analysis III: Webpage Databases (a.k.a. "Search Engines")

Human vs. Automatic Indexing

  • Both are related to the subject analysis of information resources.
  • Human indexing is used to describe the subject analysis of various periodical databases.
  • Automatic indexing is a term used for the subject analysis operations by the computer algorithms of various webpage databases (a.k.a. search engines).
    • Research from the 60s-80s were trying to get a computer to calculate what articles were about. The most frequent words, articles like a, an, the, etc., don't really tell you much about the article, neither do the least used words. The key is finding the sweet spot based  on what the author usually writes about.
Why Webpage Database?
  • It is always important to know the documentary unit of an information database. 
  • The adjective associated with database is always a cue to the documentary unit. 
  • Webpage databases are informational databases in which a webpage is the documentary unit. 
  • They are also known as search engines and discovered databases.
Analysis of Websites and their Structure
  • What are webpages? What are websites? Webpages:Websites as pages:books
  • Standards (or lack thereof) for the authoring of web sites and webpages
    • HTML and other markup languages
    • Editors
  • What are the implications of the lack of authoring standards for web-based information resources?
Location of Webpage Subject Metadata
  • In webpage headers: For individual webpages, subject metadata can be created by authors and included in HTML headers.
  • In separate metadata record databases:
    • Subject metadata can be created by intermediaries using Dublin Core schema
    • In search engines, subject metadata is inferred "automatically" by computer algorithm.
Search Engine Questions
  • For greater understanding we need to be able to answer:
    • Why do search engines produce different results the exact same query?
    • What is the principle for ranking the display of search engine records in response to a query?
The Term "Search Engine"
  • The term has become the common designation for webpage databases, However, in actuality, webpage databases have three parts:
    • Spidering/crawling software to collect webpages.
    • Indexing software to build the index of surrogate records.
    • Retrieval software to facilitate retrieval of surrogates.
Automatic Indexing in Context
  1. Obtain information resource - spidering/crawling
    • Steps for spidering/crawling:
      • Computers owned by search engine retrieve documents by clicking on all hyperlinks on each retrieved webpage
      • Determination is made whether a webpage needs to be indexed (because it is new) or reindexed (if it has already been indexed)
      • Determination is made whether reindexing is warranted
      • New webpages and those meeting criteria for reindexing are then placed in the indexing queue
  2. Describe information resource in surrogate record - read off webpages by indexing software
    • Left Side elements must be inferred by searcher:
      • Examine structure of retrieved records
      • Examine advanced search interface
      • Element sets are not standard, i.e., they will vary across search engines.
    • Right Side Content:
      • What is the source for the content?
      • Authority control?
  3. Subject analyze information resource in surrogate record - indexing software:
    • Verbal - inferred by computer algorithm
    • Classification - inferred by computer algorithm
    • Subject Indexing in Search Engines
      • The subject fields of webpage surrogate records include the words that describe what the webpage is about. 
      • Right side subject content is inferred through the application of proprietary algorithms.
      • Subject terms added to surrogate records are weighted:
        • Doc #1: SU = dogs (.99); breeding (.87);dachshund (.30)
        • Doc #2: cats (.92); dogs(.44); dachshund (.03)
        • The weights are computed by proprietary algorithm.
Retrieval from Search Engines
  • Unlike bibliographic databases, in which the ordering of retrieved surrogate records is reverse chronological, search engines use a relevance-based ranking.
  • The search engine component of a search engine takes the entered query and compares it to the terms to the index.
  • The documents that are retrieved first are those that contain a higher "relevance" score:
    • Doc #1: SU = dogs (.99); breeding (.87);dachshund (.30)
    • Doc #2: cats (.92); dogs(.44); dachshund (.03)
    • "dog" query would rank document #1 ahead of document #2
    • "breeding" query would rank document #1 ahead of document #2
    • "cats"query would rank document #2 ahead of document #1
How are Subject Weights Calculated?
  • Conventional methods (Dating from the 1950s) for automatically inferring what a document is about include the following three techniques:
    • Frequency of word occurrences
    • Location of words occurrences
    • Size of word occurrences
  • In the web era, however, these techniques did not scale well to meet the needs of databases containing billions of records:
    • Could facilitate retrieval of relevant documents, but could not distinguish between "good" and "bad" documents.
    • Were also subject to manipulation by authors desiring higher search engine retrieval (spamming)
Two responses to Early Indexing Failure
  • Yahoo! era (late 1990's)
    • Human indexing (website directories)
    • More discussion during lectures on classification. 
  • Google era (since 1999)
    • Additional criteria introduced to infer aboutness, e.g.,;
      • $ - paid submissions, such as Alta Vista
      • Quality - PageRank algorithm of Google
Google Approach to Authomatic Indexing
  • Issue addressed by Google concerns the quality problem: How to cause the "best" documents to rise to the top of a set of retrieved webpages.
  • Solution concerns identifying additional criteria to include int he subject weighting algorithm.
  • Google maintains additional metadata elements for each surrogate record in its index of webpages:
    • How many other webpages link to a given webpage
      • The more webpages (i.e. linkers) a dachshund webpage has poiting to it, the more quality it has.
      • This factors into the weight assigned to the "dachshund" descriptor inthe subject field of its surrogate record
    • Who are the linkers
      • Those linkers that have a higher quality rank are given more weight than those linkers with a lower quality rank.

Tuesday, July 5, 2016

Verbal Subject Analysis II: Periodical and Other Databases

Subject Cataloging vs. Indexing
  • Both are related to the subject analysis of resources. 
  • Subject cataloging  is a term used for the subject analysis operations in library cataloging. 
  • Indexing is a term generally used for the subject analysis operations in various other resource organization contexts, including periodical databases and search engines. 
Brief History of Periodical Indexes
  • Around the turn of the 20th century, the library community decided not to add article citations to the catalog. 
  • This development led to the growth of the commercial indexing industry. 
  • The result of this has been:
    • Split files
    • Fees for licensing database content
    • Difficulty fulfilling Cutter's 2nd objective
Analytical Cataloging
  • Analytical cataloging techniques are needed in order to provide access to the component parts of composite information resources, most commonly:
    • Book chapters
    • Proceedings articles (usually of academic meetings)
    • Journal articles
  • Definition from AARC2: Analysis is the process of preparing a bibliographic record that describes a part (or parts) of an item for which a comprehensive entry is made.
Analytical Cataloging Techniques
  • Complex entries made within the record of composite work [cheap]:
    • Analytical added entries:
      • Use 740 tag for second of two works mentioned in title of item
    • Note area for comprehensive entry of larger work:
      • Use 505 tag for structured display of table of contents.
  • Separate records created for the component parts of composite works ("In" Analytics)[expensive]:
    • Use 773 to trace the component part record to parent record
Analytical Access to Journal Content
  • Decision to not provide analytical access to journal content (i.e. directly to articles) was because of the expense:
    • Excessive number of records would have to be created.
    • Additional authority work would need to be done.
  • As a result, through the 20th century, cataloging and periodical indexing/bibliography creation techniques evolved separate approaches. 
Overview Comparison
  • Catalog
    • Authority work
    • Cataloging records represents the holdings of a library
  • Periodical indexes:
    • Subject indexes are extensive topical bibliographies (often include books and book chapters, too), usually covering large swaths of "territory"
    • Domain-wide indexes (e.g. Index Medicus) attempt to capture an entire discipline (may include book chapters, too)
    • No single library could ever own all items referred to in exhaustive bibliographies/indexes, thus leading to ILL (inter-Library Loan) services
    • Authority work nonexistent (except controlled vocabularies)
Surrogate Records in Periodical Databases
  • As is the case with library catalogs, periodical databases contain structured surrogate records. 
  • This structuring is fairly consistent across periodical databases, both in terms of stored records (two part metadata model holds) and how records are displayed
  • There is some authority control at work, but not in ways that you might think.
Collocation in Periodical Databases
  • By subject - what about vocabulary control?
  • By author - what about authority control?
  • By journal - what about authority control?
  • By language
  • By publication type
  • By date
  • Etc., etc., etc. 
In all Collocation Contexts: MATCH!
  • EXAMPLES:
    • Indexers → author name → match ← author name ← users
    • Indexers → journal name → match ← journal name ← users
    • Indexers → vocabulary → match ←vocabulary ← users
Inverted File Structures
  • How surrogate records are physically stored in the index of a database.
  • Each surrogate record has a unique identifies (also called a pointer)
  • Each word and phrase of the index has a record in the index; each record contains the UI for each surrogate record that contains that word or phrase:
    • Dog: 235, 527; 5,345,672; 117,127,923
    • Cat: 127; 2,753; 917,538; 327,543,238
How is Surrogate Information Stored?
  • Print periodical indexes and bibliographies. 
  • Online periodical databases:
  • ALWAYS KNOW THE START DATE OF YOUR ONLINE PERIODICAL DATABASE!
    • Manual search of the literature is often needed for exhaustive searches
    • Retrospective conversion of print indexes to online is not generally undertaken by database providers due to expense. 
Periodical Database Characteristics
  • They usually hold more records than a library catalog.
  • Information resources (i.e. journal articles) contain less information than the info resources in library catalogs and there are no detailed secondary navigation aids such as book indexes.
  • More fields (i.e. left side elements) available for search word qualification.
  • Important to distinguish database producing companies/organizations from database interface companies:
    • Some companies/organizations provide both (e.g. PubMed MEDLINE)
    • Other companies provide interface services, such as Dialog or EBSCO.
User Interfaces (UI) in Periodical Databases
  • Common UIs employed in periodical databases:
    • ISSN - uniquely numerical identification for individual serial publications
    • Internal numbering systems within a periodical database, such as the PMID in MEDLINE (for known item searches)
  • The most important UI is pre-Web and that is the "Address" of an article in the bibliographic universe (also known as the citation data) - also good for known item searches:
    • Journal name
    • Volume number (in some journals, the issue number also)
    • First page number of article. 
Authority Control in Periodical Databases
  • Titles - not controlled
  • Authors - somewhat controlled:
    • Indexers generally enter author name from the information resource
    • Control rests with periodical editors, who often have policies on author names that may be different than other periodical editors
  • Subjects - controlled:
    • Controlled vocabularies are imposed across periodical  and over time by indexers
    • However, subject searching is still subject to the problems associated with the "Great Pop vs. Soda Controversy"
Author Indexes in Periodical Databases
  • Examine how author data is entered into surrogate records:
    • Generally taken from the information resource in hand
    • PubMed MEDLINE is an exception
  • Some databases will provide lists of author names from which to choose:
    • Indicator of authority work?
      • Library Literature is an example
  • Subject Indexing in Periodical Databases
    • Indexing approach is "information access," therefore depth indexing is the general rule. 
    • Indexers index to the most specific, therefore, hierarchies remain important in controlled vocabularies. 
    • Pre-coordination and post-coordination are important concepts.
Subject Indexes in Periodical Databases
  • Depth Indexing:
    • General goal is to provide subject access to the information contained in an article.
    • However, this practice is not universal across database producers; therefore, determine the depth of indexing of the database you are searching by examining existing records
  • Be aware of the lag time related to subject indexing of articles
Management of Controlled Vocabularies
  • Homonymy:
    • Addressed by domain specificity of most periodical databases
    • Various forms of qualification are used, including parenthetical, hierarchy, and scope notes. 
  • New concepts:
    • Literary warrant is generally not employed in periodical database vocabularies
    • New terms are added after new concepts have established themselves in the literature
    • This poses a challenge if you are searching at a "research front" (may need to perform a keyword search strategy of the abstract field)
Some Vocabularies for Periodical Databases
Other Controlled Vocabulary Contexts
Web Content for Human Indexing
Indexing in Context
  1. Obtain information resource
  2. Describe information resource in surrogate record
  3. Subject analyze information resource in surrogate record:
    • Verbal 
    • Classification
Article Indexing Process
  • Two steps:
    • Analyze information resource to generate list of candidate concepts that describe its subject content
    • Translate those concepts into the controlled vocabulary of the database
  • ISO Standard for article indexing - special attention should be paid to certain sources of information:
    • Title
    • Abstract (when provided)
    • Introduction; opening and concluding paragraphs
    • Illustrations, diagrams, etc and their captions
    • Words or groups of words that are underlined, bold, etc.