Showing posts with label databases. Show all posts
Showing posts with label databases. Show all posts

Thursday, July 7, 2016

Verbal Subject Analysis III: Webpage Databases (a.k.a. "Search Engines")

Human vs. Automatic Indexing

  • Both are related to the subject analysis of information resources.
  • Human indexing is used to describe the subject analysis of various periodical databases.
  • Automatic indexing is a term used for the subject analysis operations by the computer algorithms of various webpage databases (a.k.a. search engines).
    • Research from the 60s-80s were trying to get a computer to calculate what articles were about. The most frequent words, articles like a, an, the, etc., don't really tell you much about the article, neither do the least used words. The key is finding the sweet spot based  on what the author usually writes about.
Why Webpage Database?
  • It is always important to know the documentary unit of an information database. 
  • The adjective associated with database is always a cue to the documentary unit. 
  • Webpage databases are informational databases in which a webpage is the documentary unit. 
  • They are also known as search engines and discovered databases.
Analysis of Websites and their Structure
  • What are webpages? What are websites? Webpages:Websites as pages:books
  • Standards (or lack thereof) for the authoring of web sites and webpages
    • HTML and other markup languages
    • Editors
  • What are the implications of the lack of authoring standards for web-based information resources?
Location of Webpage Subject Metadata
  • In webpage headers: For individual webpages, subject metadata can be created by authors and included in HTML headers.
  • In separate metadata record databases:
    • Subject metadata can be created by intermediaries using Dublin Core schema
    • In search engines, subject metadata is inferred "automatically" by computer algorithm.
Search Engine Questions
  • For greater understanding we need to be able to answer:
    • Why do search engines produce different results the exact same query?
    • What is the principle for ranking the display of search engine records in response to a query?
The Term "Search Engine"
  • The term has become the common designation for webpage databases, However, in actuality, webpage databases have three parts:
    • Spidering/crawling software to collect webpages.
    • Indexing software to build the index of surrogate records.
    • Retrieval software to facilitate retrieval of surrogates.
Automatic Indexing in Context
  1. Obtain information resource - spidering/crawling
    • Steps for spidering/crawling:
      • Computers owned by search engine retrieve documents by clicking on all hyperlinks on each retrieved webpage
      • Determination is made whether a webpage needs to be indexed (because it is new) or reindexed (if it has already been indexed)
      • Determination is made whether reindexing is warranted
      • New webpages and those meeting criteria for reindexing are then placed in the indexing queue
  2. Describe information resource in surrogate record - read off webpages by indexing software
    • Left Side elements must be inferred by searcher:
      • Examine structure of retrieved records
      • Examine advanced search interface
      • Element sets are not standard, i.e., they will vary across search engines.
    • Right Side Content:
      • What is the source for the content?
      • Authority control?
  3. Subject analyze information resource in surrogate record - indexing software:
    • Verbal - inferred by computer algorithm
    • Classification - inferred by computer algorithm
    • Subject Indexing in Search Engines
      • The subject fields of webpage surrogate records include the words that describe what the webpage is about. 
      • Right side subject content is inferred through the application of proprietary algorithms.
      • Subject terms added to surrogate records are weighted:
        • Doc #1: SU = dogs (.99); breeding (.87);dachshund (.30)
        • Doc #2: cats (.92); dogs(.44); dachshund (.03)
        • The weights are computed by proprietary algorithm.
Retrieval from Search Engines
  • Unlike bibliographic databases, in which the ordering of retrieved surrogate records is reverse chronological, search engines use a relevance-based ranking.
  • The search engine component of a search engine takes the entered query and compares it to the terms to the index.
  • The documents that are retrieved first are those that contain a higher "relevance" score:
    • Doc #1: SU = dogs (.99); breeding (.87);dachshund (.30)
    • Doc #2: cats (.92); dogs(.44); dachshund (.03)
    • "dog" query would rank document #1 ahead of document #2
    • "breeding" query would rank document #1 ahead of document #2
    • "cats"query would rank document #2 ahead of document #1
How are Subject Weights Calculated?
  • Conventional methods (Dating from the 1950s) for automatically inferring what a document is about include the following three techniques:
    • Frequency of word occurrences
    • Location of words occurrences
    • Size of word occurrences
  • In the web era, however, these techniques did not scale well to meet the needs of databases containing billions of records:
    • Could facilitate retrieval of relevant documents, but could not distinguish between "good" and "bad" documents.
    • Were also subject to manipulation by authors desiring higher search engine retrieval (spamming)
Two responses to Early Indexing Failure
  • Yahoo! era (late 1990's)
    • Human indexing (website directories)
    • More discussion during lectures on classification. 
  • Google era (since 1999)
    • Additional criteria introduced to infer aboutness, e.g.,;
      • $ - paid submissions, such as Alta Vista
      • Quality - PageRank algorithm of Google
Google Approach to Authomatic Indexing
  • Issue addressed by Google concerns the quality problem: How to cause the "best" documents to rise to the top of a set of retrieved webpages.
  • Solution concerns identifying additional criteria to include int he subject weighting algorithm.
  • Google maintains additional metadata elements for each surrogate record in its index of webpages:
    • How many other webpages link to a given webpage
      • The more webpages (i.e. linkers) a dachshund webpage has poiting to it, the more quality it has.
      • This factors into the weight assigned to the "dachshund" descriptor inthe subject field of its surrogate record
    • Who are the linkers
      • Those linkers that have a higher quality rank are given more weight than those linkers with a lower quality rank.

Tuesday, July 5, 2016

Verbal Subject Analysis II: Periodical and Other Databases

Subject Cataloging vs. Indexing
  • Both are related to the subject analysis of resources. 
  • Subject cataloging  is a term used for the subject analysis operations in library cataloging. 
  • Indexing is a term generally used for the subject analysis operations in various other resource organization contexts, including periodical databases and search engines. 
Brief History of Periodical Indexes
  • Around the turn of the 20th century, the library community decided not to add article citations to the catalog. 
  • This development led to the growth of the commercial indexing industry. 
  • The result of this has been:
    • Split files
    • Fees for licensing database content
    • Difficulty fulfilling Cutter's 2nd objective
Analytical Cataloging
  • Analytical cataloging techniques are needed in order to provide access to the component parts of composite information resources, most commonly:
    • Book chapters
    • Proceedings articles (usually of academic meetings)
    • Journal articles
  • Definition from AARC2: Analysis is the process of preparing a bibliographic record that describes a part (or parts) of an item for which a comprehensive entry is made.
Analytical Cataloging Techniques
  • Complex entries made within the record of composite work [cheap]:
    • Analytical added entries:
      • Use 740 tag for second of two works mentioned in title of item
    • Note area for comprehensive entry of larger work:
      • Use 505 tag for structured display of table of contents.
  • Separate records created for the component parts of composite works ("In" Analytics)[expensive]:
    • Use 773 to trace the component part record to parent record
Analytical Access to Journal Content
  • Decision to not provide analytical access to journal content (i.e. directly to articles) was because of the expense:
    • Excessive number of records would have to be created.
    • Additional authority work would need to be done.
  • As a result, through the 20th century, cataloging and periodical indexing/bibliography creation techniques evolved separate approaches. 
Overview Comparison
  • Catalog
    • Authority work
    • Cataloging records represents the holdings of a library
  • Periodical indexes:
    • Subject indexes are extensive topical bibliographies (often include books and book chapters, too), usually covering large swaths of "territory"
    • Domain-wide indexes (e.g. Index Medicus) attempt to capture an entire discipline (may include book chapters, too)
    • No single library could ever own all items referred to in exhaustive bibliographies/indexes, thus leading to ILL (inter-Library Loan) services
    • Authority work nonexistent (except controlled vocabularies)
Surrogate Records in Periodical Databases
  • As is the case with library catalogs, periodical databases contain structured surrogate records. 
  • This structuring is fairly consistent across periodical databases, both in terms of stored records (two part metadata model holds) and how records are displayed
  • There is some authority control at work, but not in ways that you might think.
Collocation in Periodical Databases
  • By subject - what about vocabulary control?
  • By author - what about authority control?
  • By journal - what about authority control?
  • By language
  • By publication type
  • By date
  • Etc., etc., etc. 
In all Collocation Contexts: MATCH!
  • EXAMPLES:
    • Indexers → author name → match ← author name ← users
    • Indexers → journal name → match ← journal name ← users
    • Indexers → vocabulary → match ←vocabulary ← users
Inverted File Structures
  • How surrogate records are physically stored in the index of a database.
  • Each surrogate record has a unique identifies (also called a pointer)
  • Each word and phrase of the index has a record in the index; each record contains the UI for each surrogate record that contains that word or phrase:
    • Dog: 235, 527; 5,345,672; 117,127,923
    • Cat: 127; 2,753; 917,538; 327,543,238
How is Surrogate Information Stored?
  • Print periodical indexes and bibliographies. 
  • Online periodical databases:
  • ALWAYS KNOW THE START DATE OF YOUR ONLINE PERIODICAL DATABASE!
    • Manual search of the literature is often needed for exhaustive searches
    • Retrospective conversion of print indexes to online is not generally undertaken by database providers due to expense. 
Periodical Database Characteristics
  • They usually hold more records than a library catalog.
  • Information resources (i.e. journal articles) contain less information than the info resources in library catalogs and there are no detailed secondary navigation aids such as book indexes.
  • More fields (i.e. left side elements) available for search word qualification.
  • Important to distinguish database producing companies/organizations from database interface companies:
    • Some companies/organizations provide both (e.g. PubMed MEDLINE)
    • Other companies provide interface services, such as Dialog or EBSCO.
User Interfaces (UI) in Periodical Databases
  • Common UIs employed in periodical databases:
    • ISSN - uniquely numerical identification for individual serial publications
    • Internal numbering systems within a periodical database, such as the PMID in MEDLINE (for known item searches)
  • The most important UI is pre-Web and that is the "Address" of an article in the bibliographic universe (also known as the citation data) - also good for known item searches:
    • Journal name
    • Volume number (in some journals, the issue number also)
    • First page number of article. 
Authority Control in Periodical Databases
  • Titles - not controlled
  • Authors - somewhat controlled:
    • Indexers generally enter author name from the information resource
    • Control rests with periodical editors, who often have policies on author names that may be different than other periodical editors
  • Subjects - controlled:
    • Controlled vocabularies are imposed across periodical  and over time by indexers
    • However, subject searching is still subject to the problems associated with the "Great Pop vs. Soda Controversy"
Author Indexes in Periodical Databases
  • Examine how author data is entered into surrogate records:
    • Generally taken from the information resource in hand
    • PubMed MEDLINE is an exception
  • Some databases will provide lists of author names from which to choose:
    • Indicator of authority work?
      • Library Literature is an example
  • Subject Indexing in Periodical Databases
    • Indexing approach is "information access," therefore depth indexing is the general rule. 
    • Indexers index to the most specific, therefore, hierarchies remain important in controlled vocabularies. 
    • Pre-coordination and post-coordination are important concepts.
Subject Indexes in Periodical Databases
  • Depth Indexing:
    • General goal is to provide subject access to the information contained in an article.
    • However, this practice is not universal across database producers; therefore, determine the depth of indexing of the database you are searching by examining existing records
  • Be aware of the lag time related to subject indexing of articles
Management of Controlled Vocabularies
  • Homonymy:
    • Addressed by domain specificity of most periodical databases
    • Various forms of qualification are used, including parenthetical, hierarchy, and scope notes. 
  • New concepts:
    • Literary warrant is generally not employed in periodical database vocabularies
    • New terms are added after new concepts have established themselves in the literature
    • This poses a challenge if you are searching at a "research front" (may need to perform a keyword search strategy of the abstract field)
Some Vocabularies for Periodical Databases
Other Controlled Vocabulary Contexts
Web Content for Human Indexing
Indexing in Context
  1. Obtain information resource
  2. Describe information resource in surrogate record
  3. Subject analyze information resource in surrogate record:
    • Verbal 
    • Classification
Article Indexing Process
  • Two steps:
    • Analyze information resource to generate list of candidate concepts that describe its subject content
    • Translate those concepts into the controlled vocabulary of the database
  • ISO Standard for article indexing - special attention should be paid to certain sources of information:
    • Title
    • Abstract (when provided)
    • Introduction; opening and concluding paragraphs
    • Illustrations, diagrams, etc and their captions
    • Words or groups of words that are underlined, bold, etc.

Thursday, June 11, 2015

Databases

What is the difference between data and information? 

  • Data - random content, no context
  • Information - data within a context, data with MEANING
Database Anatomy
  • Table
    • Columns - topics/headers/descriptors of data
      • Field
    • Rows - individual instances of data
      • Records
Primary key - unique identifies (each record must be unique)
Foreign Key - relating information in other table

Table Relationships

One to oneBoth tables can have only one record on either side of the relationship. Each primary key value relates to only one (or no) record in the related table. Most one-to-one relationships are forced by business rules and don't flow naturally from the data. In the absence of such a rule, you can usually combine both tables into one table without breaking any normalization rules.
       -e.g. Spouses are 1 to 1 records. You can only have 1 spouse.



One to manyThe primary key table contains only one record that relates to none, one, or many records in the related table.
       -e.g. Parent-child relationships are one-to -many, as parents can have many children

Many to many - Each record in both tables can relate to any number of records (or no records) in the other table. Many-to-many relationships require a third table, known as a bridge table, because relational systems can't directly accommodate the relationship.
         -e.g. Students can take many courses, and courses can have many students

Integrity Rules

Entity Integrity - Each entity has a unique key

Referential Integrity - Foreign key value is null or matches primary key values in related table

Database Efficiency

Normalization  - reduces repetitive entries an makes the design and structure of the database as efficient as possible. There are multiple levels to this normalization (1 NF, 2 NF, 3 NF)

1 NF 
  • Eliminates duplicative columns from the same table.
  • Identifies each row with a unique column (the primary key
    • Steps
      • Eliminate repeating groups
      • Identify primary key
      • Identify all dependencies

2 NF
  • It is in 1 NF
  • There are no partial dependencies
    • Steps
      • Start with 1 NF formula
      • Write each key component w/ partial dependency) on separate line
      • Write original (composite) key on last line
      • Each component is new table
      • Write dependent attributes after each key

3 NF

  • It is in 2 NF
  • There are no transitive dependencies
    • Steps
      • Start with 2NF format
      • Break off the TP pieces and create separate tables