Showing posts with label internet. Show all posts
Showing posts with label internet. Show all posts

Friday, September 11, 2009

Week 3: Digital Object Identifiers and Their Implications


The discussion of digital objects started off harmlessly enough.  Computer language must be converted into human language - characters (i.e. A, a, 4, $) are created using strings of 1's and 0's.   Digital documents are formatted for structure using a family of markup languages (SGML) and for format using HTML (a combination of structure and format capability) and style sheets.

Computers must be able to understand each other as they communicate over networks, so a system of identifiers such as URNs and URLs are put in place so that computers can have a universal method of keeping track of digital objects and their locations.  The problem with URLs is that they are not stable identifiers. 

The Webmaster may change or remove content on a page without warning.   If another person uses a URL to cite linkages between their work sand someone else's work, there is no guarantee that linkage of information will still be valid 6 months down the road.

A reference to a lawsuit between Microsoft and Ticketmaster in Paskin's article on Digital Object Identifiers caught my attention so I did a basic search to find out more.   Apparently, Ticketmaster sued Microsoft because Microsoft provided its users with a "deep link" to Ticketmaster's ticket purchasing page which allowed users to bypass the usual onslaught of digital advertising.

As the internet provides a creative means to generate commerce, arguments over intellectual property and "unfair competition" are bound to grow.   The intrinsic value of the Internet is in the power it gives the user to access information instantly and to follow linkages between documents or web pages simply by clicking on hypertext. 

As the concept of Digital Object Identifiers grows as a means to uniquely identify bodies of information on the internet, so does the need to evaluate the balance between the ownership of intellectual property and universal access.   Digital Object Identifiers create a standardized way of identifying information (like a serial number) and solve some problems created by the instability of URLs.

Digital Object Identifiers (DOIs) provide a way to identify bodies of information that is neither too specific (for example, an ISBN which denotes a hardcover or trade paper version of the same book) nor too generalized to have meaning.   Paskin talks about some of the benefits of DOIs. DOIs provide resolution by tying a digital object to a specific name.   They have the potential to support interoperability.  If separate digital libraries adopt the same naming rules and standards for digital objects, those objects can be identified across digital libraries.   They also supply persistence - a deleted page will no longer signify a broken link between two bodies of information as some central agency can help us keep track of our digital objects and their changing locations. DOI. 

The implications for this system are frightening as well as dazzling.  Paskin identifies some monumental concerns in his article about DOIs.   What happens when a user can no longer freely navigate between pages?  If the central agency has control of the DOI, and by implication, the means to access the content the DOI identifies, it has the ability to restrict access to content by withholding that DOI.   How much harder will it be to access content?  How much time will it take?  Does the central agency that manages the DOI database have the right to know what information we are accessing? 

As I read more I had questions of my own. To what extent will this system be used as a way to generate commerce from materials that we now have access to for free?  How does this affect research?  The ability to follow citation via hypertext when accessing research publications is pivotal to research.  If some third party regulates this access, will research suffer?   This has a huge impact on our society.

Lynch, in his article concludes with "In a very real sense, there are no bad identifiers, but it is very possible to put identifiers to bad or inappropriate uses."   I agree with this statement.  The ability to manage information efficiently so that it can be accessed easily is a a huge benefit to those who access the information.  But what is the cost?

  1. LESK sections 2.1, 2.2, 2.7, chapter 3.
  2. ARMS. Chapters 9. http://www.cs.cornell.edu/wya/DigLib/MS1999/Chapter9.html.
  3. Clifford Lynch, “Identifiers and Their Role In Networked Information Applications”. http://www.arl.org/bm~doc/identifier.pdf
  4. Norman Paskin. “Digital Object Indentifier (DOI) System”. Encyclopedia of Library and Information Sciences. http://www.doi.org/overview/080625DOI-ELIS-Paskin.pdf
Background Readings:
  1. Sam Sun, Larry Lannom, and Brian Boesch. "Handle System Overview", http://www.handle.net/rfc/rfc3650.html.    
     6.  Netlitigation.  "Ticketmaster v. Microsoft, United States District Court for the Central District of California, Civil Action Number 97-3055DPP",  http://www.netlitigation.com/netlitigation/cases/ticketmaster.htm

Friday, September 4, 2009

Week 2: Notes from Works About Interoperability, Extensibility, and the World Wide Web

In addition to establishing an accepted definition for digital libraries, there is a movement within the field to establish a protocol or set of standards for building the components of digital libraries.

Having a standard protocol for digital library development allows for both interoperability (the ability to use the same components with the same functionality across different interfaces) and extensibility (allowing for future development of the components without compromising functionality).  In order to develop a protocol, one must define and standardize the components of digital libraries.

Some of the components defined in these articles are:
  • repositories where digital objects are stored 
  • handles or identifiers for digital objects
  • disseminators or processes associated with digital objects that provide increased functionality
  • servlets or programs containing a collection of operations and associated with a specific disseminator types
Interoperability and extensibility allow developments and information to flow across multiple sources (as information flows across servers in the World Wide Web) and allows them to be dynamic and open to improving functionality as technology improves.  An established set of protocols for building components of digital libraries would allow those components to become mobile and would allow creators of digital libraries to customize a set of tools already established (instead of re-creating the wheel with each new library).  This is not a new concept - programming languages (such as java) use an established set of tools to build and customize sophisticated programs which can be used across many interfaces.

The Digital Library Research Group at Cornell University and the Corporation for National Research Initiatives (CNRI) have been working together to establish infrastructures that provide real-life examples of interoperability and extensibility through open architecture.     They tested the interoperability of their product across repositories, disseminators, and the behaviors of their digital objects.

Their tests show that:
  • Interoperability across repositories can be established using RAP (Repository Access Protocol)  interface
  • The same tasks can be accomplished in separate repositories with the same outcome
  • Extensibility can be established, which allows community standards to be integrated into existing digital libraries
Future work of the Cornell/CNRI group will focus on:
  • Allowing more variation between repositories without compromising interoperability
  • Creating a measurement system for interoperability
  • More testing for interoperability
  • Issues of security
William Arms discusses the history and components of the World Wide Web in his chapter, "Overview of the World Wide Web".   This chapter applies to the development of digital libraries because the web can serve as an example of a well-established network of information with both  interoperability and extensibility.  Like the architectural concepts of digital libraries described by Arms et al., the World Wide Web is supported by a set of well-defined components which are functional across systems and allow information to flow freely across networks.  Arms would argue that the world wide web is under-credited (because of it's perceived exclusion from the library world) as being significant to digital library development.


References:
1.Hussein Suleman and Edward A. Fox. “A Framework for Building Open Digital Libraries”, D-Lib Magazine, December 2001. Volume 7 Number 12. http://www.dlib.org/dlib/december01/suleman/12suleman.html.
2.William Y. Arms , Christophe Blanchi, Edward A. Overly. ”An Architecture for Information in Digital Libraries”. D-Lib Magazine, February 1997. http://www.dlib.org/dlib/february97/cnri/02arms1.html.
3.Sandra Payette, Christophe Blanchi, Carl Lagoze, Edward A. Overly. “Interoperability for Digital Objects and Repositories, The Cornell/CNRI Experiments”, D-Lib Magazine, May 1999, Volume 5 Issue 5. http://www.dlib.org/dlib/may99/payette/05payette.html.
4.ARMS, Chapter 2, or “An Overview of World Wide Web”, http://www.cio.com/WebMaster/sem2_home.html