Thursday, September 17, 2009

Muddiest Point (9/14/09)

My question concerns our final project.   How will copyright affect our digital library collections?   Can we use media (text, images, video) produced by others if we give them credit for their work? 

Sunday, September 13, 2009

Week 1: Defining Digital Libraries

In his work, Structure of Scientific Revolutions, Thomas Kuhn states that a science begins with the definition of problems and methods for future generations of scientists.   These definitions are further articulated as those future scientists study them and add precision.  

The DELOS Manifesto takes on the task of setting definitions and setting the foundations for future research.   It proposes models and possible objectives for research, functionality, quality, architecture, and policy.

William Arms talks about how the field of Library and Information Sciences is impacted by improvements in automation.   What will the future of librarians be as advancements in digital libraries replace the traditional roles of library science professionals?

Paepcke suggests that librarians' jobs will become more specialized, helping researchers interact with technology during the production process and by adding value to materials that computer scientists put online. Arms says that librarians can use their skill with abstract ideas and the idiosyncrasies of information mining - skills that web crawlers and automated search engines have not yet mastered.

I wonder what place the automated future will carve out for those trained in library and information sciences.   What role do they have to play in developing these systems that will eventually replace many of the responsibilities that librarians have always held?  Will their skill set become more technological?  As technology and commerce intersect with digital materials, will their focus be more rooted in philosophy of and the advocacy for access?  Or will they become highly specialized computer scientists?

I think librarians have much to offer in the development of digital information systems (and visa versa).  Now, more than ever, it's important for library science to keep in step with technology and as it does so, carve out its own place in a rapidly changing environment.   Maybe in doing so, library science can create  its own manifesto, declaring ways for the field to use its philosophical foundations and collective skill set to improve the way humans interact with information and technology.



  1.  Christine L. Borgman, “From Gutenberg to the Global Information Infrastructure: Access to Information in the Networked World”. MIT Press, 2001.
  2. Leonardo Candela et. al. (2007) Setting the Foundations of Digital Libraries.  D-Lib Magazine 13(3-4), March/April 2007. 
  3. Andreas Paepcke, Hector Garcia-Molina, Rebecca Wesley, “Dewey Meets Turing: Librarians, Computer Scientists, and the Digital Libraries Initiative” D-Lib Magazine, Volume 11 Number 7/8, July/August 2005.
  4. Christian Lupovici. (2008) The growth of the role of librarians and information officers in digital libraries. Digital Libraries, Fabrice Papy (eds). ISTE and John Wiley & Sons, Inc.
  1. William Y. Arms. “Automated Digital Libraries, How Effectively Can Computers Be Used for the Skilled Tasks of Professional Librarianship?” D-Lib Magazine July/August 2000. 6(7/8).  

Friday, September 11, 2009

The New Librarian

Image by: David Coverly
Accessed from: http://www.speedbump.com/librarian.html

Week 3: Digital Object Identifiers and Their Implications


The discussion of digital objects started off harmlessly enough.  Computer language must be converted into human language - characters (i.e. A, a, 4, $) are created using strings of 1's and 0's.   Digital documents are formatted for structure using a family of markup languages (SGML) and for format using HTML (a combination of structure and format capability) and style sheets.

Computers must be able to understand each other as they communicate over networks, so a system of identifiers such as URNs and URLs are put in place so that computers can have a universal method of keeping track of digital objects and their locations.  The problem with URLs is that they are not stable identifiers. 

The Webmaster may change or remove content on a page without warning.   If another person uses a URL to cite linkages between their work sand someone else's work, there is no guarantee that linkage of information will still be valid 6 months down the road.

A reference to a lawsuit between Microsoft and Ticketmaster in Paskin's article on Digital Object Identifiers caught my attention so I did a basic search to find out more.   Apparently, Ticketmaster sued Microsoft because Microsoft provided its users with a "deep link" to Ticketmaster's ticket purchasing page which allowed users to bypass the usual onslaught of digital advertising.

As the internet provides a creative means to generate commerce, arguments over intellectual property and "unfair competition" are bound to grow.   The intrinsic value of the Internet is in the power it gives the user to access information instantly and to follow linkages between documents or web pages simply by clicking on hypertext. 

As the concept of Digital Object Identifiers grows as a means to uniquely identify bodies of information on the internet, so does the need to evaluate the balance between the ownership of intellectual property and universal access.   Digital Object Identifiers create a standardized way of identifying information (like a serial number) and solve some problems created by the instability of URLs.

Digital Object Identifiers (DOIs) provide a way to identify bodies of information that is neither too specific (for example, an ISBN which denotes a hardcover or trade paper version of the same book) nor too generalized to have meaning.   Paskin talks about some of the benefits of DOIs. DOIs provide resolution by tying a digital object to a specific name.   They have the potential to support interoperability.  If separate digital libraries adopt the same naming rules and standards for digital objects, those objects can be identified across digital libraries.   They also supply persistence - a deleted page will no longer signify a broken link between two bodies of information as some central agency can help us keep track of our digital objects and their changing locations. DOI. 

The implications for this system are frightening as well as dazzling.  Paskin identifies some monumental concerns in his article about DOIs.   What happens when a user can no longer freely navigate between pages?  If the central agency has control of the DOI, and by implication, the means to access the content the DOI identifies, it has the ability to restrict access to content by withholding that DOI.   How much harder will it be to access content?  How much time will it take?  Does the central agency that manages the DOI database have the right to know what information we are accessing? 

As I read more I had questions of my own. To what extent will this system be used as a way to generate commerce from materials that we now have access to for free?  How does this affect research?  The ability to follow citation via hypertext when accessing research publications is pivotal to research.  If some third party regulates this access, will research suffer?   This has a huge impact on our society.

Lynch, in his article concludes with "In a very real sense, there are no bad identifiers, but it is very possible to put identifiers to bad or inappropriate uses."   I agree with this statement.  The ability to manage information efficiently so that it can be accessed easily is a a huge benefit to those who access the information.  But what is the cost?

  1. LESK sections 2.1, 2.2, 2.7, chapter 3.
  2. ARMS. Chapters 9. http://www.cs.cornell.edu/wya/DigLib/MS1999/Chapter9.html.
  3. Clifford Lynch, “Identifiers and Their Role In Networked Information Applications”. http://www.arl.org/bm~doc/identifier.pdf
  4. Norman Paskin. “Digital Object Indentifier (DOI) System”. Encyclopedia of Library and Information Sciences. http://www.doi.org/overview/080625DOI-ELIS-Paskin.pdf
Background Readings:
  1. Sam Sun, Larry Lannom, and Brian Boesch. "Handle System Overview", http://www.handle.net/rfc/rfc3650.html.    
     6.  Netlitigation.  "Ticketmaster v. Microsoft, United States District Court for the Central District of California, Civil Action Number 97-3055DPP",  http://www.netlitigation.com/netlitigation/cases/ticketmaster.htm

Friday, September 4, 2009

Week 2: Notes from Works About Interoperability, Extensibility, and the World Wide Web

In addition to establishing an accepted definition for digital libraries, there is a movement within the field to establish a protocol or set of standards for building the components of digital libraries.

Having a standard protocol for digital library development allows for both interoperability (the ability to use the same components with the same functionality across different interfaces) and extensibility (allowing for future development of the components without compromising functionality).  In order to develop a protocol, one must define and standardize the components of digital libraries.

Some of the components defined in these articles are:
  • repositories where digital objects are stored 
  • handles or identifiers for digital objects
  • disseminators or processes associated with digital objects that provide increased functionality
  • servlets or programs containing a collection of operations and associated with a specific disseminator types
Interoperability and extensibility allow developments and information to flow across multiple sources (as information flows across servers in the World Wide Web) and allows them to be dynamic and open to improving functionality as technology improves.  An established set of protocols for building components of digital libraries would allow those components to become mobile and would allow creators of digital libraries to customize a set of tools already established (instead of re-creating the wheel with each new library).  This is not a new concept - programming languages (such as java) use an established set of tools to build and customize sophisticated programs which can be used across many interfaces.

The Digital Library Research Group at Cornell University and the Corporation for National Research Initiatives (CNRI) have been working together to establish infrastructures that provide real-life examples of interoperability and extensibility through open architecture.     They tested the interoperability of their product across repositories, disseminators, and the behaviors of their digital objects.

Their tests show that:
  • Interoperability across repositories can be established using RAP (Repository Access Protocol)  interface
  • The same tasks can be accomplished in separate repositories with the same outcome
  • Extensibility can be established, which allows community standards to be integrated into existing digital libraries
Future work of the Cornell/CNRI group will focus on:
  • Allowing more variation between repositories without compromising interoperability
  • Creating a measurement system for interoperability
  • More testing for interoperability
  • Issues of security
William Arms discusses the history and components of the World Wide Web in his chapter, "Overview of the World Wide Web".   This chapter applies to the development of digital libraries because the web can serve as an example of a well-established network of information with both  interoperability and extensibility.  Like the architectural concepts of digital libraries described by Arms et al., the World Wide Web is supported by a set of well-defined components which are functional across systems and allow information to flow freely across networks.  Arms would argue that the world wide web is under-credited (because of it's perceived exclusion from the library world) as being significant to digital library development.


References:
1.Hussein Suleman and Edward A. Fox. “A Framework for Building Open Digital Libraries”, D-Lib Magazine, December 2001. Volume 7 Number 12. http://www.dlib.org/dlib/december01/suleman/12suleman.html.
2.William Y. Arms , Christophe Blanchi, Edward A. Overly. ”An Architecture for Information in Digital Libraries”. D-Lib Magazine, February 1997. http://www.dlib.org/dlib/february97/cnri/02arms1.html.
3.Sandra Payette, Christophe Blanchi, Carl Lagoze, Edward A. Overly. “Interoperability for Digital Objects and Repositories, The Cornell/CNRI Experiments”, D-Lib Magazine, May 1999, Volume 5 Issue 5. http://www.dlib.org/dlib/may99/payette/05payette.html.
4.ARMS, Chapter 2, or “An Overview of World Wide Web”, http://www.cio.com/WebMaster/sem2_home.html

Tuesday, September 1, 2009

Muddiest Point (8/31/09)

This is more of a request for more information than clarification:

I had a chance to look at the "Visualize the Collections" or "Bungee View" prototype for the Historic  Pittsburgh Image Collections website.   Are there any other visualization tools like this one that we can look at on the web?

Which aspects of Bungee Viewer are representative of emerging technology? Is it in the components or the way the capabilities are bundled into one package?

The Blog

Welcome to my blog!  All postings will be relevant to the Digital Libraries course at the University of Pittsburgh, Fall 2009.   My posts will include reflections, questions, and resources relevant to the world of digital libraries.  Enjoy!