Friday, December 4, 2009
Sunday, November 29, 2009
Reading Notes: Security and Economics
This week's readings encompassed the ideas of security, encryption, authentication, and methods of managing differing users' levels of access. Libraries are able to manage user profiles and set pre-determined rights for access. Public and private key encryption allows trusted allies and new business partners to exchange protected information over the web. One of the most obvious reasons for limiting rights for access are the economic pressures which limit the number of copies a library can carry of a certain item and requests by the rights holders to withhold access, generally to protect the right holders' financial self-interests. Journals such as D-Lib Magazine have provided models for open-access literature. Unrefereed journals are able to provide increased access but until more top universities develop tenure systems more friendly towards these economic alternatives (and it seems they are beginning to do just this), things will probably not change greatly. While advertisements on the web allow information producers to provide users with free access, this is not a model that libraries are likely to adopt. Some companies have found other methods (besides technological security methods) of encouraging users to invest in their products by value-added services such as customer support.
ARMS, chapter 7
William Arms, “Implementing Policies for Access Management”, D-Lib Magazine,1998. http://www.dlib.org/dlib/february98/arms/02arms.html.
LESK, chapter 10 “economics” lesk-ch9-economics.pdf
ARMS, chapter 6, economics part http://www.cs.cornell.edu/wya/DigLib/new/Chapter6.html
READING QUESTIONS/COMMENTS:
-I am interested in more updated information about encryption and authentication technology. How have these technologies changed in the last ten years.
-Are their digital libraries that provide information which must be protected for reasons other than financial incentives for rights holders? In other words, do digital libraries ever deal with information that is available only to a limited number of people?
-How are medical informatics systems like digital libraries?
-Can you talk about some reason economic trends in open access?
ARMS, chapter 7
William Arms, “Implementing Policies for Access Management”, D-Lib Magazine,1998. http://www.dlib.org/dlib/february98/arms/02arms.html.
LESK, chapter 10 “economics” lesk-ch9-economics.pdf
ARMS, chapter 6, economics part http://www.cs.cornell.edu/wya/DigLib/new/Chapter6.html
READING QUESTIONS/COMMENTS:
-I am interested in more updated information about encryption and authentication technology. How have these technologies changed in the last ten years.
-Are their digital libraries that provide information which must be protected for reasons other than financial incentives for rights holders? In other words, do digital libraries ever deal with information that is available only to a limited number of people?
-How are medical informatics systems like digital libraries?
-Can you talk about some reason economic trends in open access?
Friday, November 20, 2009
Muddiest Point (11/16/09)
Before we leave the class, I would like some information about some skills we should develop if we are interested in working with digital libraries once we finish this program. For example, what programming languages might we learn? Are there currently any DL projects in the area that we can volunteer to help with? What other skills might be useful to us as we enter the field?
I am particularly interested in any oppotunities there may be for someone in the medical librarianship track.
I am particularly interested in any oppotunities there may be for someone in the medical librarianship track.
Reading Notes: Social Issues in Digital Libraries
This week's readings were reflective on the past and future of digital libraries.
In Social Aspects of the Digital Library (UCLA-NSF Social Aspects of Digital Libraries Workshop
Research issues as
human-centered (human computer interactions, how humans interact with information)
artifact-centered (focused on the artifacts of human communication)
systems-centered (focus on the systems that support the products of the first two research areas)
Issues Raised:
lack of consensus on terminology
involving users in the development process / systematically determining user needs
development of an information life cycle to represent the flow of information
The infinite library: does Google's plan to digitize millions of print books spell the death of libraries--or their rebirth?
Discusses the role of libraries in the Google books projects. Some larger libraries are participating fully (Harvard) while others (Bodleian) only allow Google access to items already in the public domain.
Digitalization is a long and complicated process and will take a vast number of resources to digitize the holdings of the libraries that have opened their stacks.
Raises the issue of what roles libraries will play if/when a vast number of books become searchable online. Only snippets of the books will be available so libraries will still function to provide free full access to books. Some libraries vow to allow users access even if Google decides to charge. Discusses the reluctance of the library community to embrace tools provided by the internet.
A Viewpoint Analysis of the Digital Library
Discussing the future of DL development.
Organizational Develpment - Identifies how different major library systems can be integrated (i.e. with the Library of Congress as the center of the library world).
Technical Development - discusses interoperability research and technical standards that allow collaboration between libraries
User Development - discusses the user as the center of the library. Discusses the ideal of one seamless interface for multiple library systems, identifies a incoherence (as a result of interoperable systems) behind the current interfaces of federated systems.
Research Agenda:
Promotes the idea of evaluation as a core area of future research. Underscores the ideas of user-friendly seamless systems. Suggests the maturation of internet as a model for the possible progression of digital library development.
References:
Social Aspects of Digital Libraries. The final report of UCLA-NSF Social Aspects of Digital Libraries Workshop. http://is.gseis.ucla.edu/research/dig_libraries/index.html
The Infinite Library, Wade Roush, Technology Review, 2005. http://www.technologyreview.com/articles/05/05/issue/feature_library.asp
William Y. Arms, “A Viewpoint Analysis of the Digital Library”, D-Lib Magazine, Volume 11 Number 7/8, July/August 2005. http://www.dlib.org/dlib/july05/arms/07arms.html
Questions/Comments:
I would like to see some examples of digital collections that embody the most current research in the field.
In Social Aspects of the Digital Library (UCLA-NSF Social Aspects of Digital Libraries Workshop
Research issues as
human-centered (human computer interactions, how humans interact with information)
artifact-centered (focused on the artifacts of human communication)
systems-centered (focus on the systems that support the products of the first two research areas)
Issues Raised:
lack of consensus on terminology
involving users in the development process / systematically determining user needs
development of an information life cycle to represent the flow of information
The infinite library: does Google's plan to digitize millions of print books spell the death of libraries--or their rebirth?
Discusses the role of libraries in the Google books projects. Some larger libraries are participating fully (Harvard) while others (Bodleian) only allow Google access to items already in the public domain.
Digitalization is a long and complicated process and will take a vast number of resources to digitize the holdings of the libraries that have opened their stacks.
Raises the issue of what roles libraries will play if/when a vast number of books become searchable online. Only snippets of the books will be available so libraries will still function to provide free full access to books. Some libraries vow to allow users access even if Google decides to charge. Discusses the reluctance of the library community to embrace tools provided by the internet.
A Viewpoint Analysis of the Digital Library
Discussing the future of DL development.
Organizational Develpment - Identifies how different major library systems can be integrated (i.e. with the Library of Congress as the center of the library world).
Technical Development - discusses interoperability research and technical standards that allow collaboration between libraries
User Development - discusses the user as the center of the library. Discusses the ideal of one seamless interface for multiple library systems, identifies a incoherence (as a result of interoperable systems) behind the current interfaces of federated systems.
Research Agenda:
Promotes the idea of evaluation as a core area of future research. Underscores the ideas of user-friendly seamless systems. Suggests the maturation of internet as a model for the possible progression of digital library development.
References:
Social Aspects of Digital Libraries. The final report of UCLA-NSF Social Aspects of Digital Libraries Workshop. http://is.gseis.ucla.edu/research/dig_libraries/index.html
The Infinite Library, Wade Roush, Technology Review, 2005. http://www.technologyreview.com/articles/05/05/issue/feature_library.asp
William Y. Arms, “A Viewpoint Analysis of the Digital Library”, D-Lib Magazine, Volume 11 Number 7/8, July/August 2005. http://www.dlib.org/dlib/july05/arms/07arms.html
Questions/Comments:
I would like to see some examples of digital collections that embody the most current research in the field.
Friday, November 13, 2009
Reading Notes: Interaction and Evaluation
Evaluation involves considering the user when designing the interface. Without this consideration, the user may dismiss a powerful and useful system as too complex, irrelevant, or inconvenient. One such example of this is in medical informatics where systems are often ignored by practitioners who find them unmatched with their needs.
There is not a lot of interest in funding this field of study, although its principles and aims are structured and well-defined and it has applications in many types of systems: digital libraries, games, websites, internet applications, critical systems, and other programs in industrial, commercial systems, and personal systems.
In terms of Digital Libraries...
Usability goals include:
-the ability of users to access, control, and maintain the system
-with minimal personnel and training
-with reliable and functional interaction between hardware and software
-meeting standards for interoperability
-work is completed on time and according to budget
User Consideration:
-physical comfort, accessibility
-possible barriers to access (time, budget, environment, etc.)
-age
-ability
-expectations
-context of use
-personality type
-cultural and linguistic factors, etc.
Criteria for evaluating usability should include (Saracevic):
content
process
format
overall system (maintenence, scalability, interoperability, sharability, costs)
Studies have also considered usage patterns, use of materials, usage statistics, types of users, time factors, and intent (Saracevic)
According to Saracevic, ongoing study should also include how digital libraries ultimately affect the environment in which they are used.
Rob Kling and Margaret Elliott "Digital Library Design for Usability" http://www.csdl.tamu.edu/DL94/paper/kling.html
Tefko Saracevic, “Evaluation of digital libraries: An overview” http://www.scils.rutgers.edu/~tefko/DL_evaluation_Delos.pdf.
Ben Sheiderman, Catherine Plaisant, "Designing the user interfaces" 4ed. chapter 1. A good introduction about usability and its application in human computer interaction hci4ed-ch1.pdf
READING QUESTIONS:
- Can you show us some digital libraries that are unusual in content, context, or user group?
- Can you show us some digital libraries with advanced features, not normally utilized by the general population?
-Are there examples of DLs pertinent to medical informatics? Can you show us examples?
-What sort of features are available but aren't used due to incompatibility with users and/or the environment of their use?
-What types of jobs are available for people who study HCI?
There is not a lot of interest in funding this field of study, although its principles and aims are structured and well-defined and it has applications in many types of systems: digital libraries, games, websites, internet applications, critical systems, and other programs in industrial, commercial systems, and personal systems.
In terms of Digital Libraries...
Usability goals include:
-the ability of users to access, control, and maintain the system
-with minimal personnel and training
-with reliable and functional interaction between hardware and software
-meeting standards for interoperability
-work is completed on time and according to budget
User Consideration:
-physical comfort, accessibility
-possible barriers to access (time, budget, environment, etc.)
-age
-ability
-expectations
-context of use
-personality type
-cultural and linguistic factors, etc.
Criteria for evaluating usability should include (Saracevic):
content
process
format
overall system (maintenence, scalability, interoperability, sharability, costs)
Studies have also considered usage patterns, use of materials, usage statistics, types of users, time factors, and intent (Saracevic)
According to Saracevic, ongoing study should also include how digital libraries ultimately affect the environment in which they are used.
READING QUESTIONS:
- Can you show us some digital libraries that are unusual in content, context, or user group?
- Can you show us some digital libraries with advanced features, not normally utilized by the general population?
-Are there examples of DLs pertinent to medical informatics? Can you show us examples?
-What sort of features are available but aren't used due to incompatibility with users and/or the environment of their use?
-What types of jobs are available for people who study HCI?
Friday, October 23, 2009
Muddest Point (10/16)
This is not exactly a request for clarification but positive feedback instead. I enjoyed the XML Hand on Points and assignment. I appreciated the opportunities to practice writing XML and DTDs. I thought they were very useful and practical activities. I look forward to completing more of these practical HOPs and skill development exercises in the future.
Reading Notes: Access in Digital Libraries II
The information community appears divided on the utility of federated searching. Federated searching appeals to the typical user because it is simple, doesn't involve a lot of effort, and provides the familiar comfort of a plain text box and natural language for search queries.
Federate searching is a more user-friendly approach but does currently can not support the powerful searching of an index content-specific search. Hane points out that that federated searches often lacking in advertised coverage (due to authentication problems, for one), removing duplicates, determining relevance and ranking in a way meaningful to the query, and can run into trouble keeping up with updates of various databases they cover.
Various standards have been introduced to try to improve federating searching. Z39.50 was introduced as a protocol for information retrieval which identified structures and rules for interchange. In 1997, attempts were being made to develop linkages with other standards and improving interoperability (specifically the ability to translate queries across systems).
Metadata harvesting offers possibilities in interoperability and increased dissemination of information by providing rules and a framework for sharing descriptive data (OAI-PMH). Reliance on user-created metadata can sometimes be problematic because when humans describe their own materials they often do so subjectively and without much consideration for the use of that metadata by outside sources and the possibility of lost context. Aggregated metadata increases efficiency, uniformity, speed, performance, analytical power, and meaningfulness in organization (OAI-PMH).
As the search tools offered by the library move further from traditional and into larger and domains with varied setups, development must focus on the ability to collage information across collections. The introduction of standards is helpful but depends on users to implement these standards with other uses in mind. Whether or not developers are willing to do this is another matter.
Questions:
Is Z39.50 still in use/development?
Can you discuss some of the current trends/developments metadata harvesting?
What other strategies are being used in implemented in modern federated searching?
Are the same federated searching roadblocks discussed in the Info Today article still valid?
Federate searching is a more user-friendly approach but does currently can not support the powerful searching of an index content-specific search. Hane points out that that federated searches often lacking in advertised coverage (due to authentication problems, for one), removing duplicates, determining relevance and ranking in a way meaningful to the query, and can run into trouble keeping up with updates of various databases they cover.
Various standards have been introduced to try to improve federating searching. Z39.50 was introduced as a protocol for information retrieval which identified structures and rules for interchange. In 1997, attempts were being made to develop linkages with other standards and improving interoperability (specifically the ability to translate queries across systems).
Metadata harvesting offers possibilities in interoperability and increased dissemination of information by providing rules and a framework for sharing descriptive data (OAI-PMH). Reliance on user-created metadata can sometimes be problematic because when humans describe their own materials they often do so subjectively and without much consideration for the use of that metadata by outside sources and the possibility of lost context. Aggregated metadata increases efficiency, uniformity, speed, performance, analytical power, and meaningfulness in organization (OAI-PMH).
As the search tools offered by the library move further from traditional and into larger and domains with varied setups, development must focus on the ability to collage information across collections. The introduction of standards is helpful but depends on users to implement these standards with other uses in mind. Whether or not developers are willing to do this is another matter.
Questions:
Is Z39.50 still in use/development?
Can you discuss some of the current trends/developments metadata harvesting?
What other strategies are being used in implemented in modern federated searching?
Are the same federated searching roadblocks discussed in the Info Today article still valid?
Friday, October 16, 2009
Reading Notes: Access in Digital Libraries --I
A digital library and/or search engine much be able to handle many different kinds of multimedia, huge numbers of simultaneous requests, and attempts by some parties to manipulate automated search engine processes.
In order to index effectively, search engines must use algorithms that can:
-Determine what content should be indexed
-Divvy up tasks between servers
-Analyze content
-Collect new URLs
-Ignore dupicate pages
-Add new URLS are added to a cue
-Ignore spam
-Avoid overloading servers which host the pages they are crawling
-Save page content for indexing indexing.
-Process simplistic or ambiguous queries like "the onion" and weed out content effectively
-multiple factors should be included in this determination
Suggested focii for improving search engine quality (Henzinger): spam, content quality, webmaster deviation from web conventions, duplicate hosts, and vaguely structured data.
Some websites use text, links, or cloaking mechanisms to improve their ranking in search result sets. This often includes adding content (key words, links, false content) in an attempt to fool search engine page ranking algorithms. Spam may be structured in the form of deceptive text (white text on a white background with text that is invisible to the user but not the webcrawler), link spam (or link farms which collect links pointing to every other page in that site), and providing entirely different content for the user and the webcrawler. Work can be done towards assessing the quality of content included in a given webpage. Some sites advertise false content (like celebrity names) in an attempt to redirect users to their site.
Most search engine users do not travel beyond the first page of search results. Since many website generate income from their traffic, there is great motivation to increases a page's ranking so that it appears on the first page of a result set. Search engine developers must constantly tweak their indexing and page ranking algorithms in order to stay one step ahead of those who wish to manipulate these processes in order to increase traffic to their own page.
David Hawking , Web Search Engines: Part 1 and Part 2 IEEE Computer, June 2006.
M. Henzinger et al. challenges in Web Search Engines. ACM SIGIR 2002.
In order to index effectively, search engines must use algorithms that can:
-Determine what content should be indexed
-Divvy up tasks between servers
-Analyze content
-Collect new URLs
-Ignore dupicate pages
-Add new URLS are added to a cue
-Ignore spam
-Avoid overloading servers which host the pages they are crawling
-Save page content for indexing indexing.
-Process simplistic or ambiguous queries like "the onion" and weed out content effectively
-multiple factors should be included in this determination
Suggested focii for improving search engine quality (Henzinger): spam, content quality, webmaster deviation from web conventions, duplicate hosts, and vaguely structured data.
Some websites use text, links, or cloaking mechanisms to improve their ranking in search result sets. This often includes adding content (key words, links, false content) in an attempt to fool search engine page ranking algorithms. Spam may be structured in the form of deceptive text (white text on a white background with text that is invisible to the user but not the webcrawler), link spam (or link farms which collect links pointing to every other page in that site), and providing entirely different content for the user and the webcrawler. Work can be done towards assessing the quality of content included in a given webpage. Some sites advertise false content (like celebrity names) in an attempt to redirect users to their site.
Most search engine users do not travel beyond the first page of search results. Since many website generate income from their traffic, there is great motivation to increases a page's ranking so that it appears on the first page of a result set. Search engine developers must constantly tweak their indexing and page ranking algorithms in order to stay one step ahead of those who wish to manipulate these processes in order to increase traffic to their own page.
David Hawking , Web Search Engines: Part 1 and Part 2 IEEE Computer, June 2006.
M. Henzinger et al. challenges in Web Search Engines. ACM SIGIR 2002.
Labels:
algorithms,
Google,
issues,
reading notes,
search engines
Muddiest Point (10/16/09)
Will there be a Penapto recording for this class? I was unable to attend class due to illness. I would like to view it online.
Friday, October 9, 2009
Muddiest Point (10/28/09)
Are we supposed to submit our muddiest point and reading notes via email as well as through the blog?
Reading Notes: XML
XML is a difficult concept to understand. It is not a programming language, like Java, which uses functional elements. XML doesn't designate style elements, like HTML. In a sense, it doesn't really DO anything. It's basically a carrier and organizer of information.
XML is important because it allows different computers to interchange documents via the internet. You use XML to create your own tags which label the parts of information in a document.
< name > Jane < /name >
< role student < /student >
This way a person or a computer can identify the elements in a document easily and without ambiguity. XML is advantageous because it is a somewhat simplified version of SGML that lets you set document structure. HTML is very valuable in telling a computer how to display information on the web but it says nothing about the type of information that is the content of the document.
XML allows the user to create a list of elements tagged in a standardized way. The elements can be defined with attributes which give qualifying information about the attribute (i.e. gender, color, length, number, etc.). This can be done with quotation marks inside the original statement or in a separate statement.
The user can determine things like how to manage white space, language, and allowable characters. XML is useful for creating a detailed catalog.
DTD allows users to list and specify their own tags as well as set the order of their tags (XML does not impose a specific order on a list of elements). A DTD is a formal document which defines the roles of the elements labeled by the user. XSD (XML schema definition) can be used instead of DTD to define the structure of a document. XSD offers extensibility, consistency, power, and support for data types and namespaces (w3schools.com)
Namespaces: allows a user to differentiate between two different elements with the same name (this happens most often when XML is combined with HTML).
Questions:
*Can you please explain the use of xlms atrributes and the use of URIs ? I'm not sure I fully understand this concept.
*Can you clarify the different between terminal and non-terminal elements?
Sources:
XML is important because it allows different computers to interchange documents via the internet. You use XML to create your own tags which label the parts of information in a document.
This way a person or a computer can identify the elements in a document easily and without ambiguity. XML is advantageous because it is a somewhat simplified version of SGML that lets you set document structure. HTML is very valuable in telling a computer how to display information on the web but it says nothing about the type of information that is the content of the document.
XML allows the user to create a list of elements tagged in a standardized way. The elements can be defined with attributes which give qualifying information about the attribute (i.e. gender, color, length, number, etc.). This can be done with quotation marks inside the original statement or in a separate statement.
The user can determine things like how to manage white space, language, and allowable characters. XML is useful for creating a detailed catalog.
DTD allows users to list and specify their own tags as well as set the order of their tags (XML does not impose a specific order on a list of elements). A DTD is a formal document which defines the roles of the elements labeled by the user. XSD (XML schema definition) can be used instead of DTD to define the structure of a document. XSD offers extensibility, consistency, power, and support for data types and namespaces (w3schools.com)
Namespaces: allows a user to differentiate between two different elements with the same name (this happens most often when XML is combined with HTML).
Questions:
*Can you please explain the use of xlms atrributes and the use of URIs ? I'm not sure I fully understand this concept.
*Can you clarify the different between terminal and non-terminal elements?
Sources:
- Martin Bryan. Introducing the Extensible Markup Language (XML) http://burks.bton.ac.uk/burks/internet/web/xmlintro.htm
- Uche Ogbuji. A survey of XML standards: Part 1. January 2004. http://www-128.ibm.com/developerworks/xml/library/x-stand1.html
- Extending you Markup: a XML tutorial by Andre Bergholz http://www.pdffinder.com/pdf/extending-your-markup-an-xml-tutorial.html, or at http://xml.coverpages.org/BergholzTutorial.pdf
- XML Schema Tutorial http://www.w3schools.com/Schema/default.asp
Saturday, October 3, 2009
Flickr Photo Collection
Below is a link to my Flickr photo collection, a project for my Digital Libraries course.
The Family Zoo
The Family Zoo
Friday, October 2, 2009
Muddiest Point (9/21/09 and Syllabus)
Would you like us to email you our readings notes / muddiest point and post them in our blogs or submit them via blogs alone?
Also, on 9/21 you stated that digitization is no longer considered a means for preservation because digital objects are actually more fragile than physical ones. Does this have to do with the impermanence of URLS and changing technology, or is it something else? Can you explain?
Also, on 9/21 you stated that digitization is no longer considered a means for preservation because digital objects are actually more fragile than physical ones. Does this have to do with the impermanence of URLS and changing technology, or is it something else? Can you explain?
Reading Notes: Metadata
Subject classification has long been a standard of library science. Subject classification uses hierarchical relationships to describe objects and their correlation with similar items. In contrast, metadata is any order-of-magnitude information about information (the description of a particular object) either in digital or physical format.
Digital metadata is used to describe digital objects (i.e. text documents, images, video, etc.). Digital metadata is ideally embedded in the object it describes so that is is not displaced if the object is used. Some metadata is static and never changes while some is dynamic and is used for updating information about an object and how it changes over time. Metadata can also be used to extract information about a text: language, acronyms and their meaning, names of people, time and date stamps, email addresses, phrase hierarchy, etc.
In Objects of a Biographical System, Witten states that there is little need for subject classification in a digital world. Isn't subject classification still useful for helping the user retrieve information and find other items that may be of interest? Metadata can also meet this need but is there anything subject classification can do which metadata cannot?
Witten discussed the possible future of MPEG-7 files and their potential ability to evaluate a few notes of music and identify similar melodies, to retrieve graphics or logos from a few user-made digital brush strokes, identify the source of sounds from pitch samples, and to describe movements from actions in video files. Have there been any like developments since this work was published? Are projects aiming to discern human gestures and postures from video recordings likely done using metadata from MPEG-7 files or some other technology?
Anne Gilliland discusses metadata's long-term benefits. She says the following,
"What we do know is that the existence of many types of metadata will prove critical to the continued online and intellectual accessibility and utility of digital resources and the information objects that they contain, as well as the original objects and collections to which they relate. In this sense, metadata provides us with the Rosetta stone that will make it possible to decode information objects and their transformation into knowledge in the cultural heritage information systems of the future."
Is she referring to digital objects as the future artifacts of our culture? If so, what digital objects may be of importance to our predecessors? I wonder if they will have interoperability issues or if technology will have solved problems like interoperability by then.
Weibel's discusses difficult and unanticipated problems in the creation of a universal standard for metadata creation. He says the following,
" To borrow from the oldest joke of the Dismal Profession, put all the data modelers in the world end to end, and you won't reach a conclusion (we did, but it took ten years to manage it)."
What kind of problems is he referring to? How do data couplers help to find solutions? What problems have been resolved since this article was written (Summer 2005)?
If professionals are unable to come to a consensus on metadata standards, is it possible to create software that can identify the different standards used and accommodate them simultaneously (or even convert them to a standard format)?
Digital metadata is used to describe digital objects (i.e. text documents, images, video, etc.). Digital metadata is ideally embedded in the object it describes so that is is not displaced if the object is used. Some metadata is static and never changes while some is dynamic and is used for updating information about an object and how it changes over time. Metadata can also be used to extract information about a text: language, acronyms and their meaning, names of people, time and date stamps, email addresses, phrase hierarchy, etc.
In Objects of a Biographical System, Witten states that there is little need for subject classification in a digital world. Isn't subject classification still useful for helping the user retrieve information and find other items that may be of interest? Metadata can also meet this need but is there anything subject classification can do which metadata cannot?
Witten discussed the possible future of MPEG-7 files and their potential ability to evaluate a few notes of music and identify similar melodies, to retrieve graphics or logos from a few user-made digital brush strokes, identify the source of sounds from pitch samples, and to describe movements from actions in video files. Have there been any like developments since this work was published? Are projects aiming to discern human gestures and postures from video recordings likely done using metadata from MPEG-7 files or some other technology?
Anne Gilliland discusses metadata's long-term benefits. She says the following,
"What we do know is that the existence of many types of metadata will prove critical to the continued online and intellectual accessibility and utility of digital resources and the information objects that they contain, as well as the original objects and collections to which they relate. In this sense, metadata provides us with the Rosetta stone that will make it possible to decode information objects and their transformation into knowledge in the cultural heritage information systems of the future."
Is she referring to digital objects as the future artifacts of our culture? If so, what digital objects may be of importance to our predecessors? I wonder if they will have interoperability issues or if technology will have solved problems like interoperability by then.
Weibel's discusses difficult and unanticipated problems in the creation of a universal standard for metadata creation. He says the following,
" To borrow from the oldest joke of the Dismal Profession, put all the data modelers in the world end to end, and you won't reach a conclusion (we did, but it took ten years to manage it)."
What kind of problems is he referring to? How do data couplers help to find solutions? What problems have been resolved since this article was written (Summer 2005)?
If professionals are unable to come to a consensus on metadata standards, is it possible to create software that can identify the different standards used and accommodate them simultaneously (or even convert them to a standard format)?
- Ian H. Witten. “How to Build a Digital Library”. Morgan Kaufmann Publisher. 2002. ISBN: 1-558-60790-0.
- Anne J. Gilliland. Introduction to Metadata, pathways to Digital Information: 1: Setting the Stage http://www.getty.edu/research/conducting_research/standards/intrometadata/setting.html
- Stuart L. Weibel, “Border Crossings: Reflections on a Decade of Metadata Consensus Building”, D-Lib Magazine, Volume 11 Number 7/8, July/August 2005 http://www.dlib.org/dlib/july05/weibel/07weibel.html
Labels:
classification,
file formats,
interoperability,
metadata,
reading notes
Thursday, September 17, 2009
Muddiest Point (9/14/09)
My question concerns our final project. How will copyright affect our digital library collections? Can we use media (text, images, video) produced by others if we give them credit for their work?
Sunday, September 13, 2009
Week 1: Defining Digital Libraries
In his work, Structure of Scientific Revolutions, Thomas Kuhn states that a science begins with the definition of problems and methods for future generations of scientists. These definitions are further articulated as those future scientists study them and add precision.
The DELOS Manifesto takes on the task of setting definitions and setting the foundations for future research. It proposes models and possible objectives for research, functionality, quality, architecture, and policy.
William Arms talks about how the field of Library and Information Sciences is impacted by improvements in automation. What will the future of librarians be as advancements in digital libraries replace the traditional roles of library science professionals?
Paepcke suggests that librarians' jobs will become more specialized, helping researchers interact with technology during the production process and by adding value to materials that computer scientists put online. Arms says that librarians can use their skill with abstract ideas and the idiosyncrasies of information mining - skills that web crawlers and automated search engines have not yet mastered.
I wonder what place the automated future will carve out for those trained in library and information sciences. What role do they have to play in developing these systems that will eventually replace many of the responsibilities that librarians have always held? Will their skill set become more technological? As technology and commerce intersect with digital materials, will their focus be more rooted in philosophy of and the advocacy for access? Or will they become highly specialized computer scientists?
I think librarians have much to offer in the development of digital information systems (and visa versa). Now, more than ever, it's important for library science to keep in step with technology and as it does so, carve out its own place in a rapidly changing environment. Maybe in doing so, library science can create its own manifesto, declaring ways for the field to use its philosophical foundations and collective skill set to improve the way humans interact with information and technology.
The DELOS Manifesto takes on the task of setting definitions and setting the foundations for future research. It proposes models and possible objectives for research, functionality, quality, architecture, and policy.
William Arms talks about how the field of Library and Information Sciences is impacted by improvements in automation. What will the future of librarians be as advancements in digital libraries replace the traditional roles of library science professionals?
Paepcke suggests that librarians' jobs will become more specialized, helping researchers interact with technology during the production process and by adding value to materials that computer scientists put online. Arms says that librarians can use their skill with abstract ideas and the idiosyncrasies of information mining - skills that web crawlers and automated search engines have not yet mastered.
I wonder what place the automated future will carve out for those trained in library and information sciences. What role do they have to play in developing these systems that will eventually replace many of the responsibilities that librarians have always held? Will their skill set become more technological? As technology and commerce intersect with digital materials, will their focus be more rooted in philosophy of and the advocacy for access? Or will they become highly specialized computer scientists?
I think librarians have much to offer in the development of digital information systems (and visa versa). Now, more than ever, it's important for library science to keep in step with technology and as it does so, carve out its own place in a rapidly changing environment. Maybe in doing so, library science can create its own manifesto, declaring ways for the field to use its philosophical foundations and collective skill set to improve the way humans interact with information and technology.
- Christine L. Borgman, “From Gutenberg to the Global Information Infrastructure: Access to Information in the Networked World”. MIT Press, 2001.
- Leonardo Candela et. al. (2007) Setting the Foundations of Digital Libraries. D-Lib Magazine 13(3-4), March/April 2007.
- Andreas Paepcke, Hector Garcia-Molina, Rebecca Wesley, “Dewey Meets Turing: Librarians, Computer Scientists, and the Digital Libraries Initiative” D-Lib Magazine, Volume 11 Number 7/8, July/August 2005.
- Christian Lupovici. (2008) The growth of the role of librarians and information officers in digital libraries. Digital Libraries, Fabrice Papy (eds). ISTE and John Wiley & Sons, Inc.
- William Y. Arms. “Automated Digital Libraries, How Effectively Can Computers Be Used for the Skilled Tasks of Professional Librarianship?” D-Lib Magazine July/August 2000. 6(7/8).
Friday, September 11, 2009
Week 3: Digital Object Identifiers and Their Implications
The discussion of digital objects started off harmlessly enough. Computer language must be converted into human language - characters (i.e. A, a, 4, $) are created using strings of 1's and 0's. Digital documents are formatted for structure using a family of markup languages (SGML) and for format using HTML (a combination of structure and format capability) and style sheets.
Computers must be able to understand each other as they communicate over networks, so a system of identifiers such as URNs and URLs are put in place so that computers can have a universal method of keeping track of digital objects and their locations. The problem with URLs is that they are not stable identifiers.
The Webmaster may change or remove content on a page without warning. If another person uses a URL to cite linkages between their work sand someone else's work, there is no guarantee that linkage of information will still be valid 6 months down the road.
A reference to a lawsuit between Microsoft and Ticketmaster in Paskin's article on Digital Object Identifiers caught my attention so I did a basic search to find out more. Apparently, Ticketmaster sued Microsoft because Microsoft provided its users with a "deep link" to Ticketmaster's ticket purchasing page which allowed users to bypass the usual onslaught of digital advertising.
As the internet provides a creative means to generate commerce, arguments over intellectual property and "unfair competition" are bound to grow. The intrinsic value of the Internet is in the power it gives the user to access information instantly and to follow linkages between documents or web pages simply by clicking on hypertext.
As the concept of Digital Object Identifiers grows as a means to uniquely identify bodies of information on the internet, so does the need to evaluate the balance between the ownership of intellectual property and universal access. Digital Object Identifiers create a standardized way of identifying information (like a serial number) and solve some problems created by the instability of URLs.
Digital Object Identifiers (DOIs) provide a way to identify bodies of information that is neither too specific (for example, an ISBN which denotes a hardcover or trade paper version of the same book) nor too generalized to have meaning. Paskin talks about some of the benefits of DOIs. DOIs provide resolution by tying a digital object to a specific name. They have the potential to support interoperability. If separate digital libraries adopt the same naming rules and standards for digital objects, those objects can be identified across digital libraries. They also supply persistence - a deleted page will no longer signify a broken link between two bodies of information as some central agency can help us keep track of our digital objects and their changing locations. DOI.
The implications for this system are frightening as well as dazzling. Paskin identifies some monumental concerns in his article about DOIs. What happens when a user can no longer freely navigate between pages? If the central agency has control of the DOI, and by implication, the means to access the content the DOI identifies, it has the ability to restrict access to content by withholding that DOI. How much harder will it be to access content? How much time will it take? Does the central agency that manages the DOI database have the right to know what information we are accessing?
As I read more I had questions of my own. To what extent will this system be used as a way to generate commerce from materials that we now have access to for free? How does this affect research? The ability to follow citation via hypertext when accessing research publications is pivotal to research. If some third party regulates this access, will research suffer? This has a huge impact on our society.
Lynch, in his article concludes with "In a very real sense, there are no bad identifiers, but it is very possible to put identifiers to bad or inappropriate uses." I agree with this statement. The ability to manage information efficiently so that it can be accessed easily is a a huge benefit to those who access the information. But what is the cost?
Computers must be able to understand each other as they communicate over networks, so a system of identifiers such as URNs and URLs are put in place so that computers can have a universal method of keeping track of digital objects and their locations. The problem with URLs is that they are not stable identifiers.
The Webmaster may change or remove content on a page without warning. If another person uses a URL to cite linkages between their work sand someone else's work, there is no guarantee that linkage of information will still be valid 6 months down the road.
A reference to a lawsuit between Microsoft and Ticketmaster in Paskin's article on Digital Object Identifiers caught my attention so I did a basic search to find out more. Apparently, Ticketmaster sued Microsoft because Microsoft provided its users with a "deep link" to Ticketmaster's ticket purchasing page which allowed users to bypass the usual onslaught of digital advertising.
As the internet provides a creative means to generate commerce, arguments over intellectual property and "unfair competition" are bound to grow. The intrinsic value of the Internet is in the power it gives the user to access information instantly and to follow linkages between documents or web pages simply by clicking on hypertext.
As the concept of Digital Object Identifiers grows as a means to uniquely identify bodies of information on the internet, so does the need to evaluate the balance between the ownership of intellectual property and universal access. Digital Object Identifiers create a standardized way of identifying information (like a serial number) and solve some problems created by the instability of URLs.
Digital Object Identifiers (DOIs) provide a way to identify bodies of information that is neither too specific (for example, an ISBN which denotes a hardcover or trade paper version of the same book) nor too generalized to have meaning. Paskin talks about some of the benefits of DOIs. DOIs provide resolution by tying a digital object to a specific name. They have the potential to support interoperability. If separate digital libraries adopt the same naming rules and standards for digital objects, those objects can be identified across digital libraries. They also supply persistence - a deleted page will no longer signify a broken link between two bodies of information as some central agency can help us keep track of our digital objects and their changing locations. DOI.
The implications for this system are frightening as well as dazzling. Paskin identifies some monumental concerns in his article about DOIs. What happens when a user can no longer freely navigate between pages? If the central agency has control of the DOI, and by implication, the means to access the content the DOI identifies, it has the ability to restrict access to content by withholding that DOI. How much harder will it be to access content? How much time will it take? Does the central agency that manages the DOI database have the right to know what information we are accessing?
As I read more I had questions of my own. To what extent will this system be used as a way to generate commerce from materials that we now have access to for free? How does this affect research? The ability to follow citation via hypertext when accessing research publications is pivotal to research. If some third party regulates this access, will research suffer? This has a huge impact on our society.
Lynch, in his article concludes with "In a very real sense, there are no bad identifiers, but it is very possible to put identifiers to bad or inappropriate uses." I agree with this statement. The ability to manage information efficiently so that it can be accessed easily is a a huge benefit to those who access the information. But what is the cost?
- LESK sections 2.1, 2.2, 2.7, chapter 3.
- ARMS. Chapters 9. http://www.cs.cornell.edu/wya/DigLib/MS1999/Chapter9.html.
- Clifford Lynch, “Identifiers and Their Role In Networked Information Applications”. http://www.arl.org/bm~doc/identifier.pdf
- Norman Paskin. “Digital Object Indentifier (DOI) System”. Encyclopedia of Library and Information Sciences. http://www.doi.org/overview/080625DOI-ELIS-Paskin.pdf
Background Readings:
- Sam Sun, Larry Lannom, and Brian Boesch. "Handle System Overview", http://www.handle.net/rfc/rfc3650.html.
6. Netlitigation. "Ticketmaster v. Microsoft, United States District Court for the Central District of California, Civil Action Number 97-3055DPP", http://www.netlitigation.com/netlitigation/cases/ticketmaster.htm
Friday, September 4, 2009
Week 2: Notes from Works About Interoperability, Extensibility, and the World Wide Web
In addition to establishing an accepted definition for digital libraries, there is a movement within the field to establish a protocol or set of standards for building the components of digital libraries.
Having a standard protocol for digital library development allows for both interoperability (the ability to use the same components with the same functionality across different interfaces) and extensibility (allowing for future development of the components without compromising functionality). In order to develop a protocol, one must define and standardize the components of digital libraries.
Some of the components defined in these articles are:
The Digital Library Research Group at Cornell University and the Corporation for National Research Initiatives (CNRI) have been working together to establish infrastructures that provide real-life examples of interoperability and extensibility through open architecture. They tested the interoperability of their product across repositories, disseminators, and the behaviors of their digital objects.
Their tests show that:
References:
1.Hussein Suleman and Edward A. Fox. “A Framework for Building Open Digital Libraries”, D-Lib Magazine, December 2001. Volume 7 Number 12. http://www.dlib.org/dlib/december01/suleman/12suleman.html.
2.William Y. Arms , Christophe Blanchi, Edward A. Overly. ”An Architecture for Information in Digital Libraries”. D-Lib Magazine, February 1997. http://www.dlib.org/dlib/february97/cnri/02arms1.html.
3.Sandra Payette, Christophe Blanchi, Carl Lagoze, Edward A. Overly. “Interoperability for Digital Objects and Repositories, The Cornell/CNRI Experiments”, D-Lib Magazine, May 1999, Volume 5 Issue 5. http://www.dlib.org/dlib/may99/payette/05payette.html.
4.ARMS, Chapter 2, or “An Overview of World Wide Web”, http://www.cio.com/WebMaster/sem2_home.html
Having a standard protocol for digital library development allows for both interoperability (the ability to use the same components with the same functionality across different interfaces) and extensibility (allowing for future development of the components without compromising functionality). In order to develop a protocol, one must define and standardize the components of digital libraries.
Some of the components defined in these articles are:
- repositories where digital objects are stored
- handles or identifiers for digital objects
- disseminators or processes associated with digital objects that provide increased functionality
- servlets or programs containing a collection of operations and associated with a specific disseminator types
The Digital Library Research Group at Cornell University and the Corporation for National Research Initiatives (CNRI) have been working together to establish infrastructures that provide real-life examples of interoperability and extensibility through open architecture. They tested the interoperability of their product across repositories, disseminators, and the behaviors of their digital objects.
Their tests show that:
- Interoperability across repositories can be established using RAP (Repository Access Protocol) interface
- The same tasks can be accomplished in separate repositories with the same outcome
- Extensibility can be established, which allows community standards to be integrated into existing digital libraries
- Allowing more variation between repositories without compromising interoperability
- Creating a measurement system for interoperability
- More testing for interoperability
- Issues of security
References:
1.Hussein Suleman and Edward A. Fox. “A Framework for Building Open Digital Libraries”, D-Lib Magazine, December 2001. Volume 7 Number 12. http://www.dlib.org/dlib/december01/suleman/12suleman.html.
2.William Y. Arms , Christophe Blanchi, Edward A. Overly. ”An Architecture for Information in Digital Libraries”. D-Lib Magazine, February 1997. http://www.dlib.org/dlib/february97/cnri/02arms1.html.
3.Sandra Payette, Christophe Blanchi, Carl Lagoze, Edward A. Overly. “Interoperability for Digital Objects and Repositories, The Cornell/CNRI Experiments”, D-Lib Magazine, May 1999, Volume 5 Issue 5. http://www.dlib.org/dlib/may99/payette/05payette.html.
4.ARMS, Chapter 2, or “An Overview of World Wide Web”, http://www.cio.com/WebMaster/sem2_home.html
Labels:
definitions,
extensibility,
internet,
interoperability
Tuesday, September 1, 2009
Muddiest Point (8/31/09)
This is more of a request for more information than clarification:
I had a chance to look at the "Visualize the Collections" or "Bungee View" prototype for the Historic Pittsburgh Image Collections website. Are there any other visualization tools like this one that we can look at on the web?
Which aspects of Bungee Viewer are representative of emerging technology? Is it in the components or the way the capabilities are bundled into one package?
I had a chance to look at the "Visualize the Collections" or "Bungee View" prototype for the Historic Pittsburgh Image Collections website. Are there any other visualization tools like this one that we can look at on the web?
Which aspects of Bungee Viewer are representative of emerging technology? Is it in the components or the way the capabilities are bundled into one package?
The Blog
Welcome to my blog! All postings will be relevant to the Digital Libraries course at the University of Pittsburgh, Fall 2009. My posts will include reflections, questions, and resources relevant to the world of digital libraries. Enjoy!
Subscribe to:
Posts (Atom)
