Friday, February 4, 2011

Metadata guidelines for the UK RDTF

Andy Powell and Pete Johnston, of Eduserv, with funding from JISC, have put up some high level draft guidelines for "how metadata associated with library, museum and archival collections should be made available for the purposes of supporting resource discovery in line with the Resource Discovery Taskforce (RDTF) Vision."

See the guidelines at http://rdtfmetadata.jiscpress.org/

They are taking comments on the draft until Feb. 18.

You can read Powell's and Johnston's own commentary/announcement at their joint blog, eFoundations, at http://efoundations.typepad.com/efoundations/2011/02/metadata-guidelines-for-the-uk-rdtf.html

A few quotes from the guidelines.

"These guidelines have been developed such that they:

1. support the RDTF Vision;
2. are compatible with the outcomes of the JISC IE Technical Review meeting in London, Aug 2010;
3. are in line with Linked Data principles as far as possible;
4. are compatible with the W3C Linked Open Data Star Scheme;
5. are in line with Designing URI Sets for the UK Public Sector;
6. take into account the Europeana Data Model and ESE;
7. are informed by mainstream web practice and search engine behaviour and are broadly in line with the notion of “making better websites” across the library, museum and archives sectors."

"The guidelines are intended to help libraries, museums and archives expose existing metadata (and any new metadata that is created using existing practices) in ways that 1) supports the development of aggregator services and that 2) integrates well with the web of data. The intention is not to change existing cataloguing practice in libraries, museums and archives."

"RDTF metadata should be made openly available using one or more of three approaches, referred to below as the community formats approach, the RDF data approach and the Linked Data approach."

For what it is, it looks good. Powell and Johnston "believe that by putting this guidance in place it will be possible to create significantly more coherence in the way that metadata is created, managed and used across the library, archives and museum sectors than is currently the case."

Friday, January 14, 2011

Digital forensics

A CLIR report on digital forensics for born digital collections is out. Digital Forensics and Born-Digital Content in Cultural Heritage Collections by Matthew G. Kirschenbaum,Richard Ovenden,Gabriela Redwine with research assistance from Rachel Donahue.

The report makes a case for applying digital forensics, an applied field originating in law enforcement,computer security, and national defense, to the archives and curatorial community since libraries, special collections, etc. increasingly receive
computer storage media (and sometimes entire computers) as part of their acquisitions of "papers" from artists, writers, musicians, etc. Upwards of 90 percent of the records (i.e. personal and corporate "papers") being created today are born digital (Dow 2009, xi).

Here's a quote from the introduction: "Digital forensics therefore offers archivists, as well as an archive’s patrons, new tools, new methodologies, and new capabilities. Yet as even this brief description must suggest, digital forensics does not affect archivists’ practices solely at the level of procedures and tools. Its methods and outcomes raise important legal, ethical, and hermeneutical questions about the nature of the cultural record, the boundaries between public and private knowledge, and the roles and responsibilities of donor, archivist, and the public in a new technological era."

This report cites an earlier one that sounds good, too. "The starting place for any cultural heritage professional interested in matters of forensics, data recovery, and storage formats is a 1999 JISC/NIPO study coauthored by Seamus Ross and Ann Gow
and entitled Digital Archaeology: Rescuing Neglected and Damaged Data Resources. Although more than a decade old, the report remains invaluable."

Monday, January 10, 2011

OCLC report on managing print collections in mass-digitized library world

Malpas, Constance. 2011. Cloud-sourcing Research Collections: Managing Print in the Mass-digitized Library Environment. Dublin, Ohio: OCLC Research. http://www.oclc.org/research/publications/library/2011/2011-01.pdf.

Cloud-sourcing Research Collections is a 76 p. (pdf) analysis of the feasibility of outsourcing management of low-use print books held in academic libraries to shared service providers, including large-scale print and digital repositories.

Mass digitization projects like Google Books and shared online collections like the HathiTrust have given substance to the visions of a transformation of library use from paper to online resources. This "flip" and related demands for physical space and care of paper resources has resulted in renewed attention to print collections in academic libraries. This is the time for discussion within and among research libraries on how to construct new systems of services based on aggregations of digital resources, local paper resource collections and shared storage repositories for online and paper resources.

The report's main conclusion is:

"Based on a year-long study of data from the HathiTrust, ReCAP, and WorldCat, we concluded that our central hypothesis was successfully confirmed: there is sufficient material in the mass-digitized library collection managed by the HathiTrust to duplicate a sizeable (and growing) portion of virtually any academic library in the United States, and there is adequate duplication between the shared digital repository and large-scale print storage facilities to enable a great number of academic libraries to reconsider their local print management operations. Significantly, we also found that the combination of a relatively small number of potential shared print providers, including the Library of Congress, was sufficient to achieve more than 70% coverage of the digitized book collection, suggesting that shared service may not require a very large network of providers."

This points a way forward for academic libraries. The report might be an interesting frame for a discussion at Yale of how we think of our collections in this environment and how we move to use the environment to create services for readers. It is one of the few reports that integrates questions of online resources with paper resources. That kind of integrated approach to collections, preservation, user/reader services makes a lot more sense than digital only or print only approaches.

Friday, December 3, 2010

Robert Darnton The Library: Three Jeremiads from the NY Review of Books, Dec. 23, 2010

The Library: Three Jeremiads from the NY Review of Books, Dec. 23, 2010
http://www.nybooks.com/articles/archives/2010/dec/23/library-three-jeremiads/?pagination=false

Excellent piece on the crises in research libraries. It's all about money and the lack of it for research libraries.

3 Jeremiads, 3 problems

1st: Monographs and scholarship
" ... a vicious circle: the escalation in the price of periodicals forces libraries to cut back on their purchase of monographs; the drop in the demand for monographs makes university presses reduce their publication of them; and the difficulty in getting them published creates barriers to careers among graduate students."

"Another rule of thumb used to prevail among the better university presses. They could count on research libraries purchasing about eight hundred copies of any new monograph. By 2000 that figure had fallen to three or four hundred, often less, and not enough in most cases to cover production costs. Therefore, the presses abandoned subjects like colonial Latin America and Africa."

2nd: Journals
"A few years later, “sustainability” had become a buzz word, and the inflationary spiral of journal prices had continued unabated. In 2007 I became director of the Harvard University Library, a strategic position from which to take the full measure of the business constraints on academic life. Although economic conditions had worsened, the faculty’s understanding of them had not improved."

"How many professors in chemistry can give you even a ballpark estimate of the cost of a year’s subscription to Tetrahedron (currently $39,082)?"

"At Harvard we developed a new model. By a unanimous vote on February 12, 2008, professors in the Faculty of Arts and Sciences bound themselves to deposit all of their future scholarly articles in an open-access repository to be established by the library and also granted the university permission to distribute them."

3rd: Google Books
"The fundamental incompatibility of purpose between libraries and Google Book Search might be mitigated if Google could offer libraries access to its digitized database of books on reasonable terms. But the terms are embodied in a 368-page document known as the “settlement,” which is meant to resolve another conflict: the suit brought against Google by authors and publishers for alleged infringement of their copyrights."

"Despite its enormous complexity, the settlement comes down to an agreement about how to divide a pie—the profits to be produced by Google Book Search: 37 percent will go to Google, 63 percent to the authors and publishers. And the libraries? They are not partners to the agreement, but many of them have provided, free of charge, the books that Google has digitized. They are being asked to buy back access to those books along with those of their sister libraries, in digitized form, for an “institutional subscription” price, which could escalate as disastrously as the price of journals."

"... my happy ending: a National Digital Library—or a Digital Public Library of America (DPLA), as some prefer to call it."

Monday, November 15, 2010

Open Bibliographic Data Guide: JISC study on the business cases for Open Bibliographic Data

http://obd.jisc.ac.uk/

links to the _The Guide to Open Bibliographic Data_ that JISC developed on behalf of its partners in the Resource Discovery Task Force. It is about the business cases for Open Bibliographic Data – releasing some or all of a library’s catalog records for open use and re-use by others. The Guide uses 17 use cases to explore
* How to license the data
* Legal issues to be considered
* Potential costs and savings
* Practical implications in terms of processes, effort and skills
* Data formats and other technical options

The assumed rationale is about discoverability and is gaining in credibility the more our resources are discovered from ‘out there’ (through such as Google) and not from ‘in here’ (through the local OPAC). --most of the above quoted or modified slightly from the Guide.

A PDF version of the use cases is at http://obd.jisc.ac.uk/wp-content/uploads/2010/11/Open-Bibliographic-Data-The-Use-Cases.pdf

On a quick review, the case studies seem useful: specific, brief, comprehensive. Each use case includes sections on description, motivation, benefits, consequences, rights & licensing, practicalities and costs.

A table of use cases and examples is at http://obd.jisc.ac.uk/examples

An example for use case 1 (publish data for unspecified use) is Open Library http://openlibrary.org and another is Cambridge U. Library http://openbiblio.net/2010/10/05/jisc-openbibliography-cul-data-release/

An example for use case 2 (publish open Linked Data for unspecified use) is Libris, the joint catalogue of the Swedish academic and research libraries http://libris.kb.se/

Wednesday, November 10, 2010

Understanding linked data and its potential for libraries

I've been hearing a lot about linked data in the past year, and I find that I'm very fuzzy about what linked data is and how it matters or might matter to libraries, to organizations that have libraries and to people who may use libraries.

My first question is What is linked data? A good starting place for me is the definition at http://linkeddata.org

"Linked Data is about using the Web to connect related data that wasn't previously linked, or using the Web to lower the barriers to linking data currently linked using other methods. More specifically, Wikipedia defines Linked Data as 'a term used to describe a recommended best practice for exposing, sharing, and connecting pieces of data, information, and knowledge on the Semantic Web using URIs and RDF.'" --linkeddata.org, read Nov. 10, 2010.

I checked wikipedia for a definition and found a slightly different, more technical definition of linked data.

"Linked Data is a sub-topic of the Semantic Web. The term Linked Data is used to describe a method of exposing, sharing, and connecting data via dereferenceable URIs on the Web." --wikipedia, read on Nov. 10, 2010.

Of course, I had no idea what "dereferenceable URIs" are. Well, a dereferenceable URI is the normal and obvious way that links on the Web work: a URI refers to a page that the web server returns a copy of.

I'm in a technical vocabulary thicket,and I don't want to be. That may be useful later, but not now. I need to put it in my own words or into words I understand.

Let me try working with that definition cited by linkeddata.org: "... a recommended best practice for exposing, sharing, and connecting pieces of data, information, and knowledge on the Semantic Web using URIs and RDF.'

Linked data is a way to expose, share, or connect to data on the Web so that the data can be understood or is meaningful to other machines on the Web. This is how Linked Data is a sub-topic of the Semantic Web. Additionally, Linked Data uses URIs as names for things and RDF as the data model so that statements about resources (in particular Web resources)are made in the form of subject-predicate-object expressions, and these expressions are known as triples.

Well, that is making sense to me, but I don't know that it would make much sense to anyone else, or be seen by anyone else as an improvement over the other available definitions. It helps me, though. That is enough for now.

Wednesday, October 20, 2010

Ed Summers on the linked data release from Deutschen Nationalbibliothek

See Ed Summers' comments on the DNB release of linked library data at
http://inkdroid.org/journal/2010/10/19/linked-library-data-at-the-deutschen-nationalbibliothek/

Summers' piece is worth reading.

Here are some highlights.

The Deutschen Nationalbibliothek (DNB)has released linked library data for

■1.8 million authors from the Personennamendatei (PND)
■1.3 million corporate bodies from the Gemeinsame Körperschaftsdatei (GKD)
■187,000 subject headings from the Schlagwortnormdatei (SWD)
■51,000 Dewey Decimal Classification categories
The full dataset that the DNB has made available for download amounts to 38,849,113 individual statements (aka triples).

See the DNB announcement at http://lists.w3.org/Archives/Public/public-lod/2010Oct/0016.html

This is a huge event for _library_ uses of linked data, and exemplary behavior from DNB. Other research and national libraries should emulate the DNB.

Summers cites Herta Müller's authority information as an illustration and notes the use of RDA vocabularies, which are also available as linked data. "RDF vocabularies are explicit ways of describing resources like people, places, topics, etc. When different things are described using the same vocabulary (or the vocabularies themselves are related together in a particular way) it becomes possible to merge the descriptions, and build software on top of it."

"Another really interesting thing to note about this RDF for Herta Müller are the links to Wikipedia (http://de.wikipedia.org/wiki/Herta_M%C3%BCller), VIAF (http://viaf.org/viaf/12324250) and dbpedia (http://dbpedia.org/resource/Herta_M%C3%BCller). These are important because they contextualize the DNB record for Herta Müller by relating it to other records for her, thus allowing it to be disambiguated from records describing other people named Herta Müller."

"[Summers] did some quick and dirty analysis of the full data dump from the DNB and found: 3,569,402 links to VIAF and 40,136 links to dbpedia (the Linked Data version of Wikipedia)."

Summers goes on to talk a bit about what more needs to be done.

"What remains to be done to some extent is leveraging this contextual information around our data in Library Applications, both cataloging, metadata enrichment applications and end user facing discovery applications."


There is a lot more in his piece, and links to many related tools, projects, and activities.

Monday, August 16, 2010

Getting on the cloud

On Friday, I began renting a virtual server from linode.com. It's a linux thing. I'm using Ubuntu. Renting the virtual server is cheap and pretty easy. My colleague, Daniel Lovins, is helping me with advice, encouragement and answers to a few dumb questions. (Thanks, for the hand-holding and the example, Daniel.) Today, with Daniel's help, I installed Apache and MySQL. I'm starting to learn Vi, too. My plan is to set up drupal, an instance of vufind, and the eXtensible Catalog Metadata Toolkit (XC MST), and then I'll see what I can do with these tools in this environment. I have a lot to learn, but I feel that I am now able to actually play with the right toys.

Monday, August 9, 2010

Dempsey on supply and demand

Lorcan Dempsey's Aug. 8, 2010 blog posting on "Sorting out demand" is insightful and useful. His ideas are his 3rd top trend as presented at the 2010 LITA Top Tech Trends panel at ALA Annual Conference. http://orweblog.oclc.org/archives/002124.html

His argument about the shift (in libraries) from a focus on managing supply to "sorting out demand" is an economic one. Costs - in time, effort or money - for the user drives what services the library can provide and the "right" structure for the library. Libraries in the 20th century reduced the user's supply-based transaction costs by integrating the many sources of supply and bringing them close to the user. Libraries in the 21st century (or this early part of it)must also reduce user costs but those will be different costs as the supply transaction costs are falling due to the effect of digital format and the network. Dempsey mentions several examples of how libraries can provide services on the demand side: recommendations, contextualizing content for particular communities, connective services, tailoring content to purpose, and managing institutional assets. Each of these examples is interesting and promising, but not seem to be as compelling as the 20th century economic rationale for libraries.

I continue to think that libraries as we know them will be changed in ways very like bookstores and printed journals or newspapers; they will not serve the necessary local distribution needs as well as a globally networked provider. Some will survive because of local peculiarities or because they develop special services for their community--these may be the same thing. Research libraries will become more research museum-like, that is, more artifact-centric. But how a university, for instance, manages its institutional digital repositories and its licensed online resources may make more sense outside of the library. Yale has created a university-wide Office of Digital Assets and Infrastructure, that, though in its institutional infancy, is clearly the focal point on digital repositories at Yale, digital preservation at Yale, and discovery of resources across the university's many collections in libraries, archives, museums, etc. The changing economics will change the institutional structures. The more radical the changes in economics the more radical the resulting institutional changes.

Lastly, Dempsey closes his post with a link to Dan Chudnov's prescient 2006 post "help people build their own libraries" http://onebiglibrary.net/story/because-this-is-the-business-weve-chosen

Friday, August 6, 2010

catalogs and scholarship: a bit of sentiment

OK, it must be Paul Courant day here or something like that. I was looking at his blog post about the closing and more specifically the removal of the U. Michigan card catalog, and in his discussion of responses to the removal he spoke of his own sentimental response. These are the lines that struck me:

"... I’ll always remember the card catalog as the rich, powerful and brilliant piece of scholarship that it was, and as a place that I visited in eager anticipation of learning something new. I don’t think that I was ever disappointed."

The catalog as a work of scholarship is a view one rarely hears anymore, but it is the correct view of the card catalog of a research university. Those catalogs were the heart of the heart of the university, or the soul in the machine, or whatever lovely, sentimental phrase you most prefer to use. When one was "in" the catalog following a citation or browsing an author's works or perusing a subject, one was in more or less the active mind of the library and thus one could imagine it as the mind or memory of the university or of scholarship itself.

If the catalog was a work of scholarship, then the cataloger was a scholar. And in that I think we can feel the loss that many catalogers feel when they consider the past 30 years and look ahead to the future of cataloging and libraries: their work as scholars is at an end. It is possible to see a descending arc--the catalog as scholarship to the catalog as information repository to the catalog as a database. And the arc descends for the cataloger from scholar to information manager to data assistant. This view is a sentimental one and a depressing one; it is not objectively true, but I know that it feels true to many catalogers. It is part of the sorrow of catalogers that many lament the loss of status as scholars and don't feel any warmth for the status of a data-centric programmer.

Paul Courant's talk at OCLC: "Economic Perspectives on Academic Libraries,"

Paul Courant's recent talk at OCLC is now online. It is called, "Economic Perspectives on Academic Libraries," and it is worth a look and listen. http://www.oclc.org/research/news/2010-08-05.htm

His talk and slide take about 70 minutes. There is a QA video, too. Another 15 minutes or so.

His talk is interesting just to get his perspective as a university librarian (U. Michigan) and as a former provost (also at UM) and as an economist (on faculty at UM).

Summary: A library is a complicated institution, a big nonprofit business that supports the mission of an even bigger nonprofit business. It plays essential roles in the production and distribution of scholarship, which can be understood as an industry. For over a century, a library's focus has been almost entirely on the interests of the local institution and its local value has been almost entirely dependent on it role in sharing the costs of expensive information. However, digital information technology is radically altering the value of that focus and that role. A library is profoundly affected by both the emerging role of the network and by the fact that copying and distribution (of books, articles, etc.) are now very cheap. Courant develops these themes and shares some of his thoughts on the effective and efficient functioning of academic libraries.

There is no transcript to read, but there are 17 slides to view and both and MP3 to listen to and a Webcast to watch and hear.

His blog is at http://paulcourant.net/ and is called Au Courant.

Thursday, July 29, 2010

two surveys

Two surveys worth looking at. One by Ithaka S+R and one by BL/JISC.
The Ithaka S+R report:
Faculty Survey 2009: Key Strategic Insights for Libraries, Publishers, and Societies (April 7, 2010) at
http://www.ithaka.org/ithaka-s-r/research/faculty-surveys-2000-2009/Faculty%20Study%202009.pdf

The BL/JISC report:
Researchers of Tomorrow:A three year (BL/JISC) study tracking the research behaviour of 'Generation Y' doctoral students. Annual Report 2009-2010 (June 2010) at
http://explorationforchange.net/attachments/056_RoT%20Year%201%20report%20final%20100622.pdf

Each is excellent and worth reading, but my first impression is that nothing here is surprising. Digital technologies are transforming scholarship and communication and the relation of libraries (and archives and musuems) to scholars (faculty and graduate students) is changing. The need for libraries as direct intermediaries between scholars and local collections is lessening. Digital technologies present new opportunities for scholarship and communication, but institutions--publishers, societies, libraries, universities--have been slow to capitalize on them in any coordinated way. Researcher behaviors have adapted to the new technologies and have as yet held on to traditional attitudes, values and skills per evaluation and use of sources.