Wednesday, October 14, 2009

Carr's Big Switch redux

I'm about 1/2 done with Big Switch, and its good.

Carr's thesis is that information services (hardware and software) are now and increasingly operating on the scale of the electrical power utilities. Just as companies don't generally create their own power, they now no longer need their own IT depts. As network access speeds and reliability approach that available on one's own computer, the network itself becomes one big machine. He's look at and past things like Amazon's EC2 "elastic computing cloud" that allows companies to use Amazon's computers (H&S) as if they were their own. Another example is 3Tera's AppLogic, a cloud computing platform.The customer pays for the computing power consumed when they consume it--just like we pay for electricity. Wow!

Carr briefly highlight large-scale consequences of the electrical gridon society and suggests that similarly large-scale effects will follow from the utilitization of computing. I think he's right. Consider the consequences of large-scale, utility-style cloud computing on digital preservation. If say higher education institutions outsource their computing utility-style to global third-party providers, then preservation of the digital content (an oxymoron; what we mean is curation of digital content over time via migrations) also moves to the third party. There the scale is much larger, part of the ongoing access to content, and costs to individual institutions is amortized across the aggregate of all the institutions using the third party computing utility. In short, indiviual institutions (here colleges and universities) need not themselves directly work to preserve their digital content. They have out-sourced it to their computing utility. It (digital curation--my prefered phrase for digital preservation) still has to be done but not repeatedly at a local (say we call it--retail) level. There are many trust issues here, but then most of us trust our utilities now for water and power.

Things RDA

http://ac.bslw.com/community/blog/category/rda/

a blog on RDA from backstage folks; a good list of links to RDA materials.

Among the best:

1. http://www.rda-jsc.org/rda.html – The RDA page on the Joint Steering Committee website.

2. http://metadataregistry.org/rdabrowse.htm – The registered RDA elements and vocabularies.

3. http://tsig.wikispaces.com/Pre-conference+2009+presentation+materials – CLA power points and other materials from Montreal Pre-conference.

Tuesday, October 13, 2009

The Law of Linked Data

http://www.mkbergman.com/837/the-law-of-linked-data/

Mike Bergman at AI3 blogs about the Law of Linked Data. "The Linked Data Law: the value of a linked data network is proportional to the square of the number of links between data objects."

He argues that linking (meaningfully) the existing nodes on the 'net produces network effects for the semantic web. He makes a nice analogy with Metcalfe’s law, which "states that the value of a telecommunications network is proportional to the square of the number of users of the system.

Bergman thinks a good marshal would deliver law and order to linked data and the semantic enterprise. What would the good marshal do? Well, he doesn't say beyond "deliver law and order." What the hell does that mean? OK, aside from that the piece is worth reading and it links to more good reading, too. Are footnotes the original linked data?

Monday, October 12, 2009

Report on Large-Scale Digitization of Manuscript Collections

Extending the Reach of Southern Sources Proceeding to Large-Scale Digitization of Manuscript Collections: Final Grant Report / Prepared by the Southern Historical Collection University Library The University of North Carolina at Chapel Hill for the Andrew W. Mellon Foundation


I like the questionnaire/decision matrix and though it here applies to questions of what should be prioritized for digitization, it could be modified to capture metadata relevant to what should be prioritized for digital preservation. See Appendix G, page 57 for the questionnaire.

http://www.lib.unc.edu/mss/archivalmassdigitization/Extending_the_Reach.pdf

Friday, October 9, 2009

Google books editorial in NYTimes

Sergey Brin wrote in the NYTimes today about the Google book deal in an op-ed piece called "A Library to Last Forever. " He makes a good case for a deal of some sort and gently addresses a few of the concerns that have been raised. His argument for the deal has two aspects: one, implied by the title, that the deal will protect books forever in a new kind of library that is disaster proof and, apparently, Google has solved the digital preservation problems (forever is a long time.) The other aspect is access. The book deal will make a century's printed output easily avaiable to all. Sounds good, but I don't know much about most of what he was talking about. The one thing I know something about--access to library collections--was mentioned in one sentence that is just completely wrong.

"Today, if you want to access a typical out-of-print book, you have only one choice — fly to one of a handful of leading libraries in the country and hope to find it in the stacks."

What the hell is he talking about? There is no need to fly and hope. He must know that you have at least one other choice: use your computer to 1. look up the book in WorldCat to see what libraries have copies 2. email your local library to use its inter-library loan service to get the book for you. He can't be ignorant of this--Google has a deal with OCLC that joins Google Book Search and WorldCat--so why did he say "fly" and "hope"? One effect of this: it makes me wonder if the other things he says are just as fishy as this. I don't know anything about those other things, but seeing what he said about the one piece I do know about makes me doubt everything else he says.

Thursday, October 8, 2009

Knowledge in an age of abundance

A talk at LITA Forum by David Weinberger, author of Everything Is Miscellaneous.

http://www.al.ala.org/insidescoop/2009/10/05/lita-forum-saturday-keynote-knowledge-in-the-age-of-abundance/

"Citing the Scottish philosopher Andy Clark, Weinberger explained that the internet becomes almost a sort of extension of our mind (scaffolding, he called it) so that we think with our brains and store information elsewhere."

My reaction:

Information abundance (for the affluent or for affluent societies, anyway) does seem to be a primary characteristic of our information economy. Our institutions though are shaped by the past environment that was characterized by information scarcity. Libraries seem a prime example. When information is scarce, then collecting it creates pockets of abundance for specific sets of users in particular places--a city, a university, etc. But when information is abundant, then creating local collections of information (to overcome information's "natural" scarcity) is a waste of time. The environment libraries (and the host institutions of libraries: cities, nations, universities, etc.) thrived within is gone. Libraries and other institutions built for an information economy characterized by scarcity must re-make themselves so they fit an information abundance economy. Libraries--as we have known them--are moot.

One model for libraries that seems to be working in the abundant information economy is the library as museum. The library becomes less information-centric and more artifact-centric. Artifacts may remain scarce, so the scarcity-based model of a library as museum could work for collections of rare, unique or otherwise special materials. But libraries as information-centric institutions are ill-suited for an abundant information economy. Perhaps the most obvious characteristic of this ill-fit is that libraries are primarily _local_ institutions that serve a host organization (city, state, university, etc.) The new organizations that have grown up in the abundant information economy are global: Amazon, Google, etc., and one can see a similar pattern of movement from many small local entities to a few large global entities in other information-centric activities like banking and stock brokerage.

In an age of abundant information, the information-centric organizations we need help us find what we are looking for (search tools like Google,) share what we have found (like blogs and social tools), and use or re-use what we have found (productivity tools designed with the "cloud" in mind.) The traditional infrastructure libraries and similar collection-oriented institutions provide doesn't address the needs of users in an economy of information abundance. It seems likely that the information-centric organizations that emerge to help users navigate and manage and use abundant information will be global organizations that are not subordinate to local host institutions like Sioux City, Iowa, the US Dept. of Labor, Yale University, etc. Its a big change.

Five trends in discoverability

Lorcan Dempsey posted on his blog about a U. MN study on discoverability.

The 5 trends are

1. Users are discovering relevant resources outside of traditional library systems.

2. Users expect discovery and delivery to coincide.

3. Expanding use of portable Intenet devices.

4. Recommendation increasingly push discovery.

5. Users rely increasingly on non-traditional information objects.


This is worth reading.

Wednesday, October 7, 2009

Thingology post on Ebook economics: Are libraries screwed?

http://www.librarything.com/thingology/index.php

Tim Spaulding has a nice piece on ebook pricing for libraries. A lot of doom and gloom, but the gist is on target publishers/bundlers will rent ebooks to libraries as e-journals are now and thus price increases for ebooks on the scale and model of e-journals is likely.

The whole concept of collection development is altered when the library is a renter and not an owner of books and journals. If a library is rooted in its possession of a collection, then a library that rents is not a library.

Tuesday, October 6, 2009

iPRES '09 today and yesterday

International Conference on Preservation of Digital Objects (iPRES 2009) at Mission Bay Conference Center in San Francisco, October 5th and 6th, 2009, explores the latest trends, innovations, and practices in preserving our scientific and cultural digital heritage.
http://www.cdlib.org/iPres/

blog postings on iPRES 09 from Digital Curation blog at http://digitalcuration.blogspot.com/search/label/iPres09

Monday, October 5, 2009

LC's effort to define an extended date/time format

Library of Congress has proposed an extended date/time definition for use with ISO 8601 and possibly with W3C as an XML schema type.

http://www.loc.gov/standards/datetime/index.html

The problem: No standard date/time format meets the needs of XML metadata schemas. W3C XML Schema built-in types xs:date, xs:time, and xs:dateTime are inadequate, as is W3CDTF, and TEMPER. ISO 8601 and the W3C schema are incompatible. The LC proposal addresses that and adds BCE dates, open date ranges, and useful/necessary concepts like "uncertain" and "approximate" to the definition and the format.

The proposal could be incorporated into schemas such as MODS and METS. (Note: it is already in use within the PREMIS schema.) It may be proposed for standardization in ISO 8601 or it might be proposed to W3C for adoption as an XML schema type – the benefits of this are clear, among them: strict validation would be supported.

Wednesday, September 30, 2009

Google Book deal explained

The folks at JISC and eFoundations have been commenting on the Google book deal. This link goes to a brief with some links. http://efoundations.typepad.com/efoundations/2009/09/the-google-book-settlement.html

Friday, September 25, 2009

VIAF now available as linked data.

Thom Hickey of OCLC posted on the VIAF as linked data on his blog.

Search the VIAF beta at http://viaf.org/

Thom says,

There are some 9.5 million personae described in VIAF and have established more than 4 million links between the files. To us linked data means:
URIs for everything
HTTP 303 redirects for URIs representing the personae our metadata is about
HTTP content negotiation for different data formats
An RDF view of the data
A rich a set of internal and external links in our data