COMET (Cambridge Open METadata) project was completed this July 2011.
The COMET blog final post sums up the work done, the lessons learned, and indicates some next steps.
The CUL open data service is worth a look. (It is funded under the JISC Infrastructure for Resource Discovery program.)
I was particularly interested in the document on the ownership of MARC21 records by Hugh Taylor, Head of Collection Description and Development at Cambridge University Library. It is nice brief on the issues of intellectual property law and contracts and licences as they relate to MARC21 records in library catalogs. Hugh is very good on "reading the ownership of MARC 21 bibliographic records." The whole project is nicely documented at the COMET project blog.
Showing posts with label linked data. Show all posts
Showing posts with label linked data. Show all posts
Monday, September 12, 2011
Wednesday, June 29, 2011
The W3C Library Linked Data Incubator Group DRAFT REPORT
The W3C Library Linked Data Incubator Group has issued a draft report for commentary. Linked library data is an opportunity to bring library concepts and practices to the Web in ways that transcend any individual library and its individual limitations.
I quote here one line from the benefits section. "The Linked Data approach offers significant advantages over current practices for creating and delivering library data while providing a natural extension to the collaborative sharing models historically employed by libraries, archives, and museums ("memory institutions")." And one more from the benefits to "memory institutions." "By using Linked Data, memory institutions will create an open, global pool of shared data that can be used and re-used to describe resources, with a limited amount of redundant effort compared with current cataloguing processes." Cheaper, faster, better.
From a cataloger's viewpoint, library linked data is the ultimate cooperative cataloging environment and the ultimate user services environment.
The report
http://www.w3.org/2005/Incubator/lld/wiki/DraftReportWithTransclusion
includes these sections:
Benefits
Vocabularies and Datasets
Relevant Technologies
Implementation challenges
Recommendations
Two related parts are:
Use Cases, a survey report describing existing projects
http://www.w3.org/2005/Incubator/lld/wiki/UseCaseReport
Vocabularies and Datasets, a survey report
http://www.w3.org/2005/Incubator/lld/wiki/Vocabulary_and_Dataset
The LLD XG invite comments from interested members of the public.
Feedback can sent as comments to individual sections posted on the dedicated blog at http://blogs.ukoln.ac.uk/w3clld/ or by email to the public mailing list (public-lld@w3.org, archived at http://lists.w3.org/Archives/Public/public-lld/ ) using descriptive subject lines such as '[COMMENTS] "Benefits" section.'
Comments are especially welcome in the next four weeks (through 22 July). Reviewers should note that as with Wikipedia, the text may be revised and corrected by its editors in response to comments at any time, but that earlier versions of a document may be viewed by clicking on the History tab.
It is anticipated that the three reports will be published in final form by 31 August.
I quote here one line from the benefits section. "The Linked Data approach offers significant advantages over current practices for creating and delivering library data while providing a natural extension to the collaborative sharing models historically employed by libraries, archives, and museums ("memory institutions")." And one more from the benefits to "memory institutions." "By using Linked Data, memory institutions will create an open, global pool of shared data that can be used and re-used to describe resources, with a limited amount of redundant effort compared with current cataloguing processes." Cheaper, faster, better.
From a cataloger's viewpoint, library linked data is the ultimate cooperative cataloging environment and the ultimate user services environment.
The report
http://www.w3.org/2005/Incubator/lld/wiki/DraftReportWithTransclusion
includes these sections:
Benefits
Vocabularies and Datasets
Relevant Technologies
Implementation challenges
Recommendations
Two related parts are:
Use Cases, a survey report describing existing projects
http://www.w3.org/2005/Incubator/lld/wiki/UseCaseReport
Vocabularies and Datasets, a survey report
http://www.w3.org/2005/Incubator/lld/wiki/Vocabulary_and_Dataset
The LLD XG invite comments from interested members of the public.
Feedback can sent as comments to individual sections posted on the dedicated blog at http://blogs.ukoln.ac.uk/w3clld/ or by email to the public mailing list (public-lld@w3.org, archived at http://lists.w3.org/Archives/Public/public-lld/ ) using descriptive subject lines such as '[COMMENTS] "Benefits" section.'
Comments are especially welcome in the next four weeks (through 22 July). Reviewers should note that as with Wikipedia, the text may be revised and corrected by its editors in response to comments at any time, but that earlier versions of a document may be viewed by clicking on the History tab.
It is anticipated that the three reports will be published in final form by 31 August.
Monday, May 23, 2011
Library of Congress: Bibliographic Framework Transition Initiative
The Library of Congress has begun a "Bibliographic Framework Transition Initiative."
"A major focus of the initiative will be to determine a transition path for the MARC 21 exchange format in order to reap the benefits of newer technology while preserving a robust data exchange that has supported resource sharing and cataloging cost savings in recent decades."
"This work will be carried out in consultation with the format's formal partners -- Library and Archives Canada and the British Library -- and informal partners -- the Deutsche Nationalbibliothek and other national libraries, the agencies that provide library services and products, the many MARC user institutions, and the MARC advisory committees such as the MARBI committee of ALA, the Canadian Committee on MARC, and the BIC Bibliographic Standards Group in the UK."
This could make the RDA effort look like a piece of cake. How will the process be arranged to include these players?
Good luck, though. Sounds like fun to me. Let's get to work.
The press release has more information.
"A major focus of the initiative will be to determine a transition path for the MARC 21 exchange format in order to reap the benefits of newer technology while preserving a robust data exchange that has supported resource sharing and cataloging cost savings in recent decades."
"This work will be carried out in consultation with the format's formal partners -- Library and Archives Canada and the British Library -- and informal partners -- the Deutsche Nationalbibliothek and other national libraries, the agencies that provide library services and products, the many MARC user institutions, and the MARC advisory committees such as the MARBI committee of ALA, the Canadian Committee on MARC, and the BIC Bibliographic Standards Group in the UK."
This could make the RDA effort look like a piece of cake. How will the process be arranged to include these players?
Good luck, though. Sounds like fun to me. Let's get to work.
The press release has more information.
Labels:
FRBR,
Library of Congress,
linked data,
MARC,
RDA,
Semantic Web
Monday, November 15, 2010
Open Bibliographic Data Guide: JISC study on the business cases for Open Bibliographic Data
http://obd.jisc.ac.uk/
links to the _The Guide to Open Bibliographic Data_ that JISC developed on behalf of its partners in the Resource Discovery Task Force. It is about the business cases for Open Bibliographic Data – releasing some or all of a library’s catalog records for open use and re-use by others. The Guide uses 17 use cases to explore
* How to license the data
* Legal issues to be considered
* Potential costs and savings
* Practical implications in terms of processes, effort and skills
* Data formats and other technical options
The assumed rationale is about discoverability and is gaining in credibility the more our resources are discovered from ‘out there’ (through such as Google) and not from ‘in here’ (through the local OPAC). --most of the above quoted or modified slightly from the Guide.
A PDF version of the use cases is at http://obd.jisc.ac.uk/wp-content/uploads/2010/11/Open-Bibliographic-Data-The-Use-Cases.pdf
On a quick review, the case studies seem useful: specific, brief, comprehensive. Each use case includes sections on description, motivation, benefits, consequences, rights & licensing, practicalities and costs.
A table of use cases and examples is at http://obd.jisc.ac.uk/examples
An example for use case 1 (publish data for unspecified use) is Open Library http://openlibrary.org and another is Cambridge U. Library http://openbiblio.net/2010/10/05/jisc-openbibliography-cul-data-release/
An example for use case 2 (publish open Linked Data for unspecified use) is Libris, the joint catalogue of the Swedish academic and research libraries http://libris.kb.se/
links to the _The Guide to Open Bibliographic Data_ that JISC developed on behalf of its partners in the Resource Discovery Task Force. It is about the business cases for Open Bibliographic Data – releasing some or all of a library’s catalog records for open use and re-use by others. The Guide uses 17 use cases to explore
* How to license the data
* Legal issues to be considered
* Potential costs and savings
* Practical implications in terms of processes, effort and skills
* Data formats and other technical options
The assumed rationale is about discoverability and is gaining in credibility the more our resources are discovered from ‘out there’ (through such as Google) and not from ‘in here’ (through the local OPAC). --most of the above quoted or modified slightly from the Guide.
A PDF version of the use cases is at http://obd.jisc.ac.uk/wp-content/uploads/2010/11/Open-Bibliographic-Data-The-Use-Cases.pdf
On a quick review, the case studies seem useful: specific, brief, comprehensive. Each use case includes sections on description, motivation, benefits, consequences, rights & licensing, practicalities and costs.
A table of use cases and examples is at http://obd.jisc.ac.uk/examples
An example for use case 1 (publish data for unspecified use) is Open Library http://openlibrary.org and another is Cambridge U. Library http://openbiblio.net/2010/10/05/jisc-openbibliography-cul-data-release/
An example for use case 2 (publish open Linked Data for unspecified use) is Libris, the joint catalogue of the Swedish academic and research libraries http://libris.kb.se/
Wednesday, November 10, 2010
Understanding linked data and its potential for libraries
I've been hearing a lot about linked data in the past year, and I find that I'm very fuzzy about what linked data is and how it matters or might matter to libraries, to organizations that have libraries and to people who may use libraries.
My first question is What is linked data? A good starting place for me is the definition at http://linkeddata.org
"Linked Data is about using the Web to connect related data that wasn't previously linked, or using the Web to lower the barriers to linking data currently linked using other methods. More specifically, Wikipedia defines Linked Data as 'a term used to describe a recommended best practice for exposing, sharing, and connecting pieces of data, information, and knowledge on the Semantic Web using URIs and RDF.'" --linkeddata.org, read Nov. 10, 2010.
I checked wikipedia for a definition and found a slightly different, more technical definition of linked data.
"Linked Data is a sub-topic of the Semantic Web. The term Linked Data is used to describe a method of exposing, sharing, and connecting data via dereferenceable URIs on the Web." --wikipedia, read on Nov. 10, 2010.
Of course, I had no idea what "dereferenceable URIs" are. Well, a dereferenceable URI is the normal and obvious way that links on the Web work: a URI refers to a page that the web server returns a copy of.
I'm in a technical vocabulary thicket,and I don't want to be. That may be useful later, but not now. I need to put it in my own words or into words I understand.
Let me try working with that definition cited by linkeddata.org: "... a recommended best practice for exposing, sharing, and connecting pieces of data, information, and knowledge on the Semantic Web using URIs and RDF.'
Linked data is a way to expose, share, or connect to data on the Web so that the data can be understood or is meaningful to other machines on the Web. This is how Linked Data is a sub-topic of the Semantic Web. Additionally, Linked Data uses URIs as names for things and RDF as the data model so that statements about resources (in particular Web resources)are made in the form of subject-predicate-object expressions, and these expressions are known as triples.
Well, that is making sense to me, but I don't know that it would make much sense to anyone else, or be seen by anyone else as an improvement over the other available definitions. It helps me, though. That is enough for now.
My first question is What is linked data? A good starting place for me is the definition at http://linkeddata.org
"Linked Data is about using the Web to connect related data that wasn't previously linked, or using the Web to lower the barriers to linking data currently linked using other methods. More specifically, Wikipedia defines Linked Data as 'a term used to describe a recommended best practice for exposing, sharing, and connecting pieces of data, information, and knowledge on the Semantic Web using URIs and RDF.'" --linkeddata.org, read Nov. 10, 2010.
I checked wikipedia for a definition and found a slightly different, more technical definition of linked data.
"Linked Data is a sub-topic of the Semantic Web. The term Linked Data is used to describe a method of exposing, sharing, and connecting data via dereferenceable URIs on the Web." --wikipedia, read on Nov. 10, 2010.
Of course, I had no idea what "dereferenceable URIs" are. Well, a dereferenceable URI is the normal and obvious way that links on the Web work: a URI refers to a page that the web server returns a copy of.
I'm in a technical vocabulary thicket,and I don't want to be. That may be useful later, but not now. I need to put it in my own words or into words I understand.
Let me try working with that definition cited by linkeddata.org: "... a recommended best practice for exposing, sharing, and connecting pieces of data, information, and knowledge on the Semantic Web using URIs and RDF.'
Linked data is a way to expose, share, or connect to data on the Web so that the data can be understood or is meaningful to other machines on the Web. This is how Linked Data is a sub-topic of the Semantic Web. Additionally, Linked Data uses URIs as names for things and RDF as the data model so that statements about resources (in particular Web resources)are made in the form of subject-predicate-object expressions, and these expressions are known as triples.
Well, that is making sense to me, but I don't know that it would make much sense to anyone else, or be seen by anyone else as an improvement over the other available definitions. It helps me, though. That is enough for now.
Wednesday, October 20, 2010
Ed Summers on the linked data release from Deutschen Nationalbibliothek
See Ed Summers' comments on the DNB release of linked library data at
http://inkdroid.org/journal/2010/10/19/linked-library-data-at-the-deutschen-nationalbibliothek/
Summers' piece is worth reading.
Here are some highlights.
The Deutschen Nationalbibliothek (DNB)has released linked library data for
■1.8 million authors from the Personennamendatei (PND)
■1.3 million corporate bodies from the Gemeinsame Körperschaftsdatei (GKD)
■187,000 subject headings from the Schlagwortnormdatei (SWD)
■51,000 Dewey Decimal Classification categories
The full dataset that the DNB has made available for download amounts to 38,849,113 individual statements (aka triples).
See the DNB announcement at http://lists.w3.org/Archives/Public/public-lod/2010Oct/0016.html
This is a huge event for _library_ uses of linked data, and exemplary behavior from DNB. Other research and national libraries should emulate the DNB.
Summers cites Herta Müller's authority information as an illustration and notes the use of RDA vocabularies, which are also available as linked data. "RDF vocabularies are explicit ways of describing resources like people, places, topics, etc. When different things are described using the same vocabulary (or the vocabularies themselves are related together in a particular way) it becomes possible to merge the descriptions, and build software on top of it."
"Another really interesting thing to note about this RDF for Herta Müller are the links to Wikipedia (http://de.wikipedia.org/wiki/Herta_M%C3%BCller), VIAF (http://viaf.org/viaf/12324250) and dbpedia (http://dbpedia.org/resource/Herta_M%C3%BCller). These are important because they contextualize the DNB record for Herta Müller by relating it to other records for her, thus allowing it to be disambiguated from records describing other people named Herta Müller."
"[Summers] did some quick and dirty analysis of the full data dump from the DNB and found: 3,569,402 links to VIAF and 40,136 links to dbpedia (the Linked Data version of Wikipedia)."
Summers goes on to talk a bit about what more needs to be done.
"What remains to be done to some extent is leveraging this contextual information around our data in Library Applications, both cataloging, metadata enrichment applications and end user facing discovery applications."
There is a lot more in his piece, and links to many related tools, projects, and activities.
http://inkdroid.org/journal/2010/10/19/linked-library-data-at-the-deutschen-nationalbibliothek/
Summers' piece is worth reading.
Here are some highlights.
The Deutschen Nationalbibliothek (DNB)has released linked library data for
■1.8 million authors from the Personennamendatei (PND)
■1.3 million corporate bodies from the Gemeinsame Körperschaftsdatei (GKD)
■187,000 subject headings from the Schlagwortnormdatei (SWD)
■51,000 Dewey Decimal Classification categories
The full dataset that the DNB has made available for download amounts to 38,849,113 individual statements (aka triples).
See the DNB announcement at http://lists.w3.org/Archives/Public/public-lod/2010Oct/0016.html
This is a huge event for _library_ uses of linked data, and exemplary behavior from DNB. Other research and national libraries should emulate the DNB.
Summers cites Herta Müller's authority information as an illustration and notes the use of RDA vocabularies, which are also available as linked data. "RDF vocabularies are explicit ways of describing resources like people, places, topics, etc. When different things are described using the same vocabulary (or the vocabularies themselves are related together in a particular way) it becomes possible to merge the descriptions, and build software on top of it."
"Another really interesting thing to note about this RDF for Herta Müller are the links to Wikipedia (http://de.wikipedia.org/wiki/Herta_M%C3%BCller), VIAF (http://viaf.org/viaf/12324250) and dbpedia (http://dbpedia.org/resource/Herta_M%C3%BCller). These are important because they contextualize the DNB record for Herta Müller by relating it to other records for her, thus allowing it to be disambiguated from records describing other people named Herta Müller."
"[Summers] did some quick and dirty analysis of the full data dump from the DNB and found: 3,569,402 links to VIAF and 40,136 links to dbpedia (the Linked Data version of Wikipedia)."
Summers goes on to talk a bit about what more needs to be done.
"What remains to be done to some extent is leveraging this contextual information around our data in Library Applications, both cataloging, metadata enrichment applications and end user facing discovery applications."
There is a lot more in his piece, and links to many related tools, projects, and activities.
Friday, March 19, 2010
OCLC report: Implications of MARC tag usage on library metadata practices
OCLC has published a report called, Implications of MARC Tag Usage on Library Metadata Practices[pdf]
I've only begun to read it.
The "implications" section of the Exec. Summ. are interesting.
1. Be consistent. Splitting content across multiple fields will negatively affect indexing, retrieval, and mapping to other encoding schema.
2. Respond to local user needs whether that is counting plates in a book or adding contents notes.
3. Focus on authorized names, classifications, and controlled vocabularies that key word searching of full-text will not provide (as full text online negates some value of descriptive surrogates).
4. Use specific MARC fields for particular types of note if they are available rather than the general 500 note.
5. Map the 200 or so MARC 21 fields in use to simpler schema. (MARC data cannot continue to exist in its own discrete environment, separate from the rest of the information universe. Leverage it and use it in other domains to reach users in heir own networked environments.)
6. Accuracy of fields that are used in machine matching becomes more important in environments using linked data to leverage fuller descriptions and other related information generated from other sources.
And MARC's future:
1. MARC is a niche data communication format approaching the end of its life cycle.
2. Encoding schema will need to robust MARC crosswalks to ingest millions of legacy records.
3. How would we create, capture, structure, store, search, retrieve, and display objects and metadata if we didn’t have to use MARC and if we weren’t limited by current library systems?
4. How do we best take advantage of linked data and avoid creating the same redundant metadata in individual records?
5. How do we integrate library metadata with sources outside the traditional library environment?
6. To meet the demands of the rest of the information universe, give priority to interoperability with other encoding schema and systems.
I've only begun to read it.
The "implications" section of the Exec. Summ. are interesting.
1. Be consistent. Splitting content across multiple fields will negatively affect indexing, retrieval, and mapping to other encoding schema.
2. Respond to local user needs whether that is counting plates in a book or adding contents notes.
3. Focus on authorized names, classifications, and controlled vocabularies that key word searching of full-text will not provide (as full text online negates some value of descriptive surrogates).
4. Use specific MARC fields for particular types of note if they are available rather than the general 500 note.
5. Map the 200 or so MARC 21 fields in use to simpler schema. (MARC data cannot continue to exist in its own discrete environment, separate from the rest of the information universe. Leverage it and use it in other domains to reach users in heir own networked environments.)
6. Accuracy of fields that are used in machine matching becomes more important in environments using linked data to leverage fuller descriptions and other related information generated from other sources.
And MARC's future:
1. MARC is a niche data communication format approaching the end of its life cycle.
2. Encoding schema will need to robust MARC crosswalks to ingest millions of legacy records.
3. How would we create, capture, structure, store, search, retrieve, and display objects and metadata if we didn’t have to use MARC and if we weren’t limited by current library systems?
4. How do we best take advantage of linked data and avoid creating the same redundant metadata in individual records?
5. How do we integrate library metadata with sources outside the traditional library environment?
6. To meet the demands of the rest of the information universe, give priority to interoperability with other encoding schema and systems.
Thursday, February 18, 2010
Karen Coyle on libraries, metadata, and the Semantic Web
Karen Coyle's article "Understanding the Semantic Web: bibliographic data and metadata" may be of wide interest in libraries. It is about 30 p. and is the full content of _Library Technology Reports_ v. 46, issue 1, Jan. 2010. (This is available online at Yale via Gale Cengage Academic One File.) The first chapter (of two) is available online at http://alatechsource.metapress.com/content/p3022442071g7655/fulltext.pdf
Coyle briefly places traditional (i.e. Panizzi and since) library bibliographic data in its contemporary context (Internet, etc.) and calls for changing our traditions to fit the new context. She then discusses several steps necessary to make the change from what is more or less textual descriptions (e.g. catalog cards, MARC records) to sets of data elements that can be processed by computers (e.g. linked data in the semantic web.)
Overall, this is a nice brief on library metadata and its relation with the semantic web, RDF (Resource Description Framework), identifiers, and linked data. If you already know this, Coyle's article is worth the review. If you don't already know this, Coyle's article is a good place to begin.
Coyle briefly places traditional (i.e. Panizzi and since) library bibliographic data in its contemporary context (Internet, etc.) and calls for changing our traditions to fit the new context. She then discusses several steps necessary to make the change from what is more or less textual descriptions (e.g. catalog cards, MARC records) to sets of data elements that can be processed by computers (e.g. linked data in the semantic web.)
Overall, this is a nice brief on library metadata and its relation with the semantic web, RDF (Resource Description Framework), identifiers, and linked data. If you already know this, Coyle's article is worth the review. If you don't already know this, Coyle's article is a good place to begin.
Tuesday, October 20, 2009
More on Linked Data: Linked Data Design note
Tim Berners-Lee linked data design note.
http://www.w3.org/DesignIssues/LinkedData.html
A brief and readable note on linked data from TBL. From 2006, but this makes a nice primer on linked data.
Four rules for linked data:
1. Use URIs as names for things
2. Use HTTP URIs so that people can look up those names.
3. When someone looks up a URI, provide useful information, using the standards (RDF, SPARQL)
4. Include links to other URIs. so that they can discover more things.
http://www.w3.org/DesignIssues/LinkedData.html
A brief and readable note on linked data from TBL. From 2006, but this makes a nice primer on linked data.
Four rules for linked data:
1. Use URIs as names for things
2. Use HTTP URIs so that people can look up those names.
3. When someone looks up a URI, provide useful information, using the standards (RDF, SPARQL)
4. Include links to other URIs. so that they can discover more things.
Tuesday, October 13, 2009
The Law of Linked Data
http://www.mkbergman.com/837/the-law-of-linked-data/
Mike Bergman at AI3 blogs about the Law of Linked Data. "The Linked Data Law: the value of a linked data network is proportional to the square of the number of links between data objects."
He argues that linking (meaningfully) the existing nodes on the 'net produces network effects for the semantic web. He makes a nice analogy with Metcalfe’s law, which "states that the value of a telecommunications network is proportional to the square of the number of users of the system.
Bergman thinks a good marshal would deliver law and order to linked data and the semantic enterprise. What would the good marshal do? Well, he doesn't say beyond "deliver law and order." What the hell does that mean? OK, aside from that the piece is worth reading and it links to more good reading, too. Are footnotes the original linked data?
Mike Bergman at AI3 blogs about the Law of Linked Data. "The Linked Data Law: the value of a linked data network is proportional to the square of the number of links between data objects."
He argues that linking (meaningfully) the existing nodes on the 'net produces network effects for the semantic web. He makes a nice analogy with Metcalfe’s law, which "states that the value of a telecommunications network is proportional to the square of the number of users of the system.
Bergman thinks a good marshal would deliver law and order to linked data and the semantic enterprise. What would the good marshal do? Well, he doesn't say beyond "deliver law and order." What the hell does that mean? OK, aside from that the piece is worth reading and it links to more good reading, too. Are footnotes the original linked data?
Friday, September 25, 2009
VIAF now available as linked data.
Thom Hickey of OCLC posted on the VIAF as linked data on his blog.
Search the VIAF beta at http://viaf.org/
Thom says,
There are some 9.5 million personae described in VIAF and have established more than 4 million links between the files. To us linked data means:
URIs for everything
HTTP 303 redirects for URIs representing the personae our metadata is about
HTTP content negotiation for different data formats
An RDF view of the data
A rich a set of internal and external links in our data
Search the VIAF beta at http://viaf.org/
Thom says,
There are some 9.5 million personae described in VIAF and have established more than 4 million links between the files. To us linked data means:
URIs for everything
HTTP 303 redirects for URIs representing the personae our metadata is about
HTTP content negotiation for different data formats
An RDF view of the data
A rich a set of internal and external links in our data
Subscribe to:
Posts (Atom)
