Monday, March 30, 2009

Spring CNI Plenary

I can finally reveal the mysterious talk I referred to in this comment; it is the opening plenary at CNI's Spring Task Force meeting one week from today. In essence, the talk is a look back at Jeff Rothenberg's 1995 Scientific American article "Ensuring the Longevity of Digital Documents" which asks:
  • What led Jeff to his dire predictions?
  • Would one make the same dire predictions now?
  • If not, what dire predictions would one make, and why?
This talk arose because Cliff Lynch invited me to give a talk at UC Berkeley's School of Information last November. He liked it enough to ask me to give it at CNI and I agreed, thinking I had already written it. But once I thought about it more, and considered the very different audience, I realized that it needed to be almost completely rewritten. I'm still revising it, based on feedback from giving it to the Stanford Library staff.

CNI will post the slides after the talk, and I plan to post here a commentary on them providing links to sources and additional details. You will be able to see how the discussions here were a very valuable resource. Thank you all.

Monday, March 23, 2009

2009 FAST conference

I attended the 2009 FAST Conference in San Francisco almost a month ago. This post was delayed because, as I mentioned in this comment, I've been working on an important talk which draws from recent discussions on this blog. The papers are on the web following Usenix's commendable open access policy. Follow me below the fold for my highlights.

Thursday, January 15, 2009

Postel's Law

In RFC 793 (1981) the late, great Jon Postel laid down one of the basic design principles of the Internet, Postel's Law or the Robustness Principle:

"Be conservative in what you do; be liberal in what you accept from others."

Its important not to lose sight of the fact that digital preservation is on the "accept" side of Postel's Law, but it seems that people often do.

On the Digital Curation Centre Associates mail list, Adil Hasan started a discussion by asking:

"Does anyone know [whether] there has been a study to estimate how many PDF documents do not comply with the PDF standards?"


No-one in the subsequent discussion knew of a comprehensive study, but Sheila Morrissey reported on the results of Portico's use of JHOVE to classify the 9 million PDFs they have received from 68 publishers as one of not well formed, well-formed and not valid, and well-formed and valid. A significant proportion were classified as either not well formed or well-formed and not valid.

These results are not unexpected. It is well known that much of the HTML on the Web fails the W3C validation tests. Indeed, a 2001 study reportedly concluded that less than 1% of it was valid SGML. Alas, I couldn't retrieve the original document via this link, but our experience confirms that much HTML is poorly formed. For this very reason LOCKSS uses a crawler based on work by James Gosling at Sun Microsystems to develop techniques for extracting links from HTML that are very tolerant of malformed input; an application of Postel's Law.

Follow me below the fold to see why, although questions like Adil's are frequently asked, devoting resources to answering them or acting upon the answers is unlikely to help digital preservation.

Sunday, January 4, 2009

Are format specifications important for preservation?

On the Digital Curation Centre Associates mail list, Steven Ranking pointed to the release of Microsoft's specifications for the Office formats under their Open Specification Promise. This sparked a discussion in which two topics were confused; the suitability of Microsoft Office formats for preservation, and the value of the specifications for preservation. As regards the first, I believe that "it became necessary to change the content in order to preserve it" is a very bad idea; we should preserve what's out there without adding cost and losing information by preemptively migrating to a format we believe (normally without evidence) is less doomed. I'm a skeptic about the second; I don't think preserving the specifications contributes anything to practical digital preservation, as I explain below the fold.

Tuesday, December 30, 2008

Persistence of Poor Peer Reviewing

Another thing I've been doing during the hiatus is serving as a judge for Elsevier's Grand Challenge. Anita de Waard and her colleagues at Elsevier's research labs set up this competition with a substantial prize for the best demonstration of what could be done to improve science and scientific communication given unfettered access to Elsevier's vast database of publications. I think the reason I'm on the panel of judges is that after Anita's talk at the Spring CNI she and I had an interesting discussion. The talk described her team's work to extract information from full-text articles to help authors and readers. I asked who was building tools to help reviewers and make their reviews better. This is a bete noir of mine, both because I find doing reviews really hard work, and because I think the quality of reviews (including mine) is really poor. For an ironic example of the problem, follow me below the fold.

Sunday, December 28, 2008

Foot, meet bullet

The gap in posting since March was caused by a bad bout of RSI in my hands. I believe it was triggered by the truly terrible ergonomics of the mouse buttons on the first-generation Asus EEE (which otherwise fully justifies its reputation as a game-changing product). It took a long time to recover and even longer to catch up with all the work I couldn't do when I couldn't type for more than a few minutes.

One achievement during this enforced hiatus was to turn my series of posts on A Petabyte for a Century into a paper entitled Bit Preservation: A Solved Problem? (190KB PDF) and present it at the iPRES 2008 conference last September at the British Library.

I also attended the 4th International Digital Curation Conference in Edinburgh. As usual these days, for obvious reasons, sustainability was at the top of the agenda. Brian Lavoie of OCLC talked (461KB .ppt) about the work of the Blue Ribbon Task Force on Sustainable Digital Preservation and Access which he co-chairs with Fran Berman of the San Diego Supercomputer Center. NSF, the Andrew W. Mellon Foundation and others are sponsoring this effort; the LOCKSS team have presented to the Task Force. Their interim report has just been released.

Listening to Brian talk about the need to persuade funding organizations of the value of digital preservation efforts I came to understand the extent to which the tendency to present simply preserving the bits as a trivial, solved problem has caused the field to shoot itself in the foot.

The activities that the funders are told they need to support are curation-focused, such as generating metadata to prepare for possible format obsolescence, and finding the content for future readers. The problem is that, as a result, the funders see a view of the future in which, even if they do nothing, the bits will survive. There might possibly be problems in the distant future if formats go obsolete, but there might not be. There might be problems finding content in the future, but there might not be. After all, funders might think, if the bits survive and Google can index them, how much worse than the current state could things be? Why should they pour money into activities intended to enhance the data? After all, the future can figure out what to do with the bits when they need them; they'll be there whatever happens.

A more realistic view of the world, as I showed in my iPRES paper, would be that there are huge volumes of data that need to be preserved, that simply storing a few copies of all of it is more costly than we can currently cope with, and that even if we spend enough to use the best available technology we can't be sure the bits will be safe. If this were the view being presented to the funders, that unless they immediately provide funds important information would gradually be lost, they might be scared into actually doing something.

Thursday, March 6, 2008

More bad news on storage reliability

Last year's FAST conference had depressing news for those who think that bits are safe in modern storage systems, with two papers showing that disks in use in large-scale storage facilities are much less reliable than the manufacturers claim, and a keynote (PDF) reporting that errors in file system code are endemic. This year's FAST had more sobering news. I'll return to these papers in more detail, but here are the take-away messages.

Jiang et al from UIUC and NetApp took a detailed look at the various subsystems in modern storage system, showing that 45-75% of the apparent disk unreliability in last year's papers is probably due to the unreliability of other components in the storage system, and that the correlations between errors are even worse than last year's papers suggested.

Gunawi et al from Wisconsin analyzed one of the root causes of the incorrect response of file systems to errors in the underlying storage that was reported in earlier papers from Wisconsin (PDF) and Stanford (PDF), namely the way the file systems propogate reported errors between functions and modules, showing that correct handling of these problems is so hard that implementors often throw up their hands.

Bairavasundaram et al from Wisconsin, NetApp and Toronto presented a massive study of silent data corruption in storage systems, reinforcing the earlier study (PDF) from CERN in showing an alarming incidence of both these errors and of correlations between them.

Krioukov et al from Wisconsin and NetApp analyzed the techniques RAID-based storage systems use to tolerate silent data corruption, showing that in various ways all current systems fall short of an adequate solution to this problem.

Greenan and Wylie (PDF) from HP Labs gave a work-in-progress presentation showing that the Markov models which are pretty much the exclusive technique for analyzing failures in storage systems give results that are systematically optimistic because they depend on assumptions that are known to be untrue.

Considerable kudos is due to NetApp for their many contributions to both data and analysis and to the University of Wisconsin, which is building an impressive track record in this important area.