Showing posts with label Conference. Show all posts
Showing posts with label Conference. Show all posts

Saturday, March 28, 2009

Video about IBM ProtecTIER data deduplication

IBM has published a marketing video about their ProtecTIER data deduplication system recorded at the Pulse09 conference in February:

Key message: It is scalable. But the video contains 3 minutes of marketing stuff without much real information.

What I really find more interessing: At the SYSTOR'09 conference (one of the interessing talks I mentioned here) will be a research talk about the technology and concepts behind the ProtecTIER system, which is based on the product from the company Diligent that IBM bought April 2008. Abstract:

We describe some of the design choices that were made during the development of the IBM TS7650G ProtecTier, a fast, scalable, inline, deduplication device. The system's design goals and how they were achieved are presented. This is the first and only deduplication device that uses similarity matching. The paper provides the following original research contributions: we show how similarity signatures can serve in a deduplication scheme; a novel type of similarity signatures is presented and its advantages in the context of deduplication requirements are explained.
It is also shown how to combine similarity matching schemes with hash based identity schemes.
I really look forward to this talk. Especially how the delimit their approach in comparision to approaches like DERD, DeepStore and other.

First paper accepted

My first paper has been accepted for publication at the SYSTOR'09 conference that takes place in Haifa at May 4-6.

It is based on the first part of my master thesis, but the contents has been extended and revised afterwards:

Data deduplication systems detect redundancies between data blocks to either reduce storage needs or to reduce network traffic. A class of deduplication systems splits the data stream into data blocks (chunks) and then finds exact duplicates of these blocks.

This paper compares the influence of different chunking approaches on multiple levels. On a macroscopic level, we compare the chunking approaches based on real-live user data in a weekly full backup scenario, both at a single point in time as well as over several weeks.

In addition, we analyze how small changes affect the deduplication ratio for different file types on a microscopic level for chunking approaches and delta encoding. An intuitive assumption is that small semantic changes on documents cause only small modifications in the binary representation of files, which would imply a high ratio of deduplication. We will show that this assumption is not valid for many important file types and that application specific chunking can help to further decrease storage capacity demands.

I really look forward to that conference because surprisingly many talks in the program look really interesting and it is my first chance to meet storage researchers outside the Fürstenallee.