Repository logo

Infoscience

  • English
  • French
Log In
Logo EPFL, École polytechnique fédérale de Lausanne

Infoscience

  • English
  • French
Log In
  1. Home
  2. Academic and Research Output
  3. Journal articles
  4. A privacy-preserving solution for compressed storage and selective retrieval of genomic data
 
research article

A privacy-preserving solution for compressed storage and selective retrieval of genomic data

Huang, Zhicong  
•
Ayday, Erman  
•
Lin, Huang  
Show more
October 27, 2016
Genome Research

In clinical genomics, the continuous evolution of bioinformatic algorithms and sequencing platforms makes it beneficial to store patients' complete aligned genomic data in addition to variant calls relative to a reference sequence. Due to the large size of human genome sequence data files (varying from 30 GB to 200 GB depending on coverage), two major challenges facing genomics laboratories are the costs of storage and the efficiency of the initial data processing. In addition, privacy of genomic data is becoming an increasingly serious concern, yet no standard data storage solutions exist that enable compression, encryption, and selective retrieval. Here we present a privacy-preserving solution named SECRAM (elective retrieval on Encrypted and Compressed Reference oriented Alignment Map) for the secure storage of compressed aligned genomic data. Our solution enables selective retrieval of encrypted data and improves the efficiency of downstream analysis (e.g., variant calling). Compared with BAM, the de facto standard for storing aligned genomic data, SECRAM uses 18% less storage. Compared with CRAM, one of the most compressed nonencrypted formats (using 34% less storage than BAM), SECRAM maintains efficient compression and downstream data processing, while allowing for unprecedented levels of security in genomic data storage. Compared with previous work, the distinguishing features of SECRAM are that (1) it is position-based instead of read-based, and (2) it allows random querying of a subregion from a BAM-like file in an encrypted form. Our method thus offers a space-saving, privacy-preserving, and effective solution for the storage of clinical genomic data.

  • Files
  • Details
  • Metrics
Type
research article
DOI
10.1101/gr.206870.116
Web of Science ID

WOS:000389563000006

Author(s)
Huang, Zhicong  
Ayday, Erman  
Lin, Huang  
Aiyar, Raeka S.
Molyneaux, Adam
Xu, Zhenyu
Fellay, Jacques  
Steinmetz, Lars M.
Hubaux, Jean-Pierre  
Date Issued

2016-10-27

Publisher

Cold Spring Harbor Laboratory

Published in
Genome Research
Volume

26

Issue

12

Start page

1687

End page

1696

Note

This article is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License

Editorial or Peer reviewed

REVIEWED

Written at

EPFL

EPFL units
LDS  
UPFELLAY  
Available on Infoscience
January 24, 2017
Use this identifier to reference this record
https://infoscience.epfl.ch/handle/20.500.14299/133573
Logo EPFL, École polytechnique fédérale de Lausanne
  • Contact
  • infoscience@epfl.ch

  • Follow us on Facebook
  • Follow us on Instagram
  • Follow us on LinkedIn
  • Follow us on X
  • Follow us on Youtube
AccessibilityLegal noticePrivacy policyCookie settingsEnd User AgreementGet helpFeedback

Infoscience is a service managed and provided by the Library and IT Services of EPFL. © EPFL, tous droits réservés