BigWig and BigBed: enabling browsing of large distributed datasets

There is no author summary for this article yet. Authors can add summaries to their articles on ScienceOpen to make them more accessible to a non-specialist audience.

Abstract

Summary: BigWig and BigBed files are compressed binary indexed files containing data at several resolutions that allow the high-performance display of next-generation sequencing experiment results in the UCSC Genome Browser. The visualization is implemented using a multi-layered software approach that takes advantage of specific capabilities of web-based protocols and Linux and UNIX operating systems files, R trees and various indexing and compression tricks. As a result, only the data needed to support the current browser view is transmitted rather than the entire file, enabling fast remote access to large distributed data sets.

Availability and implementation: Binaries for the BigWig and BigBed creation and parsing utilities may be downloaded at http://hgdownload.cse.ucsc.edu/admin/exe/linux.x86_64/. Source code for the creation and visualization software is freely available for non-commercial use at http://hgdownload.cse.ucsc.edu/admin/jksrc.zip, implemented in C and supported on Linux. The UCSC Genome Browser is available at http://genome.ucsc.edu

Contact: ann@ 123456soe.ucsc.edu

Supplementary information: Supplementary byte-level details of the BigWig and BigBed file formats are available at Bioinformatics online. For an in-depth description of UCSC data file formats and custom tracks, see http://genome.ucsc.edu/FAQ/FAQformat.html and http://genome.ucsc.edu/goldenPath/help/hgTracksHelp.html

Related collections

Most cited references 5

Record: found
Abstract: found
Article: found

Is Open Access

The UCSC Genome Browser database: update 2010

Brooke Rhead, Donna Karolchik, Robert M. Kuhn … (2010)

The University of California, Santa Cruz (UCSC) Genome Browser website (http://genome.ucsc.edu/) provides a large database of publicly available sequence and annotation data along with an integrated tool set for examining and comparing the genomes of organisms, aligning sequence to genomes, and displaying and sharing users’ own annotation data. As of September 2009, genomic sequence and a basic set of annotation ‘tracks’ are provided for 47 organisms, including 14 mammals, 10 non-mammal vertebrates, 3 invertebrate deuterostomes, 13 insects, 6 worms and a yeast. New data highlights this year include an updated human genome browser, a 44-species multiple sequence alignment track, improved variation and phenotype tracks and 16 new genome-wide ENCODE tracks. New features include drag-and-zoom navigation, a Wiki track for user-added annotations, new custom track formats for large datasets (bigBed and bigWig), a new multiple alignment output tool, links to variation and protein structure tools, in silico PCR utility enhancements, and improved track configuration tools.

0 comments Cited 260 times – based on 0 reviews      Review now

Bookmark

Record: found
Abstract: not found
Article: not found

R-trees

Antonin Guttman (1984)

0 comments Cited 97 times – based on 0 reviews      Review now

Bookmark

Record: found
Abstract: found
Article: not found

Nested Containment List (NCList): a new algorithm for accelerating interval query of genome alignment and interval databases.

Alexander V. Alekseyenko, Brian J. Lee (2007)

The exponential growth of sequence databases poses a major challenge to bioinformatics tools for querying alignment and annotation databases. There is a pressing need for methods for finding overlapping sequence intervals that are highly scalable to database size, query interval size, result size and construction/updating of the interval database. We have developed a new interval database representation, the Nested Containment List (NCList), whose query time is O(n + log N), where N is the database size and n is the size of the result set. In all cases tested, this query algorithm is 5-500-fold faster than other indexing methods tested in this study, such as MySQL multi-column indexing, MySQL binning and R-Tree indexing. We provide performance comparisons both in simulated datasets and real-world genome alignment databases, across a wide range of database sizes and query interval widths. We also present an in-place NCList construction algorithm that yields database construction times that are approximately 100-fold faster than other methods available. The NCList data structure appears to provide a useful foundation for highly scalable interval database applications. NCList data structure is part of Pygr, a bioinformatics graph database library, available at http://sourceforge.net/projects/pygr

0 comments Cited 19 times – based on 0 reviews      Review now

Bookmark

All references

Author and article information

Journal

Journal ID (nlm-ta): Bioinformatics

Journal ID (publisher-id): bioinformatics

Journal ID (hwp): bioinfo

Title: Bioinformatics

Publisher: Oxford University Press

ISSN (Print): 1367-4803

ISSN (Electronic): 1367-4811

Publication date (Print): 1 September 2010

Publication date (Electronic): 17 July 2010

Publication date PMC-release: 17 July 2010

Volume: 26

Issue: 17

Pages: 2204-2207

Affiliations

Center for Biomolecular Science and Engineering, School of Engineering, University of California, Santa Cruz (UCSC), Santa Cruz, CA 95064, USA

Author notes

* To whom correspondence should be addressed.

Associate Editor: Jonathan Wren

Article

Publisher ID: btq351

DOI: 10.1093/bioinformatics/btq351

PMC ID: 2922891

PubMed ID: 20639541

SO-VID: 9157eb6c-bd31-4c82-83de-db851c0d76a5

License:

This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License ( http://creativecommons.org/licenses/by-nc/2.5), which permits unrestricted non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited.

History

Date received : 18 February 2010

Date revision received : 10 June 2010

Date accepted : 28 June 2010

Comments

Comment on this article

scite_

Cited by 560

See all cited by

Most referenced authors 328

See all reference authors

BigWig and BigBed: enabling browsing of large distributed datasets

Read this article at

Abstract

Related collections

Genetoberfest

Most cited references 5

The UCSC Genome Browser database: update 2010

R-trees

Nested Containment List (NCList): a new algorithm for accelerating interval query of genome alignment and interval databases.

Author and article information

Journal

Affiliations

Author notes

Article

History

Categories

Comments

Comment on this article

Similar content 24

Cited by 560

Most referenced authors 328