Accelerating Columnar Storage Based on Asynchronous Skipping Strategy

Big Data Research - Tập 31 - Trang 100352 - 2023
Wenhai Li1,2, Zheng Yang1,2, Lingfeng Deng1,2, Zhiling Cheng1,2, Weidong Wen1, Yanxiang He1
1Wuhan University, Wuhan, Hubei, 430072, PR China
2Pengcheng Laboratory, Shenzhen, Guangdong, PR China

Tài liệu tham khảo

Ailamaki, 2002 Vermeij, 2008, MonetDB, a novel spatial column-store DBMS Scholl, 1987, Supporting flat relations by a nested relational kernel, 137 Paul, 1987, Architecture and implementation of the Darmstadt database kernel system, 196 Behzad, 2015, Pattern-driven parallel I/O tuning, 43 Behzad, 2013, Taming parallel I/O complexity with auto-tuning, 68 Mane, 2014 Liu, 2015, Hierarchical collective I/O scheduling for high-performance computing, Big Data Res., 2, 117, 10.1016/j.bdr.2015.01.007 Broneske, 2017, Accelerating multi-column selection predicates in main-memory - the elf approach, 647 He, 2012, Efficient iceberg query evaluation using compressed bitmap index, IEEE Trans. Knowl. Data Eng., 24, 1570, 10.1109/TKDE.2011.73 Wen, 2019, CORES: towards scan-optimized columnar storage for nested records, ACM Trans. Storage, 15, 16, 10.1145/3321704 Wang, 2017, Exploiting common patterns for tree-structured data, 883 Amur, 2013, Memory-efficient GroupBy-Aggregate using Compressed Buffer Trees, 1 Shvachko, 2010, The Hadoop distributed file system, 1 Dean, 2004, MapReduce: simplified data processing on large clusters, 10 Zaharia, 2012, Resilient distributed datasets: a fault-tolerant abstraction for in-memory cluster computing, 2 Behm, 2014, Storage management in asterixDB, Proc. VLDB Endow., 7, 841, 10.14778/2732951.2732958 Alsubaiee, 2014, AsterixDB: a scalable, open source BDMS, Proc. VLDB Endow., 7, 1905, 10.14778/2733085.2733096 Yu, 2009, DryadLINQ: a system for general-purpose distributed data-parallel computing using a high-level language, 1 Melnik, 2010, Dremel: interactive analysis of web-scale datasets, Commun. ACM, 3, 114 Sun, 2014, A partitioning framework for aggressive data skipping, Proc. VLDB Endow., 7, 1617, 10.14778/2733004.2733044 Rumbold, 2018, What are data? A categorization of the data sensitivity spectrum, Big Data Res., 12, 49, 10.1016/j.bdr.2017.11.001 Tsirogiannis, 2009, Query processing techniques for solid state drives, 59 Lin, 2014, Migratory compression: coarse-grained data reordering to improve compressibility, 257 Schindler, 2011, Improving throughput for small disk requests with proximal I/O, 133 Lamb, 2012, The vertica analytic database: C-store 7 years later, Comput. Sci., 5 Borkar, 2011, Hyracks: a flexible and extensible foundation for data-intensive computing, 1151 Stonebraker, 2005, C-store: a column-oriented DBMS, 553 Tatarowicz, 2012, Lookup tables: fine-grained partitioning for distributed databases, 102 Welch, 2008, Scalable performance of the Panasas parallel file system, 2 Ahn, 2019, Cache-aware block allocation for memory-technology storage targeted file systems, 1424 Lu, 2013, An efficient and compact indexing scheme for large-scale data store, 326 Zhang, 2016, Virtual denormalization via array index reference for main memory OLAP, IEEE Trans. Knowl. Data Eng., 28, 1061, 10.1109/TKDE.2015.2499199 Beyer, 1999, Bottom-up computation of sparse and iceberg cubes, SIGMOD Rec., 28, 359, 10.1145/304181.304214 Lehner, 2001, Fast refresh using mass query optimization, 391 Stockinger, 2001, Design and implementation of bitmap indices for scientific data, 47 Kaufmann, 2013, Storing and processing temporal data in a main memory column store Liu, 2017, Graphene: fine-grained IO management for graph computing, 285 Das, 2015, Query optimization in oracle 12c database in-memory, 1770 Comer, 1987, A vertical partitioning algorithm for relational databases, 30 Afrati, 2014, Storing and querying tree-structured records in Dremel, Proc. VLDB Endow., 7, 1131, 10.14778/2732977.2732987 Maltzahn, 2010, Ceph as a scalable alternative to the Hadoop distributed file system, 38 Shanbhag, 2016, Amoeba: a shape changing storage system for big data, Proc. VLDB Endow., 9, 1569, 10.14778/3007263.3007311 Gupta, 2021, ComBI: compressed binary search tree for approximate k-NN searches in Hamming space, Big Data Res., 25, 10.1016/j.bdr.2021.100223 Harnik, 2013, To Zip or not to Zip: effective resource usage for real-time compression, 229 Lu, 2017, Canopus: enabling extreme-scale data analytics on big HPC storage via progressive refactoring Zhang, 2019, Finesse: fine-grained feature locality based fast resemblance detection for post-deduplication delta compression, 121 Samanta, 2019, Compact and power efficient SEC-DED codec for computer memory, Microsyst. Technol., 1 Zhou, 2019, Fast erasure coding for data storage: a comprehensive study of the acceleration techniques, 317