DWS-AQA: a cost effective approach for very large data warehouses

J. Bernardino1, P. Furtado2, H. Madeira2
1Institute Polytechnic of Coimbra, DEIS, ISEC, Portugal
2DEI, University of Coimbra, Portugal

Tóm tắt

Data warehousing applications typically involve massive amounts of data that push database management technology to the limit. A scalable architecture is crucial, not only to handle very large amount of data but also to assure interactive response time to the users. Large data warehouses require a very expensive setup, typically based on high-end servers or high-performance clusters. In this paper we propose and evaluate a simple but very effective method to implement a data warehouse using the computers and workstations typically available in large organizations. The proposed approach is called data warehouse striping with approximate query answering (DWS-AQA). The goal is to use the processing and disk capacity normally available in large workstation networks to implement a data warehouse with a very reduced infrastructure cost. As the data warehouse shares computers that are also being used for other purposes, most of the times only a fraction of the computers will be able to execute the partial queries in time. However, as we show in the paper, the approximated answers estimated from partial results have a very small error for most of the plausible scenarios. Moreover, as the data warehouse facts are partitioned in a strict uniform way, it is possible to calculate tight confidence intervals for the approximated answers, providing the user with a measure of the accuracy of the query results. A set of experiments on the TPC-H benchmark database is presented to show the accuracy of DWS-AQA for a large number of scenarios.

Từ khóa

#Costs #Data warehouses #Databases #Workstations #Computer networks #Delay #Time sharing computer systems #Concurrent computing #Warehousing #Technology management

Tài liệu tham khảo

cochran, 1977, Sampling Techniques 10.1109/SSDM.1997.621151 10.1145/253260.253291 1998, Informix Decision Support Indexing for the Enterprise Data Warehouses, White Paper lu, 1994, Query Processing in Parallel Relational Database Systems, IEEE Computer Society 10.1145/253260.253268 1997, Star queries in Oracle 8, White Paper ozsu, 1999, Principles of Distributed Database Systems 10.1109/2.970558 10.1145/276304.276309 10.1007/3-540-44801-2_34 agrawal, 2000, Automated Selection of Materialized Views and Indexes for SQL Databases, Proc of the 26th International Conference on Very Large Databases, 496 brobst, 1998, Starburst Grows Bright, DB2 UDB, Database Programming & Design 10.1109/IDEAS.2001.938099 10.1145/375663.375694 catozzi, 2001, Operating Systems Extensions for the Teradata Parallel VLDB, Proc Intl Conf on Very Large Databases, 676 10.1145/342009.335450 abdelguerfi, 1998, Parallel Database Techniques chaudhuri, 1997, An Efficient, Cost-Driven Index Selection Tool for Microsoft SQL Server, Proc of the 23rd VLDB Conference, 146 1998, Star Schema Processing for Complex Queries, White Paper benchmark, 1999, Transaction Processing Council