Integrating induction and deduction for noisy data mining

Information Sciences - Tập 180 Số 14 - Trang 2663-2673 - 2010
Yan Zhang1, Xindong Wu2
1Department of Computer Science, University of Vermont, Burlington, VT 05405, USA
2Department of Computer Science, University of Vermont, Burlington, VT 05405, USA and School of Computer Science and Information Engineering, Hefei University of Technology, Hefei 230009, China#TAB#

Tóm tắt

Từ khóa


Tài liệu tham khảo

Rakesh Agrawal, Ramakrishnan Srikant, Fast algorithms for mining association rules in large databases, in: Proceedings of the 20th International Conference on Very Large Data Bases, Santiago, Chile, 1994, pp. 487–499.

Rakesh Agrawal, Tomasz Imielinski, Arun N. Swami, Mining association rules between sets of items in large databases, in: The 1993 ACM SIGMOD International Conference on Management of Data, Washington, DC, USA, 1993, pp. 207–216.

Anagnostopoulos, 2008, Effective and efficient classification on a search-engine model, Knowledge and Information Systems, 16, 129, 10.1007/s10115-007-0102-6

Assent, 2008, Clustering multidimensional sequences in spatial and temporal databases, Knowledge and Information Systems, 16, 29, 10.1007/s10115-007-0121-3

Brad M. Barber, Terrance Odean, Ning Zhu, Systematic noise, in: AFA, San Diego, 2004.

Breiman, 1984

Chen, 2007, Enriching the ER model based on discovered association rules, Information Sciences, 177, 1558, 10.1016/j.ins.2006.07.001

Cheng, 2008, A survey on algorithms for mining frequent itemsets over data streams, Knowledge and Information Systems, 16, 1, 10.1007/s10115-007-0092-4

Creighton, 2003, Mining gene expression databases for association rules, Bioinformatics, 19, 79, 10.1093/bioinformatics/19.1.79

Fisher, 1996, Iterative optimization and simplification of hierarchical clusterings, Journal of Artifical Intelligence Research, 4, 147, 10.1613/jair.276

Dunnet, 1955, A multiple comparison procedure for comparing several treatments with a control, Journal of the American Statistical Association, 50, 1096, 10.2307/2281208

M. Ester, H. Kriegel, S. Jorg, X. Xu, A density-based algorithm for discovering clusters in large spatial databases with noise, in: Proceedings of the Second International Conference on Knowledge Discovery and Data Mining, 1996, pp. 226–231.

Flach, 2001, Confirmation-guided discovery of first-order rules with tertius, Machine Learning, 42, 61, 10.1023/A:1007656703224

Freund, 1997, A decision-theoretic generalization of on-line learning and an application to boosting, Journal of Computer and System Sciences, 55, 119, 10.1006/jcss.1997.1504

Greco, 2001, Combining inductive and deductive tools for data analysis, AI Communications, 14, 69

S. Guha, R. Rastogi, K. Shim, Cure: an efficient clustering algorithm for large databases, in: Proceedings of 1998 ACM-SIGMOD International Conference Management of Data, 1998, pp. 73–84.

Hamzaebi, 2008, Improving artificial neural networks performance in seasonal time series forecasting, Information Sciences, 178, 4550, 10.1016/j.ins.2008.07.024

J. Han, J. Pei, Y. Yin, Mining frequent patterns without candidate generation, in: Proceedings of the 2000 ACM SIGMOD International Conference on Management of Data, 2000, pp. 1–12.

Hand, 1996, Idiots Bayes: not so stupid after all?, International Statistical Review, 69, 385, 10.1111/j.1751-5823.2001.tb00465.x

Hartigan, 1979, A k-means clustering algorithm, Applied Statistics, 28, 100, 10.2307/2346830

Hastie, 1996, Discriminant adaptive nearest neighbor classification, IEEE Transactions on Pattern Analysis and Machine Intelligence, 18, 607, 10.1109/34.506411

S. Hettich, S.D. Bay, The UCI KDD archive, University of California, Department of Information and Computer Science, Irvine, CA, 1999. <http://kdd.ics.uci.edu>.

Janssens, 2005, Adapting the CBA algorithm by means of intensity of implication, Information Sciences, 173, 305, 10.1016/j.ins.2004.03.022

Kanevski, 2004, Environmental data mining and modeling based on machine learning algorithms and geostatistics, Environmental Modelling and Software, 19, 845, 10.1016/j.envsoft.2003.03.004

Kim, 2008, Image retrieval model based on weighted visual features determined by relevance feedback, Information Sciences, 178, 4301, 10.1016/j.ins.2008.06.025

Lin, 2008, Fast discovery of sequential patterns in large databases using effective time-indexing, Information Sciences, 178, 4228, 10.1016/j.ins.2008.07.012

B. Liu, W. Hsu, Y. Ma, Integrating classification and association rule mining, in: Proceedings of the Fourth International Conference on Knowledge Discovery and Data Mining (KDD98), 1998.

D. Luebbers, U. Grimmer, M. Jarke, Systematic development of data mining-based data quality tools, in: Proceedings of the 29th International Conference on Very Large Data Bases (VLDB-2003), Berlin, Germany, 2003, pp. 548–559.

McLachlan, 2000

Nayak, 2008, Fast and effective clustering of XML data using structural information, Knowledge and Information Systems, 14, 197, 10.1007/s10115-007-0080-8

Orr, 1998, Data quality and systems theory, Communications of the ACM, 42, 66, 10.1145/269012.269023

Quinlan, 1993

Recknagel, 2001, Applications of machine learning to ecological modeling, Ecological Modelling, 1, 10.1016/S0304-3800(01)00291-5

Vapnik, 1995

Wu, 2008, Mining with noise knowledge: error-aware data mining, IEEE Transactions on Systems, Man and Cybernetics, Part A, 38, 917, 10.1109/TSMCA.2008.923034

Wu, 2008, Top 10 algorithms in data mining, Knowledge and Information Systems, 1, 10.1007/s10115-007-0114-2

Yang, 2006, 10 challenging problems in data mining research, International Journal of Information Technology and Decision Making, 5, 597, 10.1142/S0219622006002258

You, 2008, Search structures and algorithms for personalized ranking, Information Sciences, 178, 3925, 10.1016/j.ins.2008.06.009

Zhang, 1996, Birch: an efficient data clustering method for very large databases, ACM SIGMOD Record, 25, 103, 10.1145/235968.233324

Zhu, 2007, Knowledge Discovery and Data Mining: Challenges and Realities, IGI Global

X. Zhu, Semi-supervised Learning Literature Survey, Cs tr-1530, University of Wisconsin-Madison, 2008.

Zhu, 2006, Bridging local and global data cleansing: identifying class noise in large, distributed datasets, Data Mining and Knowledge Discovery, 12, 275, 10.1007/s10618-005-0012-8

Xingquan Zhu, Xindong Wu, Error awareness data mining, in: Proceedings of the IEEE International Conference on Granular Computing (GRC-2006), Atlanta, 2006.