| 2003 | Generalized Sensitivity Analysis: A Framework for Evaluating Data Analysis Results. | Ronald K. Pearson |
| 2003 | Approximate Query Answering by Model Averaging. | Dmitry Pavlov, Padhraic Smyth |
| 2003 | Mining Temporal Databases for Subsequence Patterns. | Wen Niu, Raj Bhatnagar |
| 2003 | Efficient Unsupervised Mining from Noisy Data Sets: Application to Clustering Co-occurrence Data. | Hiroshi Mamitsuka |
| 2003 | Field-Theoretic Methods for Intractable Probabilistic Models. | Dennis Lucarelli, Cheryl Resch, I-Jeng Wang, Fernando J. Pineda |
| 2003 | Sort-Merge Feature Selection for Video Data. | Yan Liu, John R. Kender |
| 2003 | Using Low-Memory Representations to Cluster Very Large Data Sets. | David Littau, Daniel Boley |
| 2003 | An Outlier-based Data Association Method for Linking Criminal Incidents. | Song Lin, Donald E. Brown |
| 2003 | Active Sampling: An Effective Approach to Feature Selection. | Huan Li, Hongjun Lu, Lei Yu |
| 2003 | A Comparative Study of Anomaly Detection Schemes in Network Intrusion Detection. | Aleksandar Lazarevic, Levent Ertz, Vipin Kumar, Aysel Ozgur, Jaideep Srivastava |
| 2003 | On using Page Cooccurrences for Computing Clickstream Similarity. | Ravi Kothari, Parul A. Mittal, Vivek Jain, Mukesh K. Mohania |
| 2003 | ApproxMAP: Approximate Mining of Consensus Sequential Patterns. | Hye-Chung Kum, Jian Pei, Wei Wang, Dean Duncan |
| 2003 | Communication and Memory Efficient Parallel Decision Tree Construction. | Ruoming Jin, Gagan Agrawal |
| 2003 | Feature Mining Paradigms for Scientific Data. | Ming Jiang, Tat-Sang Choy, Sameep Mehta, Matt Coatney, Steve Barr, Kaden Hazzard, David Richie, Srinivasan Parthasarathy, Raghu Machiraju, David S. Thompson, John Wilkins, Boyd Gatlin |
| 2003 | Estimation of Topological Dimension. | Douglas R. Hundley, Michael J. Kirby |
| 2003 | Mixture Models and Frequent Sets: Combining Global and Local Methods for 0-1 Data. | Jaakko Hollmn, Jouni K. Seppnen, Heikki Mannila |
| 2003 | Nonparametric Density Estimation: Toward Computational Tractability. | Alexander G. Gray, Andrew W. Moore |
| 2003 | A New Gravitational Clustering Algorithm. | Jonatan Gmez, Dipankar Dasgupta, Olfa Nasraoui |
| 2003 | Hierarchical Document Clustering using Frequent Itemsets. | Benjamin C. M. Fung, Ke Wang, Martin Ester |
| 2003 | Finding Clusters of Different Sizes, Shapes, and Densities in Noisy, High Dimensional Data. | Levent Ertz, Michael S. Steinbach, Vipin Kumar |
| 2003 | PageRank: HITS and a Unified Framework for Link Analysis. | Chris H. Q. Ding, Xiaofeng He, Parry Husbands, Hongyuan Zha, Horst D. Simon |
| 2003 | Anytime Query-Tuned Kernel Machines via Cholesky Factorization. | Dennis DeCoste |
| 2003 | On the Techniques for Data Clustering with Numerical Constraints. | Bi-Ru Dai, Cheng-Ru Lin, Ming-Syan Chen |
| 2003 | Learning Bayesian Network Structure from Distributed Data. | Rong Chen, Krishnamoorthy Sivakumar, H. Khargupta |
| 2003 | The Application of Text Mining Software to Examine Coded Information. | Patricia B. Cerrito, James Cox |