| 2021 | COLT | When does gradient descent with logistic loss interpolate using deep networks with smoothed ReLU activations? | Niladri S. Chatterji, Philip M. Long, Peter L. Bartlett |
| 2020 | ALT | On the Complexity of Proper Distribution-Free Learning of Linear Classifiers. | Philip M. Long, Raphael J. Long |
| 2020 | ICLR | Generalization bounds for deep convolutional neural networks. | Philip M. Long, Hanie Sedghi |
| 2020 | ICLR | On the Global Convergence of Training Deep Linear ResNets. | Difan Zou, Philip M. Long, Quanquan Gu |
| 2019 | ICLR | The Singular Values of Convolutional Layers. | Hanie Sedghi, Vineet Gupta, Philip M. Long |
| 2018 | FOCS | Learning Sums of Independent Random Variables with Sparse Collective Support. | Anindya De, Philip M. Long, Rocco A. Servedio |
| 2018 | ICML | Gradient descent with identity initialization efficiently learns positive definite linear transformations. | Peter L. Bartlett, David P. Helmbold, Philip M. Long |
| 2017 | ALT | New bounds on the price of bandit feedback for mistake-bounded online multiclass learning. | Philip M. Long |
| 2017 | COLT | Surprising properties of dropout in deep networks. | David P. Helmbold, Philip M. Long |
| 2014 | WACV | Benchmarking large-scale Fine-Grained Categorization. | Anelia Angelova, Philip M. Long |
| 2014 | STOC | The power of localization for efficiently learning linear separators with noise. | Pranjal Awasthi, Maria-Florina Balcan, Philip M. Long |
| 2013 | COLT | Active and passive learning of linear separators under log-concave distributions. | Maria-Florina Balcan, Philip M. Long |
| 2013 | ICML | Consistency versus Realizable H-Consistency for Multiclass Classification. | Philip M. Long, Rocco A. Servedio |
| 2011 | ICML | On the Necessity of Irrelevant Variables. | David P. Helmbold, Philip M. Long |
| 2010 | ICML | Finding Planted Partitions in Nearly Linear Time using Arrested Spectral Clustering. | Nader H. Bshouty, Philip M. Long |
| 2010 | ICML | Restricted Boltzmann Machines are Hard to Approximately Evaluate or Simulate. | Philip M. Long, Rocco A. Servedio |
| 2009 | COLT | Linear Classifiers are Nearly Optimal When Hidden Variables Have Diverse Effect. | Nader H. Bshouty, Philip M. Long |
| 2009 | ICALP | Learning Halfspaces with Malicious Noise. | Adam R. Klivans, Philip M. Long, Rocco A. Servedio |
| 2008 | ICML | Random classification noise defeats all convex potential boosters. | Philip M. Long, Rocco A. Servedio |
| 2006 | AAAI | Predicting Electricity Distribution Feeder Failures Using Machine Learning Susceptibility Analysis. | Philip Gross, Albert Boulanger, Marta Arias, David L. Waltz, Philip M. Long, Charles Lawson, Roger Anderson, Matthew Koenig, Mark Mastrocinque, William Fairechio, John A. Johnson, Serena Lee, Frank Doherty, Arthur Kressner |
| 2006 | ALT | Editors' Introduction. | Jos L. Balczar, Philip M. Long, Frank Stephan |
| 2006 | COLT | Online Multitask Learning. | Ofer Dekel, Philip M. Long, Yoram Singer |
| 2006 | COLT | Discriminative Learning Can Succeed Where Generative Learning Fails. | Philip M. Long, Rocco A. Servedio |
| 2005 | COLT | Martingale Boosting. | Philip M. Long, Rocco A. Servedio |
| 2005 | ICML | Unsupervised evidence integration. | Philip M. Long, Vinay Varadan, Sarah Gilman, Mark Treshock, Rocco A. Servedio |
| 2003 | COLT | Boosting with Diverse Base Classifiers. | Sanjoy Dasgupta, Philip M. Long |
| 2002 | AAAI | Minimum Majority Classification and Boosting. | Philip M. Long |
| 2001 | COLT | Agnostic Boosting. | Shai Ben-David, Philip M. Long, Yishay Mansour |
| 2001 | COLT | A Theoretical Analysis of Query Selection for Collaborative Filtering. | Wee Sun Lee, Philip M. Long |
| 2001 | COLT | On Agnostic Learning with {0, *, 1}-Valued and Real-Valued Hypotheses. | Philip M. Long |
| 2001 | WADS | Using the Pseudo-Dimension to Analyze Approximation Algorithms for Integer Programming. | Philip M. Long |
| 2000 | COLT | On the Difficulty of Approximately Maximizing Agreements. | Shai Ben-David, Nadav Eiron, Philip M. Long |
| 2000 | SODA | Improved bounds on the sample complexity of learning. | Yi Li, Philip M. Long, Aravind Srinivasan |
| 1999 | ICML | Associative Reinforcement Learning using Linear Probabilistic Concepts. | Naoki Abe, Philip M. Long |
| 1998 | COLT | The complexity of learning according to two models of a drifting environment. | Philip M. Long |
| 1998 | COLT | On the Sample Complexity of Learning Functions with Bounded Variation. | Philip M. Long |
| 1997 | COLT | On-line Evaluation and Prediction using Linear Functions. | Philip M. Long |
| 1997 | DCC | Text Compression Via Alphabet Re-Representation. | Philip M. Long, Apostol Natsev, Jeffrey Scott Vitter |
| 1997 | STOC | Approximating Hyper-Rectangles: Learning and Pseudo-Random Sets. | Peter Auer, Philip M. Long, Aravind Srinivasan |
| 1996 | ALT | Improved Bounds about On-line Learning of Smooth Functions of a Single Variable. | Philip M. Long |
| 1996 | COLT | On the Complexity of Learning from Drifting Distributions. | Rakesh D. Barve, Philip M. Long |
| 1996 | COLT | PAC Learning Axis-Aligned Rectangles with Respect to Product Distributions from Multiple-Instance Examples. | Philip M. Long, Lei Tan |
| 1996 | DCC | Efficient Cost Measures for Motion Compensation at Low Bit Rates (Extended Abstract). | Dzung T. Hoang, Philip M. Long, Jeffrey Scott Vitter |
| 1995 | COLT | More Theorems about Scale-sensitive Dimensions and Learning. | Peter L. Bartlett, Philip M. Long |
| 1995 | DCC | Multiple-Dictionary Coding Using Partial Matching. | Dzung T. Hoang, Philip M. Long, Jeffrey Scott Vitter |
| 1995 | ICML | Learning to Make Rent-to-Buy Decisions with Systems Applications. | P. Krishnan, Philip M. Long, Jeffrey Scott Vitter |
| 1994 | COLT | Fat-Shattering and the Learnability of Real-Valued Functions. | Peter L. Bartlett, Philip M. Long, Robert C. Williamson |
| 1994 | DCC | Explicit Bit Minimization for Motion-Compensated Video Coding. | Dzung T. Hoang, Philip M. Long, Jeffrey Scott Vitter |
| 1994 | STOC | Simulating access to hidden information while learning. | Peter Auer, Philip M. Long |
| 1993 | COLT | On the Complexity of Function Learning. | Peter Auer, Philip M. Long, Wolfgang Maass, Gerhard J. Woeginger |
| 1993 | COLT | Worst-Case Quadratic Loss Bounds for a Generalization of the Widrow-Hoff Rule. | Nicol Cesa-Bianchi, Philip M. Long, Manfred K. Warmuth |
| 1993 | COLT | On-Line Learning with Linear Loss Constraints. | Nick Littlestone, Philip M. Long |
| 1992 | COLT | Characterizations of Learnability for Classes of { | Shai Ben-David, Nicol Cesa-Bianchi, Philip M. Long |
| 1992 | COLT | The Learning Complexity of Smooth Functions of a Single Variable. | Don Kimber, Philip M. Long |
| 1992 | FOCS | Apple Tasting and Nearly One-Sided Learning | David P. Helmbold, Nick Littlestone, Philip M. Long |
| 1991 | COLT | Tracking Drifting Concepts Using Random Examples. | David P. Helmbold, Philip M. Long |
| 1991 | STOC | On-Line Learning of Linear Functions | Nick Littlestone, Philip M. Long, Manfred K. Warmuth |
| 1990 | COLT | Composite Geometric Concepts and Polynomial Predictability. | Philip M. Long, Manfred K. Warmuth |