Tony Cai’s current research focuses on the statistical foundations of modern data science, particularly transfer learning, differential privacy, federated and distributed learning, causal inference, and high-dimensional machine learning. A central theme is how to learn optimally from heterogeneous, sensitive, and decentralized data: when information from related populations can improve a target analysis, how privacy and communication constraints change statistical limits, and how procedures can adapt safely without suffering negative transfer.
His recent work develops decision-theoretic and minimax frameworks for these problems, including privacy-preserving and federated estimation, testing, transfer learning, and individualized treatment decisions. These questions are increasingly important for AI, where modern systems must learn from large, distributed, and heterogeneous data sources while respecting privacy, reliability, and resource constraints. His work contributes statistical foundations for trustworthy AI by clarifying when learning is possible, what information is fundamentally required, and how optimal procedures can be designed under such constraints.
Tony Cai’s research contributions span several major areas of modern statistics. He has developed influential theory and methodology for high-dimensional covariance and precision-matrix estimation, sparse PCA, graphical models, regression, and high-dimensional testing, as well as for nonparametric estimation, adaptation, and uncertainty quantification. His work on large-scale multiple testing addresses power and false discovery control in complex high-dimensional settings, while contributions to binomial confidence intervals, singular-subspace perturbation theory, and the theoretical analysis of t-SNE have provided widely used tools and benchmarks. Many of these contributions also underpin modern AI and machine learning, particularly through their treatment of high-dimensional structure, spectral methods, dimension reduction, uncertainty, and reliable inference from complex data. A recurring theme throughout his work is the development of statistically optimal and adaptive procedures together with a precise understanding of the limits of what can be achieved under high dimensionality, structural complexity, privacy, communication, and heterogeneity.
- Arnab Auddy, Tony Cai, Abhinav Chakraborty (2026), Minimax and adaptive transfer learning for nonparametric classification under distributed differential privacy constraints, Journal of the Royal Statistical Society, Series B , 27. Abstract
This paper considers minimax and adaptive transfer learning for nonparametric classification under the posterior drift model with distributed differential privacy constraints. Our study is conducted within a heterogeneous framework, encompassing diverse sample sizes, varying privacy parameters, and data heterogeneity across different servers. We first establish the minimax misclassification rate, precisely characterizing the effects of privacy constraints, source samples, and target samples on classification accuracy. The results reveal interesting phase transition phenomena and highlight the intricate trade-offs between preserving privacy and achieving classification accuracy. We then develop a data-driven adaptive classifier that achieves the optimal rate within a logarithmic factor across a large collection of parameter spaces while satisfying the same set of differential privacy constraints. Simulation studies and real-world data applications further elucidate the theoretical analysis with numerical results.
- Tony Cai, Abhinav Chakraborty, Yichen Wang (2026), Optimal differentially private ranking from pairwise comparisons, Journal of the American Statistical Association.
- Tony Cai, Abhinav Chakraborty, Lasse Vuursteen (2026), Optimal federated learning for nonparametric regression with heterogenous distributed differential privacy constraints, Journal of the American Statistical Association.
- Tony Cai and Hongji Wei (2024), Distributed Gaussian Mean Estimation under Communication Constraints: Optimal Rates and Communication-efficient Algorithms, Journal of Machine Learning Research , 63.
- Tony Cai, Rungang Han, Anru Zhang (2022), On the Non-asymptotic Concentration of Heteroskedastic Wishart-type Matrix, Electronic Journal of Probability, 27 (), pp. 1-40.
- Anru Zhang, Tony Cai, Yihong Wu (2022), Heteroskedastic PCA: Algorithm, Optimality, and Applications, Annals of Statistics, 50 (1), pp. 53-80.
- Zijian Guo, Claude Renaux, Peter Bühlmann, Tony Cai (2021), Group Inference in High Dimensions with Applications to Hierarchical Testing, Electronic Journal of Statistics, 15 (), pp. 6633-6676.
- Tony Cai, Tiefeng Jiang, Xiaoou Li (Under Review), Asymptotic Analysis for Extreme Eigenvalues of Principal Minors of Random Matrices.
- Tony Cai and Hongji Wei (2020), Transfer Learning for Nonparametric Classification: Minimax Rate and Adaptive Classifier, Annals of Statistics, (to appear) ().
- Shulei Wang, Tony Cai, Hongzhe Li (2020), Optimal Estimation of Wasserstein Distance on A Tree with An Application to Microbiome Studies, Journal of the American Statistical Association, (to appear) ().
- All Research from Tony Cai »