∞
π Σ ∫ ∂ Δ √
Δ

Speaker:Xueqin Wang(University of Science and Technology of China)

Time:2022-06-07, 15:00

Location:Conference Room 105 at Experiment Building at Haiyun Campus

Abstract:

Statistical inference aims to use observed samples to learn the unknown properties of a population. It has become an integral step in scientific reasoning. A building block of nonparametric statistical inference is distribution function. The distribution function and samples are connected to form a directed closed loop by the correspondence theorem in measure theory and the Glivenko-Cantelli and Donsker properties in statistics, and this connection creates a paradigm for statistical inference. However, existing distribution functions are defined in Euclidean spaces. Those distribution functions are no longer convenient to use or applicable in characterizing the rapidly evolving data objects of complex nature. Thus, it is imperative to develop the concept of the distribution function in a more general space to meet emerging needs. Note that the linearity allows us to use hypercubes to define the distribution function in an Euclidean space, but without the linearity in a metric space, we must work with balls as the basis of the metric topology in defining a probability measure. We introduce a class of novel quasi-distribution functions, or ball functions, for metric space-valued random objects. A ball requires a center and a radius. The center depends on the random point of interest, and the radius is determined by the distance between the center and another random point. Working with balls in defining a probability measure is particularly challenging because unlike hypercubes, the intersection of two balls may not be a ball. We overcome this challenge to prove the correspondence theorem and the Glivenko-Cantelli theorem in metric spaces that lie the foundation for conducting rational statistical inference for metric space-valued data. Based on ball function, we develop statistical methods for homogeneity test, mutual independence test, and hierarchical clustering for non-Euclidean random objects, and present comprehensive empirical evidence to support the performance of our proposed methods.