Journal of Machine Learning Research Papers Volume 27に記載されている内容を一覧にまとめ、機械翻訳を交えて日本語化し掲載します。
目次
- 1 論文
- 1.1 The surrogate Gibbs-posterior of a corrected stochastic MALA: Towards uncertainty quantification for neural networks
- 1.2 Online Detection of Changes in Moment--Based Projections: When to Retrain Deep Learners or Update Portfolios?
- 1.3 Efficient frequent directions algorithms for approximate decomposition of matrices and higher-order tensors
- 1.4 Identifying Weight-Variant Latent Causal Models
- 1.5 Classification Under Local Differential Privacy with Model Reversal and Model Averaging
- 1.6 Stochastic Gradient Methods: Bias, Stability and Generalization
- 1.7 Extending Mean-Field Variational Inference via Entropic Regularization: Theory and Computation
- 1.8 skwdro: a library for Wasserstein distributionally robust machine learning
- 1.9 Guaranteed Nonconvex Low-Rank Tensor Estimation via Scaled Gradient Descent
- 1.10 A Data-Augmented Contrastive Learning Approach to Nonparametric Density Estimation
- 1.11 Nonlocal Techniques for the Analysis of Deep ReLU Neural Network Approximations
- 1.12 Nonlinear function-on-function regression by RKHS
- 1.13 UQLM: A Python Package for Uncertainty Quantification in Large Language Models
- 1.14 A Reinforcement Learning Approach in Multi-Phase Second-Price Auction Design
- 1.15 Error Analysis for Deep ReLU Feedforward Density-Ratio Estimation with Bregman Divergence
- 1.16 Flexible Functional Treatment Effect Estimation
- 1.17 Neural Network Parameter-optimization of Gaussian Pre-marginalized Directed Acyclic Graphs
- 1.18 Extrapolated Markov Chain Oversampling Method for Imbalanced Text Classification
- 1.19 An Anytime Algorithm for Good Arm Identification
- 1.20 Simulation-based Calibration of Uncertainty Intervals under Approximate Bayesian Estimation
- 1.21 Learning Bayesian Network Classifiers to Minimize Class Variable Parameters
- 1.22 Nonparametric Estimation of a Factorizable Density using Diffusion Models
- 1.23 The Distribution of Ridgeless Least Squares Interpolators
- 1.24 LazyDINO: Fast, Scalable, and Efficiently Amortized Bayesian Inversion via Structure-Exploiting and Surrogate-Driven Measure Transport
- 1.25 A Common Interface for Automatic Differentiation
- 1.26 Refined Risk Bounds for Unbounded Losses via Transductive Priors
- 1.27 Decorrelated Local Linear Estimator: Inference for Non-linear Effects in High-dimensional Additive Models
- 1.28 Communication-efficient Distributed Statistical Inference for Massive Data with Heterogeneous Auxiliary Information
- 1.29 Generative Bayesian Inference with GANs
- 1.30 Exploring Novel Uncertainty Quantification through Forward Intensity Function Modeling
- 1.31 Persistence Diagrams Estimation of Multivariate Piecewise Hölder-continuous Signals
- 1.32 CHANI: Correlation-based Hawkes Aggregation of Neurons with bio-Inspiration
- 1.33 Finite Neural Networks as Mixtures of Gaussian Processes: From Provable Error Bounds to Prior Selection
- 1.34 Optimization and Generalization of Gradient Descent for Shallow ReLU Networks with Minimal Width
- 1.35 Adaptive Forward Stepwise: A Method for High Sparsity Regression
- 1.36 Optimizing Attention with Mirror Descent: Generalized Max-Margin Token Selection
- 1.37 Hierarchical Causal Models
- 1.38 Reparameterized Complex-valued Neurons Can Efficiently Learn More than Real-valued Neurons via Gradient Descent
- 1.39 Unsupervised Feature Selection via Nonnegative Orthogonal Constrained Regularized Minimization
- 1.40 A causal fused lasso for interpretable heterogeneous treatment effects estimation
- 1.41 Bayesian Inference of Contextual Bandit Policies via Empirical Likelihood
- 1.42 Convergence and complexity of block majorization-minimization for constrained block-Riemannian optimization
- 1.43 Two-way Node Popularity Model for Directed and Bipartite Networks
- 1.44 A Symplectic Analysis of Alternating Mirror Descent
- 1.45 Contrasting Local and Global Modeling with Machine Learning and Satellite Data: A Case Study Estimating Tree Canopy Height in African Savannas
- 1.46 Boosted Control Functions: Distribution Generalization and Invariance in Confounded Models
- 1.47 DCatalyst: A Unified Accelerated Framework for Decentralized Optimization
- 1.48 Covariate-dependent Hierarchical Dirichlet Processes
- 1.49 Online Bernstein-von Mises theorem
- 1.50 Transformers Can Overcome the Curse of Dimensionality: A Theoretical Study from an Approximation Perspective
- 1.51 The Role of Contextual Information in Best Arm Identification
- 1.52 A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks
- 1.53 Inference with non-differentiable surrogate loss in a general high-dimensional classification framework
- 1.54 Knowledge Cascade: Reverse Knowledge Distillation on Nonparametric Multivariate Functional Estimation
- 1.55 Neural Exploitation and Exploration of Contextual Bandits
- 1.56 Causal Influences over Social Learning Networks
- 1.57 Do We Need to Penalize Variance of Losses for Learning with Label Noise?
- 1.58 Nonparametric generative modeling for time series via Schrödinger bridge
- 1.59 Sparse Topic Modeling via Spectral Decomposition and Thresholding
- 1.60 Probabilistic Rainfall Downscaling: Joint Generalized Neural Models with Censored Spatial Gaussian Copula
- 1.61 Global Fréchet Manifold Learning for Random Objects, With Application to Low-Dimensional Wasserstein Representations of Distributional Data
- 1.62 A Convex Framework for Confounding Robust Inference
- 1.63 Kernel-based Distributed Learning
- 1.64 Why "Classic" Transformers Are Shallow and A Depth-Enabling Technique
- 1.65 Bayes-Optimal Fair Classification with Linear Disparity Constraints via Pre-, In-, and Post-processing
- 1.66 Investigating the Histogram Loss in Regression
- 1.67 Optimal Approximation and Generalization Errors for Deep Convolutional Neural Networks
- 1.68 Approximations and Learning for Continuous State and Action MDPs under Average Cost Criteria
- 1.69 Minimax density estimation in the adversarial framework under local differential privacy
- 1.70 A Mean-Field Analysis of Neural Stochastic Gradient Descent-Ascent for Functional Minimax Optimization
- 1.71 Nonparametric Partial Disentanglement via Mechanism Sparsity: Sparse Actions, Interventions and Sparse Temporal Dependencies
- 1.72 Corruptions of Supervised Learning Problems: Typology and Mitigations
- 1.73 Differentially Private Best-Arm Identification
- 1.74 Vector-Valued Gaussian Processes for Approximating Divergence- or Rotation-free Vector Fields
- 1.75 Stochastic Differential Equations models for Least-Squares Stochastic Gradient Descent
- 1.76 A Fully Parameter-Free Second-Order Algorithm for Convex-Concave Minimax Problems
- 1.77 Limiting Over-Smoothing and Over-Squashing of Graph Message Passing by Deep Scattering Transforms
- 1.78 Multi-relational Network Autoregression Model with Latent Group Structures
- 1.79 Enhancing Accuracy in Generative Models via Knowledge Transfer
- 1.80 Demographic Parity in Regression and Classification Within the Unawareness Framework
- 1.81 Differentially Private Estimation and Inference in High-Dimensional Regression with FDR Control
- 1.82 Beyond Unconstrained Features: Neural Collapse for Shallow Neural Networks with General Data
- 1.83 Asymptotics of Stochastic Gradient Descent with Dropout Regularization in Linear Models
- 1.84 Almost Sure Convergence of Linear Temporal Difference Learning with Arbitrary Features
- 1.85 Convergence of Noise-Free Sampling Algorithms with Regularized Wasserstein Proximals
- 1.86 Best Arm Identification with Minimal Regret
- 1.87 On the Relevance of Byzantine Robust Optimization Against Data Poisoning
- 1.88 Node Regression on Latent Position Random Graphs via Local Averaging
- 1.89 Towards Convexity in Anomaly Detection: A New Formulation of SSLM with Unique Optimal Solutions
- 1.90 Transfer Conformal Predictive Inference for Regression
- 1.91 Generalized Resubstitution for Regression Error Estimation
- 1.92 A Natural Primal-Dual Hybrid Gradient Method for Adversarial Neural Network Training on Solving Partial Differential Equation
- 1.93 Cheap Bootstrap for Fast Uncertainty Quantification of Stochastic Gradient Descent
- 1.94 Transfer Learning via Regularized Random-effects Linear Discriminant Analysis
- 1.95 Semi-supervised learning for linear extremile regression
- 1.96 Deep Nonparametric Conditional Independence Tests for Images
- 1.97 Kernel Mean Embedding Deviation Subspace for Unsupervised Learning with Heterogeneous Data
- 1.98 Exogenous Randomness Empowering Random Forests
- 1.99 High-dimensional Parameter Transfer With Fused-Regularizer
- 1.100 Spectral Truncation Kernels: Noncommutativity in C*-algebraic Kernel Machines
- 1.101 Deconvolution in unlinked linear models
- 1.102 Statistical Learning Theory for Neural Operators
- 1.103 A Single-Loop Stochastic Proximal Quasi-Newton Method for Large-Scale Nonsmooth Convex Optimization
- 1.104 Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural Networks
- 1.105 A Unified Approach to Analysis and Design of Denoising Markov Models
- 1.106 Nested Subspace Learning with Flags
- 1.107 FLAGG: Flexible Autoregressive Graph Generation
- 1.108 A Two-Timescale Primal-Dual Framework for Reinforcement Learning via Online Dual Variable Guidance
- 1.109 Convergence of Decentralized Stochastic Subgradient-based Methods for Nonsmooth Nonconvex Optimization
- 1.110 Three Types of Calibration using Properties and their Semantic and Formal Relationships
- 1.111 Embedding Network Autoregression for Time Series Analysis and Causal Peer Effect Inference
- 1.112 Finite-Time Decoupled Convergence in Nonlinear Two-Time-Scale Stochastic Approximation
- 1.113 The Within-Orbit Adaptive Leapfrog No-U-Turn Sampler
- 1.114 STDE++: Polynomial-Time Amortization for Linear Differential Operators
- 1.115 Vecchia-Inducing-Points Full-Scale Approximations for Gaussian Processes
- 1.116 Statistical guarantees for denoising reflected diffusion models
- 1.117 Learning general conditional independence structures via the neighbourhood lattice
- 1.118 Accelerating Constrained Sampling: A Large Deviations Approach
- 1.119 Statistical Test for Attention in Transformers for Images and Time Series
- 1.120 py/cuTAGI: An Open-Source Library for Tractable Approximate Gaussian Inference in Bayesian Neural Networks
- 1.121 Gradient Span Algorithms Make Predictable Progress in High Dimension
- 1.122 Robust training of implicit generative models for multivariate and heavy-tailed distributions with an invariant statistical loss
- 1.123 Adaptive Nonparametric Perturbations of Parametric Models with Generalized Bayes
- 1.124 Minimax Optimal Convergence of Gradient Descent in Logistic Regression via Large and Adaptive Stepsizes
- 1.125 Approximation-Free Differentiable Oblique Decision Trees
- 1.126 Underdamped Langevin MCMC with third order convergence
- 1.127 Mixing times of data-augmentation Gibbs samplers for high-dimensional probit regression
- 1.128 Abstract Gradient Training: A Unified Certification Framework for Data Poisoning, Unlearning, and Differential Privacy
- 1.129 Doubly Debiased Robust Subsampling for Transfer Learning
- 1.130 Learning to Play Two-Player Perfect-Information Games without Knowledge
- 1.131 Graph-based Clustering Revisited: A Relaxation of Kernel k-Means Perspective
- 1.132 End-to-End Deep Learning for Predicting Metric Space-Valued Outputs
- 1.133 The Sample Complexity of Parameter-Free Stochastic Convex Optimization
- 1.134 Near-optimal Delta-convex Estimation of Lipschitz Functions
- 1.135 Error Analyses of Auto-Regressive Video Diffusion Models
- 1.136 High-Dimensional Analysis of Gradient Flow for Extensive-Width Quadratic Neural Networks
- 1.137 Bridging Domain Invariance and Diversity: A Fine-Grained Risk Bound for Domain Generalization
- 2 参考文献
- 3 関連情報
論文
The surrogate Gibbs-posterior of a corrected stochastic MALA: Towards uncertainty quantification for neural networks
The surrogate Gibbs-posterior of a corrected stochastic MALA: Towards uncertainty quantification for neural networks / 補正された確率的MALAの代理ギブス事後分布:ニューラルネットワークの不確実性定量化に向けて
MALA is a popular gradient-based Markov chain Monte Carlo method to access the Gibbs-posterior distribution. Stochastic MALA (sMALA) scales to large data sets, but changes the target distribution from the Gibbs-posterior to a surrogate posterior which only exploits a reduced sample size. We introduce a corrected stochastic MALA (csMALA) with a simple correction term for which distance between the resulting surrogate posterior and the original Gibbs-posterior decreases in the full sample size while retaining scalability. In a nonparametric regression model, we prove a PAC-Bayes oracle inequality for the surrogate posterior. Uncertainties can be quantified by sampling from the surrogate posterior. Focusing on Bayesian neural networks, we analyze the diameter and coverage of credible balls for shallow neural networks and we show optimal contraction rates for deep neural networks. Our credibility result is independent of the correction and can also be applied to the standard Gibbs-posterior. A simulation study in a high-dimensional parameter space demonstrates that an estimator drawn from csMALA based on its surrogate Gibbs-posterior indeed exhibits these advantages in practice.
MALAは、ギブス事後分布にアクセスするための一般的な勾配ベースのマルコフ連鎖モンテカルロ法です。確率的MALA(sMALA)は大規模なデータセットに拡張できますが、ターゲット分布をギブス事後分布から、縮小されたサンプルサイズのみを利用する代理事後分布に変更します。私たちは、スケーラビリティを維持しながら、結果として得られる代理事後分布と元のギブス事後分布の間の距離が完全なサンプルサイズで減少する単純な補正項を持つ補正確率的MALA(csMALA)を紹介します。ノンパラメトリック回帰モデルでは、代理事後分布のPAC-ベイズオラクル不等式を証明します。不確実性は、代理事後分布からサンプリングすることで定量化できます。ベイジアンニューラルネットワークに焦点を当て、浅いニューラルネットワークの信頼できるボールの直径と範囲を分析し、深いニューラルネットワークの最適な収縮率を示します。我々の信頼性結果は補正とは独立しており、標準的なギブス事後分布にも適用できます。高次元パラメータ空間におけるシミュレーション研究は、csMALAからその代理ギブス事後分布に基づいて導出された推定値が、実際にこれらの利点を示すことを実証しています。
Online Detection of Changes in Moment–Based Projections: When to Retrain Deep Learners or Update Portfolios?
Online Detection of Changes in Moment–Based Projections: When to Retrain Deep Learners or Update Portfolios? / モーメントベース射影における変化のオンライン検出:ディープラーナーの再学習とポートフォリオの更新はいつ行うべきか?
Training deep learning neural networks often requires massive amounts of computational ressources. We proposeto sequentially monitor network predictions to trigger retraining only if the predictions are no longer valid. This can reduce drastically computational costs and opens a door to green deep learning. Our approach is based on the relationship to projected second moments monitoring, a problem also arising in other areas such as computational finance. Various open-end as well as closed-end monitoring rules are studied under mild assumptions on the training sample and the observations of the monitoring period. The results allow for high-dimensional non-stationary time series data and thus, especially, non-i.i.d. training data. Asymptotics is based on Gaussian approximations of projected partial sums allowing for an estimated projection vector. Estimation of projection vectors is studied both for classical non-$\ell_0$-sparsity as well as under sparsity. For the case that the optimal projection depends on the unknown covariance matrix, hard- and soft-thresholded estimators are studied. The method is analyzed by simulations and supported by synthetic data experiments.
深層学習ニューラルネットワークの学習には、しばしば膨大な計算資源が必要です。我々は、ネットワーク予測を逐次監視し、予測がもはや有効でなくなった場合にのみ再学習をトリガーすることを提案します。これにより計算コストを大幅に削減し、グリーンな深層学習への道が開かれます。我々のアプローチは、計算金融などの他の分野でも発生する問題である、投影された秒モーメント監視との関係に基づいています。学習サンプルと監視期間の観測値に関する緩やかな仮定の下で、様々なオープンエンドおよびクローズドエンドの監視規則が研究されます。その結果、高次元の非定常時系列データ、特に非独立同値学習データが可能になります。漸近解析は、推定射影ベクトルを許容する射影部分和のガウス近似に基づいています。射影ベクトルの推定は、古典的な非$\ell_0$-スパース性とスパース性の両方について研究されています。最適な射影が未知の共分散行列に依存する場合については、ハード閾値推定値とソフト閾値推定値が研究されています。この手法はシミュレーションによって解析され、合成データ実験によって裏付けられています。
Efficient frequent directions algorithms for approximate decomposition of matrices and higher-order tensors
Efficient frequent directions algorithms for approximate decomposition of matrices and higher-order tensors / 行列と高階テンソルの近似分解のための効率的な頻繁方向アルゴリズム
In the framework of the FD (frequent directions) algorithm, we first develop two efficient algorithms for low-rank matrix approximations under the embedding matrices composed of the product of any SpEmb (sparse embedding) matrix and any standard Gaussian matrix, or any SpEmb matrix and any SRHT (subsampled randomized Hadamard transform) matrix. The theoretical results are also achieved based on the bounds of singular values of standard Gaussian matrices and the theoretical results for SpEmb and SRHT matrices. With a given Tucker-rank, we then obtain several efficient FD-based randomized variants of T-HOSVD (the truncated high-order singular value decomposition) and ST-HOSVD (sequentially T-HOSVD), which are two common algorithms for computing the approximate Tucker decomposition of any tensor with a given Tucker-rank. We also consider efficient FD-based randomized algorithms for computing the approximate TT (tensor-train) decomposition of any tensor with a given TT-rank. Finally, we illustrate the efficiency and accuracy of these algorithms using synthetic and real-world matrix (and tensor) data.
FD(頻繁方向)アルゴリズムの枠組みにおいて、まず、任意のSpEmb(スパース埋め込み)行列と任意の標準ガウス行列、または任意のSpEmb行列と任意のSRHT(部分標本化ランダム化アダマール変換)行列の積で構成される埋め込み行列の下での低ランク行列近似のための2つの効率的なアルゴリズムを開発します。また、標準ガウス行列の特異値の境界と、SpEmb行列およびSRHT行列の理論的結果に基づいて理論的結果も得られます。与えられたタッカーランクを用いて、T-HOSVD(切り捨て高階特異値分解)とST-HOSVD(逐次T-HOSVD)の効率的なFDベースのランダム化変種をいくつか得ることができます。これらは、与えられたタッカーランクを持つ任意のテンソルの近似タッカー分解を計算するための一般的な2つのアルゴリズムです。また、与えられたTTランクを持つ任意のテンソルの近似TT(テンソル列)分解を計算するための効率的なFDベースのランダム化アルゴリズムについても検討します。最後に、合成データと実世界の行列(およびテンソル)データを用いて、これらのアルゴリズムの効率と精度を示します。
Identifying Weight-Variant Latent Causal Models
Identifying Weight-Variant Latent Causal Models / 重み変数潜在因果モデルの識別
The task of causal representation learning aims to uncover latent higher-level causal variables that affect lower-level observations. Identifying the true latent causal variables from observed data, while allowing instantaneous causal relations among latent variables, remains a challenge, however. To this end, we start with the analysis of three intrinsic indeterminacies in identifying latent variables from observations: transitivity, permutation indeterminacy, and scaling indeterminacy. We find that transitivity acts as a key role in impeding the identifiability of latent causal variables. To address the unidentifiable issue due to transitivity, we introduce a novel identifiability condition where the underlying latent causal model satisfies a linear-Gaussian model, in which the causal coefficients and the distribution of Gaussian noise are modulated by an additional observed variable. Under certain assumptions, including the existence of a reference condition under which latent causal influences vanish, we can show that the latent causal variables can be identified up to trivial permutation and scaling, and that partial identifiability results can still be obtained when this reference condition is violated for a subset of latent variables. Furthermore, based on these theoretical results, we propose a novel method, termed Structural caUsAl Variational autoEncoder (SuaVE), which directly learns causal representations and causal relationships among them, together with the mapping from the latent causal variables to the observed ones. Experimental results on synthetic and real data demonstrate the identifiability and consistency results and the efficacy of SuaVE in learning causal representations.
因果表現学習の課題は、低レベルの観測に影響を与える潜在的な高レベルの因果変数を発見することを目指しています。しかしながら、潜在変数間の瞬間的な因果関係を許容しながら、観測データから真の潜在的因果変数を特定することは依然として課題です。この目的のために、我々は観測から潜在変数を特定する際の3つの本質的な不確定性、すなわち推移性、順列不確定性、およびスケーリング不確定性の分析から始めます。その結果、推移性が潜在的因果変数の識別可能性を阻害する上で重要な役割を果たしていることが分かりました。推移性による識別不能な問題に対処するため、我々は、基礎となる潜在的因果モデルが線形ガウスモデルを満たすという新たな識別可能性条件を導入します。この線形ガウスモデルでは、因果係数とガウスノイズの分布が追加の観測変数によって変調されます。潜在的因果的影響が消失する参照条件の存在を含む特定の仮定の下で、潜在的因果変数は自明な順列とスケーリングまで識別可能であり、潜在変数のサブセットでこの参照条件が破られた場合でも部分識別可能性の結果が得られることが示せます。さらに、これらの理論的結果に基づいて、構造的因果変分オートエンコーダ(SuaVE)と呼ばれる新しい手法を提案します。この手法は、因果表現とそれらの間の因果関係、および潜在的因果変数から観測された因果変数へのマッピングを直接学習します。合成データと実データを用いた実験結果は、識別可能性と一貫性の結果、そして因果表現の学習におけるSuaVEの有効性を実証しています。
Classification Under Local Differential Privacy with Model Reversal and Model Averaging
Classification Under Local Differential Privacy with Model Reversal and Model Averaging / モデル反転とモデル平均化を用いた局所差分プライバシー下での分類
Local differential privacy has become a central topic in data privacy research, offering strong privacy guarantees by perturbing user data at the source and removing the need for a trusted curator. However, the noise introduced by local differential privacy often significantly reduces data utility. To address this issue, we reinterpret private learning under local differential privacy as a transfer learning problem, where the noisy data serve as the source domain and the unobserved clean data as the target. We propose novel techniques specifically designed for local differential privacy to improve classification performance without compromising privacy: (1) a noised binary feedback-based evaluation mechanism for estimating dataset utility; (2) model reversal, which salvages underperforming classifiers by inverting their decision boundaries; and (3) model averaging, which assigns weights to multiple reversed classifiers based on their estimated utility. We provide theoretical excess risk bounds under local differential privacy and demonstrate how our methods reduce this risk. Empirical results on both simulated and real-world datasets show substantial improvements in classification accuracy.
ローカル差分プライバシーは、データプライバシー研究の中心的なトピックとなっており、ユーザーデータをソースで撹乱し、信頼できるキュレーターの必要性を排除することで強力なプライバシー保証を提供します。しかし、ローカル差分プライバシーによって導入されるノイズは、多くの場合、データの有用性を大幅に低下させます。この問題に対処するため、我々はローカル差分プライバシー下でのプライバシー学習を転移学習問題として再解釈します。転移学習問題では、ノイズの多いデータがソースドメイン、観測されていないクリーンなデータがターゲットとなります。我々は、プライバシーを損なうことなく分類性能を向上させるため、ローカル差分プライバシー向けに特別に設計された以下の新しい手法を提案します。(1)データセットの有用性を推定するためのノイズ付きバイナリフィードバックベースの評価メカニズム、(2)決定境界を反転することで性能の低い分類器を救済するモデル反転、(3)推定された有用性に基づいて複数の反転分類器に重みを割り当てるモデル平均化。我々は、ローカル差分プライバシー下での理論的な過剰リスク境界を示し、我々の手法がどのようにこのリスクを軽減するかを示す。シミュレートされたデータセットと実際のデータセットの両方で実験した結果、分類精度が大幅に向上した。
Stochastic Gradient Methods: Bias, Stability and Generalization
Stochastic Gradient Methods: Bias, Stability and Generalization / 確率的勾配法:バイアス、安定性、汎化
Recent developments of stochastic optimization often suggest biased gradient estimators to improve either the robustness, communication efficiency or computational speed. Representative biased stochastic gradient methods (BSGMs) include Zeroth-order stochastic gradient descent (SGD), Clipped-SGD and SGD with delayed gradients. The practical success of BSGMs motivates a lot of convergence analysis to explain their impressive training behaviour. As a comparison, there is far less work on their generalization analysis, which is a central topic in modern machine learning. In this paper, we present the first framework to study the stability and generalization of BSGMs for convex and smooth problems. We introduce a generalized Lipschitz-type condition on gradient estimators and bias, under which we develop a rather general stability bound to show how the bias and the gradient estimators affect the stability. We apply our general result to develop the first stability bound for Zeroth-order SGD with reasonable step size sequences, and the first stability bound for Clipped-SGD. While our stability analysis is developed for general BSGMs, the resulting stability bounds for both Zeroth-order SGD and Clipped-SGD match those of SGD under appropriate smoothing/clipping parameters. We combine the stability and convergence analysis together, and derive excess risk bounds of order $O(1/\sqrt{n})$ for both Zeroth-order SGD and Clipped-SGD, where $n$ is the sample size.
確率的最適化の最近の発展は、堅牢性、通信効率、または計算速度のいずれかを向上させるために、バイアス勾配推定器を提案することが多い。代表的なバイアス付き確率的勾配法(BSGM)には、ゼロ次確率的勾配降下法(SGD)、Clipped-SGD、遅延勾配付きSGDなどがあります。BSGMの実用的な成功により、その優れたトレーニング動作を説明する収束解析が盛んに行われています。一方、現代の機械学習の中心的なトピックである一般化解析に関する研究ははるかに少ないです。本稿では、凸問題と滑らかな問題に対するBSGMの安定性と一般化を研究する初のフレームワークを紹介します。勾配推定値とバイアスに関する一般化Lipschitz型条件を導入し、その下で、バイアスと勾配推定値が安定性にどのように影響するかを示すためのかなり一般的な安定性境界を展開します。この一般的な結果を適用して、妥当なステップ サイズ シーケンスによるゼロ次SGDの最初の安定性境界と、Clipped-SGDの最初の安定性境界を展開します。我々の安定性解析は一般的なBSGMを対象に開発されているが、結果として得られるゼロ次SGDとClipped-SGDの安定性境界は、適切な平滑化/クリッピングパラメータを適用したSGDの安定性境界と一致しています。安定性解析と収束解析を組み合わせ、ゼロ次SGDとClipped-SGDの両方について、$O(1/\sqrt{n})$のオーダーの過剰リスク境界を導出します。ここで、$n$はサンプルサイズです。
Extending Mean-Field Variational Inference via Entropic Regularization: Theory and Computation
Extending Mean-Field Variational Inference via Entropic Regularization: Theory and Computation / エントロピー正則化による平均場変分推論の拡張:理論と計算
Variational inference (VI) has emerged as a popular method for approximate inference for high-dimensional Bayesian models. In this paper, we propose a novel VI method that extends the naive mean field via entropic regularization, referred to as $\Xi$-variational inference ($\Xi$-VI). $\Xi$-VI has a close connection to the entropic optimal transport problem and benefits from the computationally efficient Sinkhorn algorithm. We show that $\Xi$-variational posteriors effectively recover the true posterior dependency, where the likelihood function is downweighted by a regularization parameter. We analyze the role of dimensionality of the parameter space on the accuracy of $\Xi$-variational approximation and the computational complexity of computing the approximate distribution, providing a rough characterization of the statistical-computational trade-off in $\Xi$-VI, where higher statistical accuracy requires greater computational effort. We also investigate the frequentist properties of $\Xi$-VI and establish results on consistency, asymptotic normality, high-dimensional asymptotics, and algorithmic stability. We provide sufficient criteria for our algorithm to achieve polynomial-time convergence. Finally, we show the inferential benefits of using $\Xi$-VI over mean-field VI and other competing methods, such as normalizing flow, on simulated and real datasets.
変分推論(VI)は、高次元ベイズモデルの近似推論手法として広く普及しています。本稿では、エントロピー正則化を介してナイーブ平均場を拡張する新しいVI手法、$\Xi$変分推論($\Xi$-VI)を提案します。$\Xi$-VIはエントロピー最適輸送問題と密接な関係があり、計算効率の高いシンクホーンアルゴリズムの恩恵を受ける。$\Xi$変分事後分布は、尤度関数が正則化パラメータによって重み付けを下げられることで、真の事後依存性を効果的に回復することを示す。パラメータ空間の次元が$\Xi$変分近似の精度と近似分布の計算量に及ぼす影響を分析し、統計精度の向上にはより大きな計算量が必要となる$\Xi$-VIにおける統計的・計算的トレードオフの大まかな特徴づけを行う。また、$\Xi$-VIの頻度論的特性を調査し、一貫性、漸近正規性、高次元漸近性、アルゴリズムの安定性に関する結果を確立します。本アルゴリズムが多項式時間収束を達成するための十分な基準を提供します。最後に、シミュレーションおよび実データセットにおいて、平均場VIや正規化フローなどの他の競合手法よりも$\Xi$-VIを使用することの推論上の利点を示す。
skwdro: a library for Wasserstein distributionally robust machine learning
skwdro: a library for Wasserstein distributionally robust machine learning / skwdro:分布的にロバストなWasserstein機械学習用ライブラリ
We present skwdro, a Python library for training robust machine learning models.The library is based on distributionally robust optimization using Wasserstein distances, popular in optimal transport and machine learnings. The goal of the library is to make the training of robust models easier for a wide audience by proposing a wrapper for PyTorch modules, enabling model loss’ robustification with minimal code changes. It comes along with scikit-learn compatible estimators for some popular objectives. The core of the implementation relies on an entropic smoothing of the original robust objective, in order to ensure maximal model flexibility. The library is available at https://github.com/iutzeler/skwdro and the documentation at https://skwdro.readthedocs.io
堅牢な機械学習モデルのトレーニングのためのPythonライブラリであるskwdroを紹介します。このライブラリは、最適輸送および機械学習で人気のあるワッサースタイン距離を用いた分布的に堅牢な最適化に基づいています。このライブラリの目標は、PyTorchモジュールのラッパーを提案し、最小限のコード変更でモデル損失の堅牢化を可能にすることで、幅広いユーザーが堅牢なモデルのトレーニングを容易にすることです。いくつかの一般的な目的関数に対するscikit-learn互換の推定量が付属しています。実装の中核は、モデルの柔軟性を最大限に高めるために、元の堅牢な目的関数のエントロピー平滑化に依存しています。ライブラリはhttps://github.com/iutzeler/skwdroで、ドキュメントはhttps://skwdro.readthedocs.ioで入手できます。
Guaranteed Nonconvex Low-Rank Tensor Estimation via Scaled Gradient Descent
Guaranteed Nonconvex Low-Rank Tensor Estimation via Scaled Gradient Descent / スケール勾配降下法による保証付き非凸低ランクテンソル推定
Tensors, which give a faithful and effective representation to deliver the intrinsic structure of multi-dimensional data, play a crucial role in an increasing number of signal processing and machine learning problems. However, tensor data are often accompanied by arbitrary signal corruptions, including missing entries and sparse noise. A fundamental challenge is to reliably extract the meaningful information from corrupted tensor data in a statistically and computationally efficient manner. This paper develops a scaled gradient descent (ScaledGD) algorithm to directly estimate the tensor factors with tailored spectral initializations under the tensor-tensor product (t-product) and tensor singular value decomposition (t-SVD) framework. With tailored variants for tensor robust principal component analysis, (robust) tensor completion and tensor regression, we theoretically show that ScaledGD achieves linear convergence at a constant rate that is independent of the condition number of the ground truth low-rank tensor, while maintaining the low per-iteration cost of gradient descent. To the best of our knowledge, ScaledGD is the first algorithm that provably has such properties for low-rank tensor estimation with the t-SVD. Finally, numerical examples are provided to demonstrate the efficacy of ScaledGD in accelerating the convergence rate of ill-conditioned low-rank tensor estimation in a number of applications.
テンソルは、多次元データの本質的な構造を忠実かつ効果的に表現するものであり、信号処理や機械学習における様々な問題において重要な役割を果たしています。しかし、テンソルデータには、欠落エントリやスパースノイズといった、様々な信号破損が伴うことがよくあります。根本的な課題は、破損したテンソルデータから統計的かつ計算効率の高い方法で、意味のある情報を確実に抽出することです。本論文では、テンソル-テンソル積(t-product)およびテンソル特異値分解(t-SVD)フレームワークに基づき、調整されたスペクトル初期化を用いてテンソル因子を直接推定する、スケールド勾配降下法(ScaledGD)アルゴリズムを開発します。テンソルロバスト主成分分析、(ロバスト)テンソル補完、テンソル回帰といった様々な手法を用いて、ScaledGDが勾配降下法の反復ごとのコストを低く抑えつつ、真の低ランクテンソルの条件数に依存しない一定速度で線形収束を達成することを理論的に示す。我々の知る限り、ScaledGDはt-SVDを用いた低ランクテンソル推定においてこのような特性を証明できる最初のアルゴリズムです。最後に、様々なアプリケーションにおいて、悪条件の低ランクテンソル推定の収束速度を加速するScaledGDの有効性を示す数値例を提示します。
A Data-Augmented Contrastive Learning Approach to Nonparametric Density Estimation
A Data-Augmented Contrastive Learning Approach to Nonparametric Density Estimation / ノンパラメトリック密度推定へのデータ拡張対照学習アプローチ
In this paper, we introduce a data-augmented nonparametric noise contrastive estimation method to density estimation using deep neural networks. By leveraging the idea of contrastive learning, our density estimator exhibits efficiency with a one-step and simulation-free evaluation process, imposes no constraints on the neural network, and is shown to be consistent and asymptotically automatically normalized. A novel data augmentation procedure allows us to mitigate the influence of the choice of reference distribution on our method. Non-asymptotic upper bounds for the expected $L_{2}$-risk and the expected total variation distance have been established, which achieve minimax optimal rates. Moreover, our new method exhibits inherent adaptivity to low dimensional structures of data with a faster convergence rate under a compositional structure assumption. Numerical experiments show the competitiveness of our new method compared with the state-of-the-art nonparametric density estimation methods.
本稿では、ディープニューラルネットワークを用いた密度推定に、データ拡張型ノンパラメトリックノイズ対照推定法を導入します。対照学習の考え方を活用することで、本密度推定法は、ワンステップかつシミュレーション不要の評価プロセスで効率性を示し、ニューラルネットワークに制約を課さず、一貫性があり漸近的に自動的に正規化されることが示されます。新たなデータ拡張手順により、参照分布の選択が本手法に与える影響を軽減することができます。期待$L_{2}$-リスクと期待総変動距離の非漸近的上限が確立されており、これによりミニマックス最適速度が達成されます。さらに、本新手法は、構成構造仮定の下でより高速な収束速度を有し、データの低次元構造に固有の適応性を示す。数値実験は、最先端のノンパラメトリック密度推定法と比較して、本新手法の競争力を示しています。
Nonlocal Techniques for the Analysis of Deep ReLU Neural Network Approximations
Nonlocal Techniques for the Analysis of Deep ReLU Neural Network Approximations / Deep ReLUニューラルネットワーク近似の解析のための非局所的手法
In recent work concerned with the approximation and expressive powers of deep neural networks, Daubechies, DeVore, Foucart, Hanin, and Petrova introduced a system of piecewise linear functions, which can be easily reproduced by artificial neural networks with the ReLU activation function, and showed that it forms a Riesz basis of $L_2([0, 1])$. Their work was subsequently generalized to the multivariate setting by Schneider and Vybíral. In the work at hand, we show that this system serves as a Riesz basis also for Sobolev spaces $W^s([0,1]^d)$ and Barron classes ${\mathbb B}^s([0,1]^d)$ with smoothness $0\lt s\lt 1$. We apply this fact to re-prove some recent results on the approximation of functions from these classes by deep neural networks. Our proof method avoids using local approximations and also allows us to track the implicit constants as well as to show that we can avoid the curse of dimension. Moreover, we also study how well one can approximate Sobolev and Barron functions by neural networks if only function values are known.
深層ニューラルネットワークの近似と表現力に関する最近の研究において、Daubechies、DeVore、Foucart、Hanin、およびPetrovaは、ReLU活性化関数を用いた人工ニューラルネットワークによって容易に再現可能な区分線形関数のシステムを導入し、それが$L_2([0, 1])$のリース基底を形成することを示した。彼らの研究はその後、SchneiderとVybíralによって多変数設定に一般化されました。本研究では、このシステムがソボレフ空間$W^s([0,1]^d)$と滑らかさ$0\lt s\lt 1$を持つバロン類${\mathbb B}^s([0,1]^d)$に対してもリース基底として機能することを示す。この事実を適用して、深層ニューラルネットワークによるこれらのクラスの関数の近似に関する最近の結果をいくつか再証明します。我々の証明法は局所近似の使用を避け、暗黙の定数を追跡するとともに次元の呪いを回避できることを示す。さらに、関数値のみが既知である場合に、ニューラルネットワークによってソボレフ関数とバロン関数をどの程度正確に近似できるかについても検証します。
Nonlinear function-on-function regression by RKHS
Nonlinear function-on-function regression by RKHS / RKHSによる非線形関数対関数回帰
We propose a nonlinear function-on-function regression model where both the covariate and the response are random functions. The nonlinear regression is carried out in two steps: we first construct Hilbert spaces to accommodate the functional covariate and the functional response, and then build a second-layer Hilbert space for the covariate to capture nonlinearity. The second-layer space is assumed to be a reproducing kernel Hilbert space, which is generated by a positive definite kernel determined by the inner product of the first-layer Hilbert space for $X$–this structure is known as the nested Hilbert spaces. We develop estimation procedures to implement the proposed method, which allows the functional data to be observed at different time points for different subjects. Furthermore, we establish the convergence rate of our estimator as well as theweak convergence of the predicted response in the Hilbert space. Numerical studies including both simulations and a data application are conducted to investigate the performance of our estimator in finite sample.
我々は、共変量と応答がともにランダム関数である非線形関数対関数回帰モデルを提案します。この非線形回帰は2段階で実行されます。まず、関数共変量と関数応答を収容するヒルベルト空間を構築し、次に非線形性を捉えるために共変量のための第2層ヒルベルト空間を構築します。第2層空間は、$X$の第1層ヒルベルト空間の内積によって決定される正定値カーネルによって生成される再生カーネルヒルベルト空間であると仮定します。この構造はネストヒルベルト空間として知られています。提案手法を実装するための推定手順を開発し、これにより、異なる被験者について異なる時点で関数データを観測することができます。さらに、ヒルベルト空間における推定値の収束率と予測応答の弱収束性を確立します。有限サンプルにおける推定値の性能を調査するため、シミュレーションとデータ適用の両方を含む数値研究を実施します。
UQLM: A Python Package for Uncertainty Quantification in Large Language Models
UQLM: A Python Package for Uncertainty Quantification in Large Language Models / UQLM: 大規模言語モデルにおける不確実性定量化のためのPythonパッケージ
Hallucinations, defined as instances where Large Language Models (LLMs) generate false or misleading content, pose a significant challenge that impacts the safety and trust of downstream applications. We introduce UQLM, a Python package for LLM hallucination detection using state-of-the-art uncertainty quantification (UQ) techniques. This toolkit offers a suite of UQ-based scorers that compute response-level confidence scores ranging from 0 to 1. This library provides an off-the-shelf solution for UQ-based hallucination detection that can be easily integrated to enhance the reliability of LLM outputs.
幻覚は、大規模言語モデル(LLM)が虚偽または誤解を招くコンテンツを生成するインスタンスとして定義され、下流のアプリケーションの安全性と信頼性に影響を与える重大な課題となります。最先端の不確実性定量化(UQ)技術を用いたLLM幻覚検出用のPythonパッケージであるUQLMを紹介します。このツールキットは、0から1の範囲の応答レベルの信頼度スコアを計算するUQベースのスコアラースイートを提供します。このライブラリは、UQベースの幻覚検出のための既製のソリューションを提供し、簡単に統合してLLM出力の信頼性を高めることができます。
A Reinforcement Learning Approach in Multi-Phase Second-Price Auction Design
A Reinforcement Learning Approach in Multi-Phase Second-Price Auction Design / 多段階セカンドプライスオークション設計における強化学習アプローチ
We study reserve price optimization in multi-phase second price auctions, where the seller’s prior actions affect the bidders’ later valuations through a Markov Decision Process (MDP). Compared to the bandit setting in existing works, the setting in ours involves three challenges.First, from the seller’s perspective, we need to efficiently explore the environment in the presence of potentially untruthful bidders who aim to manipulate the seller’s policy.Second, we want to minimize the seller’s revenue regret when the market noise distribution is unknown. Third, the seller’s per-step revenue is an unknown, nonlinear random variable, and cannot even be directly observed from the environment but realized values.We propose a mechanism addressing all three challenges. To address the first challenge, we use a combination of a new technique named “buffer periods” and inspirations from Reinforcement Learning (RL) with low switching cost to limit bidders’ surplus from untruthful bidding, thereby incentivizing approximately truthful bidding. The second one is tackled by a novel algorithm that removes the need for pure exploration when the market noise distribution is unknown. The third challenge is resolved by an extension of LSVI-UCB, where we use the auction’s underlying structure to control the uncertainty of the revenue function. The three techniques culminate in the \underline{C}ontextual-\underline{L}SVI-\underline{U}CB-\underline{B}uffer (CLUB) algorithm which achieves $\tilde{\mathcal{O}}(H^{5/2}\sqrt{K})$ revenue regret, where $K$ is the number of episodes and $H$ is the length of each episode, when the market noise is known and $\tilde{\mathcal{O}}(H^{3}\sqrt{K})$ revenue regret when the noise is unknown with no assumptions on bidders’ truthfulness.
売り手の以前の行動がマルコフ決定過程(MDP)を通じて入札者のその後の評価に影響を与える、多段階のセカンドプライスオークションにおける最低落札価格の最適化を研究します。既存研究のバンディット設定と比較して、私たちの設定には3つの課題があります。まず、売り手の視点から、売り手のポリシーを操作しようとする潜在的に不正な入札者が存在する環境で、環境を効率的に探索する必要があります。次に、市場ノイズの分布が不明な場合に、売り手の収益後悔を最小限に抑えたいと考えています。3番目に、売り手のステップごとの収益は未知の非線形ランダム変数であり、環境から直接観察することさえできず、実現値となります。私たちは、3つの課題すべてに対処するメカニズムを提案します。最初の課題に対処するために、「バッファ期間」と呼ばれる新しい手法と、スイッチング コストの低い強化学習(RL)からのインスピレーションを組み合わせて、不正な入札による入札者の余剰を制限し、それによってほぼ正直な入札を奨励します。2番目の課題については、市場ノイズの分布が不明な場合に純粋な探索の必要性を排除する新しいアルゴリズムで対処します。3つ目の課題は、LSVI-UCBの拡張によって解決されます。この拡張では、オークションの基礎構造を使用して収益関数の不確実性を制御します。3つの手法は、\underline{C}ontextual-\underline{L}SVI-\underline{U}CB-\underline{B}uffer (CLUB)アルゴリズムに集約されます。このアルゴリズムは、市場ノイズが既知の場合は$\tilde{\mathcal{O}}(H^{5/2}\sqrt{K})$の収益後悔を実現します。ここで、$K$はエピソード数、$H$は各エピソードの長さです。また、入札者の誠実さに関する仮定がなく、ノイズが未知である場合は、$\tilde{\mathcal{O}}(H^{3}\sqrt{K})$の収益後悔を実現します。
Error Analysis for Deep ReLU Feedforward Density-Ratio Estimation with Bregman Divergence
Error Analysis for Deep ReLU Feedforward Density-Ratio Estimation with Bregman Divergence / ブレグマンダイバージェンスを用いたDeep ReLUフィードフォワード密度比推定の誤差分析
We consider the problem of density-ratio estimation using Bregman Divergence with Deep ReLU feedforward neural networks (BDD). We establish non-asymptotic error bounds for BDD density-ratio estimators, which are minimax optimal up to a logarithmic factor when the data distribution has finite support. As an application of our theoretical findings, we propose an estimator for the KL-divergence that is asymptotically normal, leveraging our convergence results for the deep density-ratio estimator and a data-splitting method. We also extend our results to cases with unbounded support and unbounded density ratios. Furthermore, we show that the BDD density-ratio estimator can mitigate the curse of dimensionality when data distributions are supported on an approximately low-dimensional manifold. Our results are applied to investigate the convergence properties of the telescoping density-ratio estimator proposed by Rhodes (2020). We provide sufficient conditions under which it achieves a lower error bound than a single-ratio estimator. Moreover, we conduct simulation studies to validate our main theoretical results and assess the performance of the BDD density-ratio estimator.
Deep ReLUフィードフォワードニューラルネットワーク(BDD)を使用したBregman Divergenceを用いた密度比推定の問題を検討します。データ分布が有限のサポートを持つ場合、対数係数までミニマックス最適となるBDD密度比推定量の非漸近的誤差境界を確立します。理論的発見の応用として、深層密度比推定量とデータ分割法の収束結果を利用して、漸近的に正規なKL情報量の推定量を提案します。また、結果を無制限のサポートと無制限の密度比を持つケースに拡張します。さらに、データ分布が近似的に低次元の多様体でサポートされている場合、BDD密度比推定量は次元の呪いを軽減できることを示す。私たちの結果は、Rhodes (2020)によって提案されたテレスコーピング密度比推定量の収束特性を調査するために適用されます。私たちは、単一比率推定量よりも低い誤差境界を達成するための十分な条件を提供します。さらに、主要な理論的結果を検証し、BDD密度比推定器のパフォーマンスを評価するためのシミュレーション研究を実施します。
Flexible Functional Treatment Effect Estimation
Flexible Functional Treatment Effect Estimation / 柔軟な機能的治療効果推定
We study treatment effect estimation with functional treatments where the average potential outcome functional is a function of functions, in contrast to continuous treatment effect estimation where the target is a function of real numbers. By considering a flexible scalar-on-function marginal structural model, a weight-modified kernel ridge regression (WMKRR) is adopted for estimation. The weights are constructed by directly minimizing the uniform balancing error resulting from a decomposition of the WMKRR estimator, instead of being estimated under a particular treatment selection model. Despite the complex structure of the uniform balancing error derived under WMKRR, finite-dimensional convex algorithms can be applied to efficiently solve for the proposed weights thanks to a representer theorem. The optimal convergence rate is shown to be attainable by the proposed WMKRR estimator without any smoothness assumption on the true weight function. Corresponding empirical performance is demonstrated by a simulation study and a real data application.
我々は、目標が実数の関数である連続的な治療効果推定とは対照的に、平均潜在的結果関数が関数の関数である関数治療による治療効果推定を研究します。柔軟なスカラー・オン・ファンクション周辺構造モデルを考慮することにより、重み修正カーネルリッジ回帰(WMKRR)が推定に採用されます。重みは、特定の治療選択モデルの下で推定されるのではなく、WMKRR推定量の分解から生じる均一バランシング誤差を直接最小化することによって構築されます。WMKRRの下で導出される均一バランシング誤差の複雑な構造にもかかわらず、代表者定理のおかげで、有限次元凸アルゴリズムを適用して提案された重みを効率的に解くことができます。最適な収束率は、真の重み関数にいかなる滑らかさの仮定もなしに、提案されたWMKRR推定量によって達成可能であることが示されています。対応する経験的性能は、シミュレーション研究と実際のデータ適用によって実証されています。
Neural Network Parameter-optimization of Gaussian Pre-marginalized Directed Acyclic Graphs
Neural Network Parameter-optimization of Gaussian Pre-marginalized Directed Acyclic Graphs / ガウス事前周辺化有向非巡回グラフのニューラルネットワークパラメータ最適化
Finding the parameters of a latent variable causal model is central to causal inference and causal identification. In this article, we show that existing graphical structures that are used in causal inference are not stable under marginalization of Gaussian Bayesian networks, and present a graphical structure that faithfully represents margins of Gaussian Bayesian networks. We present the first duality between parameter optimization of a latent variable model and training a feed-forward neural network in the parameter space of the assumed family of distributions. Based on this observation, we develop an algorithm for parameter optimization of these graphical structures using the observational distribution. Then, we provide conditions for causal effect identifiability in the Gaussian setting. We propose a meta-algorithm that checks whether a causal effect is identifiable or not. Moreover, we lay a grounding for generalizing the duality between a neural network and a causal model from the Gaussian to other distributions.
潜在変数因果モデルのパラメータを見つけることは、因果推論と因果同定において中心的な役割を果たします。本稿では、因果推論で使用される既存のグラフィカル構造が、ガウスベイジアンネットワークの周辺化の下では安定しないことを示し、ガウスベイジアンネットワークの周辺を忠実に表現するグラフィカル構造を提示します。潜在変数モデルのパラメータ最適化と、想定される分布族のパラメータ空間におけるフィードフォワードニューラルネットワークのトレーニングとの間に、最初の双対性を示します。この観察に基づいて、観測分布を用いてこれらのグラフィカル構造のパラメータ最適化アルゴリズムを開発します。次に、ガウス分布の設定において因果効果を識別可能であるための条件を示します。因果効果が識別可能かどうかを確認するメタアルゴリズムを提案します。さらに、ニューラルネットワークと因果モデル間の双対性をガウス分布から他の分布に一般化するための基礎を築きます。
Extrapolated Markov Chain Oversampling Method for Imbalanced Text Classification
Extrapolated Markov Chain Oversampling Method for Imbalanced Text Classification / 不均衡テキスト分類のための外挿マルコフ連鎖オーバーサンプリング法
Text classification is the task of automatically assigning text documents correct labels from a predefined set of categories. In real-life (text) classification tasks, observations and misclassification costs are often unevenly distributed between the classes – known as the problem of imbalanced data. Synthetic oversampling is a popular approach to imbalanced classification. The idea is to generate synthetic observations in the minority class to balance the classes in the training set. Many general-purpose oversampling methods can be applied to text data; however, imbalanced text data poses a number of distinctive difficulties that stem from the unique nature of text compared to other domains. One such factor is that when the sample size of text increases, the sample vocabulary (i.e., feature space) is likely to grow as well. We introduce a novel Markov chain based text oversampling method. The transition probabilities are estimated from the minority class but also partly from the majority class, thus allowing the minority feature space to expand in oversampling. We evaluate our approach against prominent oversampling methods and show that our approach is able to produce highly competitive results against the other methods in several real data examples, especially when the imbalance is severe.
テキスト分類とは、事前定義されたカテゴリセットからテキスト文書に正しいラベルを自動的に割り当てるタスクです。実際の(テキスト)分類タスクでは、観測値と誤分類コストがクラス間で不均等に分散していることが多く、これは不均衡データ問題として知られています。合成オーバーサンプリングは、不均衡分類に対する一般的なアプローチです。その考え方は、少数クラスで合成観測値を生成し、トレーニングセット内のクラス間のバランスをとることです。多くの汎用オーバーサンプリング手法をテキストデータに適用できますが、不均衡なテキストデータは、他の領域と比較したテキストの特異な性質に起因する多くの特有の問題を引き起こします。そのような要因の1つは、テキストのサンプルサイズが増加すると、サンプル語彙(すなわち、特徴空間)も増加する可能性があることです。本稿では、マルコフ連鎖に基づく新しいテキストオーバーサンプリング手法を紹介します。遷移確率は少数クラスから推定されますが、一部は多数クラスからも推定されるため、オーバーサンプリングにおいて少数クラスの特徴空間を拡張することができます。我々は、提案手法を主要なオーバーサンプリング手法と比較して評価し、特に不均衡が深刻な場合、いくつかの実データ例において他の手法に対して非常に競争力のある結果を生み出すことができることを示す。
An Anytime Algorithm for Good Arm Identification
An Anytime Algorithm for Good Arm Identification / 良好なアーム識別のためのいつでもアルゴリズム
In good arm identification (GAI), the goal is to identify one arm whose average performance exceeds a given threshold, referred to as a good arm, if it exists. Few works have studied GAI in the fixed-budget setting when the sampling budget is fixed beforehand, or in the anytime setting, when a recommendation can be asked at any time. We propose APGAI, an anytime and parameter-free sampling rule for GAI in stochastic bandits. APGAI can be straightforwardly used in fixed-confidence and fixed-budget settings. First, we derive upper bounds on its probability of error at any time. They show that adaptive strategies can be more efficient in detecting the absence of good arms than uniform sampling in several diverse instances. Second, when APGAI is combined with a stopping rule, we prove upper bounds on the expected sampling complexity, holding at any confidence level. Finally, we show the good empirical performance of APGAI on synthetic and real-world data. Our work offers an extensive overview of the GAI problem in all settings.
良好アーム識別(GAI)における目標は、平均パフォーマンスが所定の閾値を超える1つのアーム(良好アームと呼ばれる)を特定することです。サンプリング予算が事前に固定されている固定予算設定、またはいつでも推奨を尋ねられるいつでも設定でのGAIを研究した研究はほとんどない。我々は、確率的バンディットにおけるGAIのための、いつでもパラメータフリーのサンプリング規則であるAPGAIを提案します。APGAIは、固定信頼度および固定予算設定で直接使用できます。まず、任意の時点における誤り確率の上限を導出します。その結果、適応戦略は、複数の多様なインスタンスにおいて均一サンプリングよりも良好アームの不在を検出するのに効率的であることを示す。次に、APGAIを停止規則と組み合わせた場合、あらゆる信頼度レベルで成立する、予想されるサンプリング複雑度の上限を証明します。最後に、合成データおよび実世界データにおけるAPGAIの良好な経験的性能を示す。本研究は、あらゆる設定におけるGAI問題の広範な概要を提供します。
Simulation-based Calibration of Uncertainty Intervals under Approximate Bayesian Estimation
Simulation-based Calibration of Uncertainty Intervals under Approximate Bayesian Estimation / 近似ベイズ推定における不確実性区間のシミュレーションベースのキャリブレーション
The mean field variational Bayes (VB) algorithm implemented in Stan is relatively fast and efficient, making it feasible to produce model-estimated official statistics on a rapid timeline. Yet, while consistent point estimates of parameters are achieved for continuous data models, the mean field approximation often produces inaccurate uncertainty quantification to the extent that parameters are correlated a posteriori. In this paper, we propose a simulation procedure that calibrates uncertainty intervals for model parameters estimated under approximate algorithms to achieve nominal coverages. Our procedure detects and corrects biased estimation of both first and second moments of approximate marginal posterior distributions induced by any estimation algorithm that produces consistent first moments under specification of the correct model. The method generates replicate data sets using parameters estimated in an initial model run. The model is subsequently re-estimated on each replicate data set, and we use the empirical distribution over the re-samples to formulate calibrated confidence intervals of parameter estimates of the initial model run that are guaranteed to asymptotically achieve nominal coverage. We demonstrate the performance of our procedure in Monte Carlo simulation study and apply it to real data from the Current Employment Statistics survey.
Stanに実装された平均場変分ベイズ(VB)アルゴリズムは比較的高速かつ効率的であり、モデル推定に基づく公式統計量を迅速に生成することが可能になります。しかしながら、連続データモデルではパラメータの一貫した点推定値が得られる一方で、平均場近似は、パラメータが事後的に相関する程度まで不正確な不確実性定量化をもたらすことが多い。本論文では、近似アルゴリズムを用いて推定されたモデルパラメータの不確実性区間を較正し、名義カバレッジを達成するシミュレーション手順を提案します。本手順は、正しいモデルの仕様に基づいて一貫した第一モーメントを生成する推定アルゴリズムによって引き起こされる、近似周辺事後分布の第一モーメントと第二モーメントの両方の偏った推定値を検出し、修正します。この手法は、初期モデル実行で推定されたパラメータを用いて複製データセットを生成します。その後、各複製データセットでモデルを再推定し、再標本における経験分布を用いて、初期モデル実行時のパラメータ推定値の較正済み信頼区間を定式化します。この信頼区間は、漸近的に名目カバレッジを達成することが保証されています。本手法の性能をモンテカルロシミュレーション研究で実証し、これを「雇用統計調査」の実データに適用します。
Learning Bayesian Network Classifiers to Minimize Class Variable Parameters
Learning Bayesian Network Classifiers to Minimize Class Variable Parameters / クラス変数パラメータを最小化するベイジアンネットワーク分類器の学習
This study proposes and evaluates a novel Bayesian network classifier which can asymptotically estimate the true probability distribution of the class variable with the fewest class variable parameters among all structures for which the class variable has no parent. Moreover, to search for an optimal structure of the proposed classifier, we propose (1) a depth-first search based method and (2) an integer programming based method. The proposed methods are guaranteed to obtain the true probability distribution asymptotically while minimizing the number of class variable parameters. Comparative experiments using benchmark datasets demonstrate the effectiveness of the proposed method.
本研究では、クラス変数が親を持たないすべての構造の中で、クラス変数パラメータが最も少ないクラス変数の真の確率分布を漸近的に推定できる、新しいベイジアンネットワーク分類器を提案し、評価します。さらに、提案する分類器の最適な構造を探索するために、(1)深さ優先探索に基づく手法と(2)整数計画法に基づく手法を提案します。提案手法は、クラス変数パラメータの数を最小化しながら、真の確率分布を漸近的に得ることを保証します。ベンチマークデータセットを用いた比較実験は、提案手法の有効性を実証します。
Nonparametric Estimation of a Factorizable Density using Diffusion Models
Nonparametric Estimation of a Factorizable Density using Diffusion Models / 拡散モデルを用いた因数分解可能な密度のノンパラメトリック推定
In recent years, diffusion models, and more generally score-based deep generative models, have achieved remarkable success in various applications, including image and audio generation. In this paper, we view diffusion models as an implicit approach to nonparametric density estimation and study them within a statistical framework to analyze their surprising performance. A key challenge in high-dimensional statistical inference is leveraging low-dimensional structures inherent in the data to mitigate the curse of dimensionality. We assume that the underlying density exhibits a low-dimensional structure by factorizing into low-dimensional components, a property common in examples such as Bayesian networks and Markov random fields. Under suitable assumptions, we demonstrate that an implicit density estimator constructed from diffusion models adapts to the factorization structure and achieves the minimax optimal rate with respect to the total variation distance. In constructing the estimator, we design a sparse weight-sharing neural network architecture, where sparsity and weight-sharing are key features of practical architectures such as convolutional neural networks and recurrent neural networks.
近年、拡散モデル、そしてより一般的にはスコアベースの深層生成モデルは、画像や音声の生成を含む様々なアプリケーションで目覚ましい成功を収めています。本稿では、拡散モデルを非パラメトリック密度推定への暗黙的なアプローチと捉え、統計的枠組みの中でその驚くべき性能を分析します。高次元統計的推論における重要な課題は、データに内在する低次元構造を活用して次元の呪いを軽減することです。我々は、ベイジアンネットワークやマルコフ確率場などの例でよく見られる特性である、低次元成分への因数分解によって基礎密度が低次元構造を示すと仮定します。適切な仮定の下で、拡散モデルから構築された暗黙的な密度推定量が因数分解構造に適応し、総変動距離に関してミニマックス最適速度を達成することを示す。推定量の構築にあたり、我々はスパースな重み共有ニューラルネットワークアーキテクチャを設計します。ここで、スパース性と重み共有は、畳み込みニューラルネットワークやリカレントニューラルネットワークなどの実用的なアーキテクチャの重要な特徴です。
The Distribution of Ridgeless Least Squares Interpolators
The Distribution of Ridgeless Least Squares Interpolators / リッジレス最小二乗補間器の分布
The Ridgeless minimum $\ell_2$-norm interpolator in overparametrized linear regression has attracted considerable attention in recent years in both machine learning and statistics communities. While it seems to defy conventional wisdom that overfitting leads to poor prediction, recent theoretical research on its $\ell_2$-type risks reveals that its norm minimizing property induces an `implicit regularization’ that helps prediction in spite of interpolation.This paper takes a further step that aims at understanding its precise stochastic behavior as a statistical estimator. Specifically, we characterize the distribution of the Ridgeless interpolator in high dimensions, in terms of a Ridge estimator in an associated Gaussian sequence model with positive regularization, which provides a precise quantification of the prescribed implicit regularization in the most general distributional sense. Our distributional characterizations hold for general non-Gaussian random designs and extend uniformly to positively regularized Ridge estimators.As a direct application, we obtain a complete characterization for a general class of weighted $\ell_q$ risks of the Ridge(less) estimators that are previously only known for $q=2$ by random matrix methods. These weighted $\ell_q$ risks not only include the standard prediction and estimation errors, but also include the non-standard covariate shift settings. Our uniform characterizations further reveal a surprising feature of the commonly used generalized and $k$-fold cross-validation schemes: tuning the estimated $\ell_2$ prediction risk by these methods alone lead to simultaneous optimal $\ell_2$ in-sample, prediction and estimation risks, as well as the optimal length of debiased confidence intervals.
過パラメータ化線形回帰におけるリッジレス最小$\ell_2$ノルム補間器は、近年、機械学習と統計学の両方のコミュニティで大きな注目を集めています。過剰適合は予測精度の低下につながるという通説に反するように思われるが、$\ell_2$型リスクに関する最近の理論的研究では、そのノルム最小化特性が「暗黙の正則化」を誘導し、補間にもかかわらず予測精度を向上させることが明らかにされています。本論文では、統計的推定量としてのその正確な確率的挙動を理解することを目指し、さらに一歩踏み込む。具体的には、高次元におけるリッジレス補間量の分布を、正の正則化を伴うガウス系列モデルにおけるリッジ推定量を用いて特徴付ける。このリッジ推定量は、最も一般的な分布的意味で、規定された暗黙の正則化を正確に定量化します。我々の分布的特徴付けは、一般的な非ガウスランダム設計に当てはまり、正の正則化されたリッジ推定量に一様に拡張されます。直接的な応用として、これまでランダム行列法によって$q=2$の場合のみ知られていた、リッジ(レス)推定量の重み付き$\ell_q$リスクの一般的なクラスに対する完全な特徴付けを得る。これらの重み付き$\ell_q$リスクには、標準的な予測誤差と推定誤差だけでなく、非標準的な共変量シフト設定も含まれます。我々の均一な特性評価は、一般的に用いられる一般化クロスバリデーションと$k$分割クロスバリデーションの驚くべき特徴をさらに明らかにします。これらの手法のみで推定された$\ell_2$予測リスクを調整すると、標本内リスク、予測リスク、推定リスク、そしてバイアス除去信頼区間の最適な長さが同時に得られます。
LazyDINO: Fast, Scalable, and Efficiently Amortized Bayesian Inversion via Structure-Exploiting and Surrogate-Driven Measure Transport
LazyDINO: Fast, Scalable, and Efficiently Amortized Bayesian Inversion via Structure-Exploiting and Surrogate-Driven Measure Transport / LazyDINO: 構造利用とサロゲート駆動の測度輸送による高速、スケーラブル、かつ効率的な償却ベイジアン逆変換
We present LazyDINO, a transport map variational inference method for fast, scalable, and efficiently amortized solutions of high-dimensional nonlinear Bayesian inverse problems with expensive parameter-to-observable (PtO) maps. Our method consists of an offline phase, in which we construct a derivative-informed neural surrogate of the PtO map using joint samples of the PtO map and its Jacobian as training data. During the online phase, when given observational data, we rapidly approximate the posterior using surrogate-driven training of a lazy map, i.e., a structure-exploiting transport map with low-dimensional nonlinearity. Our surrogate construction is optimized for amortized Bayesian inversion using lazy map variational inference. We show that (i) the derivative-based reduced basis architecture minimizes an upper bound on the expected error in surrogate posterior approximation, and (ii) the derivative-informed surrogate training minimizes the expected error due to surrogate-driven variational inference. Our numerical results demonstrate that LazyDINO is highly efficient in cost amortization for Bayesian inversion. We observe a reduction of one to two orders of magnitude in offline cost for accurate online posterior approximation, compared to amortized simulation-based inference via conditional transport and to conventional surrogate-driven transport. In particular, LazyDINO consistently outperforms Laplace approximation using fewer than 1000 offline PtO map evaluations, while competing methods struggle and sometimes fail at 16,000 evaluations.
我々は、高次元非線形ベイズ逆問題(パラメータ観測値(PtO)マップを扱う)に対し、高速かつスケーラブルで効率的な償却解を求める輸送マップ変分推論手法LazyDINOを提案します。本手法は、PtOマップとそのヤコビアンを結合したサンプルを学習データとして用い、PtOマップの微分情報に基づくニューラルサロゲートを構築するオフラインフェーズから構成されます。オンラインフェーズでは、観測データが与えられた際に、サロゲート駆動型の遅延マップ(低次元非線形性を持つ構造利用輸送マップ)の学習を用いて事後分布を迅速に近似します。このサロゲート構築は、遅延マップ変分推論を用いた償却ベイズ逆問題に最適化されています。(i)微分ベースの縮小基底アーキテクチャが代理事後近似における期待誤差の上限を最小化し、(ii)微分情報に基づく代理学習が代理駆動変分推論による期待誤差を最小化することを示します。数値結果は、LazyDINOがベイズ逆変換のコスト償却において非常に効率的であることを示しています。条件付きトランスポートを介した償却シミュレーションベースの推論や従来の代理駆動トランスポートと比較して、正確なオンライン事後近似のオフラインコストが1~2桁削減されることが確認できました。特に、LazyDINOは1,000回未満のオフラインPtOマップ評価でラプラス近似よりも一貫して優れた性能を発揮しますが、競合手法は16,000回の評価で苦戦し、失敗することもあります。
A Common Interface for Automatic Differentiation
A Common Interface for Automatic Differentiation / 自動微分のための共通インターフェース
For scientific machine learning tasks with a lot of custom code, picking the right Automatic Differentiation (AD) system matters. Our Julia package DifferentiationInterface.jl provides a common frontend to a dozen AD backends, unlocking easy comparison and modular development. In particular, its built-in preparation mechanism leverages the strengths of each backend by amortizing one-time computations. This is key to enabling sophisticated features like sparsity handling without putting additional burdens on the user.
カスタムコードを多く含む科学的機械学習タスクでは、適切な自動微分(AD)システムを選択することが重要です。JuliaパッケージDifferentiationInterface.jlは、12種類のADバックエンドに共通のフロントエンドを提供し、容易な比較とモジュール開発を可能にします。特に、組み込みの準備メカニズムは、1回限りの計算を償却することで各バックエンドの長所を活用します。これは、ユーザーに追加の負担をかけることなく、スパース性処理などの高度な機能を実現するための鍵となります。
Refined Risk Bounds for Unbounded Losses via Transductive Priors
Refined Risk Bounds for Unbounded Losses via Transductive Priors / トランスダクティブ事前分布を用いた無制限損失のリスク境界の精緻化
We revisit the sequential variants of linear regression with the squared loss, classification problems with hinge loss, and logistic regression, all characterized by unbounded losses in the setup where no assumptions are made on the magnitude of design vectors and the norm of the optimal vector of parameters. The key distinction from existing results lies in our assumption that the set of design vectors is known in advance (though their order is not), a setup sometimes referred to as transductive online learning. While this assumption might seem similar to fixed design regression or denoising, we demonstrate that the sequential nature of our algorithms allows us to convert our bounds into statistical ones with random design without making any additional assumptions about the distribution of the design vectors-an impossibility for standard denoising results. Our key tools are based on the exponential weights algorithm with carefully chosen transductive (design-dependent) priors, which exploit the full horizon of the design vectors, as well as additional aggregation tools that address the possibly unbounded norm of the vector of the optimal solution.Our classification regret bounds have a feature that is only attributed to bounded losses in the literature: they depend solely on the dimension of the parameter space and on the number of rounds, independent of the design vectors or the norm of the optimal solution. For linear regression with squared loss, we further extend our analysis to the sparse case, providing sparsity regret bounds that depend additionally only on the magnitude of the response variables. We argue that these improved bounds are specific to the transductive setting and unattainable in the worst-case sequential setup.Our algorithms, in several cases, have polynomial-time approximations and reduce to sampling with respect to log-concave measures instead of aggregating over hard-to-construct epsilon-covers of classes.
二乗損失を伴う線形回帰のシーケンシャルバリアント、ヒンジ損失を伴う分類問題、ロジスティック回帰を再検討します。これらはすべて、設計ベクトルの大きさと最適なパラメータベクトルのノルムについて仮定を設けない設定において、無制限の損失を特徴とします。既存の結果との主な違いは、設計ベクトルのセットが事前にわかっている(ただし、順序はわからない)という仮定にあります。この設定は、トランスダクティブオンライン学習と呼ばれることもあります。この仮定は固定設計回帰やノイズ除去に似ているように見えるかもしれませんが、本アルゴリズムの逐次的な性質により、設計ベクトルの分布に関する追加の仮定を立てることなく、境界をランダム設計による統計的な境界に変換できることを実証します。これは標準的なノイズ除去結果では不可能なことです。主要ツールは、設計ベクトルの全範囲を活用する、慎重に選択されたトランスダクティブ(設計依存)事前分布を用いた指数重みアルゴリズムと、最適解のベクトルのノルムが非有界である可能性に対処する追加の集約ツールに基づいています。本分類リグレット境界には、文献では有界損失にのみ関連付けられている特徴があります。それは、パラメータ空間の次元とラウンド数のみに依存し、設計ベクトルや最適解のノルムとは独立しているということです。損失を2乗した線形回帰については、スパースケースに分析をさらに拡張し、応答変数の大きさのみにも依存するスパースリグレット境界を提供します。私たちは、これらの改善された境界はトランスダクティブ設定に特有のものであり、最悪の場合の順次設定では達成できないと主張します。私たちのアルゴリズムは、いくつかのケースで多項式時間の近似を持ち、構築が難しいクラスのイプシロンカバーにわたって集約するのではなく、対数凹測度に関するサンプリングに簡略化されます。
Decorrelated Local Linear Estimator: Inference for Non-linear Effects in High-dimensional Additive Models
Decorrelated Local Linear Estimator: Inference for Non-linear Effects in High-dimensional Additive Models / 無相関局所線形推定量: 高次元加法モデルにおける非線形効果の推論
Additive models play an essential role in studying non-linear relationships. Despite many recent advances in estimation, there is a lack of methods and theories for inference in high-dimensional additive models, including confidence interval construction and hypothesis testing. Motivated by inference for non-linear treatment effects, we consider the high-dimensional additive model and make inferences for the function derivative. We propose a novel decorrelated local linear estimator and establish its asymptotic normality. The main novelty is the construction of the decorrelation weights, which is instrumental in reducing the error inherited from estimating the nuisance functions in the high-dimensional additive model. We construct the confidence interval for the function derivative and conduct the related hypothesis testing. We demonstrate our proposed method over large-scale simulation studies and apply it to identify non-linear effects in the motif regression problem. Our proposed method is implemented in the R package DLL available from CRAN.
加法モデルは、非線形関係の研究において重要な役割を果たします。近年、推定法は大きく進歩していますが、高次元加法モデルにおける推論、特に信頼区間の構築や仮説検定といった手法や理論は未だ確立されていません。非線形治療効果の推論を目的として、本研究では高次元加法モデルを考察し、関数微分について推論を行います。本研究では、新たな非相関局所線形推定量を提案し、その漸近正規性を確立します。主な新規性は、高次元加法モデルにおけるニューサンス関数の推定から生じる誤差を低減する上で有効な、非相関重みの構築にあります。本研究では、関数微分の信頼区間を構築し、関連する仮説検定を実施します。提案手法を大規模シミュレーション研究で実証し、モチーフ回帰問題における非線形効果の特定に適用します。提案手法は、CRANから入手可能なRパッケージDLLに実装されています。
Communication-efficient Distributed Statistical Inference for Massive Data with Heterogeneous Auxiliary Information
Communication-efficient Distributed Statistical Inference for Massive Data with Heterogeneous Auxiliary Information / 異質な補助情報を持つ大規模データに対する通信効率の高い分散統計推論
Heterogeneous auxiliary information commonly arises in big data due to diverse study settings and privacy constraints. Excluding such indirect evidence often results in a substantial loss of statistical inference efficiency. This article proposes a novel framework for integrating a mixture of individual-level data and multiple external heterogeneous summary statistics by multiplying likelihood functions and confidence densities. Theoretically, we show that the proposed method possesses desirable properties and can achieve statistical efficiency comparable to that of the individual participant data (IPD) estimator, which uses all available individual-level data. Furthermore, we develop a communication-efficient distributed inference procedure for massive datasets with heterogeneous auxiliary information. We demonstrate that the proposed iterative algorithm achieves linear convergence under general conditions or generalized linear models. Finally, extensive simulations and real data applications are conducted to illustrate the performance of the proposed methods.
ビッグデータでは、多様な研究設定やプライバシー制約により、異質な補助情報が一般的に発生します。このような間接的な証拠を除外すると、統計的推論の効率が大幅に低下することがよくあります。本稿では、尤度関数と信頼密度を乗算することにより、個人レベルのデータと複数の外部の異種要約統計量の混合物を統合するための新しいフレームワークを提案します。理論的には、提案手法が望ましい特性を備え、利用可能なすべての個人レベルのデータを使用する個人参加者データ(IPD)推定量に匹敵する統計効率を達成できることを示します。さらに、異種の補助情報を含む大規模データセットに対して、通信効率の高い分散推論手順を開発します。提案する反復アルゴリズムは、一般的な条件または一般化線形モデルの下で線形収束を達成することを実証します。最後に、提案手法の性能を示すために、広範なシミュレーションと実際のデータへの適用を実施します。
Generative Bayesian Inference with GANs
Generative Bayesian Inference with GANs / GANを用いた生成ベイジアン推論
In the absence of explicit or tractable likelihoods, Bayesians often resort to approximate Bayesian computation (ABC) for inference. Our work bridges ABC with deep neural implicit samplers based on generative adversarial networks (GANs) and adversarial variational Bayes. Both ABC and GANs compare aspects of observed and fake data to simulate from posteriors and likelihoods, respectively. We develop a Bayesian GAN (B-GAN) sampler that directly targets the posterior by solving an adversarial optimization problem. B-GAN is driven by a deterministic mapping learned on the ABC reference by conditional GANs.Once the mapping has been trained, iid posterior samples are obtained by filtering noise at a negligible additional cost. We propose two post-processing local refinements using (1) data-driven proposals with importance reweighting, and (2) variational Bayes. We support our findings with frequentist-Bayesian results, showing that the typical total variation distance between the true and approximate posteriors converges to zero for certain neural network generators and discriminators. Our findings on simulated data show highly competitive performance relative to some of the most recent likelihood-free posterior simulators.
明示的または扱いやすい尤度がない場合、ベイズ主義者は推論のために近似ベイズ計算(ABC)に頼ることがよくあります。本研究では、ABCと、生成的敵対ネットワーク(GAN)および敵対的変分ベイズに基づくディープニューラル暗黙的サンプラーを橋渡しします。ABCとGANはどちらも、観測データと偽データの側面を比較し、それぞれ事後分布と尤度からシミュレートします。我々は、敵対的最適化問題を解決することで事後分布を直接ターゲットとするベイジアンGAN(B-GAN)サンプラーを開発しました。B-GANは、条件付きGANによってABC参照で学習された決定論的マッピングによって駆動されます。マッピングがトレーニングされると、わずかな追加コストでノイズをフィルタリングすることで、iid事後分布サンプルが得られます。我々は、(1)重要度の再重み付けを伴うデータ駆動型提案、および(2)変分ベイズを使用した2つの後処理ローカル改良を提案します。我々は、頻度主義ベイジアンの結果によって調査結果を裏付け、特定のニューラルネットワークジェネレーターと識別器について、真の事後分布と近似事後分布間の典型的な総変動距離がゼロに収束することを示しました。シミュレートされたデータに関する調査結果は、最新の尤度フリー事後分布シミュレータのいくつかと比較して、非常に競争力のあるパフォーマンスを示しています。
Exploring Novel Uncertainty Quantification through Forward Intensity Function Modeling
Exploring Novel Uncertainty Quantification through Forward Intensity Function Modeling / 順方向強度関数モデリングによる新たな不確実性定量化の探求
Predicting future time-to-event outcomes is a foundational task in statistical learning. While various methods exist for generating point predictions, quantifying the associated uncertainties poses a more substantial challenge. In this study, we introduce an innovative approach specifically designed to address this challenge, accommodating dynamic predictors that may manifest as stochastic processes. Our investigation harnesses the forward intensity function in a novel way, providing a fresh perspective on this intricate problem. The framework we propose demonstrates remarkable computational efficiency, enabling efficient analyses of large-scale investigations. We validate its soundness with theoretical guarantees, and our in-depth analysis establishes the weak convergence of function-valued parameter estimations. We illustrate the effectiveness of our framework with two comprehensive real examples and extensive simulation studies.
将来の事象発生までの時間を予測することは、統計学習における基礎的な課題です。点予測を生成するための様々な手法が存在しますが、関連する不確実性を定量化することは、より大きな課題となります。本研究では、この課題に対処するために特別に設計された革新的なアプローチを導入し、確率過程として現れる可能性のある動的な予測変数に対応します。本研究では、前向き強度関数を新たな方法で利用することで、この複雑な問題に新たな視点を提供します。提案するフレームワークは、優れた計算効率を示し、大規模な調査の効率的な分析を可能にします。その健全性は理論的な保証によって検証され、詳細な分析によって関数値パラメータ推定の弱収束が確立されます。2つの包括的な実例と広範なシミュレーション研究を用いて、このフレームワークの有効性を示します。
Persistence Diagrams Estimation of Multivariate Piecewise Hölder-continuous Signals
Persistence Diagrams Estimation of Multivariate Piecewise Hölder-continuous Signals / 多変量区分ヘルダー連続信号のパーシスタンス図推定
To our knowledge, the analysis of convergence rates for persistence diagrams estimation from noisy signals has predominantly relied on lifting signal estimation results through sup-norm (or other functional norm) stability theorems. We believe that moving forward from this approach can lead to considerable gains. We illustrate it in the setting of nonparametric regression. From a minimax perspective, we examine the inference of persistence diagrams (for the sublevel sets filtration). We show that for piecewise Hölder-continuous functions, with control over the reach of the set of discontinuities, taking the persistence diagram coming from a simple histogram estimator of the signal permits achieving the minimax rates known for Hölder-continuous functions. The key novelty lies in our use of algebraic stability instead of sup-norm stability, directly targeting the bottleneck distance through the underlying interleaving. This allows us to incorporate deformation retractions of sublevel sets to accommodate boundary discontinuities that cannot be handled by sup-norm based stability analyses.
私たちの知る限り、ノイズの多い信号からのパーシスタンス図推定の収束率の分析は、主に、ノルム超(またはその他の関数ノルム)安定性定理を用いて信号推定結果をリフティングすることに頼ってきました。このアプローチから前進することで、大きな成果が得られると考えています。これをノンパラメトリック回帰の設定で説明します。ミニマックスの観点から、パーシスタンス図(サブレベルセットのフィルタリング用)の推論を検討します。区分ヘルダー連続関数において、不連続点の到達範囲を制御しながら、信号の単純なヒストグラム推定値からパーシスタンス図を得ることで、ヘルダー連続関数で知られているミニマックス率を達成できることを示す。重要な新規性は、超ノルム安定性ではなく代数的安定性を用い、基礎となるインターリーブを通してボトルネック距離を直接ターゲットにしている点にあります。これにより、超ノルムベースの安定性解析では扱えない境界不連続点に対応するために、サブレベルセットの変形後退を組み込むことができます。
CHANI: Correlation-based Hawkes Aggregation of Neurons with bio-Inspiration
CHANI: Correlation-based Hawkes Aggregation of Neurons with bio-Inspiration / CHANI: 生物に着想を得たニューロンの相関に基づくHawkes集約
The present work aims at proving mathematically that a neural network inspired by biology can learn a classification task thanks to local transformations only. In this purpose, we propose a spiking neural network named CHANI (Correlation-based Hawkes Aggregation of Neurons with bio-Inspiration), whose neurons activity is modeled by Hawkes processes. Synaptic weights are updated thanks to an expert aggregation algorithm, providing a local and simple learning rule. We were able to prove that our network can learn on average and asymptotically. Moreover, we demonstrated that it automatically produces neuronal assemblies in the sense that the network can encode several classes and that a same neuron in the intermediate layers might be activated by more than one class, and we provided numerical simulations on synthetic datasets. This theoretical approach contrasts with the traditional empirical validation of biologically inspired networks and paves the way for understanding how local learning rules enable neurons to form assemblies able to represent complex concepts.
本研究の目的は、生物学に着想を得たニューラルネットワークが、局所的な変換のみによって分類タスクを学習できることを数学的に証明することです。この目的のため、我々はCHANI(Correlation-based Hawkes Aggregation of Neurons with bio-Inspiration)という名のスパイキングニューラルネットワークを提案します。このネットワークのニューロン活動はHawkes過程によってモデル化されます。シナプス荷重は、局所的かつシンプルな学習規則を提供するエキスパート集約アルゴリズムによって更新されます。我々は、このネットワークが平均的かつ漸近的に学習できることを証明できた。さらに、ネットワークが複数のクラスをエンコードでき、中間層の同じニューロンが複数のクラスによって活性化される可能性があるという意味で、ニューロンアセンブリを自動的に生成することを実証し、合成データセットを用いた数値シミュレーションを提供した。この理論的アプローチは、生物学的に着想を得たネットワークの従来の経験的検証とは対照的であり、局所的な学習規則によってニューロンが複雑な概念を表現できるアセンブリを形成できるようにする仕組みを理解するための道を開くものです。
Finite Neural Networks as Mixtures of Gaussian Processes: From Provable Error Bounds to Prior Selection
Finite Neural Networks as Mixtures of Gaussian Processes: From Provable Error Bounds to Prior Selection / ガウス過程の混合体としての有限ニューラルネットワーク:証明可能な誤差限界から事前選択へ
Infinitely wide or deep neural networks (NNs) with independent and identically distributed (i.i.d.) parameters have been shown to be equivalent to Gaussian processes. Because of the favorable properties of Gaussian processes, this equivalence is commonly employed to analyze neural networks and has led to various breakthroughs over the years. However, neural networks and Gaussian processes are equivalent only in the limit; in the finite case there are currently no methods available to approximate a trained neural network with a Gaussian model with bounds on the approximation error. In this work, we present an algorithmic framework to approximate a neural network of finite width and depth, and with not necessarily i.i.d. parameters, with a mixture of Gaussian processes with bounds on the approximation error. In particular, we consider the Wasserstein distance to quantify the closeness between probabilistic models and, by relying on tools from optimal transport and Gaussian processes, we iteratively approximate the output distribution of each layer of the neural network as a mixture of Gaussian processes. Crucially, for any NN and $\epsilon >0$ our approach is able to return a mixture of Gaussian processes that is $\epsilon$-close to the NN at a finite set of input points. Furthermore, we rely on the differentiability of the resulting error bound to show how our approach can be employed to tune the parameters of a NN to mimic the functional behavior of a given Gaussian process, e.g., for prior selection in the context of Bayesian inference. We empirically investigate the effectiveness of our results on both regression and classification problems with various neural network architectures. Our experiments highlight how our results can represent an important step towards understanding neural network predictions and formally quantifying their uncertainty.
独立かつ同一に分布する(i.i.d.)パラメータを持つ、無限に広く深いニューラルネットワーク(NN)は、ガウス過程と等価であることが示されています。ガウス過程の好ましい特性のため、この等価性はニューラルネットワークの解析に広く用いられ、長年にわたり様々なブレークスルーをもたらしてきました。しかし、ニューラルネットワークとガウス過程は極限においてのみ等価であり、有限の場合、近似誤差に境界を持つガウスモデルで訓練済みニューラルネットワークを近似する方法は現在のところ存在しません。本研究では、有限の幅と深さを持ち、必ずしもi.i.d.パラメータを持たないニューラルネットワークを、近似誤差に境界を持つガウス過程の混合で近似するアルゴリズムフレームワークを提示します。特に、確率モデル間の近似度を定量化するためにワッサースタイン距離を考慮し、最適輸送過程とガウス過程のツールを用いて、ニューラルネットワークの各層の出力分布をガウス過程の混合として反復的に近似します。重要な点は、任意のニューラルネットワークと$\epsilon >0$に対して、本手法は有限の入力点集合においてニューラルネットワークに$\epsilon$近似するガウス過程の混合を返すことができることです。さらに、得られた誤差境界の微分可能性を用いて、本手法を用いてニューラルネットワークのパラメータを調整し、例えばベイズ推論における事前選択など、特定のガウス過程の機能的挙動を模倣する方法を示す。様々なニューラルネットワークアーキテクチャを用いて、回帰問題と分類問題の両方において、本研究結果の有効性を実証的に検証します。本実験は、本研究結果がニューラルネットワーク予測の理解とその不確実性の形式的定量化に向けた重要な一歩となり得ることを示しています。
Optimization and Generalization of Gradient Descent for Shallow ReLU Networks with Minimal Width
Optimization and Generalization of Gradient Descent for Shallow ReLU Networks with Minimal Width / 最小幅の浅いReLUネットワークに対する勾配降下法の最適化と一般化
Understanding the generalization and optimization of neural networks is a longstanding problem in modern learning theory. The prior analysis often leads to risk bounds of order $1/\sqrt{n}$ for ReLU networks, where $n$ is the sample size. In this paper, we present a general optimization and generalization analysis for gradient descent applied to shallow ReLU networks. We develop convergence rates of the order $1/T$ for gradient descent with $T$ iterations, and show that the gradient descent iterates fall inside local balls around either an initialization point or a reference point. Then we develop improved Rademacher complexity estimates by using the activation pattern of the ReLU function in these local balls. We apply our general result to NTK-separable data with a margin $\gamma$, and develop an almost optimal risk bound of the order $1/(n\gamma^2)$ for the ReLU network with a polylogarithmic width.
ニューラルネットワークの一般化と最適化を理解することは、現代学習理論における長年の課題です。事前分析では、ReLUネットワークのリスク境界はしばしば$1/\sqrt{n}$オーダーに導かれる(ここで$n$はサンプルサイズ)。本論文では、浅いReLUネットワークに適用される勾配降下法の一般的な最適化および一般化分析を提示します。$T$回の反復を伴う勾配降下法の収束速度が$1/T$オーダーであることを明らかにし、勾配降下法の反復が初期化点または参照点の周囲の局所球内に収まることを示す。次に、これらの局所球におけるReLU関数の活性化パターンを用いることで、改良されたRademacher複雑度推定値を開発します。この一般的な結果を、マージン$\gamma$を持つNTK分離可能なデータに適用し、多重対数幅を持つReLUネットワークのほぼ最適なリスク境界を$1/(n\gamma^2)$オーダーで開発します。
Adaptive Forward Stepwise: A Method for High Sparsity Regression
Adaptive Forward Stepwise: A Method for High Sparsity Regression / 適応的前向きステップワイズ法:高スパース回帰のための手法
This paper proposes a sparse regression method that continuously interpolates between Forward Stepwise selection (FS) and the LASSO. When tuned appropriately, our solutions are much sparser than typical LASSO fits but, unlike FS fits, benefit from the stabilizing effect of shrinkage. Our method, Adaptive Forward Stepwise Regression (AFS) addresses the need for sparser models with shrinkage. We show its connection with boosting via a soft-thresholding viewpoint and demonstrate the ease of adapting the method to classification tasks. In both simulations and real data, our method has lower mean squared error and fewer selected features across multiple settings compared to popular sparse modeling procedures.
本論文では、前向きステップワイズ選択法(FS)とLASSOを連続的に補間するスパース回帰法を提案します。適切に調整することで、本手法は典型的なLASSOフィッティングよりもはるかにスパース性が高くなるが、FSフィッティングとは異なり、シュリンクによる安定化効果の恩恵を受ける。本手法である適応型前向きステップワイズ回帰法(AFS)は、シュリンクを伴うよりスパースなモデルの必要性に応える。ソフト閾値の観点からブースティングとの関連性を示し、本手法が分類タスクに容易に適応できることを実証します。シミュレーションと実データの両方において、本手法は一般的なスパースモデリング手法と比較して、平均二乗誤差が低く、複数の設定において選択される特徴量が少ないことが示されました。
Optimizing Attention with Mirror Descent: Generalized Max-Margin Token Selection
Optimizing Attention with Mirror Descent: Generalized Max-Margin Token Selection / ミラー降下法によるアテンションの最適化:一般化最大マージントークン選択
Attention mechanisms have revolutionized several domains of artificial intelligence, such as natural language processing and computer vision, by enabling models to selectively focus on relevant parts of the input data. While recent work has characterized the optimization dynamics of gradient descent (GD) in attention-based models and the structural properties of its preferred solutions, less is known about more general optimization algorithms such as mirror descent (MD). In this paper, we investigate the convergence properties and implicit biases of a family of MD algorithms tailored for softmax attention mechanisms, with the potential function chosen as the $p$-th power of the $\ell_p$-norm. Specifically, we show that these algorithms converge in direction to a generalized hard-margin SVM with an $\ell_p$-norm objective when applied to a classification problem using a softmax attention model. Notably, our theoretical results reveal that the convergence rate is comparable to that of traditional GD in simpler models, despite the highly nonlinear and nonconvex nature of the present problem. Additionally, we delve into the joint optimization dynamics of the key-query matrix and the decoder, establishing conditions under which this complex joint optimization converges to their respective hard-margin SVM solutions. Lastly, our numerical experiments on real data demonstrate that MD algorithms improve generalization over standard GD and excel in optimal token selection.
注意メカニズムは、モデルが入力データの関連部分に選択的に焦点を絞ることを可能にすることで、自然言語処理やコンピュータビジョンといった人工知能のいくつかの領域に革命をもたらしました。近年の研究では、注意ベースモデルにおける勾配降下法(GD)の最適化ダイナミクスとその推奨解の構造的特性が明らかにされていますが、ミラー降下法(MD)のようなより一般的な最適化アルゴリズムについてはあまり知られていません。本論文では、ポテンシャル関数として$\ell_p$ノルムの$p$乗を選択した、ソフトマックス注意メカニズム向けに調整されたMDアルゴリズム群の収束特性と暗黙的なバイアスを調査します。具体的には、これらのアルゴリズムをソフトマックス注意モデルを用いた分類問題に適用した場合、$\ell_p$ノルムを目的関数とする一般化ハードマージンSVMへと収束することを示す。特に、本問題が高度に非線形かつ非凸であるにもかかわらず、我々の理論的結果は、より単純なモデルにおいて、収束率が従来のGDと同等であることを示しています。さらに、キークエリ行列とデコーダーの同時最適化ダイナミクスを詳細に検討し、この複雑な同時最適化がそれぞれのハードマージンSVM解に収束する条件を確立します。最後に、実データを用いた数値実験により、MDアルゴリズムは標準的なGDよりも汎化性能が向上し、最適なトークン選択に優れていることが実証されました。
Hierarchical Causal Models
Hierarchical Causal Models / 階層的因果モデル
Causal questions often arise in settings where data are hierarchical: subunits are nested within units. Consider students in schools, cells in patients, or cities in states. In these settings, unit-level variables (e.g., a school’s budget) may affect subunit-level outcomes (e.g., student test scores), and subunit-level characteristics may aggregate to influence unit-level outcomes. In this paper, we show how to analyze hierarchical data for causal inference. We introduce hierarchical causal models, which extend structural causal models and graphical models by incorporating inner plates to represent nested data structures. We develop a graphical identification technique for these models that generalizes do-calculus. We show that hierarchical data can enable causal identification even when it would be impossible with non-hierarchical data–for example, when only unit-level summaries are available. We develop estimation strategies, including using hierarchical Bayesian models. We illustrate our results in simulation and through a reanalysis of the classic “eight schools” study.
因果関係に関する疑問は、データが階層的である設定でしばしば生じます。つまり、サブユニットはユニット内にネストされています。学校の生徒、患者の細胞、州の都市などを考えてみましょう。このような設定では、ユニットレベルの変数(例:学校の予算)がサブユニットレベルの結果(例:生徒のテストの点数)に影響を与える可能性があり、サブユニットレベルの特性が集約されてユニットレベルの結果に影響を与える可能性があります。本稿では、階層データを分析して因果推論を行う方法を示します。階層型因果モデルを紹介します。これは、ネストされたデータ構造を表現するために内部プレートを組み込むことで、構造因果モデルとグラフィカルモデルを拡張したものです。我々は、これらのモデルに対し、do計算を一般化するグラフィカル識別手法を開発します。階層型データを用いることで、非階層型データでは不可能な場合(例えば、単位レベルの要約しか利用できない場合)でも因果関係の識別が可能となることを示す。階層型ベイズモデルを用いた推定戦略も開発します。シミュレーションと、古典的な「8つの流派」研究の再分析によって、我々の結果を示す。
Reparameterized Complex-valued Neurons Can Efficiently Learn More than Real-valued Neurons via Gradient Descent
Reparameterized Complex-valued Neurons Can Efficiently Learn More than Real-valued Neurons via Gradient Descent / 再パラメータ化された複素数値ニューロンは、勾配降下法によって実数値ニューロンよりも効率的に学習できる
Complex-valued neural networks potentially possess better representations and performance than real-valued counterparts when dealing with some complicated tasks such as acoustic analysis, radar image classification, etc. Despite empirical successes, it remains unknown theoretically when and to what extent complex-valued neural networks outperform real-valued ones. We take one step in this direction by comparing the learnability of real-valued neurons and complex-valued neurons via gradient descent. We theoretically show that a complex-valued neuron can learn functions expressed by any one real-valued neuron and any one complex-valued neuron with convergence rates $O(t^{-3})$ and $O(t^{-1})$ where $t$ is the iteration index of gradient descent, respectively, whereas a two-layer real-valued neural network with finite width cannot learn a single non-degenerate complex-valued neuron. We prove that a complex-valued neuron learns a real-valued neuron with rate $\Omega (t^{-3})$, exponentially slower than the linear convergence rate of learning one real-valued neuron using a real-valued neuron. We then reparameterize the phase parameter of the complex-valued neuron and prove that a reparameterized complex-valued neuron can efficiently learn a real-valued neuron with a linear convergence rate. We further verify and extend these results via simulation experiments in more general settings.
複素値ニューラルネットワークは、音響解析、レーダー画像分類といった複雑なタスクを扱う際に、実数値ニューラルネットワークよりも優れた表現力と性能を持つ可能性があります。実験的な成功例があるにもかかわらず、複素値ニューラルネットワークが実数値ニューラルネットワークをどの程度、そしていつ、どのように凌駕するのかは理論的には未解明です。我々は、勾配降下法を用いて実数値ニューロンと複素値ニューロンの学習可能性を比較することで、この方向への一歩を踏み出す。我々は、複素ニューロンが、任意の1つの実数ニューロンと任意の1つの複素ニューロンで表される関数を、それぞれ収束速度$O(t^{-3})$と$O(t^{-1})$で学習できることを理論的に示す。ここで、$t$は勾配降下法の反復インデックスです。一方、有限幅の2層実数ニューラルネットワークは、単一の非退化複素ニューロンを学習することはできない。我々は、複素ニューロンが実数ニューロンを、実数ニューロンを用いて1つの実数ニューロンを学習する場合の線形収束速度よりも指数的に遅い速度$\Omega (t^{-3})$で学習することを証明した。次に、複素ニューロンの位相パラメータを再パラメータ化し、再パラメータ化された複素ニューロンが線形収束速度で実数ニューロンを効率的に学習できることを証明します。我々は、より一般的な設定でのシミュレーション実験により、これらの結果をさらに検証し、拡張します。
Unsupervised Feature Selection via Nonnegative Orthogonal Constrained Regularized Minimization
Unsupervised Feature Selection via Nonnegative Orthogonal Constrained Regularized Minimization / 非負直交制約付き正則化最小化による教師なし特徴選択
Unsupervised feature selection has drawn wide attention in the era of big data, since it serves as a fundamental technique for dimensionality reduction. However, many existing unsupervised feature selection models and solution methods are primarily designed for practical applications, and often lack rigorous theoretical support, such as convergence guarantees. In this paper, we first establish a novel unsupervised feature selection model based on regularized minimization with nonnegative orthogonality constraints, which has advantages of embedding feature selection into the nonnegative spectral clustering and preventing overfitting. To solve the proposed model, we develop an effective inexact augmented Lagrangian multiplier method, in which the subproblems are addressed using a proximal alternating minimization approach. We rigorously prove the algorithm’s sequence converges to a stationary point of the model. Extensive numerical experiments on popular datasets demonstrate the stability and robustness of our method. Moreover, comparative results show that our method outperforms some existing state-of-the-art methods in terms of clustering evaluation metrics. The code is available at https://github.com/liyan-amss/NOCRM_code.
教師なし特徴選択は、次元削減の基本的な手法として、ビッグデータ時代に広く注目を集めています。しかし、既存の多くの教師なし特徴選択モデルと解法は、主に実用アプリケーション向けに設計されており、収束保証などの厳密な理論的裏付けが不足していることが多いです。本稿では、まず、非負直交性制約を伴う正則化最小化に基づく新しい教師なし特徴選択モデルを確立します。このモデルは、特徴選択を非負スペクトルクラスタリングに組み込み、過学習を防ぐという利点があります。提案モデルを解くために、効果的な不正確拡張ラグランジュ乗数法を開発し、この方法では、部分問題に近似交代最小化アプローチを用いて対処します。アルゴリズムのシーケンスがモデルの定常点に収束することを厳密に証明します。一般的なデータセットを用いた広範な数値実験により、この方法の安定性と堅牢性が実証されています。さらに、比較結果から、本手法はクラスタリング評価指標の点で既存の最先端手法のいくつかよりも優れていることが示されています。コードはhttps://github.com/liyan-amss/NOCRM_codeで入手できます。
A causal fused lasso for interpretable heterogeneous treatment effects estimation
A causal fused lasso for interpretable heterogeneous treatment effects estimation / 解釈可能な異種治療効果推定のための因果融合Lasso
We propose a novel method for estimating heterogeneous treatment effects based on the fused lasso. By first ordering samples based on the propensity or prognostic score, we match units from the treatment and control groups. We then run the fused lasso to obtain piecewise constant treatment effects with respect to the ordering defined by the score. Similar to the existing methods based on discretizing the score, our methods yield interpretable subgroup effects. However, existing methods fixed the subgroup a priori, but our causal fused lasso forms data-adaptive subgroups. We show that the estimator consistently estimates the treatment effects conditional on the score under very general conditions on the covariates and treatment. We demonstrate the performance of our procedure using extensive experiments that show that it can be interpretable and competitive with state-of-the-art methods.
本手法では、融合Lassoに基づいて異質な治療効果を推定する新しい手法を提案します。まず、傾向スコアまたは予後スコアに基づいてサンプルを順序付けすることで、治療群と対照群のユニットをマッチングさせます。次に、融合Lassoを実行して、スコアによって定義された順序付けに関して区分的に一定の治療効果を取得します。スコアの離散化に基づく既存の手法と同様に、本手法は解釈可能なサブグループ効果をもたらします。ただし、既存の手法はサブグループを事前に固定していましたが、本手法の因果的融合Lassoはデータ適応型サブグループを形成します。本手法は、共変量と治療に関する非常に一般的な条件下で、スコアを条件として治療効果を一貫して推定できることを示します。我々は、広範な実験を用いて本手法の性能を実証し、それが解釈可能であり、最先端の手法と競合可能であることを示す。
Bayesian Inference of Contextual Bandit Policies via Empirical Likelihood
Bayesian Inference of Contextual Bandit Policies via Empirical Likelihood / ベイズ推論経験的尤度による文脈的バンディットポリシー
Policy inference plays an essential role in the contextual bandit problem. In this paper, we use empirical likelihood to develop a Bayesian inference method for the joint analysis of multiple contextual bandit policies in finite sample regimes. The proposed inference method is robust to small sample sizes and is able to provide accurate uncertainty measurements for policy value evaluation. In addition, it allows for flexible inferences on policy comparison with full uncertainty quantification. We demonstrate the effectiveness of the proposed inference method using Monte Carlo simulations and its application to an adolescent body mass index data set.
政策推論は、文脈バンディット問題において重要な役割を果たす。本稿では、経験尤度を用いて、有限サンプルレジームにおける複数の文脈バンディット政策の共同分析のためのベイズ推論手法を開発します。提案する推論手法は、小規模なサンプルサイズに対してロバストであり、政策の価値評価のための正確な不確実性測定を提供することができます。さらに、完全な不確実性定量化を伴う政策比較に関する柔軟な推論を可能にします。提案する推論手法の有効性は、モンテカルロシミュレーションと、青年期のBMIデータセットへの適用によって実証します。
Convergence and complexity of block majorization-minimization for constrained block-Riemannian optimization
Convergence and complexity of block majorization-minimization for constrained block-Riemannian optimization / 制約付きブロックリーマン最適化におけるブロック主要化最小化の収束と計算量
Block majorization-minimization (BMM) is a simple iterative algorithm for nonconvex optimization that sequentially minimizes a majorizing surrogate of the objective function in each block coordinate while the other block coordinates are held fixed. We consider a family of BMM algorithms for minimizing nonsmooth nonconvex objectives, where each parameter block is constrained within a subset of a Riemannian manifold. We establish that this algorithm converges asymptotically to the set of stationary points, and attains an $\epsilon$-stationary point within $\widetilde{O}(\epsilon^{-2})$ iterations. In particular, the assumptions for our complexity results are completely Euclidean when the underlying manifold is a product of Euclidean or Stiefel manifolds, although our analysis makes explicit use of the Riemannian geometry. Our general analysis applies to a wide range of algorithms with Riemannian constraints: Riemannian MM, block projected gradient descent, Bures-JKO scheme for Wasserstein variational inference, optimistic likelihood estimation, geodesically constrained subspace tracking, robust PCA, and Riemannian CP-dictionary-learning. We experimentally validate that our algorithm converges faster than standard Euclidean algorithms applied to the Riemannian setting.
ブロック・マジョライゼーション・ミニマライゼーション(BMM)は、非凸最適化のための単純な反復アルゴリズムであり、他のブロック座標を固定したまま、各ブロック座標における目的関数のマジョライゼーション代理関数を順次最小化します。本稿では、各パラメータブロックがリーマン多様体のサブセット内に制約される、非滑らかな非凸目的関数を最小化するBMMアルゴリズム群を検討します。このアルゴリズムは、定常点の集合に漸近収束し、$\widetilde{O}(\epsilon^{-2})$回の反復で$\epsilon$定常点に到達することを確立します。特に、本解析ではリーマン幾何学を明示的に用いているが、基礎となる多様体がユークリッド多様体またはシュティーフェル多様体の積である場合、複雑性の結果に対する仮定は完全にユークリッド的です。我々の一般的な分析は、リーマン制約を持つ幅広いアルゴリズム、すなわちリーマンMM、ブロック射影勾配降下法、ワッサーシュタイン変分推論のためのBures-JKO法、楽観的尤度推定、測地学的制約付き部分空間追跡、ロバストPCA、リーマンCP辞書学習に適用できます。我々のアルゴリズムは、リーマン制約に適用される標準的なユークリッドアルゴリズムよりも収束が速いことを実験的に検証しました。
Two-way Node Popularity Model for Directed and Bipartite Networks
Two-way Node Popularity Model for Directed and Bipartite Networks / 有向ネットワークおよび二部ネットワークにおける双方向ノード人気モデル
There has been increasing research attention on community detection in directed and bipartite networks. However, these studies often fail to consider the popularity of nodes in different communities, which is a common phenomenon in real-world networks. To address this issue, we propose a new probabilistic framework called the Two-Way Node Popularity Model (TNPM). The TNPM also accommodates edges from different distributions within a general sub-Gaussian family. We introduce the Delete-One-Method (DOM) for model fitting and community structure identification, and provide a comprehensive theoretical analysis with novel technical skills dealing with sub-Gaussian generalization. Additionally, we propose the Two-Stage Divided Cosine Algorithm (TSDC) to handle large-scale networks more efficiently. Our proposed methods offer multi-folded advantages in terms of estimation accuracy and computational efficiency, as demonstrated through extensive numerical studies. We apply our methods to two real-world applications, uncovering interesting findings.
有向ネットワークおよび二部ネットワークにおけるコミュニティ検出に関する研究はますます注目されています。しかし、これらの研究では、異なるコミュニティにおけるノードの人気度が考慮されていないことが多く、これは現実世界のネットワークでよく見られる現象です。この問題に対処するために、我々は双方向ノード人気度モデル(TNPM)と呼ばれる新しい確率的フレームワークを提案します。TNPMは、一般的なサブガウス族内の異なる分布からのエッジも受け入れます。モデルフィッティングとコミュニティ構造同定のためのDelete-One-Method(DOM)を導入し、サブガウス汎化を扱う新しい技術を用いた包括的な理論分析を提供します。さらに、大規模ネットワークをより効率的に処理するためのTwo-Stage Divided Cosine Algorithm(TSDC)を提案します。提案手法は、広範な数値研究によって実証されているように、推定精度と計算効率の点で多面的な利点を提供します。我々は、この手法を2つの実世界アプリケーションに適用し、興味深い知見を得た。
A Symplectic Analysis of Alternating Mirror Descent
A Symplectic Analysis of Alternating Mirror Descent / 交代ミラー降下法のシンプレクティック解析
Motivated by understanding the behavior of the Alternating Mirror Descent (AMD) algorithm for bilinear zero-sum games, we study the discretization of continuous-time Hamiltonian flow via the symplectic Euler method. We provide a framework for analysis using results from Hamiltonian dynamics and symplectic numerical integrators, with an emphasis on the existence and properties of a conserved quantity, the modified Hamiltonian (MH), for the symplectic Euler method.We compute the MH in closed-form when the original Hamiltonian is a quadratic function, and show that it generally differs from the other conserved quantity known previously in the literature. We derive new error bounds on the MH when truncated at orders in the stepsize in terms of the number of iterations, $K$, and use these bounds to show an improved $\mathcal{O}(K^{1/5})$ total regret bound and an $\mathcal{O}(K^{-4/5})$ duality gap of the average iterates for AMD. Finally, we propose a conjecture which, if true, would imply that the total regret for AMD scales as $\mathcal{O}\left(K^{\varepsilon}\right)$ and the duality gap of the average iterates as $\mathcal{O}\left(K^{-1+\varepsilon}\right)$ for any $\varepsilon>0$, and we can take $\varepsilon=0$ upon certain convergence conditions for the MH.
双線形零和ゲームにおけるAlternating Mirror Descent(AMD)アルゴリズムの挙動を理解することを目的に、シンプレクティックオイラー法を用いて連続時間ハミルトンフローの離散化を研究します。我々は、ハミルトニアンダイナミクスとシンプレクティック数値積分器の結果を用いた解析の枠組みを提示し、特にシンプレクティックオイラー法における保存量である修正ハミルトニアン(MH)の存在と特性に着目します。元のハミルトニアンが2次関数である場合にMHを閉じた形で計算し、それが文献で以前知られている他の保存量とは一般に異なることを示す。反復回数$K$に関して、ステップサイズのオーダーで切り捨てられた場合のMHの新しい誤差境界を導出し、この境界を用いてAMDの平均反復回数の改良された合計後悔境界$\mathcal{O}(K^{1/5})$と双対性ギャップ$\mathcal{O}(K^{-4/5})$を示す。最後に、我々は、もし真であれば、AMDの総後悔は$\mathcal{O}\left(K^{\varepsilon}\right)$に比例し、平均の双対性ギャップは任意の$\varepsilon>0$に対して$\mathcal{O}\left(K^{-1+\varepsilon}\right)$に比例し、MHの特定の収束条件では$\varepsilon=0$を取ることができるという予想を提案します。
Contrasting Local and Global Modeling with Machine Learning and Satellite Data: A Case Study Estimating Tree Canopy Height in African Savannas
Contrasting Local and Global Modeling with Machine Learning and Satellite Data: A Case Study Estimating Tree Canopy Height in African Savannas / 機械学習と衛星データを用いた局所モデリングとグローバルモデリングの対比:アフリカのサバンナにおける樹冠高推定のケーススタディ
While advances in machine learning with satellite imagery (SatML) are facilitating environmental monitoring at a global scale, developing SatML models that are accurate and useful for local regions remains critical to understanding and acting on an ever-changing planet. As increasing attention and resources are being devoted to training SatML models with global data, it is important to understand when improvements in global models will make it easier to train or fine-tune models that are accurate in specific regions. To explore this question, we design the first study that explicitly contrasts local and global training paradigms for SatML, through a case study of tree canopy height (TCH) mapping in the Karingani Game Reserve, Mozambique. We find that recent advances in global TCH mapping do not necessarily translate to better local modeling abilities in our study region. Specifically, small models trained only with locally-collected data outperform published global TCH maps, and even outperform globally pretrained models that we fine-tune using local data. Analyzing these results further, we identify specific points of conflict and synergy between local and global modeling paradigms that can inform future research toward aligning local and global performance objectives in geospatial machine learning.
衛星画像を用いた機械学習(SatML)の進歩により、地球規模の環境モニタリングが容易になっている一方で、常に変化する地球を理解し、対応していくためには、正確で地域に有用なSatMLモデルの開発が依然として重要です。地球規模のデータを用いたSatMLモデルのトレーニングにますます注目とリソースが投入されるようになるにつれ、地球規模のモデルの改善によって、特定の地域で正確なモデルのトレーニングや微調整が容易になるのはいつなのかを理解することが重要です。この問題を探るため、モザンビークのカリンガニ動物保護区における樹冠高(TCH)マッピングのケーススタディを通して、SatMLのローカルおよびグローバルのトレーニングパラダイムを明示的に対比する初の研究を設計しました。その結果、近年の地球規模のTCHマッピングの進歩が、必ずしも研究対象地域におけるローカルモデリング能力の向上につながるわけではないことがわかりました。具体的には、ローカルで収集されたデータのみでトレーニングされた小規模なモデルは、公開されている地球規模のTCHマップよりも性能が高く、ローカルデータを使用して微調整した地球規模で事前トレーニングされたモデルよりも性能が優れています。これらの結果をさらに分析することで、局所的モデリングパラダイムと全体的モデリングパラダイム間の具体的な対立点と相乗効果を特定し、地理空間機械学習における局所的および全体的パフォーマンス目標の整合に向けた将来の研究に役立てることができます。
Boosted Control Functions: Distribution Generalization and Invariance in Confounded Models
Boosted Control Functions: Distribution Generalization and Invariance in Confounded Models / ブースト制御関数:交絡モデルにおける分布の汎化と不変性
Modern machine learning methods and the availability of large-scale data have significantly advanced our ability to predict target quantities from large sets of covariates. However, these methods often struggle under distributional shifts, particularly in the presence of hidden confounding. While the impact of hidden confounding is well-studied in causal effect estimation, e.g., instrumental variables, its implications for prediction tasks under shifting distributions remain underexplored. This work addresses this gap by introducing a strong notion of invariance that, unlike existing weaker notions, allows for distribution generalization even in the presence of nonlinear, non-identifiable structural functions. Central to this framework is the Boosted Control Function (BCF), a novel, identifiable target of inference that satisfies the proposed strong invariance notion and is provably worst-case optimal under distributional shifts. The theoretical foundation of our work lies in Simultaneous Equation Models for Distribution Generalization (SIMDGs), which bridge machine learning with econometrics by describing data-generating processes under distributional shifts. To put these insights into practice, we propose the ControlTwicing algorithm to estimate the BCF using nonparametric machine-learning techniques and study its generalization performance on synthetic and real-world datasets compared to robust and empirical risk minimization approaches.
現代の機械学習手法と大規模データの利用可能性は、大規模な共変量セットから目標量を予測する能力を飛躍的に向上させた。しかし、これらの手法は分布シフト、特に隠れた交絡が存在する状況下ではしばしば困難をきたす。隠れた交絡の影響は、操作変数などの因果効果推定においては十分に研究されているが、シフトする分布下での予測タスクへの影響については十分に検討されていない。本研究では、既存の弱い概念とは異なり、非線形で識別不可能な構造関数が存在する場合でも分布の一般化を可能にする強い不変性の概念を導入することで、このギャップを解消します。この枠組みの中心となるのは、ブーステッド制御関数(BCF)です。これは、提案された強い不変性の概念を満たし、分布シフト下で最悪ケース最適であることが証明できる、新しい識別可能な推論対象です。本研究の理論的基礎は、分布一般化のための同時方程式モデル(SIMDG)にあります。SIMDGは、分布シフト下におけるデータ生成プロセスを記述することで、機械学習と計量経済学の橋渡しをします。これらの洞察を実践に移すため、ノンパラメトリック機械学習技術を用いてBCFを推定するControlTwicingアルゴリズムを提案し、合成データセットと実世界のデータセットにおけるその一般化性能を、堅牢な経験的リスク最小化アプローチと比較して研究します。
DCatalyst: A Unified Accelerated Framework for Decentralized Optimization
DCatalyst: A Unified Accelerated Framework for Decentralized Optimization / DCatalyst:分散最適化のための統合高速化フレームワーク
We study decentralized optimization over a network of agents, modeled as an undirected graph and operating without a central server. The objective is to minimize a composite function $f+r$, where $f$ is a (strongly) convex function representing the average of the agents’ losses, and $r$ is a convex, extended-value function (regularizer).We introduce DCatalyst, a unified black-box framework that injects Nesterov-type acceleration into decentralized optimization algorithms. At its core, DCatalyst is an inexact, momentum-accelerated proximal scheme (outer loop) that seamlessly wraps around a given decentralized method (inner loop). We show that DCatalyst attains optimal (up to logarithmic factors) communication and computational complexity across a broad family of decentralized algorithms and problem instances. In particular, it delivers accelerated rates for problem classes that previously lacked accelerated decentralized methods, thereby broadening the effectiveness of decentralized methods.On the technical side, our framework introduces inexact estimating sequences–an extension of Nesterov’s classical estimating sequences, tailored to decentralized, composite optimization. This construction systematically accommodates consensus errors and inexact solutions of local subproblems, addressing challenges that existing estimating-sequence-based analyses cannot handle while retaining a black-box, plug-and-play character.
本研究では、無向グラフとしてモデル化され、中央サーバーなしで動作するエージェントネットワーク上の分散最適化を研究します。目的は、合成関数$f+r$を最小化することです。ここで、$f$はエージェントの損失の平均を表す(強い)凸関数、$r$は凸拡張値関数(正則化子)です。分散最適化アルゴリズムにネステロフ型加速を注入する統合ブラックボックスフレームワークであるDCatalystを紹介します。DCatalystの核となるのは、不正確な運動量加速型近似スキーム(外側のループ)で、与えられた分散手法(内側のループ)をシームレスにラップします。DCatalystは、幅広い分散アルゴリズムと問題インスタンスにわたって、最適な(対数係数まで)通信と計算複雑性を実現することを示します。特に、これまで加速分散手法が不足していた問題クラスに対して加速率を実現し、分散手法の有効性を高めます。技術面では、私たちのフレームワークは、ネステロフの古典的な推定シーケンスを分散型の複合最適化に合わせて拡張した、不正確な推定シーケンスを導入します。この構築は、ブラックボックスのプラグアンドプレイの特性を維持しながら、コンセンサスエラーとローカルサブ問題の不正確な解決を体系的に受け入れ、既存の推定シーケンスベースの分析では処理できない課題に対処します。
Covariate-dependent Hierarchical Dirichlet Processes
Covariate-dependent Hierarchical Dirichlet Processes / 共変量依存階層的ディリクレ過程
Bayesian hierarchical modeling is a natural framework to effectively integrate data and borrow information across groups. In this paper, we address problems related to density estimation and identifying clusters across related groups, by proposing a hierarchical Bayesian approach that incorporates additional covariate information. To achieve flexibility, our approach builds on ideas from Bayesian nonparametrics, combining the hierarchical Dirichlet process with dependent Dirichlet processes. The proposed model is widely applicable, accommodating multiple and mixed covariate types through appropriate kernel functions as well as different output types through suitable component-specific likelihoods. This extends our ability to discern the relationship between covariates and clusters, while also effectively borrowing information and quantifying differences across groups. By employing a data augmentation trick, we are able to tackle the intractable normalized weights and construct a Markov chain Monte Carlo algorithm for posterior inference. The proposed method is illustrated on simulated data and two real data sets on single-cell RNA sequencing (scRNA-seq) and calcium imaging. For scRNA-seq data, we show that the incorporation of cell dynamics facilitates the discovery of additional cell subgroups. On calcium imaging data, our method identifies interpretable clusters of time frames with similar neural activity, aligning with the observed behavior of the animal.
ベイズ階層モデリングは、グループ間でデータを効果的に統合し、情報を借用するための自然なフレームワークです。本稿では、追加の共変量情報を組み込んだ階層ベイズアプローチを提案することで、関連グループ間の密度推定とクラスターの特定に関する問題に取り組みます。柔軟性を実現するために、私たちのアプローチはベイズノンパラメトリックスのアイデアに基づき、階層ディリクレ過程と従属ディリクレ過程を組み合わせています。提案モデルは幅広い適用が可能で、適切なカーネル関数によって複数の混合共変量タイプに対応し、適切なコンポーネント固有の尤度によって異なる出力タイプに対応します。これにより、共変量とクラスターの関係を識別する能力が拡張されると同時に、グループ間で効果的に情報を借用し、差異を定量化できます。データ拡張トリックを用いることで、扱いにくい正規化重みに対処し、事後推論のためのマルコフ連鎖モンテカルロアルゴリズムを構築できます。提案手法は、シミュレーションデータと、シングルセルRNAシーケンシング(scRNA-seq)およびカルシウムイメージングに関する2つの実データセットを用いて実証されています。scRNA-seqデータでは、細胞ダイナミクスを組み込むことで、新たな細胞サブグループの発見が容易になることを示します。カルシウムイメージングデータでは、本手法は、動物の観察された行動と一致する、類似した神経活動を示す解釈可能な時間フレームのクラスターを識別します。
Online Bernstein-von Mises theorem
Online Bernstein-von Mises theorem / オンライン・バーンスタイン=フォン・ミーゼス定理
Online learning is an inferential paradigm in which parameters are updated incrementally from sequentially available data, in contrast to batch learning, where the entire dataset is processed at once. In this paper, we assume that mini-batches from the full dataset become available sequentially. The Bayesian framework, which updates beliefs about unknown parameters after observing each mini-batch, is naturally suited for online learning. At each step, we update the posterior distribution using the current prior and new observations, with the updated posterior serving as the prior for the next step. However, this recursive Bayesian updating is rarely computationally tractable unless the model and prior are conjugate. When the model is regular, the updated posterior can be approximated by a normal distribution, as justified by the Bernstein-von Mises theorem. We adopt a variational approximation at each step and investigate the frequentist properties of the final posterior obtained through this sequential procedure. Under mild assumptions, we show that the accumulated approximation error becomes negligible once the mini-batch size exceeds a threshold depending on the parameter dimension. As a result, the sequentially updated posterior is asymptotically indistinguishable from the full posterior.
オンライン学習は、データセット全体を一度に処理するバッチ学習とは対照的に、順次利用可能なデータからパラメータを段階的に更新する推論パラダイムです。本稿では、データセット全体からのミニバッチが順次利用可能になると仮定します。各ミニバッチを観測した後に未知のパラメータに関する確信を更新するベイズフレームワークは、オンライン学習に自然に適しています。各ステップで、現在の事前分布と新しい観測値を使用して事後分布を更新し、更新された事後分布を次のステップの事前分布として使用します。しかし、この再帰的なベイズ更新は、モデルと事前分布が共役でない限り、計算的に扱いにくいことがほとんどです。モデルが正則である場合、更新された事後分布は、ベルンシュタイン-フォン・ミーゼス定理によって正当化されるように、正規分布で近似できます。各ステップで変分近似を採用し、この逐次的な手順によって得られる最終的な事後分布の頻度主義的特性を調査します。軽い仮定の下で、ミニバッチサイズがパラメータ次元に依存する閾値を超えると、累積近似誤差は無視できることを示します。その結果、逐次更新された事後分布は、完全な事後分布と漸近的に区別できなくなります。
Transformers Can Overcome the Curse of Dimensionality: A Theoretical Study from an Approximation Perspective
Transformers Can Overcome the Curse of Dimensionality: A Theoretical Study from an Approximation Perspective / トランスフォーマーは次元の呪いを克服できる:近似の観点からの理論的研究
The Transformer model is widely used in various application areas of machine learning, such as natural language processing. This paper investigates the approximation of the Hölder continuous function class $\mathcal{H}_{Q}^{\beta}\left([0,1]^{d\times n},\mathbb{R}^{d\times n}\right)$ by Transformers and constructs several Transformers that can overcome the curse of dimensionality. These Transformers consist of one self-attention layer with one head and the softmax function as the activation function, along with several feedforward layers. For example, to achieve an approximation accuracy of $\epsilon$, if the activation functions of the feedforward layers in the Transformer are ReLU and floor, only $\mathcal{O}\left(\log\frac{1}{\epsilon}\right)$ layers of feedforward layers are needed, with widths of these layers not exceeding $\mathcal{O}\left(\frac{1}{\epsilon^{2/\beta}}\log\frac{1}{\epsilon}\right)$. If other activation functions are allowed in the feedforward layers, the width of the feedforward layers can be further reduced to a constant. These results demonstrate that Transformers have a strong expressive capability. The construction in this paper is based on the Kolmogorov-Arnold Superposition Theorem and does not require the concept of contextual mapping, hence our proof is more intuitively clear compared to previous Transformer approximation works. Additionally, the translation technique proposed in this paper helps to apply the previous approximation results of feedforward neural networks to Transformer research.
Transformerモデルは、自然言語処理など、機械学習のさまざまな応用分野で広く使用されています。本論文では、Hölder連続関数クラス$\mathcal{H}_{Q}^{\beta}\left([0,1]^{d\times n},\mathbb{R}^{d\times n}\right)$のTransformerによる近似を検証し、次元の呪いを克服できる複数のTransformerを構築します。これらのTransformerは、1つのヘッドを持つ1つの自己注意層と、活性化関数としてソフトマックス関数、そして複数のフィードフォワード層から構成されます。例えば、近似精度$\epsilon$を達成するには、Transformerのフィードフォワード層の活性化関数がReLUとfloorである場合、フィードフォワード層は$\mathcal{O}\left(\log\frac{1}{\epsilon}\right)$層のみ必要であり、これらの層の幅は$\mathcal{O}\left(\frac{1}{\epsilon^{2/\beta}}\log\frac{1}{\epsilon}\right)$を超えません。フィードフォワード層で他の活性化関数が許可されている場合、フィードフォワード層の幅はさらに定数まで削減できます。これらの結果は、Transformerが強力な表現能力を持っていることを示しています。本論文の構成はKolmogorov-Arnoldの重ね合わせ定理に基づいており、コンテキスト マッピングの概念を必要としないため、これまでのTransformer近似研究と比較して、証明はより直感的に明確です。さらに、本論文で提案された変換技術は、フィードフォワードニューラルネットワークの以前の近似結果をTransformer研究に適用するのに役立ちます。
The Role of Contextual Information in Best Arm Identification
The Role of Contextual Information in Best Arm Identification / 最良アーム特定における文脈情報の役割
We study the best-arm identification problem with fixed confidence when contextual (covariate) information is available in stochastic bandits. In each round, we observe contextual information before selecting an arm. The distribution of the reward associated with the selected arm depends on the observed contextual information. We are interested in finding the arm with the maximum mean reward marginalized over the contextual distribution and not the mean reward conditioned on contexts. Our goal is to identify the best arm with a minimal number of samples under a given error probability. First, we derive the instance-specific sample-complexity lower bounds under the contextual information. Then, we propose a context-aware version of the Track-and-Stop strategy, wherein the proportions of arm draws track the set of optimal allocations, and prove that the expected number of arm draws asymptotically matches the lower bound. We demonstrate that the contextual information can be used to improve the efficiency of the identification of the best marginalized mean reward when compared with the results of Garivier and Kaufmann(2016). Furthermore, we experimentally confirm that contextual information contributes to faster best-arm identification.
本研究では、確率的バンディット設定においてコンテキスト(共変量)情報が利用可能な場合の「最良アーム特定問題」を扱います。各ラウンドにおいて、アームを選択する前にコンテキスト情報を観測します。選択されたアームに関連する報酬の分布は、観測されたコンテキスト情報に依存します。我々の関心対象は、コンテキストで条件付けられた平均報酬ではなく、コンテキスト分布に関して周辺化した平均報酬が最大となるアームを特定することにあります。目標は、所与の誤り確率の下で、最小限のサンプル数で最良のアームを特定することです。まず、コンテキスト情報が存在する場合の、インスタンスごとのサンプル複雑度に関する下界を導出します。次に、アームの選択割合が最適配分集合を追跡するように制御する「Track-and-Stop」戦略のコンテキスト対応版を提案し、アーム選択の期待回数が漸近的に下界と一致することを証明します。Garivier and Kaufmann (2016)の結果と比較して、コンテキスト情報を利用することで、周辺化平均報酬が最大となるアームの特定効率を向上できることを示します。さらに、コンテキスト情報が最良アームの特定を高速化することに寄与することを実験的に確認します。
A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks
A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks / 部分的に学習された3層ニューラルネットワークの関数空間平均場理論
To understand the training dynamics of neural networks, prior studies have considered the mean-field (MF) limit of two-layer NNs as the width tends to infinity, establishing theoretical guarantees for its convergence under gradient flow training as well as approximation and generalization capabilities. In this work, we study the infinite-width limit of a type of three-layer neural network where the first-layer weights are untrained. To rigorously define the limiting model, we extend the MF theory by lifting the representation of neurons from Euclidean to functional spaces. This allows us to establish the MF training dynamics as a functional gradient flow with a time-varying kernel that remains positive-definite under suitable assumptions, thus proving a linear-rate convergence of its training loss. Furthermore, we define novel function spaces that contain the solutions obtained through the MF training dynamics and prove Rademacher complexity bounds for these spaces. Notably, our analysis applies to a range of scaling choices of the model, resulting in two distinct regimes of the MF limit that both exhibit feature learning through training.
ニューラルネットワークの学習ダイナミクスを理解するために、先行研究では、幅を無限大にした際の2層ニューラルネットワークの平均場(MF)極限が検討されてきました。これにより、勾配流による学習における収束性や、近似能力・汎化能力に関する理論的保証が確立されています。本研究では、第1層の重みが学習されないタイプの3層ニューラルネットワークについて、幅を無限大にした極限を研究します。極限モデルを厳密に定義するために、ニューロンの表現をユークリッド空間から関数空間へと持ち上げることで、MF理論を拡張します。これにより、MF学習ダイナミクスを、適切な仮定の下で正定値性を維持する時間変化カーネルを伴う関数勾配流として定式化でき、学習損失が線形レートで収束することを証明します。さらに、MF学習ダイナミクスを通じて得られる解を含む新たな関数空間を定義し、それらの空間に対するラデマッハー複雑度の限界を証明します。特筆すべき点として、我々の解析はモデルのスケーリングに関する様々な選択に適用可能であり、その結果、いずれも学習を通じて特徴量学習を示す2つの異なるMF極限のレジームが導かれます。
Inference with non-differentiable surrogate loss in a general high-dimensional classification framework
Inference with non-differentiable surrogate loss in a general high-dimensional classification framework / 一般的な高次元分類の枠組みにおける非微分可能な代理損失を用いた推論
Penalized empirical risk minimization with a surrogate loss function is often used to learn a high-dimensional linear decision rule in classification problems. Although much of the literature focus on the generalization error, there is a lack of inference procedures for identifying the driving factors of the estimated decision rule, especially when the surrogate loss is non-differentiable. We propose a kernel-smoothed decorrelated score to construct hypothesis tests and interval estimators for a linear decision rule estimated using a piece-wise linear surrogate loss, which has a discontinuous gradient and non-regular Hessian. Specifically, we adopt kernel approximations to smooth the discontinuous gradient near discontinuity points and approximate the non-regular Hessian of the surrogate loss. In applications where additional nuisance parameters are involved, we propose a novel cross-fitted version to accommodate flexible nuisance estimates and kernel approximations. We establish the limiting distribution of the kernel-smoothed decorrelated score and its cross-fitted version in a high-dimensional setup. Simulation and real data analysis are conducted to demonstrate the validity and the superiority of the proposed method.
分類問題において高次元の線形決定規則を学習する際には、代理損失関数を用いたペナルティ付き経験リスク最小化が頻繁に利用されます。既存研究の多くは汎化誤差に焦点を当てていますが、特に代理損失(surrogate loss)が微分不可能である場合、推定された決定規則の主要な要因を特定するための推論手法は不足しています。本研究では、勾配が不連続かつヘッセ行列が非正則である区分線形代理損失を用いて推定された線形決定規則に対し、仮説検定や区間推定を行うための「カーネル平滑化された非相関化スコア」を提案します。具体的には、カーネル近似を用いて不連続点近傍の勾配を平滑化し、代理損失の非正則なヘッセ行列を近似します。追加のnuisanceパラメータ(局外パラメータ)を伴う応用においては、柔軟なnuisance推定やカーネル近似に対応可能な、新規のクロスフィッティング版手法を提案します。また、高次元設定におけるカーネル平滑化非相関化スコアおよびそのクロスフィッティング版の極限分布を導出します。さらに、シミュレーションおよび実データ解析を通じて、提案手法の妥当性と優位性を示します。
Knowledge Cascade: Reverse Knowledge Distillation on Nonparametric Multivariate Functional Estimation
Knowledge Cascade: Reverse Knowledge Distillation on Nonparametric Multivariate Functional Estimation / Knowledge Cascade:ノンパラメトリック多変量関数推定における逆知識蒸留
As machine learning models and datasets continue to grow, developing complex models has become increasingly computationally demanding. Knowledge distillation reduces deployment cost by compressing a large, well-trained teacher model into a compact student model, but it does not address settings where constructing the teacher itself is the bottleneck. Motivated by this challenge, we introduce Knowledge Cascade, a reverse knowledge distillation framework that uses information from a small, inexpensive student model to guide the development of a more complex teacher model. Although this direction is counterintuitive because the teacher typically has greater representational capacity, we show that student-to-teacher transfer can be principled when supported by statistical scaling relationships. We first develop Knowledge Cascade for nonparametric multivariate functional estimation in reproducing kernel Hilbert spaces via smoothing splines, where selecting multiple smoothing parameters is a major computational bottleneck. Knowledge Cascade transfers student-selected smoothing parameters to the full-sample regime through asymptotic scaling laws, substantially reducing computational cost for high-dimensional and large-scale datasets while retaining theoretical guarantees. Beyond smoothing splines, we illustrate the same principle through kernel density estimation and deep learning hyperparameter transfer. Simulations and real-data experiments show that Knowledge Cascade achieves substantial computational savings while maintaining strong statistical performance, and can sometimes outperform the corresponding full-sample procedure.
機械学習モデルやデータセットの増大に伴い、複雑なモデルの開発には膨大な計算資源が必要とされるようになっています。知識蒸留(Knowledge Distillation)は、十分に学習された大規模な「教師モデル」をコンパクトな「生徒モデル」に圧縮することでデプロイ時のコストを削減しますが、教師モデル自体の構築そのものがボトルネックとなる状況には対応していません。こうした課題に対処するため、我々は「Knowledge Cascade(知識カスケード)」という手法を提案します。これは、小規模で低コストな生徒モデルからの情報を利用して、より複雑な教師モデルの開発を導く、いわば「逆向きの知識蒸留」とも言える枠組みです。一般的に教師モデルの方が高い表現能力を持つため、生徒から教師へと情報を伝達するというこのアプローチは直感に反するように思えるかもしれません。しかし、我々は、統計的なスケーリング関係(規模の拡大に伴う法則性)に基づけば、生徒から教師への知識伝達が理論的に妥当なものとなり得ることを示します。まず、我々は、複数の平滑化パラメータの選択が計算上の大きなボトルネックとなる「再生核ヒルベルト空間における平滑化スプラインを用いたノンパラメトリック多変量関数推定」に対して、Knowledge Cascadeを適用しました。Knowledge Cascadeは、漸近的なスケーリング則を利用して生徒モデルが選択した平滑化パラメータを全サンプル規模の推定へと転用します。これにより、理論的な保証を維持しつつ、高次元かつ大規模なデータセットにおける計算コストを大幅に削減することが可能になります。さらに、平滑化スプラインにとどまらず、カーネル密度推定やディープラーニングにおけるハイパーパラメータの転用といった事例を通じても、同様の原理が有効であることを示します。シミュレーションおよび実データを用いた実験の結果、Knowledge Cascadeは優れた統計的性能を維持しながら大幅な計算コストの削減を実現し、場合によっては全サンプルを用いた従来の手法を上回る性能を示すことさえあることが明らかになりました。
Neural Exploitation and Exploration of Contextual Bandits
Neural Exploitation and Exploration of Contextual Bandits / 文脈付きバンディットにおけるニューラルな活用と探索
In this paper, we study the neural exploration strategy for contextual bandits.The dilemma of exploitation and exploration widely exists in real-world applications such as recommender systems, online advertising, and clinical trials.Contextual bandits provide principled methods to solve this dilemma, including two prevalent techniques: Thompson Sampling (TS), and Upper Confidence Bound (UCB).Neural contextual bandits have been studied to adapt to the non-linear reward function, combined with TS or UCB strategies for exploration.In this paper, we introduce, EE-Net, which is a novel framework to utilize another neural network to learn the potential gain of exploitation neural network for exploration, different from UCB-based and TS-based approaches that rely on the large-deviation-based statistical confidence bound. In addition, we provide an instance-based $\widetilde{\mathcal{O}}(\sqrt{T})$ regret upper bound for EE-Net with a new proof workflow. Empirically, we show that EE-Net outperforms related linear and neural contextual bandit baselines on real-world datasets.
本論文では、コンテクスト・バンディット(contextual bandits)におけるニューラルな探索戦略について研究します。推薦システム、オンライン広告、臨床試験などの実世界の応用において、「活用(exploitation)」と「探索(exploration)」のジレンマは広く存在します。コンテクスト・バンディットは、このジレンマを解決するための原理的な手法を提供しており、その代表的なものとしてトンプソン・サンプリング(TS)やUCB(Upper Confidence Bound:信頼区間上限)法が挙げられます。非線形な報酬関数に適応するために、TSやUCBといった探索戦略を組み合わせた「ニューラル・コンテクスト・バンディット」の研究も行われてきました。本論文では、UCBやTSのような「大偏差原理に基づく統計的信頼区間」に依存する手法とは異なり、別のニューラルネットワークを用いて「活用用ニューラルネットワーク」の潜在的な利得を学習し、それを探索に利用する新しいフレームワーク「EE-Net」を提案します。さらに、新たな証明手順を用いて、EE-Netに対するインスタンス依存の$\widetilde{\mathcal{O}}(\sqrt{T})$というリグレット(後悔)の上界を導出します。実証実験では、実世界のデータセットにおいて、EE-Netが既存の線形およびニューラル・コンテクスト・バンディットの手法(ベースライン)よりも優れた性能を示すことを明らかにします。
Causal Influences over Social Learning Networks
Causal Influences over Social Learning Networks / 社会的学習ネットワークにおける因果的影響
This paper investigates causal influences between agents linked by a social graph and interacting over time. In particular, the work examines the dynamics of social learning models and distributed decision-making protocols, and derives expressions that reveal the causal relations between pairs of agents and explain the flow of influence over the network. The results turn out to be dependent on the graph topology and the level of information that each agent has about the inference problem they are trying to solve. Using these conclusions, the paper proposes an algorithm to rank the overall influence between agents to discover highly influential agents. It also provides a method to learn the necessary model parameters from raw observational data. The results and the proposed algorithm are illustrated by considering both synthetic data and real social media data.
本論文では、ソーシャルグラフで結ばれ、時間の経過とともに相互作用するエージェント間の因果的影響について調査します。特に、ソーシャル・ラーニング・モデルや分散意思決定プロトコルのダイナミクスを検討し、エージェントのペア間の因果関係を明らかにし、ネットワーク上での影響の流れを説明する式を導出します。その結果、影響の度合いは、グラフのトポロジーや、各エージェントが解決しようとしている推論問題について保有する情報のレベルに依存することが示されます。これらの結論に基づき、本論文では、エージェント間の全体的な影響力をランク付けし、影響力の大きいエージェントを特定するためのアルゴリズムを提案します。また、生の観測データから必要なモデルパラメータを学習する方法も提示します。提案された結果やアルゴリズムについては、合成データと実際のソーシャルメディアデータの双方を用いて具体的に示します。
Do We Need to Penalize Variance of Losses for Learning with Label Noise?
Do We Need to Penalize Variance of Losses for Learning with Label Noise? / ラベルノイズを含む学習において損失の分散にペナルティを科す必要はあるか?
Statistically consistent algorithms have been widely employed for dealing with noisy labels. Their objective functions are designed so that minimizing the expected risk on noisy data leads to the same minimizer as minimizing the expected risk on clean data. From the weak law of large numbers, penalizing the variance of losses would reduce the discrepancy between the average loss and the expected risk on the clean data when there is a finite training sample, and the estimation error in the model’s parameters can be reduced. Interestingly, we found that the variance of losses needs to be encouraged for label-noise learning. Specifically, encouraging a large variance of losses would boost the memorization effect and reduce the harmfulness of incorrect labels. Regularizers can be easily designed to encourage a large variance of losses and be plugged into many existing algorithms. Empirically, the proposed method by encouraging a large variance of losses could improve the generalization ability of baselines on both synthetic and real-world datasets.
ノイズを含むラベル(ノイズありラベル)を扱う手法として、統計的に整合性のあるアルゴリズムが広く採用されています。これらのアルゴリズムの目的関数は、ノイズありデータにおける期待リスクを最小化することが、クリーンなデータにおける期待リスクを最小化することと同じ最適解(最小化解)を導くように設計されています。大数の弱法則に基づくと、損失の分散にペナルティを課すことで、有限の学習サンプルが存在する場合において、平均損失とクリーンなデータにおける期待リスクとの乖離を低減でき、モデルパラメータの推定誤差を抑えることが可能になります。興味深いことに、ラベルノイズを含むデータでの学習においては、損失の分散を大きくすることが有益であると分かりました。具体的には、損失の分散を大きくすることで、データの「記憶(memorization)」効果が促進され、誤ったラベルによる悪影響を低減できます。損失の分散を大きくするような正則化手法は容易に設計でき、既存の多くのアルゴリズムに組み込むことが可能です。実験の結果、提案手法を用いて損失の分散を大きくすることで、合成データセットと実データセットの双方において、ベースライン手法の汎化性能を向上させることができました。
Nonparametric generative modeling for time series via Schrödinger bridge
Nonparametric generative modeling for time series via Schrödinger bridge / シュレーディンガー・ブリッジを用いた時系列のノンパラメトリック生成モデリング
We propose a novel generative model for time series based on Schrödinger bridge (SB) approach. This consists in the entropic interpolation via optimal transport between a reference probability measure on path space and a target measure consistent with the joint data distribution of the time series. The solution is characterized by a stochastic differential equation on finite horizon with a path-dependent drift function, hence respec\-ting the temporal dynamics of the time series distribution.We estimate the drift function from data samples by nonparametric, e.g. kernel regression methods, and the simulation of the SB diffusion yields new synthetic data samples of the time series. The performance of our generative model is evaluated through a series of numerical experiments. First, we test with autoregressive models, a GARCH Model, and the example of fractional Brownian motion, and measure the accuracy of our algorithm with marginal, temporal dependencies metrics, and predictive scores. Next, we use our SB generated synthetic samples for the application to deep hedging on real-data sets.
本研究では、シュレーディンガー・ブリッジ(SB)アプローチに基づく時系列用の新しい生成モデルを提案します。これは、パス空間上の基準確率測度と、時系列の結合データ分布に整合する目標測度との間で、最適輸送を介したエントロピー補間を行うものです。その解は、パスに依存するドリフト関数を持つ有限時間区間上の確率微分方程式によって特徴付けられ、時系列分布の時間的ダイナミクスを保持します。我々は、カーネル回帰などのノンパラメトリック手法を用いてデータサンプルからドリフト関数を推定し、SB拡散のシミュレーションを行うことで、時系列の新たな合成データサンプルを生成します。提案する生成モデルの性能は、一連の数値実験を通じて評価されます。まず、自己回帰モデル、GARCHモデル、およびフラクショナル・ブラウン運動の例を用いて検証を行い、周辺分布や時間的依存性に関する指標、および予測スコアを用いてアルゴリズムの精度を測定します。次に、SBによって生成された合成サンプルを、実データセットに対するディープ・ヘッジ(deep hedging)の適用に利用します。
Sparse Topic Modeling via Spectral Decomposition and Thresholding
Sparse Topic Modeling via Spectral Decomposition and Thresholding / スペクトル分解と閾値処理によるスパース・トピック・モデリング
In probabilistic Latent Semantic Indexing (pLSI), word frequencies across document corpora are modeled through a low-rank factorization of the expected document-term matrix into topic-word and topic-document components. In this paper, we study the estimation of the topic-word matrix under a sparsity structure motivated by Zipf’s law: word frequencies within each topic exhibit a rapid empirical decay, with most probability mass concentrated on a small subset of words. Motivated by this observation, we introduce a spectral estimator that adaptively thresholds rare words prior to factorization. We show that the resulting estimator achieves an $\ell_1$-error rate whose dependence on the vocabulary size $p$ is only logarithmic. Our error bounds hold across parameter regimes, including high-dimensional settings with extremely large vocabularies, a practically important scenario that has received limited theoretical attention. Unlike many existing methods, our approach does not require the separability (or anchor-word) assumption. Synthetic and real-data experiments demonstrate that the proposed procedure is computationally efficient, statistically reliable, and effective across domains with widely varying dimensions, sparsity levels, and document lengths.
確率的潜在意味インデックス法(pLSI)では、文書コーパス全体における単語の出現頻度が、期待文書・単語行列をトピック・単語成分とトピック・文書成分へと低ランク分解することによってモデル化されます。本論文では、ジップの法則に動機づけられたスパース性構造の下でのトピック・単語行列の推定について検討します。この構造では、各トピック内の単語頻度が経験的に急激に減衰し、確率質量の大部分が少数の単語サブセットに集中するという特徴があります。この観察に基づき、分解を行う前に稀な単語に対して適応的に閾値処理を施すスペクトル推定量を導入します。その結果得られる推定量は、語彙サイズ$p$への依存性が対数的オーダーにとどまる$\ell_1$誤差率を達成することを示します。この誤差限界は、極めて大きな語彙を持つ高次元設定を含む様々なパラメータ領域で成立します。高次元設定は実用上重要であるにもかかわらず、理論的な研究は限られていました。多くの既存手法とは異なり、本アプローチでは分離可能性(またはアンカーワード)の仮定を必要としません。合成データおよび実データを用いた実験により、提案手法が計算効率に優れ、統計的に信頼性が高く、次元数、スパース性、文書長が大きく異なる様々なドメインにおいて有効であることが示されます。
Probabilistic Rainfall Downscaling: Joint Generalized Neural Models with Censored Spatial Gaussian Copula
Probabilistic Rainfall Downscaling: Joint Generalized Neural Models with Censored Spatial Gaussian Copula / 確率的降雨ダウンスケーリング:打ち切り空間ガウス・コピュラを伴う結合一般化ニューラルモデル
A novel approach for generating conditional probabilistic rainfall downscaling at finer scales from deterministic weather variables at coarser scales with temporal and spatial dependence is introduced. A two-step procedure is employed. Firstly, marginal location-specific distributions are jointly modelled conditional on the deterministic coarse weather variables. Secondly, a spatial dependency structure is learned to ensure spatial coherence among these distributions.To learn marginal distributions over rainfall values, we introduce joint generalised neural models that expand generalised linear models with a deep neural network architecture to jointly fit parameters of the distributions.The spatial dependency structure is modelled using a censored latent Gaussian copula leveraging the underlying spatial structure. We construct a distance matrix between locations, transformed into a correlation matrix by a Gaussian Process Kernel depending on a small set of parameters. To estimate these parameters, we propose a general framework for the estimation of latent Gaussian copulas employing scoring rules as a measure of divergence between distributions. Uniting our two contributions, namely the joint generalised neural model and the censored latent Gaussian copulas into a single model, our probabilistic approach provides downscaled rainfall. We demonstrate its efficacy using a large UK data set, outperforming existing methods.
時間的・空間的依存性を持つ粗いスケールの決定論的気象変数から、より細かいスケールでの条件付き確率的降雨ダウンスケーリングを行うための新しいアプローチを導入します。ここでは2段階の手順が採用されます。まず、決定論的な粗い気象変数を条件として、地点ごとの周辺分布を統合的にモデル化します。次に、これらの分布間の空間的な整合性を確保するために、空間的依存構造を学習させます。降雨量の周辺分布を学習させるために、我々は「統合型一般化ニューラルモデル」を導入します。これは、一般化線形モデルを深層ニューラルネットワークのアーキテクチャで拡張し、分布のパラメータを統合的に適合させるものです。空間的依存構造のモデル化には、対象地域の空間構造を活用した「打ち切り型潜在ガウス・コピュラ」を用います。具体的には、地点間の距離行列を構築し、少数のパラメータに依存するガウス過程カーネルを用いて相関行列へと変換します。これらのパラメータを推定するために、分布間の乖離尺度としてスコアリング・ルール(scoring rules)を採用した、潜在ガウス・コピュラ推定のための汎用的な枠組みを提案します。これら二つの手法(統合型一般化ニューラルモデルと打ち切り型潜在ガウス・コピュラ)を単一のモデルに統合することで、我々の確率的アプローチは、空間解像度を高めた(ダウンスケールされた)降雨量推定を実現します。英国の大規模データセットを用いた検証において、本手法が既存の手法を上回る性能を示すことを実証しました。
Global Fréchet Manifold Learning for Random Objects, With Application to Low-Dimensional Wasserstein Representations of Distributional Data
Global Fréchet Manifold Learning for Random Objects, With Application to Low-Dimensional Wasserstein Representations of Distributional Data / ランダムオブジェクトのための大域的フレシェ多様体学習 — 分布データの低次元ワッサースタイン表現への応用
We study manifold learning with multidimensional scaling for samples of metric space valued data. By adopting a global version of ISOMAP we obtain low-dimensional Euclidean representations. A key innovation is that we demonstrate that global Fréchet regressioncan be utilized for mapping the elements of a convex set in the Euclidean representation space back to the metric space where the objects reside. We refer to this approach as Fréchet manifold learning and showcase it with one-dimensional distributions as random objects, equipped with the Wasserstein metric, which is an important special case of our general approach. The resulting low-dimensional representations mimic the parametric representation in a parametric family of distributions but are entirely learned from the data without postulating any parametric model. These Wasserstein representations of distributional data can be viewed as an empirical parametrization of a sample of distributions. The utility of these representations rests on the map from the low-dimensional Euclidean representation space to the space of distributions, which is obtained with global Fréchet regression. We illustrate the proposed approach with distributional data for baby names, bike rentals and age pyramids and further demonstrate how it can be applied fora novel distributional regression method that features one-dimensional distributions as predictors.
本研究では、距離空間に値をとるデータサンプルを対象に、多次元尺度構成法(MDS)を用いた多様体学習について検討します。ISOMAPの「大域的(global)」な手法を採用することで、低次元ユークリッド空間における表現を得ます。重要な貢献として、ユークリッド表現空間内の凸集合の要素を、元のデータが存在する距離空間へと写像するために「大域的フレシェ回帰(global Fréchet regression)」が利用可能であることを示します。我々はこのアプローチを「フレシェ多様体学習」と呼び、その重要な特殊ケースとして、ワッサースタイン距離を備えた1次元分布をランダムな対象(random objects)とする事例を用いて実証します。得られる低次元表現は、分布のパラメトリックな族におけるパラメトリック表現と類似していますが、特定のパラメトリックモデルを仮定することなく、データから完全に学習されたものです。分布データのこうしたワッサースタイン表現は、分布サンプルの経験的なパラメタリゼーション(パラメータ化)とみなすことができます。これらの表現の有用性は、低次元ユークリッド表現空間から分布空間への写像(大域的フレシェ回帰によって得られるもの)に依拠しています。提案手法を、赤ちゃんの名前、自転車レンタル、年齢構成(人口ピラミッド)に関する分布データを用いて例示し、さらに、1次元分布を予測変数とする新しい分布回帰手法への応用可能性も示します。
A Convex Framework for Confounding Robust Inference
A Convex Framework for Confounding Robust Inference / 交絡に対して頑健な推論のための凸最適化の枠組み
We study policy evaluation of offline contextual bandits subject to unobserved confounders. Sensitivity analysis methods are commonly used to estimate the policy value under the worst-case confounding scenario within a given uncertainty set. However, existing work often resorts to some coarse relaxation of the uncertainty set for the sake of tractability, leading to overly conservative estimation of the policy value. In this paper, we propose a general estimator that provides a sharp lower bound of the policy value using convex programming. The generality of our estimator enables various extensions such as sensitivity analysis using f-divergence, model selection with cross validation and information criterion, and robust policy learning with the sharp lower bound. Furthermore, our estimation method can be reformulated as an empirical risk minimization problem thanks to the strong duality, which enables us to provide strong theoretical guarantees of the proposed estimator using M-estimation techniques.
本研究では、観測されない交絡因子が存在する状況下でのオフライン・コンテクスト・バンディットのポリシー評価について検討します。不確実性集合内における最悪の交絡シナリオ下でのポリシー価値を推定するために、感度分析手法が一般的に用いられています。しかし、既存の研究では、計算の容易さを優先して不確実性集合を粗く緩和してしまうことが多く、その結果、ポリシー価値の推定が過度に保守的になるという問題がありました。本論文では、凸計画法を用いてポリシー価値の鋭い下界を与える一般的な推定量を提案します。この推定量の汎用性により、f-ダイバージェンスを用いた感度分析、交差検証や情報量規準によるモデル選択、鋭い下界を用いたロバストなポリシー学習など、様々な拡張が可能になります。さらに、強双対性(strong duality)を利用することで、提案する推定手法を経験リスク最小化問題として再定式化でき、M推定の技法を用いて強力な理論的保証を与えることが可能となります。
Kernel-based Distributed Learning
Kernel-based Distributed Learning / カーネルに基づく分散学習
We consider one-shot distributed learning problems in a reproducing kernel Hilbert space framework. Current results are limited to the least-squares loss and extensions beyond this meet with some significant technical challenges. We establish the optimal rate of distributed learning for some general class of convex loss functions satisfying mild assumptions, using a novel empirical process on the Bregman divergence induced by the loss, which is essential for carrying out a quadratic approximation in the infinite-dimensional space. The empirical process is bounded by relating the Bregman divergence induced by the loss to the supremum norm and the $L^2$-norm of the functions. This framework incorporates many commonly used losses, including strongly smooth loss functions as well as Lipschitz continuous losses such as the quantile loss.
本研究では、再生核ヒルベルト空間の枠組みにおけるワンショット分散学習問題を扱います。既存の研究成果は最小二乗損失に限られており、それを超える拡張には大きな技術的課題が伴います。我々は、緩やかな仮定を満たすある一般的な凸損失関数のクラスに対し、分散学習における最適収束レートを確立しました。この際、当該損失関数から導かれるブレグマン・ダイバージェンス(Bregman divergence)に関する新たな経験過程を用いています。この経験過程は、無限次元空間における二次近似を行う上で不可欠なものです。具体的には、損失関数に由来するブレグマン・ダイバージェンスを関数の上限ノルム(supremum norm)および$L^2$ノルムと関連付けることで、当該経験過程の有界性を評価しています。この枠組みは、強滑らか(strongly smooth)な損失関数や、分位点損失(quantile loss)のようなリプシッツ連続な損失関数を含め、広く用いられている多くの損失関数を包含するものです。
Why “Classic” Transformers Are Shallow and A Depth-Enabling Technique
Why “Classic” Transformers Are Shallow and A Depth-Enabling Technique / なぜ「古典的」Transformerは浅いのか — 深層化を可能にする手法
Since its introduction in 2017, the Transformer has emerged as the leading neural network architecture, catalyzing revolutionary advancements in many AI disciplines. The key innovation in Transformer is a Self-Attention (SA) mechanism designed to capture contextual information. However, stacking up more layers of the same design has failed to produce trainable deeper Transformers. Thus far, various architectural modifications to the original design have been proposed to enable deeper depths for Transformer models, but a thorough understanding of this depth issue remains lacking. In this paper, we conduct a comprehensive investigation to substantiate the claim that the depth problem is caused by a phenomenon called token similarity escalation; that is, tokens grow increasingly alike after repeated applications of the SA mechanism. Our analysis reveals that, driven by the invariant leading eigenspace and large spectral gaps of attention matrices, token similarity provably escalates at a linear rate as the depth increases. This insight suggests a simple technique that surgically removes excessive token similarity without reducing the overall role of the SA mechanism, as is done by existing approaches. We perform a set of proof- of-concept, small-scale experiments to show the viability of the proposed depth-enabling technique.
2017年に登場して以来、Transformerは主要なニューラルネットワーク・アーキテクチャとしての地位を確立し、多くのAI分野において革新的な進歩を促してきました。Transformerの重要な革新点は、文脈情報を捉えるために設計された自己注意(Self-Attention: SA)メカニズムにあります。しかし、同じ設計の層を単に積み重ねるだけでは、学習可能な、より深いTransformerを実現することはできませんでした。これまで、Transformerモデルの深さを増すために様々なアーキテクチャ上の改良が提案されてきましたが、この「深さ」に関する問題についての十分な理解は依然として得られていません。本論文では、「トークン類似度の増大(token similarity escalation)」、すなわちSAメカニズムの適用を繰り返すにつれてトークン同士がますます似通ってくるという現象が、深さに関する問題の原因であるという主張を裏付けるため、包括的な調査を行います。我々の分析により、アテンション行列の不変な主固有空間と大きなスペクトルギャップに起因して、深さが増すにつれてトークン類似度が線形な割合で増大することが証明されました。この知見は、既存のアプローチのようにSAメカニズムの全体的な役割を損なうことなく、過剰なトークン類似度をピンポイントで除去する単純な手法を示唆しています。我々は、提案する「深さ対応(depth-enabling)」手法の実現可能性を示すため、小規模な概念実証実験を行います。
Bayes-Optimal Fair Classification with Linear Disparity Constraints via Pre-, In-, and Post-processing
Bayes-Optimal Fair Classification with Linear Disparity Constraints via Pre-, In-, and Post-processing / 線形不公平性制約下でのベイズ最適公平分類:前処理・学習時処理・後処理によるアプローチ
Machine learning algorithms may have disparate impacts on protected groups. To address this, we develop methods for Bayes-optimal fair classification, aiming to minimize classification error subject to given group fairness constraints. We introduce the notion of linear disparity measures, which are linear functions of a probabilistic classifier; and bilinear disparity measures, which are also linear in the group-wise regression functions. We show that several popular disparity measures—the deviations from demographic parity, equality of opportunity, and predictive equality—are bilinear.We find the form of Bayes-optimal fair classifiers under a single linear disparity measure, by uncovering a connection with the Neyman-Pearson lemma. For bilinear disparity measures, we are able to find the explicit form of Bayes-optimal fair classifiers as group-wise thresholding rules with explicitly characterized thresholds. We develop similar algorithms for when the protected attribute cannot be used at the prediction phase. Moreover, we obtain analogous theoretical characterizations of optimal classifiers for a multi-class protected attribute and for equalized odds. Leveraging our theoretical results, we design methods that learn fair Bayes-optimal classifiers under bilinear disparity constraints. Our methods cover three popular approaches to fairness-aware classification, via pre-processing (Fair Up- and Down-Sampling), in-processing(Fair cost-sensitive Classification) and post-processing (a Fair Plug-In Rule). Our methods control disparity directly while achieving near-optimal fairness-accuracy tradeoffs. We show empirically that our methods have state-of-the-art performance compared to existing algorithms. In particular, our pre-processing method can reach a higher accuracy than prior pre-processing methods at low disparity levels.
機械学習アルゴリズムは、保護対象となる特定のグループに対して不均等な影響を及ぼす可能性があります。これに対処するため、我々はベイズ最適かつ公平な分類を行う手法を開発し、所定のグループ公平性制約の下で分類誤差を最小化することを目指します。ここで我々は、確率的分類器の線形関数である「線形不均衡尺度(linear disparity measures)」と、グループごとの回帰関数に対しても線形である「双線形不均衡尺度(bilinear disparity measures)」という概念を導入します。人口統計的パリティ(demographic parity)、機会の平等(equality of opportunity)、予測の平等(predictive equality)からの乖離といった一般的な不均衡尺度のいくつかが、双線形不均衡尺度であることを示します。単一の線形不均衡尺度におけるベイズ最適かつ公平な分類器の形式を、ネイマン・ピアソンの補題との関連性を明らかにすることで導き出します。双線形不均衡尺度については、閾値が明示的に規定された「グループごとの閾値処理ルール」として、ベイズ最適かつ公平な分類器の具体的な形式を導出します。また、予測段階で保護属性を使用できない場合についても同様のアルゴリズムを開発します。さらに、多クラスの保護属性や「機会の均等化(equalized odds)」の場合についても、最適分類器に関する同様の理論的特性を導き出します。これらの理論的成果を活用し、双線形不均衡制約の下で公平かつベイズ最適な分類器を学習する手法を設計します。本研究で提案する手法は、公平性を考慮した分類(fairness-aware classification)における主要な3つのアプローチ、すなわち前処理(公平性を考慮したアップサンプリングおよびダウンサンプリング)、学習時処理(公平性を考慮したコスト感応型分類)、および後処理(公平性を考慮したプラグイン・ルール)を網羅しています。これらの手法は、公平性と精度のトレードオフにおいて最適に近い水準を達成しつつ、不公平性(disparity)を直接的に制御します。既存のアルゴリズムと比較して、提案手法が最先端(SOTA)の性能を示すことを実証的に明らかにしました。特に、前処理手法に関しては、低い不公平性のレベルにおいて、従来の前処理手法よりも高い精度を達成できることが示されました。
Investigating the Histogram Loss in Regression
Investigating the Histogram Loss in Regression / 回帰におけるヒストグラム損失の検討
It is becoming increasingly common in regression to train neural networks that model the entire distribution even if only the mean is required for prediction. This additional modeling often comes with performance gains, and the reasons behind the improvement are not fully known. This paper investigates a recent approach to regression, the histogram loss, which involves learning the conditional distribution of the target variable by minimizing the cross-entropy between a target distribution and a flexible histogram prediction. The resulting loss corresponds to a classification loss: a cross-entropy between the outputs and a smoothed label vector. We design theoretical and empirical analyses to determine why and when this performance gain appears and how different components of the loss contribute to it. Our results suggest that the benefits of learning distributions in this setup come from improvements in optimization rather than modeling extra information. We then demonstrate the viability of the histogram loss in common deep learning applications without the need for costly hyperparameter tuning.
回帰タスクにおいて、予測に必要なのが平均値のみである場合でも、分布全体をモデル化するニューラルネットワークを学習させることが一般的になりつつあります。このような追加的なモデル化はしばしば性能向上をもたらしますが、その改善の理由は完全には解明されていません。本論文では、回帰に対する近年のアプローチである「ヒストグラム損失(histogram loss)」について調査します。この手法は、ターゲット分布と柔軟なヒストグラム予測との間の交差エントロピーを最小化することで、目的変数の条件付き分布を学習するものです。その結果生じる損失は分類タスクの損失、すなわちネットワークの出力と平滑化されたラベルベクトルとの間の交差エントロピーに対応します。我々は、なぜ、そしてどのような状況でこの性能向上が生じるのか、また損失の各構成要素がどのように寄与しているのかを明らかにするため、理論的および実証的な分析を行いました。その結果、この設定における分布学習の利点は、追加的な情報をモデル化することではなく、最適化プロセスの改善に由来することが示唆されました。さらに、コストのかかるハイパーパラメータの調整を必要とすることなく、一般的なディープラーニングの応用においてヒストグラム損失が有効であることを実証しました。
Optimal Approximation and Generalization Errors for Deep Convolutional Neural Networks
Optimal Approximation and Generalization Errors for Deep Convolutional Neural Networks / 深層畳み込みニューラルネットワークの最適近似と汎化誤差
This paper focuses on approximation and learning performances of deep convolutional neural networks with zero-padding and max-pooling. We prove that, to approximate $r$-smooth function, the approximation rates of deep convolutional neural networks with depth $L$ are of order $ (L/\log L)^{-2r/d} $, which is optimal up to a logarithmic factor. Furthermore, we deduce almost optimal generalization errors for implementing empirical risk minimization over deep convolutional neural networks. Our theoretical results are verified by several numerical experiments to show the power of the convolutional structure, zero-padding and max-pooling.
本論文では、ゼロパディングとマックスプーリングを用いた深層畳み込みニューラルネットワークの近似および学習性能に焦点を当てる。我々は、$r$階滑らかな関数を近似する際、深さ$L$の深層畳み込みニューラルネットワークの近似誤差がオーダー$(L/\log L)^{-2r/d}$であることを証明します。これは対数因子を除いて最適です。さらに、深層畳み込みニューラルネットワーク上での経験リスク最小化の実行に伴う、ほぼ最適な汎化誤差を導出します。我々の理論的結果は、いくつかの数値実験によって検証され、畳み込み構造、ゼロパディング、およびマックスプーリングの有効性が示されます。
Approximations and Learning for Continuous State and Action MDPs under Average Cost Criteria
Approximations and Learning for Continuous State and Action MDPs under Average Cost Criteria / 平均コスト基準下での連続状態・行動MDPの近似と学習
In this paper, for Markov Decision Processes (MDPs) with standard Borel spaces, (i) we first provide a discretization based approximation method for MDPs with continuous spaces under average cost criteria, and provide error bounds for approximations when the dynamics are only weakly continuous (for asymptotic convergence of errors as the grid sizes vanish) or Wasserstein continuous (with a rate in approximation as the grid sizes vanish) under certain ergodicity assumptions. In particular, we relax the total variation condition given in prior work to weak continuity or Wasserstein continuity. (ii) We provide synchronous and asynchronous (quantized) Q-learning algorithms for continuous spaces via quantization (where the quantized state is taken to be the actual state in corresponding Q-learning algorithms presented in the paper), and establish their convergence. (iii) We finally show that the convergence is to the optimal Q values of a finite approximate model constructed via quantization, which implies near optimality of the arrived solution.
本論文では、標準ボレル空間上のマルコフ決定過程(MDP)について、以下の点を示す。(i)まず、平均コスト基準の下で連続空間上のMDPに対する離散化に基づく近似手法を提示します。また、特定のエルゴード性仮定の下で、ダイナミクスが弱連続である場合(グリッドサイズがゼロに近づく際の誤差の漸近収束について)やワッサースタイン連続である場合(グリッドサイズがゼロに近づく際の近似レートについて)の近似誤差限界を導出します。特に、先行研究における全変動に関する条件を、弱連続性またはワッサースタイン連続性へと緩和します。(ii)量子化を介して連続空間向けの同期型および非同期型(量子化)Q学習アルゴリズムを提示し(ここで、量子化された状態は、本論文で提示されるQ学習アルゴリズムにおける実際の状態として扱われる)、それらの収束性を確立します。(iii)最後に、その収束先が量子化によって構築された有限近似モデルの最適Q値であることを示し、これにより得られた解が準最適であることを明らかにします。
Minimax density estimation in the adversarial framework under local differential privacy
Minimax density estimation in the adversarial framework under local differential privacy / 局所差分プライバシー下、敵対的枠組みにおけるミニマックス密度推定
We consider the problem of nonparametric density estimation under privacy constraints in an adversarial framework. To this end, we study minimax rates over Sobolev spaces under local differential privacy. We first obtain a lower bound which allows us to quantify the impact of privacy compared with the classical framework. Next, we introduce a new Coordinate block privacy mechanism that guarantees local differential privacy, which, coupled with a projection estimator, achieves the minimax optimal rates. Finally, we develop an adaptive procedure which is optimal in the minimax sense up to logarithmic terms.
本論文では、敵対的枠組みにおけるプライバシー制約下でのノンパラメトリック密度推定問題を考察します。そのために、局所差分プライバシーの下でのソボレフ空間上のミニマックスレートを研究します。まず、従来の枠組みと比較してプライバシーの影響を定量化できる下界を導出します。次に、局所差分プライバシーを保証する新しい「座標ブロック・プライバシー・メカニズム」を導入し、これを射影推定量と組み合わせることで、ミニマックス最適レートを達成します。最後に、対数項を除いてミニマックスの意味で最適となる適応的手法を開発します。
A Mean-Field Analysis of Neural Stochastic Gradient Descent-Ascent for Functional Minimax Optimization
A Mean-Field Analysis of Neural Stochastic Gradient Descent-Ascent for Functional Minimax Optimization / 関数型ミニマックス最適化におけるニューラル確率的勾配降下・上昇法の平均場解析
This paper studies minimax optimization problems defined over infinite-dimensional function classes of over-parameterized two-layer neural networks. In particular, we consider the minimax optimization problem stemming from estimating linear functional equations defined by conditional expectations, where the objective functions are quadratic in the functional spaces. We address (i) the convergence of the stochastic gradient descent-ascent algorithm and (ii) the representation learning of the neural networks. We establish convergence in the mean-field regime by considering the continuous-time, infinite-width limit of the optimization dynamics.Under this regime, stochastic gradient descent-ascent corresponds to a Wasserstein gradient flow over the space of probability measures defined over the space of neural network parameters. We prove that the Wasserstein gradient flow converges globally to a stationary point of the minimax objective at a $\mathcal{O}(T^{-1} + \alpha^{-1} ) $ sublinear rate, and additionally finds the solution to the functional equation when the regularizer of the minimax objective is strongly convex. Here $T$ denotes the time and $\alpha$ is a scaling parameter of the neural networks. In terms of representation learning, our results show that the feature representation induced by the neural networks may deviate from the initial representation by a factor of $\mathcal{O}(\alpha^{-1})$, measured by the Wasserstein distance. Finally, we apply our general results to concrete examples, including policy evaluation, nonparametric instrumental variable regression, and asset pricing.
本論文では、過剰パラメータ化された2層ニューラルネットワークの無限次元関数クラス上で定義されるミニマックス最適化問題を研究します。特に、条件付き期待値によって定義される線形関数方程式の推定に由来するミニマックス最適化問題を考察します。ここでは、関数空間上の目的関数が二次形式となります。我々は、(i)確率的勾配降下・上昇アルゴリズムの収束性と、(ii)ニューラルネットワークによる表現学習、という2つの点に取り組みます。最適化ダイナミクスの連続時間かつ無限幅の極限を考えることで、平均場レジームにおける収束性を確立します。このレジームにおいて、確率的勾配降下・上昇法は、ニューラルネットワークのパラメータ空間上の確率測度空間におけるワッサースタイン勾配流に対応します。我々は、このワッサースタイン勾配流がミニマックス目的関数の定常点へ$\mathcal{O}(T^{-1} + \alpha^{-1})$の劣線形レートで大域的に収束すること、さらにミニマックス目的関数の正則化項が強凸である場合には関数方程式の解を求めることを証明します。ここで、$T$は時間、$ \alpha $はニューラルネットワークのスケーリングパラメータを表します。表現学習の観点からは、ニューラルネットワークによって誘導される特徴表現が、ワッサースタイン距離で測った際に初期表現から$\mathcal{O}(\alpha^{-1})$のオーダーで乖離し得ることを示します。最後に、我々の一般的な結果を、方策評価、ノンパラメトリック操作変数回帰、資産価格決定などの具体例に適用します。
Nonparametric Partial Disentanglement via Mechanism Sparsity: Sparse Actions, Interventions and Sparse Temporal Dependencies
Nonparametric Partial Disentanglement via Mechanism Sparsity: Sparse Actions, Interventions and Sparse Temporal Dependencies / メカニズムのスパース性によるノンパラメトリックな部分的非絡み合い表現学習:スパースなアクション、介入、および時間的依存関係
This work introduces a novel principle for disentanglement we call mechanism sparsity regularization, which applies when the latent factors of interest depend sparsely on observed auxiliary variables and/or past latent factors. We propose a representation learning method that induces disentanglement by simultaneously learning the latent factors and the sparse causal graphical model that explains them. We develop a nonparametric identifiability theory that formalizes this principle and shows that the latent factors can be recovered by regularizing the learned causal graph to be sparse, under some assumptions such as the absence of instantaneous causal effects between latent factors. More precisely, we show identifiability up to a novel equivalence relation we call consistency, which allows some latent factors to remain entangled (hence the term partial disentanglement). To describe the structure of this entanglement, we introduce the notions of entanglement graphs and graph preserving functions. We further provide a graphical criterion which guarantees complete disentanglement, that is identifiability up to permutations and element-wise transformations. We demonstrate the scope of the mechanism sparsity principle as well as the assumptions it relies on with several worked out examples. For instance, the framework shows how one can leverage multi-node interventions with unknown targets on the latent factors to disentangle them. We further draw connections between our nonparametric results and the now popular exponential family assumption. Lastly, we propose an estimation procedure based on variational autoencoders and a sparsity constraint and demonstrate it on various synthetic datasets. This work is meant to be a significantly extended version of a work published at CLeaR 2022.
本研究では、「メカニズムのスパース性正則化(mechanism sparsity regularization)」と呼ぶ、新たな分離(disentanglement)の原理を導入します。これは、関心対象である潜在因子が、観測された補助変数や過去の潜在因子にスパースに依存する場合に適用されるものです。我々は、潜在因子とそれらを説明するスパースな因果グラフィカルモデルを同時に学習することで分離を誘導する、表現学習手法を提案します。この原理を定式化するノンパラメトリックな識別可能性理論を展開し、潜在因子間の同時的な因果効果が存在しないといった仮定の下で、学習された因果グラフをスパースにするよう正則化することで潜在因子を復元できることを示します。より正確には、「整合性(consistency)」と呼ぶ新たな同値関係の範囲内での識別可能性を示します。この関係は、一部の潜在因子が分離されずに絡み合ったまま残ることを許容するものです(そのため「部分的分離(partial disentanglement)」と呼びます)。この絡み合いの構造を記述するために、「絡み合いグラフ(entanglement graphs)」および「グラフ保存関数(graph preserving functions)」という概念を導入します。さらに、完全な分離(すなわち、置換および要素ごとの変換を除いた識別可能性)を保証するグラフ的基準を提示します。具体的な例を用いて、「メカニズムのスパース性(稀少性)」という原理の適用範囲と、それが依拠する仮定について説明します。例えば、本フレームワークは、潜在因子に対する介入の対象が未知である場合でも、多ノード介入を活用してそれらを分離する方法を示しています。また、我々のノンパラメトリックな結果と、現在広く用いられている指数型分布族の仮定との関連性についても論じます。最後に、変分オートエンコーダとスパース性制約に基づく推定手法を提案し、様々な合成データセットを用いてその有効性を実証します。本研究は、CLeaR 2022で発表された研究を大幅に拡張したものです。
Corruptions of Supervised Learning Problems: Typology and Mitigations
Corruptions of Supervised Learning Problems: Typology and Mitigations / 教師あり学習問題における汚染(Corruptions):分類と緩和策
Corruption is notoriously widespread in data collection. Despite extensive research, the existing literature predominantly focuses on specific settings and learning scenarios, lacking aunified view of corruption modelization and mitigation. In this work, we develop a generaltheory of corruption, which incorporates all modifications to a supervised learning problem,including changes in model class and loss. Focusing on changes to the underlying probability distributions via Markov kernels, our approach leads to three novel opportunities.First, it enables the construction of a novel, provably exhaustive corruption framework,distinguishing among different corruption types. This serves to unify existing models andestablish a consistent nomenclature. Second, it facilitates a systematic analysis of corruption consequences on learning tasks, by considering Bayes risks in the clean and corruptedscenarios. Notably, while label corruptions affect only the loss function, attribute corruptions additionally influence the hypothesis class. Third, building upon these results, weinvestigate mitigations for various corruption types. We expand existing loss-correctionmethods for label corruption to handle dependent corruption types. Our findings highlightthe necessity to generalize this classical corruption-corrected learning framework to a newparadigm with weaker requirements to encompass more corruption types. We provide sucha paradigm as well as loss correction formulas in the attribute and joint corruption cases.
データ収集における「汚染(corruption)」は、広く蔓延している問題として知られています。広範な研究が行われているものの、既存の文献は特定の環境や学習シナリオに焦点を当てたものが大半であり、汚染のモデル化や緩和策に関する統一的な見解は欠如しています。本研究では、モデルクラスや損失関数の変更を含む、教師あり学習問題へのあらゆる変更を包含する、汚染に関する一般理論を構築します。マルコフカーネルを介した基礎となる確率分布の変化に着目する本アプローチにより、3つの新たな可能性が開かれます。第一に、異なる汚染タイプを区別しつつ、網羅的であることが証明可能な新しい汚染フレームワークを構築できます。これは、既存モデルの統合や一貫した用語体系の確立に寄与します。第二に、クリーンなシナリオと汚染されたシナリオにおけるベイズリスクを考慮することで、学習タスクに対する汚染の影響を体系的に分析できます。特筆すべき点として、ラベル汚染は損失関数にのみ影響を及ぼすのに対し、属性汚染は仮説クラスにも影響を与えることが挙げられます。第三に、これらの結果に基づき、様々な汚染タイプに対する緩和策を検討します。ラベル汚染に対する既存の損失補正手法を拡張し、依存関係のある汚染タイプにも対応できるようにします。我々の知見は、より多くの汚染タイプを包含するために、この古典的な汚染補正学習フレームワークを、より緩やかな要件を持つ新しいパラダイムへと一般化する必要性を示唆しています。本研究では、そのようなパラダイムを提示するとともに、属性汚染および複合的な汚染(joint corruption)のケースにおける損失補正式を導出します。
Differentially Private Best-Arm Identification
Differentially Private Best-Arm Identification / 差分プライバシーを保証する最適アーム特定
Best Arm Identification (BAI) problems are progressively used for data-sensitive applications, such as designing adaptive clinical trials, tuning hyper-parameters, and conducting user studies. Motivated by the data privacy concerns invoked by these applications, we study the problem of BAI with fixed confidence in both the local and central models, i.e. under $\epsilon$-local and $\epsilon$-global Differential Privacy (DP). First, to quantify the cost of privacy,we derive lower bounds on the sample complexity of any $\delta$-correct BAI algorithm satisfying$\epsilon$-global DP or $\epsilon$-local DP. Our lower bounds suggest the existence of two privacy regimes. In the high-privacy regime, the hardness depends on a coupled effect of privacy and novel information-theoretic quantities involving the Total Variation distance. In the low-privacy regime, the lower bounds reduce to the non-private lower bounds. We propose $\epsilon$-local DP and $\epsilon$-global DP variants of a Top Two algorithm, namely CTB-TT and AdaP-TT$\star$, respectively. For $\epsilon$-local DP, CTB-TT is asymptotically optimal by plugging in a private estimator of the means based on Randomised Response. For $\epsilon$-global DP, our private estimator of the mean runs in arm-dependent adaptive episodes and adds Laplace noise toensure a good privacy-utility trade-off. By adapting the transportation costs, the expected sample complexity of AdaP-TT$\star$ reaches the asymptotic lower bound in the asymptotic high-privacy regime, up to a small multiplicative constant.
「最良アーム特定(Best Arm Identification: BAI)」問題は、適応的臨床試験の設計、ハイパーパラメータの調整、ユーザー調査の実施など、データへの配慮が求められる応用分野でますます利用されるようになっています。こうした応用に伴うデータプライバシーへの懸念を背景に、我々は、ローカルモデルおよびセントラルモデルの双方において、すなわち$\epsilon$-ローカル差分プライバシー(DP)および$\epsilon$-グローバルDPの下で、固定信頼度(fixed confidence)によるBAI問題を研究します。まず、プライバシー保護に伴うコストを定量化するために、$\epsilon$-グローバルDPまたは$\epsilon$-ローカルDPを満たす任意の$\delta$-正当なBAIアルゴリズムのサンプル複雑さに対する下界を導出します。導出された下界は、2つのプライバシー領域(regime)の存在を示唆しています。プライバシー保護が強力な領域(high-privacy regime)では、問題の難易度は、プライバシー保護と、全変動距離(Total Variation distance)を含む新たな情報理論的量との複合的な影響に依存します。一方、プライバシー保護が緩やかな領域(low-privacy regime)では、下界はプライバシー保護を考慮しない場合の下界へと回帰します。我々は、「Top Two」アルゴリズムの$\epsilon$-ローカルDP版および$\epsilon$-グローバルDP版として、それぞれCTB-TTおよびAdaP-TT$\star$を提案します。$\epsilon$-ローカル差分プライバシー(DP)に関しては、ランダム化回答(Randomized Response)に基づく平均のプライベート推定量を組み込むことで、CTB-TTは漸近的に最適となります。$\epsilon$-グローバルDPに対しては、我々の平均のプライベート推定量は、アームに依存する適応的なエピソードで動作し、プライバシーと有用性の良好なトレードオフを確保するためにラプラスノイズを付加します。輸送コストを適切に調整することで、AdaP-TT$\star$の期待サンプル複雑度は、漸近的な高プライバシー領域において、小さな乗法定数を除き、漸近的な下限に到達します。
Vector-Valued Gaussian Processes for Approximating Divergence- or Rotation-free Vector Fields
Vector-Valued Gaussian Processes for Approximating Divergence- or Rotation-free Vector Fields / 発散フリーまたは回転フリーなベクトル場の近似のためのベクトル値ガウス過程
In this paper, we discuss vector-valued Gaussian processes for the approximation of divergence- or rotation-free functions. We establish the theory for such Gaussian processes, then link the theory to multivariate approximation theory, and finally give error estimates forthe predictive mean in various situations.
本論文では、発散フリー(divergence-free)または回転フリー(rotation-free)な関数の近似のための、ベクトル値ガウス過程について論じます。我々は、そのようなガウス過程に関する理論を構築し、その理論を多変数近似理論と結びつけ、最終的に様々な状況下での予測平均に対する誤差評価を提示します。
Stochastic Differential Equations models for Least-Squares Stochastic Gradient Descent
Stochastic Differential Equations models for Least-Squares Stochastic Gradient Descent / 最小二乗確率的勾配降下法のための確率微分方程式モデル
We study the dynamics of a continuous-time model of stochastic gradient descent (SGD) for the least-square problem. Indeed, pursuing the work of, we analyze stochastic differential equations (SDEs) that model SGD either in the case of the training loss (finite samples) or the population one (online setting). A key qualitative feature of the dynamics is the existence of a perfect interpolator of the data, irrespective of the sample size. In both scenarios, we provide precise, non-asymptotic rates of convergence to the (possibly degenerate) stationary distribution. Additionally, we describe this asymptotic distribution, offering estimates of its mean, deviations from it, and a proof of the emergence of heavy-tails related to the step-size magnitude. Numerical simulations supporting our findings are also presented.
我々は、最小二乗問題に対する確率的勾配降下法(SGD)の連続時間モデルのダイナミクスを研究します。具体的には、既存の研究を発展させ、学習損失(有限サンプル)の場合と母集団損失(オンライン設定)の場合の双方において、SGDをモデル化する確率微分方程式(SDE)を解析します。このダイナミクスの重要な定性的特徴は、サンプルサイズに関わらず、データに対する完全な補間(perfect interpolator)が存在することです。いずれのシナリオにおいても、(縮退している可能性のある)定常分布への収収束について、漸近的ではない精密な収束レートを提示します。さらに、この漸近分布について、その平均や平均からの偏差の推定、およびステップサイズの大きさに起因するヘビーテール(裾の重い分布)の出現に関する証明を提示します。また、我々の知見を裏付ける数値シミュレーションも示します。
A Fully Parameter-Free Second-Order Algorithm for Convex-Concave Minimax Problems
A Fully Parameter-Free Second-Order Algorithm for Convex-Concave Minimax Problems / 凸凹(Convex-Concave)ミニマックス問題に対する完全パラメータフリーな二次アルゴリズム
In this paper, we study second-order algorithms for the convex-concave minimax problem, which has attracted much attention in many fields such as machine learning in recent years. We propose a Lipschitz-free cubic regularization (LF-CR) algorithm for solving the convex-concave minimax optimization problem without knowing the Lipschitz constant. It can be shown that the iteration complexity of the LF-CR algorithm to obtain an $\epsilon$-optimal solution with respect to the restricted primal-dual gap is upper bounded by $\mathcal{O}(\rho^{2/3}\|z_0-z^*\|^2\epsilon^{-2/3})$ , where $z_0=(x_0,y_0)$ is a pair of initial points, $z^*=(x^*,y^*)$ is a pair of optimal solutions, and $\rho$ is the Lipschitz constant. We further propose a fully parameter-free cubic regularization (FF-CR) algorithm that does not require any parameters of the problem, including the Lipschitz constant and the upper bound of the distance from the initial point to the optimal solution. We also prove that the iteration complexity of the FF-CR algorithm to obtain an $\epsilon$-optimal solution with respect to the gradient norm is upper bounded by $\mathcal{O}(\rho^{2/3}\|z_0-z^*\|^{4/3}\epsilon^{-2/3}) $. Numerical experiments show the efficiency of both algorithms. To the best of our knowledge, the proposed FF-CR algorithm is a completely parameter-free second-order algorithm, and its iteration complexity is currently the best in terms of $\epsilon$ under the termination criterion of the gradient norm.
本論文では、近年機械学習などの多くの分野で注目を集めている凸・凹ミニマックス問題に対する二次アルゴリズムについて研究します。我々は、リプシッツ定数を事前に知ることなく凸・凹ミニマックス最適化問題を解くための、リプシッツ定数不要な(Lipschitz-free)三次正則化(LF-CR)アルゴリズムを提案します。LF-CRアルゴリズムを用いて、制限付き主双対ギャップ(restricted primal-dual gap)に関して$\epsilon$最適解を得るための反復計算量は、$\mathcal{O}(\rho^{2/3}\|z_0-z^*\|^2\epsilon^{-2/3})$で上界評価できることが示されます。ここで、$z_0=(x_0,y_0)$は初期点の組、$z^*=(x^*,y^*)$は最適解の組、$\rho$はリプシッツ定数です。さらに我々は、リプシッツ定数や初期点から最適解までの距離の上界など、問題に関するいかなるパラメータも必要としない、完全パラメータフリーな三次正則化(FF-CR)アルゴリズムを提案します。また、FF-CRアルゴリズムを用いて勾配ノルムに関して$\epsilon$最適解を得るための反復計算量が、$\mathcal{O}(\rho^{2/3}\|z_0-z^*\|^{4/3}\epsilon^{-2/3})$で上界評価できることも証明します。数値実験により、両アルゴリズムの有効性が示されます。我々の知る限り、提案するFF-CRアルゴリズムは完全にパラメータフリーな二次アルゴリズムであり、その反復計算量は、勾配ノルムに基づく停止基準の下で、$\epsilon$に関する依存性において現在最良のものです。
Limiting Over-Smoothing and Over-Squashing of Graph Message Passing by Deep Scattering Transforms
Limiting Over-Smoothing and Over-Squashing of Graph Message Passing by Deep Scattering Transforms / Deep Scattering Transformによるグラフメッセージパッシングの過剰平滑化(Over-smoothing)および過剰圧縮(Over-squashing)の抑制
Graph neural networks (GNNs) have become pivotal tools for processing graph-structured data, leveraging the message passing scheme as their core mechanism. However, traditional GNNs often grapple with issues such as instability, over-smoothing, and over-squashing, which can degrade performance and create a trade-off dilemma. In this paper, we introduce a discriminatively trained, multi-layer Deep Scattering Message Passing (DSMP) neural network designed to overcome these challenges. By harnessing spectral transformation, the DSMP model aggregates neighboring nodes with global information, thereby enhancing the precision and accuracy of graph signal processing. We provide theoretical proofs demonstrating the DSMP’s effectiveness in mitigating these issues under specific conditions. Additionally, we support our claims with empirical evidence and thorough frequency analysis, showcasing the DSMP’s superior ability to address instability, over-smoothing, and over-squashing.
グラフニューラルネットワーク(GNN)は、メッセージパッシングの仕組みを中核として活用し、グラフ構造データの処理において極めて重要なツールとなっています。しかし、従来のGNNは、不安定性、過剰平滑化(over-smoothing)、過剰圧縮(over-squashing)といった問題に直面することが多く、これらは性能低下やトレードオフのジレンマを招く可能性があります。本論文では、これらの課題を克服するために設計された、識別的に学習される多層Deep Scattering Message Passing(DSMP)ニューラルネットワークを導入します。DSMPモデルはスペクトル変換を活用することで、近傍ノードの情報を大域的な情報と統合し、それによってグラフ信号処理の精度と正確性を向上させます。我々は、特定の条件下においてDSMPがこれらの問題を緩和する有効性を示す理論的証明を提示します。さらに、実証的な証拠と詳細な周波数解析によって主張を裏付け、不安定性、過剰な平滑化(over-smoothing)、および過剰な圧縮(over-squashing)に対処するDSMPの優れた能力を実証します。
Multi-relational Network Autoregression Model with Latent Group Structures
Multi-relational Network Autoregression Model with Latent Group Structures / 潜在的なグループ構造を伴う多重関係ネットワーク自己回帰モデル
Multi-relational networks among entities are frequently observed in the era of big data. Quantifying the effects of multiple networks has attracted significant research interest recently. In this work, we model multiple network effects through an autoregressive framework for tensor-valued time series. To characterize the potential heterogeneity of the networks and handle the high dimensionality of the time series data simultaneously, we assume a separate group structure for entities in each network and estimate all group memberships in a data-driven fashion. Specifically, we propose a group tensor network autoregression (GTNAR) model, which assumes that within each network, entities in the same group share the same set of model parameters, and the parameters differ across networks. An iterative algorithm is developed to estimate the model parameters and the latent group memberships simultaneously. Theoretically, we show that the group-wise parameters and group memberships can be consistently estimated when the group numbers are correctly or possibly over-specified. An information criterion for estimating the group number for each network is also provided to consistently select the group numbers. Lastly, we apply the GTNAR method to a Yelp dataset to illustrate its usefulness.
ビッグデータ時代において、エンティティ間の多重関係ネットワークは頻繁に見られます。複数のネットワークが及ぼす影響を定量化することは、近年、大きな研究関心を集めています。本研究では、テンソル値時系列に対する自己回帰の枠組みを用いて、多重ネットワークの影響をモデル化します。ネットワークが持つ潜在的な異質性を特徴づけつつ、時系列データの高次元性にも同時に対処するため、各ネットワーク内のエンティティに対して個別のグループ構造を仮定し、データ駆動型の手法ですべてのグループ所属を推定します。具体的には、グループ・テンソル・ネットワーク自己回帰(GTNAR)モデルを提案します。このモデルは、各ネットワーク内において同一グループのエンティティが同じモデルパラメータを共有し、かつパラメータがネットワーク間で異なることを仮定するものです。モデルパラメータと潜在的なグループ所属を同時に推定するための反復アルゴリズムも開発しました。理論面では、グループ数が正しく指定された場合、あるいは過剰に指定された場合であっても、グループごとのパラメータとグループ所属を一致性を持って推定できることを示します。また、各ネットワークのグループ数を一致性を持って選択するための情報量基準も提示します。最後に、YelpのデータセットにGTNAR手法を適用し、その有用性を示します。
Enhancing Accuracy in Generative Models via Knowledge Transfer
Enhancing Accuracy in Generative Models via Knowledge Transfer / 知識転移による生成モデルの精度向上
This paper investigates the accuracy of generative models and the impact of knowledge transfer on their generation precision. Specifically, we examine a generative model for a target task, fine-tuned using a pre-trained model from a source task. Building on the “Shared Embedding” concept, which bridges the source and target tasks, we introduce a novel framework for transfer learningunder distribution metrics such as the Kullback-Leibler divergence. This framework underscores the importance of leveraging inherent similarities between diverse tasks despite their distinct data distributions. Our theory suggests that the shared structures can augment the generation accuracy for a target task, reliant on the capability of a source model to identify shared structures and effective knowledge transfer from source to target learning. To demonstrate the practical utility of this framework, we explore the theoretical implications for two specific generative models: diffusion and normalizing flows. The results show enhanced performance in both models over their non-transfer counterparts, indicating advancements for diffusion models and providing fresh insights into normalizing flows in transfer and non-transfer settings. These results highlight the significant contribution of knowledge transfer in boosting the generation capabilities of these models.
本論文では、生成モデルの精度と、知識転移が生成精度に及ぼす影響について調査します。具体的には、ソースタスクで事前学習されたモデルを用いてファインチューニングされた、ターゲットタスク向けの生成モデルを検討します。ソースタスクとターゲットタスクを橋渡しする「共有埋め込み(Shared Embedding)」の概念に基づき、カルバック・ライブラー・ダイバージェンスなどの分布間距離指標の下での転移学習のための新しいフレームワークを導入します。このフレームワークは、データ分布が異なる多様なタスク間であっても、それらが持つ本質的な類似性を活用することの重要性を強調するものです。我々の理論によれば、ソースモデルが共有構造を特定する能力と、ソースからターゲットへの効果的な知識転移が行われることを前提として、共有構造はターゲットタスクにおける生成精度を向上させることができます。このフレームワークの実用的な有用性を示すため、拡散モデル(diffusion models)と正規化流(normalizing flows)という2つの具体的な生成モデルについて、理論的な含意を検討します。その結果、両モデルとも転移学習を行わない場合と比較して性能が向上することが示されました。これは拡散モデルにおける進歩を示すとともに、転移学習および非転移学習の各設定における正規化流について新たな知見を提供するものです。これらの結果は、これらモデルの生成能力を向上させる上で、知識転移が重要な役割を果たすことを浮き彫りにしています。
Demographic Parity in Regression and Classification Within the Unawareness Framework
Demographic Parity in Regression and Classification Within the Unawareness Framework / 「非認識(Unawareness)」の枠組みにおける回帰および分類での人口統計学的パリティ(Demographic Parity)
This paper explores the theoretical foundations of fair regression under the constraint of demographic parity within the unawareness framework, where disparate treatment is prohibited, extending existing results where such treatment is permitted. Specifically, we aim to characterize the optimal fair regression function when minimizing the quadratic loss. Our results reveal that this function is given by the solution to a barycenter problem with optimal transport costs. Additionally, we study the connection between optimal fair cost-sensitive classification, and optimal fair regression. We demonstrate that nestedness of the decision sets of the classifiers is both necessary and sufficient to establish a form of equivalence between classification and regression. Under this nestedness assumption, the optimal classifiers can be derived by applying thresholds to the optimal fair regression function; conversely, the optimal fair regression function is characterized by the family of cost-sensitive classifiers.
本論文では、差別的な扱いが許容される既存の研究を拡張し、差別的な扱いが禁止される「無知(unawareness)」の枠組みの下、人口統計学的パリティ(demographic parity)の制約を課した公平な回帰分析の理論的基礎を探求します。具体的には、二乗損失を最小化する際の最適な公平回帰関数を特定することを目的とします。その結果、この関数は最適輸送コストを伴う重心問題(barycenter problem)の解として与えられることが明らかになった。さらに、最適な公平コスト敏感型分類(cost-sensitive classification)と最適な公平回帰との関連性についても検討します。分類器の決定集合が包含関係(nestedness)にあることが、分類と回帰の間に一種の等価性を成立させるための必要十分条件であることを示す。この包含関係の仮定の下では、最適な公平回帰関数にしきい値を適用することで最適な分類器を導出でき、逆に、最適な公平回帰関数はコスト敏感型分類器の族によって特徴付けられます。
Differentially Private Estimation and Inference in High-Dimensional Regression with FDR Control
Differentially Private Estimation and Inference in High-Dimensional Regression with FDR Control / FDR制御を伴う高次元回帰における差分プライバシー保護下の推定と推論
This paper proposes new methodologies for conducting practical differentially private (DP) estimation and inference in high-dimensional linear regression. We first introduce a DP Bayesian Information Criterion (DP-BIC) for selecting the unknown sparsity parameter in differentially private sparse linear regression (DP-SLR), eliminating the need for prior knowledge of model sparsity, which is a requisite in the existing literature. Next, we develop the DP debiased algorithm that enables privacy-preserving inference on a particular subset of regression parameters. Our proposed method enables privacy-preserving inference on the regression parameters by leveraging the inherent sparsity of high-dimensional linear regression models. Additionally, we address private feature selection by considering multiple testing in high-dimensional linear regression by introducing a DP multiple testing procedure that controls the false discovery rate (FDR). This allows for accurate and privacy-preserving identification of significant predictors in the regression model. Through extensive simulations and real data analyses, we demonstrate the effectiveness of our proposed methods in conducting inference for high-dimensional linear models while safeguarding privacy and controlling the FDR.
本論文では、高次元線形回帰において実用的な差分プライバシー(DP)に基づく推定および推論を行うための新しい手法を提案します。まず、差分プライバシー・スパース線形回帰(DP-SLR)における未知のスパース性パラメータを選択するためのDPベイズ情報量規準(DP-BIC)を導入します。これにより、既存の研究では必須とされていたモデルのスパース性に関する事前知識が不要となります。次に、回帰パラメータの特定のサブセットに対してプライバシーを保護しつつ推論を行うことを可能にする、DP(差分プライバシー)対応の非バイアス化アルゴリズムを開発します。提案手法は、高次元線形回帰モデルが本来持つスパース性を活用することで、回帰パラメータに関するプライバシー保護推論を実現します。さらに、偽発見率(FDR)を制御するDP多重検定手順を導入し、高次元線形回帰における多重検定を考慮することで、プライバシーを保護した特徴量選択にも取り組みます。これにより、回帰モデルにおける有意な予測変数を、精度を維持しつつプライバシーを保護して特定することが可能になります。広範なシミュレーションおよび実データ解析を通じて、プライバシーを保護しFDRを制御しつつ高次元線形モデルの推論を行う上で、提案手法が有効であることを実証します。
Beyond Unconstrained Features: Neural Collapse for Shallow Neural Networks with General Data
Beyond Unconstrained Features: Neural Collapse for Shallow Neural Networks with General Data / 制約のない特徴量を超えて:一般的なデータを用いた浅いニューラルネットワークにおけるニューラル・コラプス(Neural Collapse)
Neural collapse (${\cal NC}$) is a phenomenon that emerges at the terminal phase of the training (TPT) of deep neural networks (DNNs). The features of the data in the same class collapse to their respective sample means and the sample means exhibit a simplex equiangular tight frame (ETF). In the past few years, there has been a surge of works that focus on explaining why the ${\cal NC}$ occurs and how it affects generalization. Since the DNNs are notoriously difficult to analyze, most works mainly focus on the unconstrained feature model (UFM). While the UFM explains the ${\cal NC}$ to some extent, it fails to provide a complete picture of how the network architecture and the dataset affect ${\cal NC}$. In this work, we focus on shallow ReLU neural networks and try to understand how the width, depth, data dimension, and statistical property of the training dataset influence the neural collapse. We provide a complete characterization of when the ${\cal NC}$ occurs for two or three-layer neural networks. For two-layer ReLU neural networks, a sufficient condition on when the global minimizer of the regularized empirical risk function exhibits the ${\cal NC}$ configuration depends on the data dimension, sample size, and the signal-to-noise ratio in the data instead of the network width. For three-layer neural networks, we show that the ${\cal NC}$ occurs as long as the first layer is sufficiently wide. Regarding the connection between ${\cal NC}$ and generalization, we show the generalization heavily depends on the SNR (signal-to-noise ratio) in the data: even if the ${\cal NC}$ occurs, the generalization can still be bad provided that the SNR in the data is too low. Our results significantly extend the state-of-the-art theoretical analysis of the ${\cal NC}$ under the UFM by characterizing the emergence of the ${\cal NC}$ under shallow nonlinear networks and showing how it depends on data properties and network architecture.
ニューラル・コラプス(Neural Collapse: NC)は、深層ニューラルネットワーク(DNN)の学習の最終段階(TPT)で生じる現象です。この現象では、同一クラスのデータ特徴量がそれぞれのサンプル平均へと収束(コラプス)し、それらのサンプル平均がシンプレックス等角タイトフレーム(ETF)を形成します。近年、なぜNCが生じるのか、そしてそれが汎化性能にどのような影響を与えるのかを解明する研究が急増しています。DNNの解析は極めて困難であるため、多くの研究は主に「制約なし特徴量モデル(UFM)」に焦点を当ててきました。UFMはある程度NCを説明できるものの、ネットワーク構造やデータセットがNCに与える影響の全体像を捉えるには至っていません。本研究では、浅いReLUニューラルネットワークに焦点を当て、ネットワークの幅、深さ、データの次元、および学習データセットの統計的性質がNCにどのような影響を及ぼすかを解明しようと試みます。我々は、2層または3層のニューラルネットワークにおいてNCが生じる条件を完全に特定します。2層ReLUニューラルネットワークの場合、正則化経験的リスク関数の大域的最小化解がNCの構成を示すための十分条件は、ネットワークの幅ではなく、データの次元、サンプルサイズ、およびデータの信号対雑音比(SNR)に依存することが示されます。3層ニューラルネットワークについては、第1層の幅が十分に広ければNCが生じることを示します。NCと汎化性能の関連性に関しては、汎化性能がデータのSNRに大きく依存することを明らかにします。すなわち、データのSNRが低すぎる場合には、NCが生じていたとしても汎化性能が悪化し得るのです。本研究の結果は、浅い非線形ネットワークにおけるNCの発生を特徴づけ、それがデータの性質やネットワーク構造にどのように依存するかを示すことで、UFM下でのNCに関する最先端の理論的解析を大きく拡張するものです。
Asymptotics of Stochastic Gradient Descent with Dropout Regularization in Linear Models
Asymptotics of Stochastic Gradient Descent with Dropout Regularization in Linear Models / 線形モデルにおけるドロップアウト正則化付き確率的勾配降下法の漸近挙動
This paper proposes an asymptotic theory for online inference of the stochastic gradient descent (SGD) iterates with dropout regularization in linear regression. Specifically, we establish the geometric-moment contraction (GMC) for constant step-size SGD dropout iterates to show the existence of a unique stationary distribution of the dropout recursive function. Based on the GMC property, we use the functional dependence measure to provide quenched central limit theorems (CLT) for the gradient descent iterates with dropout regularization. Moreover, we obtain CLTs for the Ruppert-Polyak averaged GD (AGD) and averaged SGD (ASGD) iterates with dropout. Based on these asymptotic normality results, we further introduce an online estimator for the long-run covariance matrix of ASGD dropout to facilitate inference in a recursive manner with efficiency in computational time and memory. The numerical experiments demonstrate that for large samples, the proposed confidence intervals for ASGD with dropout achieve the nominal coverage probability.
本論文では、線形回帰におけるドロップアウト正則化を伴う確率的勾配降下法(SGD)の反復系列のオンライン推論に関する漸近理論を提案します。具体的には、一定ステップサイズのSGDドロップアウト反復系列に対して幾何学的モーメント縮約(GMC)を確立し、ドロップアウト再帰関数の唯一の定常分布の存在を示します。GMC(幾何学的混合条件)の性質に基づき、関数的依存性尺度を用いて、ドロップアウト正則化を伴う勾配降下法(GD)の反復系列に対するクエンチド中心極限定理(CLT)を導出します。さらに、ドロップアウトを伴うRuppert-Polyak平均化GD(AGD)および平均化SGD(ASGD)の反復系列についてもCLTを導きます。これらの漸近正規性の結果を踏まえ、計算時間とメモリの効率性を保ちつつ再帰的な推論を可能にする、ASGD(ドロップアウトあり)の長期共分散行列に対するオンライン推定量を提案します。数値実験により、大規模なサンプルにおいて、提案されたASGD(ドロップアウトあり)の信頼区間が所定の被覆確率を達成することが示されます。
Almost Sure Convergence of Linear Temporal Difference Learning with Arbitrary Features
Almost Sure Convergence of Linear Temporal Difference Learning with Arbitrary Features / 任意の(一般の)特徴量を用いた線形時間差分学習の概収束性
Temporal difference (TD) learning with linear function approximation (linear TD) is a classic and powerful prediction algorithm in reinforcement learning. While it is well-understood that linear TD converges almost surely to a unique point, this convergence traditionally requires the assumption that the features used by the approximator are linearly independent. However, this linear independence assumption does not hold in many practical scenarios. This work is the first to establish the almost sure convergence of linear TD without requiring linearly independent features. We prove that the weight iterates of linear TD converge to a bounded set, and that the value estimates derived from the weights in that set are the same almost everywhere. We also establish a notion of local stability of the weight iterates. Importantly, we do not impose assumptions tailored to feature dependence and do not modify the linear TD algorithm. Key to our analysis is a novel characterization of bounded invariant sets of the mean ODE of linear TD.
線形関数近似を用いた時間差分(TD)学習(線形TD)は、強化学習における古典的かつ強力な予測アルゴリズムです。線形TDが一意の点に概収束することはよく知られていますが、この収束には従来、近似器で使用される特徴量が線形独立であるという仮定が必要とされてきました。しかし、多くの実用的なシナリオでは、この線形独立性の仮定は成り立ちません。本研究は、特徴量の線形独立性を仮定することなく、線形TDの概収束を確立した最初のものです。我々は、線形TDの重み反復系列がある有界集合に収束すること、およびその集合内の重みから導出される価値推定値が至る所(almost everywhere)で一致することを証明します。また、重み反復系列の局所安定性という概念も確立します。重要な点として、特徴量の依存性に特化した仮定を課すことはなく、線形TDアルゴリズムの変更も行いません。我々の解析の鍵となるのは、線形TDの平均常微分方程式(mean ODE)における有界不変集合の新たな特徴付けです。
Convergence of Noise-Free Sampling Algorithms with Regularized Wasserstein Proximals
Convergence of Noise-Free Sampling Algorithms with Regularized Wasserstein Proximals / 正則化Wasserstein近接作用素を用いたノイズフリー・サンプリングアルゴリズムの収束性
In this work, we investigate the convergence properties of the backward regularized Wasserstein proximal (BRWP) method for sampling a target distribution. The BRWP approach can be shown as a semi-implicit time discretization for a probability flow ODE with the score function whose density satisfies the Fokker-Planck equation of the overdamped Langevin dynamics. Specifically, the evolution of the density-hence the score function-is approximated via a kernel representation derived from the regularized Wasserstein proximal operator. By applying the dual formulation and a localized Taylor series to obtain the asymptotic expansion of this kernel formula, we establish guaranteed convergence in terms of the Kullback-Leibler divergence for the BRWP method towards a strongly log-concave target distribution. Our analysis also identifies the optimal and maximum step sizes for convergence. Furthermore, we demonstrate that the deterministic and semi-implicit BRWP scheme outperforms many classical Langevin Monte Carlo methods, such as the Unadjusted Langevin Algorithm (ULA), by offering faster convergence and reduced bias. Numerical experiments further validate the convergence analysis of the BRWP method.
本研究では、ターゲット分布からサンプリングを行うための後退正則化ワッサースタイン近接(BRWP)法の収束特性を調査します。BRWP法は、過減衰ランジュバン力学のフォッカー・プランク方程式を満たす密度を持つスコア関数を用いた確率流常微分方程式(ODE)に対する、半陰的な時間離散化手法として解釈できます。具体的には、密度(ひいてはスコア関数)の時間発展を、正則化ワッサースタイン近接作用素から導出されるカーネル表現を用いて近似します。双対定式化と局所的なテイラー展開を適用してこのカーネル式の漸近展開を導くことで、強対数凹なターゲット分布に対するBRWP法のカルバック・ライブラー(KL)ダイバージェンスの意味での収束を保証します。また、我々の解析により、収束のための最適ステップサイズおよび最大ステップサイズも特定されます。さらに、決定論的かつ半陰的なBRWPスキームは、非調整ランジュバンアルゴリズム(ULA)などの多くの古典的なランジュバン・モンテカルロ法と比較して、より速い収束と低減されたバイアスを実現し、それらを凌駕することを示します。数値実験によっても、BRWP法の収束解析の妥当性が裏付けられます。
Best Arm Identification with Minimal Regret
Best Arm Identification with Minimal Regret / 最小リグレットでの最適アーム特定
Motivated by real-world applications that necessitate responsible experimentation, we introduce the problem of best arm identification (BAI) with minimal regret. This variant of the multi-armed bandit problem elegantly amalgamates two of its most ubiquitous objectives: regret minimization and BAI. More precisely, the agent’s goal is to identify the best arm with a prescribed confidence level $\delta$, while minimizing the cumulative regret up to the stopping time. Focusing on single-parameter exponential families of distributions, we leverage information-theoretic techniques to establish an instance-dependent lower bound on the expected cumulative regret. Moreover, we present an impossibility result that underscores the tension between cumulative regret and sample complexity in fixed-confidence BAI. Complementarily, we design and analyze the Double KL-UCB algorithm, which achieves asymptotic optimality as the confidence level tends to zero. Notably, this algorithm employs two distinct confidence bounds to guide arm selection in a randomized manner. Our findings elucidate a fresh perspective on the inherent connections between regret minimization and BAI.
責任ある実験が求められる現実世界の応用事例に動機づけられ、我々は最小リグレットでの「最良アーム特定(BAI)」問題を導入します。多腕バンディット問題のこの変種は、同問題における最も一般的な二つの目的、すなわちリグレット最小化とBAIを巧みに融合させたものです。より正確には、エージェントの目標は、停止時刻までの累積リグレットを最小化しつつ、所定の信頼水準$\delta$で最良のアームを特定することにあります。単一パラメータ指数型分布族に焦点を当て、情報理論的手法を活用することで、期待累積リグレットに関するインスタンス依存の下界を確立します。さらに、固定信頼水準BAIにおける累積リグレットとサンプル複雑性の間のトレードオフを浮き彫りにする不可能性の結果を提示します。これと相補的に、信頼水準がゼロに近づくにつれて漸近的最適性を達成する「Double KL-UCB」アルゴリズムを設計・解析します。特筆すべき点として、このアルゴリズムは二つの異なる信頼限界を用い、確率的な手法でアーム選択を行います。我々の知見は、リグレット最小化とBAIの間に存在する本質的な関連性について、新たな視点を提示するものです。
On the Relevance of Byzantine Robust Optimization Against Data Poisoning
On the Relevance of Byzantine Robust Optimization Against Data Poisoning / データポイズニングに対するビザンチン耐性最適化の意義について
The success of machine learning (ML) has been intimately linked with the availability oflarge amounts of data, typically collected from heterogeneous sources and processed onvast networks of computing devices (also called workers). Beyond accuracy, the use of MLin critical domains such as healthcare and autonomous driving calls for robustness againstdata poisoning and some faulty workers. The problem of Byzantine ML formalizes theserobustness issues by considering a distributed ML environment in which workers (storinga portion of the global dataset) can deviate arbitrarily from the prescribed algorithm.Although the problem has attracted a lot of attention from a theoretical point of view, itspractical importance for addressing realistic faults (where the behavior of any worker islocally constrained) remains unclear. It has been argued that the seemingly weaker threatmodel where only workers’ local datasets get poisoned is more reasonable. We prove that,while tolerating a wider range of faulty behaviors, Byzantine ML yields solutions that are,in a precise sense, optimal even under the weaker data poisoning threat model. Then, westudy a generic data poisoning model wherein some workers have fully-poisonous local data, i.e., their datasets are entirely corruptible, and the remainders have partially-poisonous local data, i.e., only a fraction of their local datasets is corruptible. We prove that Byzantine-robust schemes yield optimal solutions against both these forms of data poisoning, and that the former is more harmful when workers have heterogeneous local data.
機械学習(ML)の成功は、大量のデータの利用可能性と密接に結びついてきました。これらのデータは通常、異種混合のソースから収集され、計算デバイス(ワーカーとも呼ばれる)からなる巨大なネットワーク上で処理されます。医療や自動運転といった重要な領域で機械学習(ML)を利用する場合、単なる予測精度だけでなく、データポイズニングや一部の不具合のある(あるいは悪意ある)ワーカーに対する堅牢性が求められます。「ビザンチンML」の問題設定は、ワーカー(グローバルデータセットの一部を保持)が所定のアルゴリズムから恣意的に逸脱し得る分散ML環境を想定することで、こうした堅牢性の課題を定式化するものです。この問題は理論的な観点からは大きな注目を集めてきましたが、現実的な障害(ワーカーの挙動が局所的に制約されるような状況)に対処する上での実用的な重要性については、依然として不明確な点があります。ワーカーのローカルデータセットのみが汚染されるという、一見すると脅威の度合いが低いモデルの方が、より現実的で妥当であるという指摘もなされています。本研究では、より広範な不具合挙動を許容しつつも、ビザンチンMLの手法が、この「より弱い」データポイズニングの脅威モデルの下であっても、厳密な意味で最適な解をもたらすことを証明します。さらに、一部のワーカーが完全に汚染されたローカルデータ(すなわちデータセット全体が改ざんされ得る)を持ち、残りのワーカーが部分的に汚染されたローカルデータ(すなわちデータセットの一部のみが改ざんされ得る)を持つような、一般的なデータポイズニングモデルについても検討します。そして、ビザンチン堅牢性を持つ手法がこれら双方のデータポイズニングに対して最適解をもたらすこと、また、ワーカーが不均一なローカルデータを持つ場合には前者(完全な汚染)の方がより深刻な影響を及ぼすことを証明します。
Node Regression on Latent Position Random Graphs via Local Averaging
Node Regression on Latent Position Random Graphs via Local Averaging / 局所平均化を用いた潜在位置ランダムグラフ上のノード回帰
Node regression consists in predicting the value of a graph label at a node, given observations at the other nodes. We perform a theoretical study where the graph is generated bya Latent Position Model: each node has a latent position and the probability of connectiondepends on the distance between latent positions.We begin by studying the simplest estimator: averaging the label at all neighboringnodes. We show that in Latent Position Models this estimator tends to a Nadaraya-Watsonestimator in the latent space, with the same rate of convergence.One issue with this estimator is that it averages over all neighbors of a node, which maybe too large or too small a region depending on the graph model. An alternative consistsin first estimating the “true” distances between latent positions, then injecting these intoa classical Nadaraya-Watson estimator. This enables averaging in regions either smalleror larger than the typical graph neighborhood. We show that this method can achievestandard nonparametric rates even when the graph neighborhood is too large or too small.
ノード回帰とは、グラフ上の他のノードでの観測値に基づいて、あるノードにおけるグラフ・ラベルの値を予測する手法です。本研究では、潜在位置モデル(Latent Position Model)によって生成されるグラフを対象に理論的な検討を行います。このモデルでは、各ノードが潜在位置を持ち、ノード間の接続確率はそれらの潜在位置間の距離に依存します。まず、最も単純な推定量を検討します。これは、すべての近傍ノードのラベルを平均化するものです。潜在位置モデルにおいて、この推定量は潜在空間上のナダラヤ・ワトソン(Nadaraya-Watson)推定量に収束し、その収束速度も同等であることを示します。この推定量の課題の一つは、ノードのすべての近傍を平均化の対象とすることですが、グラフモデルによっては、その領域が広すぎたり狭すぎたりする可能性があります。代替案として、まず潜在位置間の「真の」距離を推定し、それを従来のナダラヤ・ワトソン推定量に組み込む手法があります。これにより、グラフ上の典型的な近傍領域よりも狭い範囲や広い範囲での平均化が可能になります。本手法を用いれば、グラフ上の近傍領域が広すぎる場合や狭すぎる場合であっても、標準的なノンパラメトリック推定における収束速度を達成できることを示します。
Towards Convexity in Anomaly Detection: A New Formulation of SSLM with Unique Optimal Solutions
Towards Convexity in Anomaly Detection: A New Formulation of SSLM with Unique Optimal Solutions / 異常検知における凸性の実現に向けて:一意な最適解を持つSSLMの新たな定式化
An unsolved issue in widely used methods such as Support Vector Data Description (SVDD) and Small Sphere and Large Margin SVM (SSLM) for anomaly detection is their nonconvexity, which hampers the analysis of optimal solutions in a manner similar to SVMs and limits their applicability in large-scale scenarios. In this paper, we introduce a novel convex SSLM formulation which has been demonstrated to revert to a convex quadratic programming problem for hyperparameter values of interest. Leveraging the convexity of our method, we derive numerous results that are unattainable with traditional nonconvex approaches. We conduct a thorough analysis of how hyperparameters influence the optimal solution, pointing out scenarios where optimal solutions can be trivially found and identifying instances of illposedness. Most notably, we establish connections between our method and traditional approaches, providing a clear determination of when the optimal solution is unique-a task unachievable with traditional nonconvex methods. We also derive the $\nu$-property to elucidate the interactions between hyperparameters and the fractions of support vectors and margin errors in both positive and negative classes.
異常検知に広く用いられているSVDD(Support Vector Data Description)やSSLM(Small Sphere and Large Margin SVM)といった手法には、非凸性という未解決の課題があります。この非凸性は、SVMの場合と同様に最適解の解析を困難にし、大規模なシナリオへの適用を制限する要因となっています。本論文では、特定のハイパーパラメータ値において凸二次計画問題に帰着することが示されている、新しい凸SSLMの定式化を提案します。提案手法の凸性を活用することで、従来の非凸なアプローチでは得られなかった数多くの知見を導き出します。ハイパーパラメータが最適解に与える影響について詳細な解析を行い、最適解が自明に求められるシナリオや、問題が適切に定義されない(ill-posedな)ケースを特定します。特に重要な点として、提案手法と従来手法との関連性を明らかにし、最適解が一意に定まる条件を明確に示します。これは、従来の非凸な手法では達成できなかったことです。さらに、$\nu$特性を導出することで、ハイパーパラメータと、正・負の各クラスにおけるサポートベクターやマージン誤差の割合との相互関係を明らかにします。
Transfer Conformal Predictive Inference for Regression
Transfer Conformal Predictive Inference for Regression / 回帰のための転移共形予測推論
Conformal prediction, a powerful framework for constructing prediction intervals for response variables using any regression function estimators, often faces the challenge of producing overly broad intervals with limited target data. In this paper, we study the transfer learning problem in conformal prediction, aiming to improve the precision of the prediction interval of the target data with insufficient data by leveraging related auxiliary source datasets. Allowing for the potential non-exchangeability between source and target datasets, we propose two transfer conformal prediction algorithms designed for scenarios where knowledge of informative source data is either present or absent. Our approach uses conditional Kullback-Leibler divergence to effectively identify relevant source datasets for transfer. A comprehensive theoretical analysis of the non-asymptotic properties of the proposed algorithms is provided, including lower and upper bounds, and the prediction interval width. These results illustrate the potential to achieve more efficient, narrower intervals without compromising coverage accuracy. Empirical results from extensive simulations and real-world data confirm the efficacy of our methods, demonstrating significant improvements in prediction interval precision by leveraging source data, achieving narrower intervals while maintaining desired coverage levels.
任意の回帰関数推定量を用いて応答変数の予測区間を構築する強力な枠組みであるコンフォーマル予測(Conformal prediction)は、対象となるデータが限られている場合、予測区間が過度に広くなってしまうという課題に直面することがよくあります。本論文では、データが不足しているターゲットデータに対する予測区間の精度を向上させることを目指し、関連する補助的なソースデータセットを活用するコンフォーマル予測における転移学習の問題を検討します。ソースデータセットとターゲットデータセットの間に交換可能性が成り立たない可能性を考慮し、有益なソースデータに関する知識の有無に応じた2つの転移コンフォーマル予測アルゴリズムを提案します。我々のアプローチでは、条件付きカルバック・ライブラー・ダイバージェンスを用いて、転移に適したソースデータセットを効果的に特定します。提案アルゴリズムの非漸近的な性質(下限・上限や予測区間の幅など)について、包括的な理論的解析を行います。これらの結果は、カバレッジ精度を損なうことなく、より効率的で幅の狭い予測区間を実現できる可能性を示しています。広範なシミュレーションおよび実データを用いた実験結果により、本手法の有効性が裏付けられました。具体的には、ソースデータを活用することで予測区間の精度が大幅に向上し、所望のカバレッジ水準を維持しつつ、より幅の狭い区間を実現できることが示されました。
Generalized Resubstitution for Regression Error Estimation
Generalized Resubstitution for Regression Error Estimation / 回帰誤差推定のための一般化再代入法
We propose generalized resubstitution error estimators for regression. Each error estimator in this class corresponds to a choice of an empirical probability measure and a loss function. The standard empirical probability measure and the quadratic loss lead to the standard sum of squares error estimator. Other choices of empirical probability measure lead to more general estimators with superior bias and variance properties. We prove that these error estimators are consistent under broad assumptions. In addition, procedures for choosing the empirical measure based on the method of moments and maximum pseudo-likelihood are proposed and investigated. Detailed experimental results using polynomial regression demonstrate empirically the superior finite-sample bias and variance properties of the proposed estimators. The R code for the experiments is provided.
本論文では、回帰分析のための一般化された再代入誤差推定量を提案します。このクラスに属する各誤差推定量は、経験確率測度と損失関数の選択に対応しています。標準的な経験確率測度と二乗損失を用いると、通常の「二乗和誤差」推定量が得られます。一方、経験確率測度を適切に選択することで、バイアスや分散の特性においてより優れた、より一般的な推定量を得ることが可能です。我々は、広範な仮定の下でこれらの誤差推定量が一致性を持つことを証明します。さらに、モーメント法や擬似最尤法に基づいた経験測度の選択手法を提案・検討します。多項式回帰を用いた詳細な実験結果により、提案する推定量が有限サンプルにおけるバイアスおよび分散の特性において優れていることが実証されています。実験に使用したRコードも提供します。
A Natural Primal-Dual Hybrid Gradient Method for Adversarial Neural Network Training on Solving Partial Differential Equation
A Natural Primal-Dual Hybrid Gradient Method for Adversarial Neural Network Training on Solving Partial Differential Equation / 偏微分方程式を解くための敵対的ニューラルネットワーク学習に向けた、自然な主双対ハイブリッド勾配法
We propose a scalable preconditioned primal-dual hybrid gradient algorithm for solving partial differential equations (PDEs). We multiply the PDE with a dual test function to obtain an inf-sup problem whose loss functional involves lower-order differential operators. The Primal-Dual Hybrid Gradient (PDHG) algorithm is then leveraged for this saddle point problem. By introducing suitable precondition operators to the proximal steps in the PDHG algorithm, we obtain an alternative natural gradient ascent-descent optimization scheme for updating the neural network parameters. We apply the Krylov subspace method (MINRES) to evaluate the natural gradients efficiently. Such treatment readily handles the inversion of precondition matrices via matrix-vector multiplication. An a posteriori convergence analysis is established for the time-continuous version of the proposed algorithm for general linear PDEs. By incorporating appropriate boundary loss terms, we further obtain a refined a priori convergence result for elliptic equations in divergence form. The algorithm is tested on various types of PDEs with dimensions ranging from $1$ to $50$, including linear and nonlinear elliptic equations, reaction-diffusion equations, and Monge-Ampere equations stemming from the $L^2$ optimal transport problems. We compare the performance of the proposed method with several commonly used deep learning algorithms such as physics-informed neural networks (PINNs), the DeepRitz method and weak adversarial networks (WANs) using either the Adam or the L-BFGS optimizer. The numerical results suggest that the proposed method performs efficiently and robustly and converges more stably with higher accuracy.
本論文では、偏微分方程式(PDE)を解くための、スケーラブルな前処理付き主双対ハイブリッド勾配(PDHG)アルゴリズムを提案します。PDEに双対テスト関数を乗じることでinf-sup問題(最小最大問題)を導き出し、その損失汎関数に低階の微分作用素が含まれるように定式化します。そして、この鞍点問題に対してPDHGアルゴリズムを適用します。PDHGアルゴリズムの近接作用素(proximal step)に適切な前処理作用素を導入することで、ニューラルネットワークのパラメータを更新するための、自然勾配を用いた新たな勾配上昇・下降最適化スキームが得られます。自然勾配を効率的に評価するために、クリロフ部分空間法(MINRES法)を適用します。この手法により、行列とベクトルの積を計算するだけで前処理行列の逆演算を容易に扱うことができます。一般的な線形PDEに対する提案アルゴリズムの時間連続版について、事後的な収束解析を確立します。さらに、適切な境界損失項を組み込むことで、発散形楕円型方程式に対する洗練された事前収束結果も導き出します。提案アルゴリズムは、線形および非線形の楕円型方程式、反応拡散方程式、$L^2$最適輸送問題に由来するモンジュ・アンペール方程式など、次元数1から50に及ぶ様々なタイプのPDEに対して検証されます。提案手法の性能を、PINN(Physics-Informed Neural Networks)、DeepRitz法、WAN(Weak Adversarial Networks)といった、AdamやL-BFGSオプティマイザを用いる一般的な深層学習アルゴリズムと比較します。数値実験の結果、提案手法は効率的かつ堅牢に動作し、より高い精度で安定して収束することが示唆されました。
Cheap Bootstrap for Fast Uncertainty Quantification of Stochastic Gradient Descent
Cheap Bootstrap for Fast Uncertainty Quantification of Stochastic Gradient Descent / 確率的勾配降下法の高速な不確実性定量化のための「安価な(Cheap)」ブートストラップ法
Stochastic gradient descent (SGD) or stochastic approximation has been widely used in model training and stochastic optimization. While there is a huge literature on analyzing its convergence, inference on the obtained solutions from SGD has only been recently studied, yet it is important due to the growing need for uncertainty quantification. We investigate two computationally cheap resampling-based methods to construct confidence intervals for SGD solutions. One uses multiple, but few, SGDs in parallel via resampling with replacement from the data, and another operates this in an online fashion. Our methods can be regarded as enhancements of established bootstrap schemes to substantially reduce the computation effort in terms of resampling requirements, while bypassing the intricate mixing conditions in existing batching methods. We achieve these via a recent so-called cheap bootstrap idea and refinement of a Berry-Esseen-type bound for SGD.
確率的勾配降下法(SGD)や確率近似法は、モデルの学習や確率的最適化において広く利用されています。SGD(確率的勾配降下法)の収束性に関する研究は膨大に存在しますが、SGDによって得られた解に対する統計的推論の研究はまだ始まったばかりです。しかし、不確実性の定量化に対するニーズの高まりを背景に、この推論は重要な課題となっています。本研究では、SGDの解に対する信頼区間を構築するための、計算コストの低い2つのリサンプリング手法を検討します。一つは、データからの復元抽出(リサンプリング)を用いて少数のSGDを並列に実行する手法であり、もう一つはこれをオンライン形式で実行する手法です。これらの手法は、既存のバッチ処理手法に伴う複雑な混合条件(mixing conditions)を回避しつつ、リサンプリングに伴う計算負荷を大幅に低減させるものであり、確立されたブートストラップ法を改良したものと位置づけられます。こうした手法の実現にあたっては、近年提案された「安価なブートストラップ(cheap bootstrap)」という概念と、SGDに対するベリー・エッセン(Berry-Esseen)型評価の精緻化を活用しています。
Transfer Learning via Regularized Random-effects Linear Discriminant Analysis
Transfer Learning via Regularized Random-effects Linear Discriminant Analysis / 正則化ランダム効果線形判別分析による転移学習
Linear discriminant analysis is a widely used method for classification. However, the high dimensionality of predictors combined with small sample sizes often results in large classification errors. To address this challenge, it is crucial to leverage data from related source models to enhance the classification performance of a target model. This paper proposes a transfer learning approach via regularized random-effects linear discriminant analysis, where the discriminant direction is estimated as a weighted combination of ridge estimates obtained from both the target and source models. Multiple strategies for determining these weights are introduced and evaluated, including one that minimizes the estimation risk of the discriminant vector and another that minimizes the classification error. Utilizing results from random matrix theory, we explicitly derive the asymptotic values of these weights and the associated classification error rates in the high-dimensional setting, where the aspect ratio $\gamma := p/n$ as $p, n\rightarrow \infty$, with $p$ representing the predictor dimension and $n$ the sample size. Extensive numerical studies, including simulations, the analysis of the proteomics-based cardiovascular disease risk classification and the lipid traits classification problem with genotype data, demonstrate the effectiveness of the proposed approach.
線形判別分析は、分類において広く用いられている手法です。しかし、予測変数の次元数が高い一方でサンプルサイズが小さい場合、大きな分類誤差が生じることがよくあります。この課題に対処するためには、関連するソースモデルのデータを活用し、ターゲットモデルの分類性能を向上させることが重要です。本論文では、正則化ランダム効果線形判別分析を用いた転移学習アプローチを提案します。この手法では、ターゲットモデルとソースモデルの双方から得られるリッジ推定量の加重結合として、判別方向を推定します。これらの重みを決定するための複数の戦略(判別ベクトルの推定リスクを最小化するものや、分類誤差を最小化するものなど)が導入・評価されています。ランダム行列理論の知見を用い、予測変数の次元数$p$とサンプルサイズ$n$が共に無限大に発散する際のアスペクト比$\gamma := p/n$を考慮した高次元設定において、これらの重みの漸近値およびそれに関連する分類誤り率を明示的に導出します。シミュレーションに加え、プロテオミクスに基づく心血管疾患リスクの分類や、遺伝子型データを用いた脂質形質の分類問題の解析を含む広範な数値実験により、提案手法の有効性が実証されています。
Semi-supervised learning for linear extremile regression
Semi-supervised learning for linear extremile regression / 線形エクストリマイル回帰のための半教師あり学習
Extremile regression, as a least squares analog of quantile regression, is potentially a useful tool for modeling and understanding the extreme tails of a distribution. However, existing extremile regression methods, as nonparametric approaches, may face challenges in high-dimensional settings due to data sparsity, computational inefficiency, and the risk of overfitting. While linear regression, particularly in high-dimensional settings, serves as the foundation for many other statistical and machine learning models due to its simplicity, interpretability, and relatively easy implementation, this paper introduces a novel definition of linear extremile regression along with an accompanying estimation methodology. The regression coefficient estimators of this method achieve root n consistency, which nonparametric extremile regression may not provide. In particular, while semi-supervised learning can leverage unlabeled data to make more accurate predictions and avoid overfitting to small labeled datasets in high-dimensional spaces, we propose a semi-supervised learning to enhance estimation efficiency, even when the specified linear extremile regression model may be misspecified. Both simulation studies and real data analyses demonstrate the finite sample performance of our proposed methods.
極値回帰(extremile regression)は、分位点回帰の最小二乗法版とも言える手法であり、分布の極端な裾野(テール)をモデル化し理解するための有用なツールとなり得ます。しかし、既存の極値回帰手法はノンパラメトリックなアプローチであるため、高次元設定においては、データの疎性、計算効率の低さ、過学習のリスクといった課題に直面する可能性があります。線形回帰は、その単純さ、解釈の容易さ、実装の簡便さから、特に高次元設定において他の多くの統計モデルや機械学習モデルの基礎となっていますが、本論文では、線形極値回帰の新たな定義と、それに伴う推定手法を提案します。本手法による回帰係数の推定量は$\sqrt{n}$整合性を達成しますが、これはノンパラメトリックな極値回帰では必ずしも得られない性質です。特に、半教師あり学習は、ラベルなしデータを活用することで、高次元空間においてより正確な予測を行い、少数のラベル付きデータセットへの過学習を回避することが可能ですが、本論文では、指定された線形極値回帰モデルが誤設定(ミススペシフィケーション)されている場合であっても推定効率を向上させるための半教師あり学習手法を提案します。シミュレーション研究および実データ解析の双方において、提案手法の有限標本における性能が実証されています。
Deep Nonparametric Conditional Independence Tests for Images
Deep Nonparametric Conditional Independence Tests for Images / 画像に対する深層ノンパラメトリック条件付き独立性検定
Conditional independence tests (CITs) test for conditional dependence between random variables given a vector of conditioning or confounder variables. As existing CITs are limited in their applicability to complex, high-dimensional variables such as images, we introduce deep nonparametric CITs (DNCITs). The DNCITs combine embedding maps, which extract feature representations of high-dimensional variables, with nonparametric CITs applicable to these feature representations. For the embedding maps, we derive general properties on their parameter estimators to obtain valid DNCITs and show that these properties include embedding maps learned through (conditional) unsupervised or transfer learning. For the nonparametric CITs, appropriate tests are selected and adapted to be applicable to feature representations. Through simulations, we investigate the performance of the DNCITs for different embedding maps and nonparametric CITs under varying confounder dimensions and confounder relationships. We apply the DNCITs to brain MRI scans and behavioral traits, given confounders, of healthy individuals from the UK Biobank, confirming null results from a number of ambiguous personality neuroscience studies, now with a larger data set and with our more powerful tests. In addition, in a confounder control study, we apply the DNCITs to brain MRI scans and a confounder set to test for sufficient confounder control. We provide an R package implementing the proposed DNCITs.
条件付き独立性検定(CIT)は、条件付け変数または交絡変数のベクトルが与えられた条件下での、確率変数間の条件付き依存性を検定するものです。既存の条件付き独立性検定(CIT)は、画像のような複雑かつ高次元の変数への適用が限られているため、我々は「深層ノンパラメトリックCIT(DNCIT)」を提案します。DNCITは、高次元変数の特徴表現を抽出する埋め込み写像と、それらの特徴表現に適用可能なノンパラメトリックCITを組み合わせたものです。埋め込み写像に関しては、妥当なDNCITを実現するためのパラメータ推定量の一般的性質を導出し、その性質が(条件付き)教師なし学習や転移学習によって学習された埋め込み写像にも当てはまることを示します。ノンパラメトリックCITについては、適切な検定手法を選定し、特徴表現に適用できるよう調整を行います。シミュレーションを通じて、様々な交絡因子の次元や関係性の下で、異なる埋め込み写像やノンパラメトリックCITを用いた場合のDNCITの性能を検証します。さらに、UKバイオバンクの健常者のデータを用い、脳MRI画像と行動特性(交絡因子を考慮)に対してDNCITを適用しました。これにより、従来の人格神経科学研究において結果が不明瞭であったいくつかの点について、より大規模なデータセットと我々の強力な検定手法を用いて、帰無仮説が棄却されないこと(有意差がないこと)を確認しました。加えて、交絡因子制御に関する研究として、脳MRI画像と交絡因子セットに対してDNCITを適用し、交絡因子の制御が十分に行われているかを検証しました。提案手法であるDNCITを実装したRパッケージも提供します。
Kernel Mean Embedding Deviation Subspace for Unsupervised Learning with Heterogeneous Data
Kernel Mean Embedding Deviation Subspace for Unsupervised Learning with Heterogeneous Data / 異種データを用いた教師なし学習のためのカーネル平均埋め込み偏差部分空間
This paper proposes a method for dimension reduction that preserves information in unsupervised learning with high-dimensional heterogeneous data, specifically targeting change point detection and clustering analysis. Our main strategy is to apply a Corrected Kernel Principal Component Analysis (CKPCA) method to construct the so-called kernel mean embedding deviation subspace. The approach efficiently identifies distributional changes in these dimension reduction subspaces for unsupervised dimension reduction.For change point detection, we demonstrate that the locations and number of change points in the dimension-reduced subspaces are identical to those in the original data.Furthermore, we extend this approach to clustering by embedding the original data into nonlinear lower-dimensional spaces, providing enhanced capabilities for clustering analysis.Additionally, we explain the necessity of using CKPCA, as the classical KPCA fails to identify the kernel mean embedding deviation subspace in these problems.Numerical studies on synthetic and real data sets suggest that the dimension reduction versions of existing methods for change point detection and clustering significantly improve the performance of current approaches in finite sample scenarios.
本論文では、高次元かつ異質なデータを用いた教師なし学習において、情報を保持しつつ次元を削減する手法を提案します。特に、変化点検知とクラスタリング分析を対象としています。主な戦略は、修正カーネル主成分分析(CKPCA)を適用し、いわゆる「カーネル平均埋め込みの偏差部分空間」を構築することです。このアプローチにより、教師なし次元削減において、次元削減された部分空間内での分布の変化を効率的に特定できます。変化点検知に関しては、次元削減された部分空間における変化点の位置と数が、元のデータにおけるそれらと一致することを示します。さらに、元のデータを非線形な低次元空間に埋め込むことでこの手法をクラスタリングに拡張し、クラスタリング分析の能力を向上させます。加えて、従来のKPCAではこれらの問題においてカーネル平均埋め込みの偏差部分空間を特定できないため、CKPCAを使用する必要性についても説明します。合成データおよび実データセットを用いた数値実験の結果、変化点検知やクラスタリングのための既存手法を次元削減版へと拡張することで、有限サンプル環境における現在の手法の性能が大幅に向上することが示唆されました。
Exogenous Randomness Empowering Random Forests
Exogenous Randomness Empowering Random Forests / ランダムフォレストを強化する外生的ランダム性
We offer theoretical and empirical insights into the impact of exogenous randomness on the effectiveness of random forests with tree-building rules independent of training data. We formally introduce the concept of exogenous randomness which can come from feature subsampling or tie-breaking in tree-building processes. We develop non-asymptotic expansions for the mean squared error (MSE) for both individual trees and forests and establish sufficient and necessary conditions for their consistency. In the special example of the linear regression model with independent features, our MSE expansions are more explicit, providing more understanding of the random forests’ mechanisms. It also allows us to derive an upper bound on the MSE with explicit consistency rates for trees and forests. Guided by our theoretical findings, we conduct simulations to further explore how exogenous randomness enhances random forests performance. Our findings unveil that feature subsampling reduces both the bias and variance of random forests compared to individual trees, serving as an adaptive mechanism to balance bias and variance. Furthermore, our results reveal an intriguing phenomenon: the presence of noise features can act as a “blessing” in enhancing the performance of random forests thanks to feature subsampling.
本論文では、学習データに依存しない木構築ルールを持つランダムフォレストの有効性に対し、外生的なランダム性が及ぼす影響について、理論的および実証的な知見を提示します。まず、特徴量のサブサンプリングや木構築プロセスにおける同順位の解消(タイブレーク)に由来しうる「外生的なランダム性」という概念を定式化します。次に、個々の木およびフォレストの双方について、平均二乗誤差(MSE)の非漸近展開を導出し、それらの整合性(コンシステンシー)に関する必要十分条件を確立します。特徴量が独立である線形回帰モデルという特殊な例においては、MSEの展開がより明示的な形となり、ランダムフォレストのメカニズムへの理解を深めることができます。また、これにより、木およびフォレストのMSEの上界を、明示的な整合性レートと共に導出することが可能になります。理論的な知見に基づき、外生的なランダム性がどのようにランダムフォレストの性能を向上させるかをさらに探るためのシミュレーションを行います。その結果、特徴量のサブサンプリングは、個々の木と比較してランダムフォレストのバイアスと分散の双方を低減させ、バイアスと分散のバランスをとる適応的なメカニズムとして機能することが明らかになりました。さらに、興味深い現象も明らかになりました。それは、特徴量のサブサンプリングのおかげで、ノイズとなる特徴量の存在が、ランダムフォレストの性能向上において「恩恵(blessing)」として作用しうるということです。
High-dimensional Parameter Transfer With Fused-Regularizer
High-dimensional Parameter Transfer With Fused-Regularizer / 融合正則化を用いた高次元パラメータ転移
Parameter transfer aims to improve parameter estimation accuracy by leveraging knowledge from related sources. This paper studies the parameter transfer problem from heterogeneous sources for high-dimensional M-estimators. Specifically, we propose a novel one-step estimator with a fused-regularizer and a target-data-oriented constraint, which can robustly capture parameter knowledge from source data in the presence of different types of data distribution shifts. Nonasymptotic bound is provided for the estimation error of target parameter, showing the proposed estimator could achieve effective parameter transfer under distribution shifts, and is guaranteed to perform no worse than any estimators learned only from the target data. We further show that the proposed estimator can achieve the minimax-optimal rate under much weaker conditions than existing methods. In addition, we extend the method to a distributed setting, requiring just one round of communication with source parameter estimators, while retaining the estimation accuracy of the centralized version. Extensive simulations and real data analysis further verify the effectiveness of the method.
パラメータ転移は、関連するソースからの知識を活用することで、パラメータ推定の精度を向上させることを目的としています。本論文では、高次元M推定量を対象に、異質なソースからのパラメータ転移の問題を研究します。具体的には、我々は、フューズド正則化項とターゲットデータ指向の制約を組み込んだ新しいワンステップ推定量を提案します。この推定量は、様々な種類のデータ分布シフトが存在する状況下でも、ソースデータからパラメータに関する知識を頑健に抽出することが可能です。ターゲットパラメータの推定誤差に対する非漸近的な評価境界を提示し、提案推定量が分布シフト下で効果的なパラメータ転移を実現できること、およびターゲットデータのみから学習されたいかなる推定量よりも性能が劣らないことが保証されることを示す。さらに、提案推定量は既存手法よりもはるかに緩やかな条件下でミニマックス最適レートを達成できることを示す。加えて、本手法を分散型設定へと拡張します。この拡張版では、集中型と同等の推定精度を維持しつつ、ソースパラメータ推定量との通信は1ラウンドのみで済む。広範なシミュレーションおよび実データ解析により、本手法の有効性をさらに実証します。
Spectral Truncation Kernels: Noncommutativity in C*-algebraic Kernel Machines
Spectral Truncation Kernels: Noncommutativity in C*-algebraic Kernel Machines / スペクトル切断カーネル:C*代数カーネルマシンにおける非可換性
A central question in vector- and function-valued learning is how to design kernels that capture both local and non-local interactions while remaining computationally tractable. Existing operator-valued kernels offer only partial answers: separable kernels are efficient but fail to model interactions across the function domain, while commutative kernels capture only pointwise structure. To address this, we propose spectral truncation kernels, a new class of positive definite kernels for vector- and function-valued learning based on spectral truncation and C*-algebra. By allowing noncommutative products in the kernel construction, the proposed kernels induce interactions across the data function domain and fill the gap between existing separable and commutative kernels. In addition, by using the C*-algebraic framework, we reduce the computational cost compared to the existing vector-valued RKHS framework with operator-valued kernels.
ベクトル値および関数値学習における中心的な課題は、計算の実行可能性を維持しつつ、局所的および非局所的な相互作用の双方を捉えるカーネルをいかに設計するかという点にあります。既存の作用素値カーネルは部分的な解決策しか提示していません。すなわち、分離可能カーネルは効率的ですが関数領域全体にわたる相互作用をモデル化できず、一方で可換カーネルは点ごとの構造しか捉えられません。これに対処するため、我々はスペクトル切断とC*代数に基づく、ベクトル値および関数値学習のための新しい正定値カーネルのクラスである「スペクトル切断カーネル」を提案します。カーネルの構成において非可換積を許容することで、提案カーネルはデータ関数領域全体にわたる相互作用を誘起し、既存の分離可能カーネルと可換カーネルの間のギャップを埋めるものとなります。さらに、C*代数の枠組みを用いることで、作用素値カーネルを用いた既存のベクトル値RKHS(再生核ヒルベルト空間)の枠組みと比較して、計算コストを低減させています。
Deconvolution in unlinked linear models
Deconvolution in unlinked linear models / 非連結線形モデルにおけるデコンボリューション
Unlinked regression, in which covariates and responses are observed separately without known correspondence, has recently gained increasing attention.Deconvolution, on the other hand, is a fundamental and challenging problem in nonparametric statistics with the aim of estimating the distribution of a latent random variable $Z$ based on observations contaminated by some additive noise. The complexity of this task is heavily influenced by the smoothness of the noise distribution and often leads to slow estimation rates.In this paper, we combine the recent unlinked linear regression problem with the classical deconvolution framework. Specifically, we study nonparametric deconvolution under the assumption that $Z$ is a linear function of an observable multidimensional covariate. This structural constraint allows us to introduce a nonparametric estimator of the distribution of $Z$ which achieves the parametric rate of convergence in the Wasserstein distance of order 1, where the smoothness of the noise does not affect the rate.Furthermore, we introduce nonparametric estimators for the unconditional density of $Z$ and the conditional density of $Z$ given an observed response. This allows us to study the problem of estimating the value of the latent linear predictor, whose link to the observed response is not accessible. Through several simulations, we illustrate the fast convergence rate of our deconvolution estimator and the performance of the proposed conditional estimators of the latent predictor in different simulation scenarios.
共変量と応答変数が対応関係不明のまま別々に観測される「非連結回帰(unlinked regression)」は、近年ますます注目を集めています。一方、デコンボリューション(逆畳み込み)は、加法的ノイズが混入した観測値に基づいて潜在確率変数$Z$の分布を推定することを目的とする、ノンパラメトリック統計における基本的かつ困難な問題です。このタスクの複雑さはノイズ分布の滑らかさに大きく依存し、しばしば推定レートの低下を招きます。本論文では、近年の非連結線形回帰問題と古典的なデコンボリューションの枠組みを組み合わせます。具体的には、$Z$が観測可能な多次元共変量の線形関数であるという仮定の下で、ノンパラメトリック・デコンボリューションを検討します。この構造的制約により、$Z$の分布に関するノンパラメトリック推定量が導入可能となり、この推定量は1次Wasserstein距離においてパラメトリックな収束レートを達成します(この際、ノイズの滑らかさは収束レートに影響しません)。さらに、我々は$Z$の無条件密度および観測された応答変数が与えられた条件の下での$Z$の条件付き密度のためのノンパラメトリック推定量も導入します。これにより、観測された応答変数との直接的な関連が不明な潜在線形予測子の値を推定する問題を検討することが可能になります。いくつかのシミュレーションを通じて、我々のデコンボリューション推定量の高速な収束レートと、異なるシミュレーション・シナリオにおける潜在予測子の条件付き推定量の性能を実証します。
Statistical Learning Theory for Neural Operators
Statistical Learning Theory for Neural Operators / ニューラルオペレータのための統計的学習理論
We present statistical convergence results for the learning of (possibly) non-linear mappings in infinite-dimensional spaces. Specifically, given a map $G_0:\mathcal X\to\mathcal Y$ between two separable Hilbert spaces, we analyze the problem of recovering $G_0$ from $n\in\mathbb{N}$ noisy input-output pairs $(x_i, y_i)_{i=1}^n$ with $y_i = G_0 (x_i)+\varepsilon_i$; here the $x_i\in\mathcal{X}$ represent randomly drawn “design” points, and the $\varepsilon_i$ are assumed to be either i.i.d. white noise processes or subgaussian random variables in $\mathcal{Y}$. We provide general convergence results for least-squares-type empirical risk minimizers over compact regression classes $\mathbf{G}\subseteq L^{\infty}(\mathcal{X},\mathcal{Y})$, in terms of their approximation properties and metric entropy bounds, which are derived using empirical process techniques. This generalizes classical results from finite-dimensional nonparametric regression to an infinite-dimensional setting. As a concrete application, we study an encoder-decoder based neural operator architecture termed FrameNet. Assuming $G_0$ to be holomorphic, we prove algebraic (in the sample size $n$) convergence rates in this setting, thereby overcoming the curse of dimensionality. To illustrate the wide applicability, as a prototypical example we discuss the learning of the non-linear solution operator to a parametric elliptic partial differential equation.
我々は、無限次元空間における(非線形である可能性のある)写像の学習に関する統計的収束結果を提示します。具体的には、2つの可分ヒルベルト空間間の写像$G_0:\mathcal X\to\mathcal Y$が与えられたとき、$y_i = G_0 (x_i)+\varepsilon_i$という形の$n$個のノイズを含む入出力ペア$(x_i, y_i)_{i=1}^n$から$G_0$を復元する問題を解析します。ここで、$x_i\in\mathcal{X}$はランダムに抽出された「設計点」を表し、$\varepsilon_i$は$\mathcal{Y}$におけるi.i.d.(独立同分布)なホワイトノイズ過程またはサブガウス確率変数のいずれかであると仮定します。我々は、コンパクトな回帰クラス$\mathbf{G}\subseteq L^{\infty}(\mathcal{X},\mathcal{Y})$上の最小二乗型経験リスク最小化法について、経験過程の手法を用いて導出された近似特性およびメトリックエントロピーの評価に基づく、一般的な収束結果を提示します。これは、有限次元のノンパラメトリック回帰における古典的な結果を無限次元の設定へと一般化するものです。具体的な応用として、「FrameNet」と呼ばれるエンコーダ・デコーダ型のニューラルオペレータ・アーキテクチャを検討します。$G_0$が正則(ホロモルフィック)であると仮定することで、この設定において(サンプルサイズ$n$に関する)代数的な収束レートを証明し、それによって「次元の呪い」を克服します。その幅広い適用性を示すため、典型的な例として、パラメータ依存の楕円型偏微分方程式に対する非線形解作用素の学習について論じます。
A Single-Loop Stochastic Proximal Quasi-Newton Method for Large-Scale Nonsmooth Convex Optimization
A Single-Loop Stochastic Proximal Quasi-Newton Method for Large-Scale Nonsmooth Convex Optimization / 大規模非平滑凸最適化のためのシングルループ確率的近接準ニュートン法
We propose a new stochastic proximal quasi-Newton method for minimizing the sum of two convex functions in the particular context that one of the functions is the average of a large number of smooth functions and the other one is nonsmooth. The new method integrates a simple single-loop SVRG (L-SVRG) technique for sampling the gradient and a stochastic limited-memory BFGS (L-BFGS) scheme for approximating the Hessian of the smooth function components. The globally linear convergence rate of the new method is proved under mild assumptions. It is also shown that the new method covers a proximal variant of the L-SVRG as a special case, and it allows for various generalization through the integration with other variance reduction methods. For example, the L-SVRG can be replaced with the SAGA or SEGA in the proposed new method and thus other new stochastic proximal quasi-Newton methods with rigorously guaranteed convergence can be proposed accordingly. Moreover, we meticulously analyze the resulting nonsmooth subproblem at each iteration and leverage a compact representation of the L-BFGS matrix with the storage of some auxiliary matrices. As a result, we propose a very efficient and easily implementable semismooth Newton solver for solving the involved subproblems, whose arithmetic operations per iteration are merely order of O(d), where d denotes the dimensionality of the problem. With this efficient inner solver, the new method performs well and its numerical efficiency is validated through extensive experiments on a regularized logistic regression problem.
本研究では、2つの凸関数の和を最小化する問題を対象に、新しい確率的近接準ニュートン法を提案します。ここで想定する特定の状況は、一方の関数が多数の滑らかな関数の平均であり、もう一方が非滑らかな関数であるというものです。この新しい手法は、勾配をサンプリングするための単純なシングルループSVRG(L-SVRG)技術と、滑らかな関数成分のヘッセ行列を近似するための確率的限定メモリBFGS(L-BFGS)スキームを統合したものです。緩やかな仮定の下で、この手法が大域的な線形収束率を持つことが証明されています。また、この手法はL-SVRGの近接型(proximal variant)を特殊なケースとして包含しており、他の分散低減手法との統合を通じて様々な形で一般化できることも示されています。例えば、提案手法におけるL-SVRGをSAGAやSEGAに置き換えることで、収束性が厳密に保証された新たな確率的近接準ニュートン法を同様に提案することが可能です。さらに、各反復で生じる非滑らかな部分問題を詳細に解析し、補助的な行列を保持することでL-BFGS行列のコンパクトな表現を活用しています。その結果、部分問題を解くための非常に効率的かつ実装が容易な半滑らかニュートン法ソルバーを提案しました。このソルバーの反復あたりの演算量は、問題の次元を$d$とすると$O(d)$のオーダーに留まります。このような効率的な内部ソルバーを備えた提案手法は良好な性能を示し、正則化ロジスティック回帰問題を用いた広範な実験を通じて、その数値的な効率性が実証されています。
Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural Networks
Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural Networks / マージンの諸相:斉次ニューラルネットワークにおける最急降下法の暗黙的バイアス
We study the implicit bias of the general family of steepest descent algorithms with infinitesimal learning rate in deep homogeneous neural networks. We show that: (a) an algorithm-dependent geometric margin starts increasing once the networks reach perfect training accuracy, and (b) any limit point of the training trajectory corresponds to a KKT point of the corresponding margin-maximization problem. We experimentally zoom into the trajectories of neural networks optimized with various steepest descent algorithms, highlighting connections to the implicit bias of popular adaptive methods (Adam and Shampoo).
本研究では、深層同次ニューラルネットワーク(deep homogeneous neural networks)において、無限小の学習率を持つ最急降下法アルゴリズムの一般的なファミリーが示す「暗黙のバイアス(implicit bias)」について考察します。具体的には、以下のことを示します。(a)ネットワークが完全な学習精度に達すると、アルゴリズムに依存する幾何学的マージンが増加し始めること、および(b)学習軌道の任意の極限点が、対応するマージン最大化問題のKKT点に対応すること。また、様々な最急降下法で最適化されたニューラルネットワークの軌道を実験的に詳細に分析し、広く使われている適応型手法(AdamやShampoo)の暗黙のバイアスとの関連性を明らかにします。
A Unified Approach to Analysis and Design of Denoising Markov Models
A Unified Approach to Analysis and Design of Denoising Markov Models / ノイズ除去マルコフモデルの解析と設計への統一的アプローチ
Probabilistic generative models based on measure transport, such as diffusion and flow-based models, are often formulated in the language of Markovian stochastic dynamics, where the choice of the underlying process impacts both algorithmic design choices and theoretical analysis. In this paper, we aim to establish a rigorous mathematical foundation for denoising Markov models, a broad class of generative models that postulate a forward process transitioning from the target distribution to a simple, easy-to-sample distribution, alongside a backward process particularly constructed to enable efficient sampling in the reverse direction. Leveraging deep connections with nonequilibrium statistical mechanics and generalized Doob’s $h$-transform, we propose a minimal set of assumptions that ensure: (1) explicit construction of the backward generator, (2) a unified variational objective directly minimizing the measure transport discrepancy, and (3) adaptations of the classical score-matching approach across diverse dynamics. Our framework unifies existing formulations of continuous and discrete diffusion models, identifies the most general form of denoising Markov models under certain regularity assumptions on forward generators, and provides a systematic recipe for designing denoising Markov models driven by arbitrary Lévy-type processes. We illustrate the versatility and practical effectiveness of our approach through novel denoising Markov models employing geometric Brownian motion and jump processes as forward dynamics, highlighting the framework’s potential flexibility and capability in modeling complex distributions.
拡散モデルやフローベースモデルなど、測度輸送(measure transport)に基づく確率的生成モデルは、しばしばマルコフ的確率力学の枠組みで定式化されます。この際、基礎となるプロセスの選択は、アルゴリズムの設計上の選択と理論的解析の双方に影響を及ぼします。本論文では、「ノイズ除去マルコフモデル(denoising Markov models)」の厳密な数学的基礎の確立を目指します。これは、ターゲット分布からサンプリングが容易な単純な分布へと移行する「順方向過程」と、その逆方向への効率的なサンプリングを可能にするよう特別に構築された「逆方向過程」を想定する、広範な生成モデルのクラスです。非平衡統計力学や一般化されたDoobの$h$-変換との深い関連性を活用し、我々は以下の3点を保証する最小限の仮定を提案します:(1)逆方向生成作用素の明示的な構成、(2)測度輸送の不一致(measure transport discrepancy)を直接最小化する統一的な変分目的関数、および(3)多様な力学系への古典的なスコアマッチング手法の適用。我々の枠組みは、連続および離散拡散モデルの既存の定式化を統合し、順方向生成作用素に関する特定の正則性仮定の下でノイズ除去マルコフモデルの最も一般的な形式を特定するとともに、任意のレヴィ型過程によって駆動されるノイズ除去マルコフモデルを設計するための体系的な手法を提供します。幾何ブラウン運動やジャンプ過程を順方向の力学系として採用した新規のノイズ除去マルコフモデルを通じて、本手法の汎用性と実用的な有効性を示し、複雑な分布をモデリングする上での本枠組みの潜在的な柔軟性と能力を強調します。
Nested Subspace Learning with Flags
Nested Subspace Learning with Flags / フラグ(旗)を用いた入れ子状部分空間学習
Many machine learning methods look for low-dimensional representations of the data. The underlying subspace can be estimated by first choosing a dimension q and then optimizing a certain objective function over the space of q-dimensional subspaces (the Grassmannian). Trying different q generally yields non-nested subspaces, which raises an important issue of consistency between the data representations. In this paper, we propose a simple and easily implementable principle to enforce nestedness in subspace learning methods. It consists in lifting Grassmannian optimization criteria to flag manifolds (the space of nested subspaces of increasing dimension) via nested projectors. We apply the flag trick to several classical machine learning methods and show that it successfully addresses the nestedness issue.
多くの機械学習手法は、データの低次元表現を探索します。基底となる部分空間は、まず次元数$q$を選択し、次いで$q$次元部分空間の空間(グラスマン多様体)上で特定の目的関数を最適化することによって推定できます。異なる$q$を試すと、一般に包含関係にない(入れ子構造を持たない)部分空間が得られることになり、データ表現間の整合性という重要な問題が生じます。本論文では、部分空間学習手法において包含関係(入れ子構造)を強制するための、単純かつ実装が容易な原理を提案します。この原理は、入れ子状の射影行列(nested projectors)を介して、グラスマン多様体上の最適化基準をフラッグ多様体(次元が増大する入れ子状の部分空間の空間)へと拡張するものです。我々は、この「フラッグ・トリック(flag trick)」をいくつかの古典的な機械学習手法に適用し、それが包含関係の問題をうまく解決することを示します。
FLAGG: Flexible Autoregressive Graph Generation
FLAGG: Flexible Autoregressive Graph Generation / FLAGG:柔軟な自己回帰グラフ生成
The Deep Graph Generation’s panorama spans two extremes: one-shot and sequential models. The former generates nodes and edges jointly, while the latter samples them autoregressively. Each method performs better in different graph domains depending on size and topology, but neither is applicable to all graph categories. For instance, one-shot methods struggle with generating large graphs, while sequential methods underperform on smaller graphs. A possible way to overcome these limitations is to flexibly combine the two methods in a unique system. In this work, we propose the FLAGG (Flexible Autoregressive Graph Generation) framework, which sequentially generates portions of graphs with one-shot models. FLAGG can apply any one-shot model to make it autoregressive, allowing flexibility in choosing the sequential policy. This policy is specified through a stochastic node removal process, which an Insertion Model learns to reverse. We evaluate FLAGG with the DiGress one-shot model on several data sets of different graph sizes and domains. We show that the approach outperforms both one-shot and autoregressive baselines in terms of sampling quality.
深層グラフ生成の分野は、ワンショットモデルと逐次生成モデルという両極端な手法によって構成されています。前者はノードとエッジを同時に生成するのに対し、後者はそれらを自己回帰的にサンプリングします。各手法はグラフのサイズやトポロジーに応じて異なる領域で優れた性能を発揮しますが、いずれもすべてのグラフカテゴリに適用できるわけではありません。例えば、ワンショット手法は大規模なグラフの生成に苦労する一方、逐次生成手法は小規模なグラフでは性能が低下します。これらの限界を克服する一つの方法は、これら二つの手法を単一のシステム内で柔軟に組み合わせることです。本研究では、ワンショットモデルを用いてグラフの一部を逐次的に生成する、FLAGG(Flexible Autoregressive Graph Generation)フレームワークを提案します。FLAGGは任意のワンショットモデルを自己回帰的に適用可能であり、逐次生成の方針(ポリシー)を柔軟に選択できます。この方針は確率的なノード削除プロセスによって規定され、挿入モデル(Insertion Model)がその逆プロセスを学習します。我々は、異なるグラフサイズや領域を持つ複数のデータセットにおいて、ワンショットモデルであるDiGressとFLAGGを組み合わせて評価を行いました。その結果、サンプリング品質の面で、本手法がワンショットおよび自己回帰ベースラインの双方を上回る性能を示すことを明らかにしました。
A Two-Timescale Primal-Dual Framework for Reinforcement Learning via Online Dual Variable Guidance
A Two-Timescale Primal-Dual Framework for Reinforcement Learning via Online Dual Variable Guidance / オンライン双対変数ガイダンスによる強化学習のための2時間スケール主双対フレームワーク
We study reinforcement learning by combining recent advances in regularized linear programming formulations with the classical theory of stochastic approximation. Motivated by the challenge of designing algorithms that leverage off-policy data while maintaining on-policy exploration, we propose PGDA-RL, a novel primal-dual projected gradient descent-ascent algorithm for solving regularized Markov decision processes (MDPs). PGDA-RL integrates experience replay-based gradient estimation with a two-timescale decomposition of the underlying nested optimization problem. The algorithm operates asynchronously, interacts with the environment through a single trajectory of correlated data, and updates its policy online in response to the dual variable associated with the occupancy measure of the underlying MDP. We prove that PGDA-RL converges almost surely to the optimal value function and policy of the regularized MDP. Our convergence analysis relies on tools from stochastic approximation theory and holds under weaker assumptions than those required by existing primal-dual RL approaches, notably removing the need for a simulator or a fixed behavioral policy. Under a strengthened ergodicity assumption on the underlying Markov chain, we establish a last-iterate finite-time guarantee with $\widetilde{\mathcal O}(k^{-2/3})$ mean-square convergence, aligning with the best-known rates for two-timescale stochastic approximation methods under Markovian sampling and biased gradient estimates.
本研究では、正則化線形計画法における近年の進展と、確率近似の古典的理論を組み合わせることで、強化学習について検討します。オンポリシーでの探索を維持しつつオフポリシーデータを活用するアルゴリズムを設計するという課題を背景に、正則化マルコフ決定過程(MDP)を解くための新しい主双対射影勾配降下・上昇アルゴリズムであるPGDA-RLを提案します。PGDA-RLは、経験再生(experience replay)に基づく勾配推定と、基底となる入れ子状最適化問題の2つの時間スケールへの分解を統合した手法です。このアルゴリズムは非同期的に動作し、相関のあるデータの単一の軌跡を通じて環境と相互作用し、対象となるMDP(マルコフ決定過程)の占有測度に関連する双対変数に応じてオンラインで方策を更新します。我々は、PGDA-RLが正則化されたMDPの最適価値関数および最適方策に概収束することを証明します。この収束解析は確率近似理論の手法に基づいており、既存の主双対型強化学習(RL)アプローチよりも緩やかな仮定の下で成立します。特に、シミュレータや固定された行動方策を必要としない点が特徴です。対象となるマルコフ連鎖に関するより強いエルゴード性仮定の下で、我々は最終反復(last-iterate)における有限時間での保証を確立しました。具体的には、$\widetilde{\mathcal O}(k^{-2/3})$の平均二乗収束を示しており、これはマルコフ的サンプリングおよびバイアスのある勾配推定を伴う2タイムスケール確率近似法において知られている最良の収束レートと整合するものです。
Convergence of Decentralized Stochastic Subgradient-based Methods for Nonsmooth Nonconvex Optimization
Convergence of Decentralized Stochastic Subgradient-based Methods for Nonsmooth Nonconvex Optimization / 非平滑非凸最適化のための分散型確率的劣勾配法の収束性
In this paper, we focus on the decentralized stochastic subgradient-based methods in minimizing nonsmooth nonconvex functions without Clarke regularity, especially in the decentralized training of nonsmooth neural networks. We propose a general framework thatunifies various decentralized subgradient-based methods, such as decentralized stochastic subgradient descent (DSGD), DSGD with gradient-tracking technique (DSGD-T), and DSGD with momentum (DSGD-M). To establish the convergence properties of our proposed framework, we relate the discrete iterates to the trajectories of a continuous-timedifferential inclusion, which is assumed to have a coercive Lyapunov function with a stable set A. We prove the asymptotic convergence of the iterates to the stable set A with sufficiently small and diminishing step-sizes. These results provide first convergence guarantees for some well-recognized of decentralized stochastic subgradient-based methodswithout Clarke regularity of the objective function. Preliminary numerical experiments demonstrate that our proposed framework yields highly efficient decentralized stochastic subgradient-based methods with convergence guarantees in the training of nonsmoothneural networks.
本論文では、Clarke正則性を仮定しない非平滑かつ非凸な関数の最小化、特に非平滑ニューラルネットワークの分散学習における、分散型確率的劣勾配法に焦点を当てる。我々は、分散型確率的劣勾配降下法(DSGD)、勾配追跡(gradient-tracking)技術を用いたDSGD(DSGD-T)、およびモメンタムを用いたDSGD(DSGD-M)といった、様々な分散型劣勾配法を統合する一般的な枠組みを提案します。提案する枠組みの収束特性を明らかにするため、離散的な反復点列を、安定集合$A$を持つ強制(coercive)なリアプノフ関数を備えた連続時間微分包含(differential inclusion)の軌跡と関連付ける。そして、十分に小さく、かつ減少するステップサイズを用いることで、反復点列が安定集合$A$へ漸近的に収束することを証明します。これらの結果は、目的関数がClarke正則性を満たさない場合における、広く知られた分散型確率的劣勾配法に対して、初めて収束の保証を与えるものです。予備的な数値実験により、提案する枠組みを用いることで、非平滑ニューラルネットワークの学習において、収束保証を伴う極めて効率的な分散型確率的劣勾配法が実現できることが示されました。
Three Types of Calibration using Properties and their Semantic and Formal Relationships
Three Types of Calibration using Properties and their Semantic and Formal Relationships / 特性を用いた3種類のキャリブレーションと、それらの意味的・形式的関係
Fueled by discussions around “trustworthiness” and algorithmic fairness, calibration of predictive systems has regained scholars’ attention. The vanilla definition and understanding of calibration is, simply put, on all days on which the rain probability has been predicted to be $p$, the actual frequency of rain days was $p$. However, the increased attention has led to an immense variety of new notions of “calibration”. Some of the notions are incomparable, serve different purposes, or imply each other. In this work, we provide two accounts which motivate calibration: self-realization of forecasted properties and precise estimation of incurred losses of the decision makers relying on forecasts. We substantiate the former via the reflection principle and the latter by actuarial fairness. For both accounts we formulate prototypical definitions via properties $\Gamma$ of outcome distributions, e.g., the mean or median. The prototypical definition for self-realization, which we call $\Gamma$-calibration, is equivalent to a certain type of swap regret under certain conditions. These implications are strongly connected to the omniprediction learning paradigm. The prototypical definition for precise loss estimation is a modification of decision calibration adopted from Zhao et al., 2021. For binary outcome sets both prototypical definitions coincide under appropriate choices of reference properties. For higher-dimensional outcome sets, both prototypical definitions can be subsumed by a natural extension of the binary definition, called distribution calibration with respect to a property. We conclude by commenting on the role of groupings in both accounts of calibration often used to obtain multicalibration. In sum, this work provides a semantic map of calibration in order to navigate a fragmented terrain of notions and definitions.
「信頼性」やアルゴリズムの公平性をめぐる議論の高まりを受け、予測システムのキャリブレーション(較正)が再び研究者の注目を集めています。キャリブレーションの基本的な定義や理解は、端的に言えば、「降水確率が$p$と予測された日において、実際に雨が降った日の頻度が$p$である」ことです。しかし、関心の高まりとともに、「キャリブレーション」に関する多種多様な新しい概念が生まれています。それらの中には、互いに比較不可能なもの、異なる目的を持つもの、あるいは互いに含意関係にあるものが存在します。本研究では、キャリブレーションを動機づける2つの視点を提示します。それは、予測された特性の「自己実現(self-realization)」と、予測に依存する意思決定者が被る損失の「正確な推定」です。前者は「反射原理(reflection principle)」を用いて、後者は「保険数理的な公平性(actuarial fairness)」を用いて具体化します。これら双方の視点について、結果分布の特性$\Gamma$(例:平均や中央値)を用いた典型的な定義を定式化します。自己実現に関する典型的な定義(我々はこれを$\Gamma$-キャリブレーションと呼ぶ)は、特定の条件下において、ある種の「スワップ・リグレット(swap regret)」と等価です。これらの含意は、「オムニプレディクション(omniprediction)」学習パラダイムと密接に関連しています。正確な損失推定に関する典型的な定義は、Zhaoら(2021)で採用された「意思決定キャリブレーション(decision calibration)」を修正したものです。二値の結果集合においては、基準となる特性を適切に選択することで、これら2つの典型的な定義は一致します。高次元の出力集合に対しては、これら2つの典型的な定義はいずれも、「ある特性に関する分布キャリブレーション(distribution calibration with respect to a property)」と呼ばれる、二値の場合の定義の自然な拡張に包含されます。最後に、マルチキャリブレーション(multicalibration)の実現によく用いられるこれら2つのキャリブレーションの枠組みにおいて、グループ化(grouping)が果たす役割について論じます。要約すると、本研究は、概念や定義が断片化している現状を整理し、キャリブレーションに関する概念的全体像(セマンティック・マップ)を提示するものです。
Embedding Network Autoregression for Time Series Analysis and Causal Peer Effect Inference
Embedding Network Autoregression for Time Series Analysis and Causal Peer Effect Inference / 時系列分析および因果的ピア効果推論のための埋め込みネットワーク自己回帰
We propose an Embedding Network Autoregressive Model for multivariate networked longitudinal data. We assume the network is generated from a latent variable model, and these unobserved variables are included in a structural peer effect model or a time series network autoregressive model. This approach takes a unified view of two related yet different problems: (1) modeling and predicting multivariate networked time series data and (2) causal peer influence estimation in the presence of confounding due to homophily from finite-time longitudinal data. Our estimation strategy comprises estimating latent variables from the observed network, followed by least squares estimation of the network autoregressive model. We show that the momentum and peer effect parameters estimated with our method are consistent and asymptotically normally distributed in setups with a growing number of network vertices ($N$) while considering both a growing number of time points $T$ (for the time series problem) and finite $T$ cases (for the peer effect problem). We allow the number of latent vectors $K$ to grow at appropriate rates. We also develop a selection criterion when $K$ is unknown that provably does not under-select. We show that the theoretical guarantees hold with the selected number for $K$, and study the bias rates when $K$ is misspecified. With the new methods, we study peer effects in conflict and school climate perception using data on more than 7000 students from 23 schools.
本研究では、多変量ネットワーク・縦断データ(longitudinal data)を対象とした「埋め込みネットワーク自己回帰モデル(Embedding Network Autoregressive Model)」を提案します。ネットワークは潜在変数モデルから生成されると仮定し、それらの観測されない変数を、構造的ピア効果モデルまたは時系列ネットワーク自己回帰モデルに組み込みます。このアプローチは、関連しつつも異なる2つの問題、すなわち(1)多変量ネットワーク時系列データのモデリングと予測、および(2)ホモフィリー(同類選択)に起因する交絡が存在する状況下での、有限期間の縦断データを用いた因果的ピア効果の推定、を統一的な視点で扱います。推定手法としては、まず観測されたネットワークから潜在変数を推定し、続いてネットワーク自己回帰モデルの最小二乗推定を行います。ネットワークの頂点数($N$)が増大する設定において、時系列問題($T$も増大する場合)とピア効果問題($T$が有限の場合)の双方を考慮しつつ、本手法で推定されたモメンタム・パラメータおよびピア効果パラメータが一致性を持ち、かつ漸近的に正規分布に従うことを示します。また、潜在ベクトルの数$K$が適切な速度で増大することを許容しています。さらに、$K$が未知の場合に、過小選択(under-selection)を起こさないことが理論的に保証された選択基準を開発しました。選択された$K$の値を用いても理論的保証が成立することを示し、$K$が誤って指定された場合のバイアス率についても検討します。これらの新しい手法を用い、23校の7,000人以上の生徒のデータに基づき、紛争や学校の雰囲気(スクール・クライメート)の認識におけるピア効果について分析を行います。
Finite-Time Decoupled Convergence in Nonlinear Two-Time-Scale Stochastic Approximation
Finite-Time Decoupled Convergence in Nonlinear Two-Time-Scale Stochastic Approximation / 非線形2時間スケール確率近似における有限時間での非結合収束
In two-time-scale stochastic approximation (SA), two iterates are updated at varying speeds using different step sizes, with each update influencing the other. Previous studies on linear two-time-scale SA have shown that the convergence rates of the mean-square errors for these updates depend solely on their respective step sizes,a phenomenon termed decoupled convergence. However, achieving decoupled convergence in nonlinear SA remains less understood. Our research investigates the potential for finite-time decoupled convergence in nonlinear two-time-scale SA. We demonstrate that, under a nested local linearity assumption, finite-time decoupled convergence rates can be achieved with suitable step size selection. To derive this result, we conduct a convergence analysis of the matrix cross term between the iterates and leverage fourth-order moment convergence rates to control the higher-order error terms induced by local linearity. To further investigate the necessity of local linearity for decoupled convergence, we also construct an example showing that, even when the fast-time-scale update is linear, the nonlinearity of the slow-time-scale update alone can destroy decoupled convergence.
2つの時間スケールを持つ確率近似(SA)では、異なるステップサイズを用いて2つの反復変数が異なる速度で更新され、それぞれの更新が互いに影響を及ぼし合います。線形な2時間スケールSAに関する先行研究では、これらの更新における平均二乗誤差の収束速度がそれぞれのステップサイズのみに依存することが示されており、この現象は「非結合収束(decoupled convergence)」と呼ばれています。しかし、非線形SAにおいて非結合収束を実現することについては、まだ十分に解明されていません。本研究では、非線形2時間スケールSAにおける有限時間での非結合収束の可能性を調査します。我々は、「入れ子状の局所線形性(nested local linearity)」という仮定の下で、適切なステップサイズを選択すれば、有限時間での非結合収束速度が達成可能であることを示します。この結果を導き出すために、反復変数間の行列交差項の収束解析を行い、局所線形性に起因する高次の誤差項を制御するために4次モーメントの収束速度を利用します。さらに、非結合収束における局所線形性の必要性を探るため、高速時間スケールの更新が線形であっても、低速時間スケールの更新が非線形であるだけで非結合収束が損なわれ得ることを示す例を構築します。
The Within-Orbit Adaptive Leapfrog No-U-Turn Sampler
The Within-Orbit Adaptive Leapfrog No-U-Turn Sampler / 軌道内適応型リープフロッグNo-U-Turnサンプラー
Locally adapting parameters within Markov chain Monte Carlo methods while preserving reversibility is notoriously difficult. The success of the No-U-Turn Sampler (NUTS) largely stems from its clever local adaptation of the integration time in Hamiltonian Monte Carlo via a geometric U-turn condition. However, posterior distributions frequently exhibit multiscale geometries with extreme variations in scale, making it necessary to also adapt the leapfrog integrator’s step size locally and dynamically. Despite its practical importance, this problem has remained largely open since the introduction of NUTS by Hoffman and Gelman (2014).To address this issue, we introduce the Within-Orbit Adaptive Leapfrog No-U-Turn Sampler (WALNUTS), a generalization of NUTS that adapts the leapfrog step size at fixed intervals of simulated time as the orbit evolves. At each interval, the algorithm selects the largest step size from a dyadic schedule that keeps the energy error below a user-specified threshold. Like NUTS, WALNUTS employs biased progressive state selection to favor states with positions that are further from the initial point along the orbit. Empirical evaluations on multiscale target distributions, including Neal’s funnel and the Stock-Watson stochastic volatility time-series model, demonstrate that WALNUTS achieves substantial improvements in sampling efficiency and robustness compared to NUTS.
マルコフ連鎖モンテカルロ法において、可逆性を維持しつつパラメータを局所的に適応させることは、極めて困難であることが知られています。No-U-Turn Sampler(NUTS)の成功は、主に、幾何学的なUターン条件を用いてハミルトニアン・モンテカルロ法の積分時間を巧みに局所適応させたことに起因しています。しかし、事後分布はしばしばスケールが極端に異なるマルチスケールな幾何学的構造を示すため、リープフロッグ積分器のステップサイズも局所的かつ動的に適応させる必要があります。その実用上の重要性にもかかわらず、この問題はHoffmanとGelman(2014)によるNUTSの導入以来、未解決のままでした。この課題に対処するため、我々は「Within-Orbit Adaptive Leapfrog No-U-Turn Sampler(WALNUTS)」を提案します。これはNUTSを一般化したものであり、軌道の進行に伴い、シミュレーション時間の一定間隔ごとにリープフロッグのステップサイズを適応させます。各間隔において、アルゴリズムは、エネルギー誤差をユーザー指定の閾値以下に抑える2進(dyadic)スケジュールの中から、最大のステップサイズを選択します。NUTSと同様に、WALNUTSは、軌道上で初期点からより遠い位置にある状態を優先する「偏りを持たせた漸進的状態選択(biased progressive state selection)」を採用しています。Nealの「漏斗(funnel)」分布やStock-Watsonの確率的ボラティリティ時系列モデルを含む、マルチスケールなターゲット分布を用いた実証評価により、WALNUTSはNUTSと比較してサンプリング効率と堅牢性が大幅に向上することが示されました。
STDE++: Polynomial-Time Amortization for Linear Differential Operators
STDE++: Polynomial-Time Amortization for Linear Differential Operators / STDE++:線形微分作用素に対する多項式時間での償却計算
Optimizing neural networks with losses that contain high-dimensional and high-order differential operators is expensive to evaluate with backpropagation due to $\mathcal{O}(d^{k})$ scaling of the derivative tensor size and the $\mathcal{O}(2^{k-1}L)$ scaling in the computation graph, where $d$ is the domain dimension, $L$ is the number of ops in the forward computation graph and $k$ is the derivative order. Previous works addressed the polynomial scaling in $d$ by amortizing the computation over the optimization process via randomization. Separately, the exponential scaling in $k$ for univariate functions ($d=1$) was addressed with high-order auto-differentiation (AD). In this work, we show how to efficiently perform arbitrary contractions of the derivative tensor of arbitrary order for multivariate functions by properly constructing the input tangents to univariate high-order AD, which can be used to randomize any differential operator efficiently. When applied to Physics-Informed Neural Networks (PINNs) and compared against the original PyTorch implementation of SDGD, our method yields about $1.34\times 10^{3}$ average speedup and $31.8\times$ average memory reduction across the three inseparable 100K-dimensional PDEs in our benchmark; the best case is $1.59\times 10^{3}$ speedup and $33.8\times$ memory reduction on Allen-Cahn. We can now solve 1-million-dimensional PDEs in 8 minutes on a single NVIDIA A100 GPU. Furthermore, we proposed new methods for computing mixed partial derivatives using Taylor mode AD, which scales polynomially with the derivative order. This work opens the possibility of using high-order differential operators in large-scale problems.
高次元かつ高階の微分演算子を含む損失関数を用いてニューラルネットワークを最適化する場合、逆伝播法(バックプロパゲーション)による評価には多大なコストがかかります。これは、微分テンソルのサイズが$\mathcal{O}(d^{k})$、計算グラフの規模が$\mathcal{O}(2^{k-1}L)$で増大するためです(ここで$d$は領域の次元、$L$は順伝播計算グラフにおける演算数、$k$は微分の階数)。従来の研究では、ランダム化を用いて最適化プロセス全体で計算を分散(償却)させることで、$d$に関する多項式的な計算量増大の問題に対処してきました。また、一変数関数($d=1$)における$k$に関する指数関数的な計算量増大については、高階自動微分(AD)を用いることで対処されてきました。本研究では、一変数高階ADへの入力接ベクトルを適切に構成することで、多変数関数の任意の階数の微分テンソルに対する任意の縮約演算を効率的に実行する手法を提示します。この手法により、あらゆる微分演算子を効率的にランダム化することが可能になります。Physics-Informed Neural Networks(PINNs)に本手法を適用し、SDGDのオリジナルのPyTorch実装と比較したところ、ベンチマーク対象とした3つの分離不可能な10万次元偏微分方程式(PDE)において、平均で約1,340倍の高速化と31.8倍のメモリ削減を達成しました。特にAllen-Cahn方程式においては、最大で1,590倍の高速化と33.8倍のメモリ削減を実現しています。これにより、単一のNVIDIA A100 GPUを用いて、8分間で100万次元のPDEを解くことが可能になりました。さらに、微分階数に対して多項式的な計算量で済むTaylorモードADを用いた、混合偏微分の計算手法も提案しました。本研究は、大規模な問題において高階微分演算子を利用する可能性を切り拓くものです。
Vecchia-Inducing-Points Full-Scale Approximations for Gaussian Processes
Vecchia-Inducing-Points Full-Scale Approximations for Gaussian Processes / ガウス過程のためのVecchia誘導点(Inducing-Points)を用いたフルスケール近似
Gaussian processes are flexible, probabilistic, non-parametric models widely used in machine learning and statistics. However, their scalability to large data sets is limited by computational constraints. To overcome these challenges, we propose Vecchia-inducing-points full-scale (VIF) approximations combining the strengths of global inducing points and local Vecchia approximations. Vecchia approximations excel in settings with low-dimensional inputs and moderately smooth covariance functions, while inducing point methods are better suited to high-dimensional inputs and smoother covariance functions. Our VIF approach bridges these two regimes by using an efficient correlation-based neighbor-finding strategy for the Vecchia approximation of the residual process, implemented via a modified cover tree algorithm. We further extend our framework to non-Gaussian likelihoods by introducing iterative methods that substantially reduce computational costs for training and prediction by several orders of magnitude compared to Cholesky-based computations when using a Laplace approximation. In particular, we propose and compare novel preconditioners and provide theoretical convergence results. Extensive numerical experiments on simulated and real-world data sets show that VIF approximations are both computationally efficient as well as more accurate and numerically stable than state-of-the-art alternatives. All methods are implemented in the open-source C++ library GPBoost with high-level Python and R interfaces.
ガウス過程は、機械学習や統計学で広く利用されている、柔軟かつ確率的なノンパラメトリックモデルです。しかし、大規模データセットへの適用可能性(スケーラビリティ)は、計算上の制約によって制限されています。こうした課題を克服するため、我々は、グローバルな誘導点(inducing points)法と局所的なヴェッキア(Vecchia)近似法の長所を組み合わせた「Vecchia誘導点フルスケール(VIF)近似」を提案します。ヴェッキア近似は低次元入力や適度に滑らかな共分散関数を持つ設定で優れた性能を発揮する一方、誘導点法は高次元入力やより滑らかな共分散関数に適しています。我々のVIFアプローチは、改良型カバーツリーアルゴリズムを用いて実装された、相関に基づく効率的な近傍探索戦略を「残差過程のヴェッキア近似」に採用することで、これら二つの領域を橋渡しします。さらに、ラプラス近似を用いる際にコレスキー分解に基づく計算と比較して学習および予測の計算コストを数桁削減できる反復法を導入し、非ガウス尤度へと枠組みを拡張しました。特に、新規の事前調整子(preconditioner)を提案・比較し、理論的な収束結果を提示します。シミュレーションデータおよび実データを用いた広範な数値実験により、VIF近似は計算効率に優れているだけでなく、最先端の代替手法と比較してより高い精度と数値的安定性を備えていることが示されました。すべての手法は、PythonおよびRの高度なインターフェースを備えたオープンソースC++ライブラリ「GPBoost」に実装されています。
Statistical guarantees for denoising reflected diffusion models
Statistical guarantees for denoising reflected diffusion models / 反射拡散モデルのノイズ除去に対する統計的保証
In recent years, denoising diffusion models have become a crucial area of research due to their abundance in the rapidly expanding field of generative AI. While recent statistical advances have delivered explanations for the generation ability of idealised denoising diffusion models for high-dimensional target data, implementations introduce thresholding procedures for the generating process to overcome issues arising from the unbounded state space of such models. This mismatch between theoretical design and implementation of diffusion models has been addressed empirically by using a reflected diffusion process as the driver of noise instead. In this paper, we study statistical guarantees of these denoising reflected diffusion models. In particular, under Sobolev smoothness assumptions, we establish rates of convergence in total variation which, up to a polylogarithmic factor, match the minimax lower bound. Our main contributions include the statistical analysis of this novel class of denoising reflected diffusion models and a refined score approximation method in both time and space, leveraging spectral decomposition and rigorous neural network analysis.
近年、急速に拡大する生成AI分野において、ノイズ除去拡散モデル(denoising diffusion models)は極めて重要な研究領域となっています。高次元のターゲットデータに対する理想化されたノイズ除去拡散モデルの生成能力については、近年の統計学の進展により理論的な説明がなされていますが、実際のモデル実装においては、状態空間が非有界であることに起因する問題を克服するために、生成過程に閾値処理(thresholding)が導入されています。拡散モデルの理論的設計と実装との間にあるこのような不整合に対し、ノイズの駆動源として反射拡散過程(reflected diffusion process)を用いることで、経験的な解決が図られてきました。本論文では、これら「ノイズ除去反射拡散モデル」に関する統計的保証について研究します。特に、ソボレフ空間における滑らかさの仮定の下で、全変動(total variation)の意味での収束レートを確立しました。このレートは、ポリログ(polylogarithmic)な因子を除き、ミニマックス下界と一致するものです。我々の主な貢献は、この新しいクラスのノイズ除去反射拡散モデルの統計的解析に加え、スペクトル分解と厳密なニューラルネットワーク解析を活用した、時間および空間の両面における洗練されたスコア近似手法の提示です。
Learning general conditional independence structures via the neighbourhood lattice
Learning general conditional independence structures via the neighbourhood lattice / 近傍格子を用いた一般的な条件付き独立構造の学習
We study the problem of learning multivariate dependencies in nonparametric and high-dimensional settings. This includes but is not limited to graphical models. Our approach effectively combines several features that are missing from previous work on this problem: We show how the entire dependence structure can be learned nonparametrically while simultaneously evading the curse of dimensionality and relaxing common assumptions such as faithfulness. To this end, we introduce and study the neighbourhood lattice decomposition of a distribution, which is a compact, non-graphical representation of conditional independence (CI) that is valid in the absence of a faithful graphical representation. We show that the neighbourhood lattice decomposition exists in any graphical model and can be computed efficiently, nonparametrically, and consistently in high-dimensions without paying the usual curse of dimensionality. This gives a way to learn all of the independence relations implied by any graphical model, without requiring a priori knowledge of the graph or even the graph type. As a special case, our results provide a general solution to the problem of nonparametric estimation of high-dimensional CI structures over any graphical model.
我々は、ノンパラメトリックかつ高次元な設定における多変量依存関係の学習問題について研究します。これにはグラフィカルモデルも含まれますが、それらに限定されるものではありません。我々のアプローチは、この問題に関する従来の研究には欠けていたいくつかの特徴を効果的に組み合わせています。具体的には、「次元の呪い」を回避しつつ、「忠実性(faithfulness)」などの一般的な仮定を緩和しながら、依存関係の構造全体をノンパラメトリックに学習する方法を提示します。そのために、我々は分布の「近傍格子分解(neighbourhood lattice decomposition)」を導入・検討します。これは、忠実なグラフィカル表現が存在しない場合でも有効な、条件付き独立性(CI)のコンパクトかつ非グラフィカルな表現です。我々は、近傍格子分解があらゆるグラフィカルモデルにおいて存在すること、そして、通常の「次元の呪い」に悩まされることなく、高次元空間において効率的かつノンパラメトリックに、かつ整合性(consistency)を保ちながら計算可能であることを示します。これにより、グラフの構造や種類に関する事前知識を必要とせずに、任意のグラフィカルモデルが示唆するすべての独立関係を学習することが可能になります。特別なケースとして、我々の成果は、任意のグラフィカルモデルにおける高次元CI構造のノンパラメトリック推定問題に対する一般的な解決策を提供するものです。
Accelerating Constrained Sampling: A Large Deviations Approach
Accelerating Constrained Sampling: A Large Deviations Approach / 制約付きサンプリングの高速化:大偏差原理に基づくアプローチ
The problem of sampling a target probability distribution on a constrained domain arises in many applications including machine learning. For constrained sampling, various Langevin algorithms such as projected Langevin Monte Carlo (PLMC), based on the discretization of reflected Langevin dynamics (RLD) and more generally skew-reflected non-reversible Langevin Monte Carlo (SRNLMC), based on the discretization of skew-reflected non-reversible Langevin dynamics (SRNLD), have been proposed and studied in the literature. This work focuses on the long-time behavior of SRNLD, where a skew-symmetric matrix is added to RLD. Although acceleration for SRNLD has been studied, it is not clear how one should design the skew-symmetric matrix in the dynamics to achieve good performance in practice. We establish a large deviation principle (LDP) for the empirical measure of SRNLD when the skew-symmetric matrix is chosen such that its product with the outward unit normal vector field on the boundary is zero. By explicitly characterizing the rate functions, we show that this choice of the skew-symmetric matrix accelerates the convergence to the target distribution compared to RLD and reduces the asymptotic variance. Numerical experiments for SRNLMC based on the proposed skew-symmetric matrix show superior performance, which validate the theoretical findings from the large deviations theory.
制約付き領域上の目標確率分布からサンプリングを行う問題は、機械学習を含む多くの応用分野で生じます。制約付きサンプリングに関しては、反射ランジュバン力学(RLD)の離散化に基づく射影ランジュバン・モンテカルロ法(PLMC)や、より一般的には、斜め反射非可逆ランジュバン力学(SRNLD)の離散化に基づく斜め反射非可逆ランジュバン・モンテカルロ法(SRNLMC)など、様々なランジュバン・アルゴリズムが提案・研究されてきました。本研究では、RLDに歪対称行列を加えた力学系であるSRNLDの長時間挙動に焦点を当てます。SRNLDの加速手法については研究されてきましたが、実用上で良好な性能を達成するために、力学系における歪対称行列をどのように設計すべきかは明らかではありませんでした。本研究では、境界上の外向き単位法線ベクトル場との積がゼロになるように歪対称行列を選んだ場合について、SRNLDの経験測度に関する大偏差原理(LDP)を確立しました。レート関数を明示的に特徴付けることで、このような歪対称行列の選択が、RLDと比較して目標分布への収束を加速させ、かつ漸近分散を低減させることを示します。提案された歪対称行列に基づくSRNLMCを用いた数値実験においても優れた性能が確認され、大偏差理論に基づく理論的知見が裏付けられました。
Statistical Test for Attention in Transformers for Images and Time Series
Statistical Test for Attention in Transformers for Images and Time Series / 画像および時系列用Transformerにおけるアテンションの統計的検定
Transformer models have achieved exceptional performance in various domains, including computer vision and time-series analysis. Their core attention mechanism is widely used to interpret model decisions by assigning importance weights to input regions, such as image patches or time series intervals. However, the reliability of these interpretations remains a major concern. High-attention weights do not necessarily indicate genuinely significant features; they may instead be artifacts of the model’s computation, undermining their reliabilities in high-stakes applications such as medical diagnostics. To address this, we propose a novel statistical framework designed to quantify the significance of high-attention regions in Transformer models. Our framework is built on selective inference (SI) to correct for the inherent selection bias that arises from testing regions chosen through the complex attention computation of the Transformer models. A key contribution of this work is a novel computational method that extends SI to the complex non-linearity of self-attention, enabling the computation of valid $p$-values for high-attention regions. These $p$-values serve as a reliable measure of significance, strengthening the interpretability of Transformer decisions. The validity and effectiveness of our approach are demonstrated through numerical experiments and applications to brain image diagnosis and electroencephalography (EEG) data analysis.
Transformerモデルは、コンピュータビジョンや時系列解析など、様々な分野で卓越した性能を達成しています。その中核となるアテンション機構は、画像パッチや時系列区間などの入力領域に重要度(重み)を割り当てることで、モデルの判断根拠を解釈するために広く利用されています。しかし、こうした解釈の信頼性については依然として大きな懸念が残っています。高いアテンション重みが必ずしも真に重要な特徴を示しているとは限らず、モデルの計算過程で生じたアーティファクトである可能性もあり、医療診断のような重要な判断が求められる応用分野において、その信頼性を損なう恐れがあります。この課題に対処するため、我々はTransformerモデルにおける高アテンション領域の有意性を定量化するための新しい統計的枠組みを提案します。本枠組みは、Transformerモデルの複雑なアテンション計算を通じて選択された領域を検定する際に生じる固有の選択バイアスを補正するために、選択的推論(SI)の手法に基づいています。本研究の重要な貢献は、自己アテンション(self-attention)の複雑な非線形性に対してSIを拡張する新しい計算手法を開発した点にあり、これにより高アテンション領域に対する妥当な$p$値の算出が可能になります。これらの$p$値は信頼性の高い有意性の指標として機能し、Transformerの判断の解釈可能性を高めます。本手法の妥当性と有効性は、数値実験、および脳画像診断や脳波(EEG)データ解析への適用を通じて実証されています。
py/cuTAGI: An Open-Source Library for Tractable Approximate Gaussian Inference in Bayesian Neural Networks
py/cuTAGI: An Open-Source Library for Tractable Approximate Gaussian Inference in Bayesian Neural Networks / py/cuTAGI:ベイズニューラルネットワークにおける扱いやすい近似ガウス推論のためのオープンソースライブラリ
This paper introduces pyTAGI, a Python wrapper, and cuTAGI, its high-performance C++/CUDA backend, implementing Tractable Approximate Gaussian Inference (TAGI) for neural networks. TAGI treats all network quantities as Gaussian random variables and derives closed-form expressions for prior/posterior expected values, variances, and covariances, enabling analytic Bayesian learning without relying on gradient descent or backpropagation. The libraries mimic PyTorch’s sequential interface, allowing users to define models by stacking layers in order and performing uncertainty-aware Bayesian inference. Beyond epistemic uncertainty, it also allows quantifying heteroscedastic aleatoric uncertainty. cuTAGI’s custom CPU/GPU kernels and distributed-data-parallel support via NCCL/MPI deliver competitive runtimes, while pyTAGI’s pip-installable frontend and MIT-licensed GitHub repo facilitate community adoption and extension. Version 0.2.1 already supports a comprehensive suite of layers and activations; future work will add eager execution, further kernel optimizations, attention mechanisms, and advanced covariance factorization. Together, py/cuTAGI offer an efficient, open-source foundation for the analytic treatment of Bayesian deep learning.
本論文では、ニューラルネットワーク向けのTractable Approximate Gaussian Inference(TAGI:扱いに適した近似ガウス推論)を実装した、PythonラッパーであるpyTAGIと、その高性能なC++/CUDAバックエンドであるcuTAGIを紹介します。TAGIはネットワーク内のすべての量をガウス確率変数として扱い、事前・事後の期待値、分散、共分散を閉形式(解析的)な式で導出するため、勾配降下法や誤差逆伝播法に依存しない解析的なベイズ学習を可能にします。これらのライブラリはPyTorchのシーケンシャル・インターフェースを模倣しており、ユーザーは層を順に積み重ねてモデルを定義し、不確実性を考慮したベイズ推論を実行できます。認識論的不確実性(epistemic uncertainty)に加え、不均一分散性を持つ偶然的不確実性(heteroscedastic aleatoric uncertainty)の定量化も可能です。cuTAGIは、カスタムCPU/GPUカーネルやNCCL/MPIによる分散データ並列処理のサポートにより、競争力のある実行速度を実現しています。一方、pyTAGIはpipでインストール可能なフロントエンドとMITライセンスのGitHubリポジトリを提供し、コミュニティによる採用や拡張を容易にしています。バージョン0.2.1の時点で、すでに多種多様な層や活性化関数がサポートされており、今後の開発では、Eager実行、さらなるカーネルの最適化、アテンション機構、高度な共分散行列の分解機能などが追加される予定です。これらpyTAGIとcuTAGIは、ベイズ深層学習を解析的に扱うための、効率的かつオープンソースな基盤を提供します。
Gradient Span Algorithms Make Predictable Progress in High Dimension
Gradient Span Algorithms Make Predictable Progress in High Dimension / 勾配スパンアルゴリズムによる高次元空間での予測可能な進捗
We prove that all ‘gradient span algorithms’ have asymptotically deterministic behavior on scaled Gaussian random functions as the dimension tends to infinity. This is a functional generalization of similar results for random quadratic functions and spin glasses. They explain the counterintuitive phenomenon that different training runs of many large machine learning models result in approximately equal cost curves despite random initialization on a complicated non-convex landscape. This ‘predictable progress’ phenomenon is exploited by the AutoML community: Since the optimization progress of a single run is already representative, multiple retries with the same hyperparameters are not necessary.
我々は、次元数が無限大に向かうにつれ、あらゆる「勾配スパンアルゴリズム(gradient span algorithms)」が、スケーリングされたガウス型ランダム関数に対して漸近的に決定論的な挙動を示すことを証明します。これは、ランダムな二次関数やスピンガラスに関する同様の結果を関数空間へと一般化したものです。この結果は、複雑で非凸な損失曲面上でランダムに初期化された大規模機械学習モデルにおいて、異なる学習試行を行ってもほぼ同じコスト曲線が得られるという、直感に反する現象を説明するものです。この「予測可能な進捗(predictable progress)」という現象は、AutoMLコミュニティにおいて活用されています。すなわち、単一の試行における最適化の進捗状況がすでに全体を代表するものであるため、同じハイパーパラメータで何度も再試行する必要がないからです。
Robust training of implicit generative models for multivariate and heavy-tailed distributions with an invariant statistical loss
Robust training of implicit generative models for multivariate and heavy-tailed distributions with an invariant statistical loss / 不変統計的損失を用いた、多変量かつ裾の重い分布に対する暗黙的生成モデルのロバスト学習
Implicit generative models are often trained adversarially, which can yield unstable dynamics and mode collapse. The invariant statistical loss (ISL) offers a fully sample-based alternative by comparing empirical ranks of real and generated samples. In this work, we formally characterize ISL as a proper divergence over continuous distributions and establish key regularity properties, showing that it is continuous and differentiable, thereby enabling stable gradient-based optimization without adversarial games. We further enhance ISL along two practical axes. First, to better model heavy-tailed data, where Gaussian latent priors can limit tail expressivity, we introduce Pareto-ISL, which replaces Gaussian noise with a generalized Pareto latent distribution to improve the representation of both typical and extreme events. Second, to handle multivariate data at scale, we propose ISL-slicing: a computationally efficient procedure that projects samples onto random one-dimensional subspaces, computes rank-based losses per projection, and averages them to capture high-dimensional structure. Experiments demonstrate improved tail fidelity with Pareto-ISL and show that ISL-slicing scales effectively to high dimensions. Specifically, in high dimensional settings we show that ISL can be used either as a standalone criterion or as a strong pretraining objective for subsequent adversarial fine-tuning.
陰的生成モデル(implicit generative models)はしばしば敵対的学習によって訓練されるが、これには動的な不安定性やモード崩壊(mode collapse)を招く恐れがあります。不変統計的損失(ISL: Invariant Statistical Loss)は、実データと生成データのサンプルの経験的順位を比較することで、完全にサンプルベースの代替手法を提供します。本研究では、ISLを連続分布上の適切なダイバージェンス(divergence)として定式化し、その主要な正則性(連続性や微分可能性)を確立します。これにより、敵対的ゲームを用いることなく、勾配に基づく安定した最適化が可能となります。さらに、我々は2つの実用的な観点からISLを拡張します。第一に、ガウス分布の潜在事前分布では裾野(tail)の表現力が制限されがちな「裾の重い(heavy-tailed)」データを適切に扱うため、「Pareto-ISL」を導入します。これはガウスノイズを一般化パレート分布の潜在分布に置き換えることで、典型的な事象と極端な事象の双方の表現力を向上させるものです。第二に、大規模な多変量データを扱うために「ISL-slicing」を提案します。これは、サンプルをランダムな1次元部分空間に射影し、各射影における順位ベースの損失を計算して平均化することで、高次元構造を捉える計算効率の良い手法です。実験により、Pareto-ISLを用いることで裾野の再現性が向上すること、およびISL-slicingが高次元データに対しても効果的にスケールすることが示されました。具体的には、高次元設定において、ISLを単独の評価基準として使用できるだけでなく、その後の敵対的ファインチューニングに向けた強力な事前学習の目的関数としても利用できることを示す。
Adaptive Nonparametric Perturbations of Parametric Models with Generalized Bayes
Adaptive Nonparametric Perturbations of Parametric Models with Generalized Bayes / 一般化ベイズを用いたパラメトリックモデルの適応的ノンパラメトリック摂動
Parametric Bayesian modeling offers a powerful and flexible toolbox for machine learning. Yet the model, however detailed, may still be wrong, and this can make inferences untrustworthy. In this paper we introduce a new class of semiparametric corrections for parametric Bayesian models, when the target of inference is a functional of the true data distribution.Our starting point is a fully Bayesian modeling approach, which explicitly accounts for the possibility that the parametric model is wrong.Asymptotic analysis shows that this approach is both robust to model misspecification and data efficient, achieving fast convergence when the parametric model is close to true. However, the fully Bayesian approach is limited in its practical usefulness by the challenges of conducting inference and computing a Bayes factor for a nonparametric model. We therefore propose a novel model correction based on generalized Bayes, which entirely avoids the need to compute a nonparametric Bayes factor, but preserves the robustness and efficiency of the fully Bayesian approach. We demonstrate our method by estimating causal effects of gene expression from single cell RNA sequencing data. Overall, we offer a new efficient approach to robust Bayesian inference with parametric models.
パラメトリックベイズモデリングは、機械学習において強力かつ柔軟な手法を提供します。しかし、モデルがいかに詳細に構築されていても誤っている可能性があり、その場合、推論結果の信頼性が損なわれる恐れがあります。本論文では、真のデータ分布の汎関数を推論の対象とする場合を想定し、パラメトリック・ベイズ・モデルに対する新しいクラスの半パラメトリック補正手法を提案します。我々のアプローチの出発点は、パラメトリック・モデルが誤っている可能性を明示的に考慮に入れた、完全ベイズ・モデリングの手法です。漸近解析により、この手法はモデルの誤設定に対して頑健(ロバスト)であると同時にデータ効率も高く、パラメトリック・モデルが真の分布に近い場合には高速な収束を実現することが示されています。しかし、完全ベイズ・アプローチは、ノンパラメトリック・モデルにおける推論やベイズ因子の計算に伴う困難さゆえに、実用上の有用性が制限されていました。そこで我々は、一般化ベイズ(generalized Bayes)に基づく新しいモデル補正手法を提案します。この手法は、ノンパラメトリック・ベイズ因子の計算を完全に回避しつつ、完全ベイズ・アプローチの持つ頑健性と効率性を維持するものです。本手法の有効性を示すため、シングルセルRNAシーケンシング・データを用いて遺伝子発現の因果効果を推定する事例を提示します。総じて、我々はパラメトリック・モデルを用いた頑健なベイズ推論のための、効率的かつ新しいアプローチを提案するものです。
Minimax Optimal Convergence of Gradient Descent in Logistic Regression via Large and Adaptive Stepsizes
Minimax Optimal Convergence of Gradient Descent in Logistic Regression via Large and Adaptive Stepsizes / 大きく適応的なステップサイズを用いたロジスティック回帰における勾配降下法のミニマックス最適収束
We study gradient descent (GD) for logistic regression on linearly separable data with stepsizes that adapt to the current risk, scaled by a constant hyperparameter \(\eta\). We show that after at most \(1/\gamma^2\) burn-in steps, GD achieves a risk upper bounded by \(\exp(-\Theta(\eta))\), where \(\gamma\) is the margin of the dataset. As \(\eta\) can be arbitrarily large, GD attains an arbitrarily small risk immediately after the burn-in steps, though the risk evolution may be non-monotonic.We further construct hard datasets with margin \(\gamma\), where any batch (or online) first-order method requires \(\Omega(1/\gamma^2)\) steps to find a linear separator. Thus, GD with large, adaptive stepsizes matches the worst-case $1/\gamma^2$ dependence when the sample size is unrestricted. Notably, the classical Perceptron, a first-order online method, also achieves a step complexity of \(1/\gamma^2\), matching GD even in constants.Finally, our GD analysis extends to a broad class of loss functions and certain two-layer networks.
本研究では、線形分離可能なデータに対するロジスティック回帰の勾配降下法(GD)を扱います。ここでは、現在のリスクに適応し、かつ定数ハイパーパラメータ$\eta$でスケーリングされたステップサイズを用います。データセットのマージンを$\gamma$とすると、最大でも$1/\gamma^2$回の「バーンイン」ステップを経た後、GDは$\exp(-\Theta(\eta))$で上界されるリスクを達成することを示します。$\eta$は任意に大きく設定できるため、リスクの推移が非単調になる可能性はあるものの、GDはバーンインステップ直後に任意に小さなリスクを達成可能です。さらに、マージン$\gamma$を持つ「困難な」データセットを構築し、そこでは線形分離面を見つけるために、あらゆるバッチ(またはオンライン)一次最適化手法が$\Omega(1/\gamma^2)$回のステップを必要とすることを示します。したがって、大きな適応的ステップサイズを用いるGDは、サンプルサイズに制限がない場合、最悪ケースにおける$1/\gamma^2$という依存性と整合します。特筆すべき点として、古典的なパーセプトロン(一次オンライン手法)もまた$1/\gamma^2$のステップ計算量を達成し、定数倍を含めてGDと一致します。最後に、本研究におけるGDの解析は、広範なクラスの損失関数や特定の2層ネットワークにも拡張可能です。
Approximation-Free Differentiable Oblique Decision Trees
Approximation-Free Differentiable Oblique Decision Trees / 近似を必要としない微分可能斜め決定木
Decision Trees (DTs) are widely used in safety-critical domains such as medical diagnosis, valued for their interpretability and effectiveness on tabular data. However, training accurate oblique DTs is challenging due to complex optimization landscapes and overfitting risks, particularly in regression. Recent advances have introduced differentiable formulations that enable gradient-based training and joint optimization of decision boundaries and leaf regressors. Yet, existing approaches typically rely on approximations, either through probabilistic softening of boundaries (soft DTs) or quantized gradients such as the Straight-Through Estimator (STE). To overcome these limitations, we propose DTSemNet, a novel, semantically equivalent, and invertible representation of hard oblique DTs as neural networks. DTSemNet enables end-to-end training with standard gradient descent, eliminating the need for approximations in both classification and regression. While classification aligns naturally with this formulation, regression remains challenging due to the joint optimization of internal nodes and leaf regressors. To address this, we analyze the limitations of STE and introduce an annealed Top-$k$ method that provides accurate gradient signals without approximation. Extensive experiments on classification and regression benchmarks show that DTSemNet-trained oblique DTs outperform state-of-the-art differentiable DTs. Furthermore, we demonstrate that DTSemNet can serve as programmatic DT policies in reinforcement learning environments, thereby broadening their applicability.
決定木(DT)は、その解釈可能性と表形式データに対する有効性から、医療診断などの安全性が極めて重要な領域で広く利用されています。しかし、特に回帰タスクにおいて、複雑な最適化の地形(ランドスケープ)や過学習のリスクがあるため、精度の高い斜め決定木(oblique DT)を学習させることは困難です。近年の進展により、勾配に基づく学習や、決定境界と葉ノードの回帰器の同時最適化を可能にする微分可能な定式化が導入されています。しかし、既存の手法は通常、境界の確率的な軟化(ソフト決定木)や、Straight-Through Estimator(STE)のような量子化された勾配といった近似に依存しています。これらの限界を克服するため、我々は「DTSemNet」を提案します。これは、ハードな斜め決定木をニューラルネットワークとして表現するものであり、意味的に等価かつ可逆的な表現となっています。DTSemNetは標準的な勾配降下法によるエンドツーエンドの学習を可能にし、分類と回帰の双方において近似を必要としません。分類タスクはこの定式化と自然に適合しますが、回帰タスクでは内部ノードと葉ノードの回帰器を同時に最適化する必要があるため、依然として課題が残ります。これに対処するため、我々はSTEの限界を分析し、近似なしに正確な勾配信号を提供する「アニール型Top-$k$法(annealed Top-$k$ method)」を導入します。分類および回帰のベンチマークを用いた広範な実験により、DTSemNetで学習された斜め決定木(oblique DT)が、最先端の微分可能決定木を凌駕することが示されました。さらに、DTSemNetを強化学習環境におけるプログラム可能な決定木ポリシーとして利用できることを実証し、その適用範囲を拡大しました。
Underdamped Langevin MCMC with third order convergence
Underdamped Langevin MCMC with third order convergence / 3次収束性を有する減衰型ランジュバンMCMC
In this paper, we propose a new numerical method for the underdamped Langevin diffusion (ULD) and present a non-asymptotic analysis of its sampling error in the 2-Wasserstein distance when the $d$-dimensional target distribution $p(x)\propto e^{-f(x)}$ is strongly log-concave and has varying degrees of smoothness. Precisely, under the assumptions that the gradient and Hessian of $f$ are Lipschitz continuous, our algorithm achieves a 2-Wasserstein error of $\varepsilon$ in $\mathcal{O}\big(\sqrt{d}/\varepsilon\big)$ and $\mathcal{O}\big(\sqrt{d}/\sqrt{\varepsilon}\big)$ steps respectively. Therefore, our algorithm has a similar complexity as other popular Langevin MCMC algorithms under matching assumptions. However, if we additionally assume that the third derivative of $f$ is Lipschitz continuous, then our algorithm achieves a 2-Wasserstein error of $\varepsilon$ in $\mathcal{O}\big(\sqrt{d}/\varepsilon^{\frac{1}{3}}\big)$ steps. To the best of our knowledge, this is the first gradient-only method for ULD with third order convergence. To support our theory, we perform Bayesian logistic regression across a range of real-world datasets, where our algorithm achieves competitive performance compared to an existing underdamped Langevin MCMC algorithm and the popular No U-Turn Sampler (NUTS).
本論文では、低減衰ランジュバン拡散(ULD)のための新しい数値解法を提案し、$d$次元の目標分布$p(x)\propto e^{-f(x)}$が強対数凹性(strongly log-concave)を持ち、かつ様々な程度の滑らかさを有する場合における、2-Wasserstein距離でのサンプリング誤差の非漸近的解析を提示します。具体的には、$f$の勾配とヘッセ行列がリプシッツ連続であるという仮定の下で、我々のアルゴリズムはそれぞれ$\mathcal{O}\big(\sqrt{d}/\varepsilon\big)$ステップおよび$\mathcal{O}\big(\sqrt{d}/\sqrt{\varepsilon}\big)$ステップで$\varepsilon$の2-Wasserstein誤差を達成します。したがって、同様の仮定の下では、我々のアルゴリズムは他の一般的なランジュバンMCMCアルゴリズムと同等の計算量を有します。しかし、$f$の3階導関数もリプシッツ連続であると仮定を強めると、我々のアルゴリズムは$\mathcal{O}\big(\sqrt{d}/\varepsilon^{\frac{1}{3}}\big)$ステップで$\varepsilon$の2-Wasserstein誤差を達成します。我々の知る限り、これは3次の収束性を有するULD向けの、勾配情報のみを用いる初のアルゴリズムです。理論を裏付けるため、様々な実データセットを用いてベイズロジスティック回帰を実施したところ、我々のアルゴリズムは、既存の低減衰ランジュバンMCMCアルゴリズムや広く利用されているNo U-Turn Sampler (NUTS)と比較して、競争力のある性能を達成しました。
Mixing times of data-augmentation Gibbs samplers for high-dimensional probit regression
Mixing times of data-augmentation Gibbs samplers for high-dimensional probit regression / 高次元プロビット回帰におけるデータ拡張ギブスサンプラーの混合時間
We investigate the convergence properties of popular data-augmentation samplers for Baye\-sian probit regression. Leveraging recent results on Gibbs samplers for log-concave targets, we provide simple and explicit non-asymptotic bounds on the associated mixing times (in Kullback-Leibler divergence). The bounds depend explicitly on the design matrix and the prior precision, while they hold uniformly over the vector of responses. We specialize the results for different regimes of statistical interest, when both the number of data points $n$ and parameters $p$ are large: in particular we identify scenarios where the mixing times remain bounded as $n,p\to\infty$, and ones where they do not. The results are shown to be tight (in the worst case with respect to the responses) and provide guidance on choices of prior distributions that provably lead to fast mixing. An empirical analysis based on coupling techniques suggests that the bounds are effective in predicting practically observed behaviours.
本研究では、ベイズ・プロビット回帰に用いられる一般的なデータ拡張サンプラーの収束特性を調査します。対数凹型(log-concave)のターゲット分布に対するギブスサンプラーに関する最近の知見を活用し、関連する混合時間(カルバック・ライブラー・ダイバージェンスの観点から)について、単純かつ明示的な非漸近的境界を提示します。これらの境界は、計画行列や事前分布の精度行列に明示的に依存する一方で、応答変数のベクトルに対して一様に成立します。データ点数$n$とパラメータ数$p$が共に大きいという、統計的に重要な異なる状況に対して結果を具体化します。特に、$n, p \to \infty$となる極限において混合時間が有界に留まるシナリオと、そうならないシナリオを特定します。これらの境界は(応答変数に関して最悪の場合において)タイト(厳密)であることが示されており、高速な混合を保証する事前分布の選択に関する指針を提供します。カップリング手法に基づく実証分析により、これらの境界が実際に観測される挙動を予測する上で有効であることが示唆されます。
Abstract Gradient Training: A Unified Certification Framework for Data Poisoning, Unlearning, and Differential Privacy
Abstract Gradient Training: A Unified Certification Framework for Data Poisoning, Unlearning, and Differential Privacy / 抽象勾配学習:データポイズニング、アンラーニング、差分プライバシーのための統一的認証フレームワーク
The impact of inference-time data perturbation (e.g., adversarial attacks) has been extensively studied in machine learning, leading to well-established certification techniques for adversarial robustness. In contrast, certifying models against training data perturbations remains a relatively under-explored area. These perturbations can arise in three critical contexts: adversarial data poisoning, where an adversary manipulates training samples to corrupt model performance; machine unlearning, which requires certifying model behavior under the removal of specific training data; and differential privacy, where guarantees must be given with respect to substituting individual data points. This work introduces Abstract Gradient Training (AGT), a unified framework for certifying robustness of a given model and training procedure to training data perturbations, including bounded perturbations, the removal of data points, and the addition of new samples. By bounding the reachable set of parameters, i.e., establishing provable parameter-space bounds, AGT provides a formal approach to analyzing the behavior of models trained via first-order optimization methods.
機械学習において、推論時のデータ摂動(敵対的攻撃など)の影響は広範に研究されており、敵対的ロバスト性に関する認証手法が確立されています。対照的に、学習データの摂動に対するモデルの認証は、依然として十分に研究されていない分野です。こうした摂動は、主に3つの重要な文脈で生じます。すなわち、攻撃者が学習サンプルを操作してモデル性能を低下させる「敵対的データポイズニング」、特定の学習データを除去した際のモデル挙動の保証が求められる「マシン・アンラーニング」、そして個々のデータ点の置き換えに関して保証が必要となる「差分プライバシー」です。本研究では、有界な摂動、データ点の除去、新規サンプルの追加など、学習データに対する摂動に対してモデルおよび学習手順のロバスト性を認証するための統一的枠組みである「抽象勾配学習(AGT)」を提案します。AGTは、到達可能なパラメータ集合を制限すること、すなわちパラメータ空間における証明可能な境界を確立することを通じて、一次最適化手法によって学習されたモデルの挙動を解析するための形式的なアプローチを提供します。
Doubly Debiased Robust Subsampling for Transfer Learning
Doubly Debiased Robust Subsampling for Transfer Learning / 転移学習のための二重バイアス補正ロバストサブサンプリング
This paper develops a general framework for doubly debiased robust subsampling for transfer learning. The setting arises when massive source datasets are computationally infeasible to use in full, while naive or heuristic subsampling leads to biased estimators that further inherit transfer bias under source-target distributional shifts. We resolve these challenges through two complementary debiasing mechanisms. Inverse probability weighting removes subsampling bias by ensuring that subsample-based estimators represent the full source distribution, while a target-based one-step refinement recenters estimators towards the target distribution, thereby mitigating transfer bias. These corrections are embedded within a distributionally robust optimization design that simultaneously controls worst-case target risk and enforces source-target alignment through maximum mean discrepancy. To optimize subsampling distributions, we propose a scalarized particle swarm algorithm that efficiently explores the robustness-alignment frontier by adjusting a single tuning parameter. We establish theoretical properties, including asymptotic normality, generalization bounds, oracle inequalities, and minimax optimality under distributional uncertainty. Simulation studies and empirical applications in text sentiment and image recognition demonstrate that the proposed method consistently improves prediction accuracy and robustness compared with uniform subsampling, target-only training, and alignment-only approaches, and that both debiasing mechanisms are essential for reliable transfer.
本論文では、転移学習における二重バイアス補正(doubly debiased)を伴うロバストな部分サンプリングのための一般的な枠組みを構築します。この設定は、大規模なソースデータセットを完全な形で使用することが計算コストの観点から困難である一方で、単純な手法やヒューリスティックな手法による部分サンプリングでは、バイアスのある推定量が生じ、さらにソース・ターゲット間の分布のずれに起因する転移バイアスも引き継いでしまうような状況で生じます。我々は、互いに補完し合う2つのバイアス補正メカニズムを用いることで、これらの課題を解決します。逆確率重み付け(Inverse probability weighting)は、サブサンプルに基づく推定量が元の分布全体を代表するようにすることでサブサンプリングによるバイアスを除去する一方、ターゲット分布に基づいたワンステップの修正は推定量をターゲット分布へと再調整し、それによって転移バイアスを緩和します。これらの修正は、最悪ケースのターゲットリスクを制御しつつ、最大平均不一致(Maximum Mean Discrepancy)を用いてソースとターゲットの整合性を確保する、分布的にロバストな最適化の枠組みに組み込まれています。サブサンプリング分布を最適化するために、我々は単一の調整パラメータを変化させることでロバスト性と整合性のトレードオフ境界(フロンティア)を効率的に探索する、スカラー化粒子群最適化アルゴリズムを提案します。また、漸近正規性、汎化誤差限界、オラクル不等式、および分布の不確実性下でのミニマックス最適性を含む理論的性質を確立します。テキストの感情分析や画像認識におけるシミュレーションおよび実データへの適用を通じて、提案手法が一様サブサンプリング、ターゲットのみを用いた学習、整合性確保のみを行う手法と比較して、予測精度とロバスト性を一貫して向上させること、そして信頼性の高い転移を実現するにはこれら二つのバイアス除去メカニズムの双方が不可欠であることを実証します。
Learning to Play Two-Player Perfect-Information Games without Knowledge
Learning to Play Two-Player Perfect-Information Games without Knowledge / 事前知識なしでの2人完全情報ゲームのプレイ学習
This paper introduces a set of techniques for learning game state evaluation functions through reinforcement learning. First, we generalize tree bootstrapping, i.e. learning the values of states encountered during search rather than restricting updates to states observed during matches, to the setting of reinforcement learning with non-linear function approximation. Second, we modifies Unbounded Best-First Minimax by extending best action sequences to terminal states. Third, we replace the traditional binary game outcome $+1/-1$ with richer reinforcement signals, including quick wins, delayed losses, and scoring. Fourth, we propose a completion mechanism that exploits state resolution.Finally, we introduce a novel action-selection distribution, referred to as the ordinal distribution.Experimental results show that each of these techniques contributes to substantial improvements in playing strength. We integrate them into a unified algorithm, Athénan, and compare it against ExIt, a leading self-play reinforcement learning approach without prior knowledge.Our results demonstrate that Athénan consistently outperforms ExIt.We further evaluate Athénan on the games Hex, Othello, and Arimaa, where it surpasses state-of-the-art performance without relying on domain-specific knowledge. In addition, we consider the single-player game Morpion Solitaire, in which Athénan again reaches state-of-the-art results under the same constraint.Overall, these results show that reinforcement learning, when combined with the proposed techniques, can achieve state-of-the-art performance across a diverse range of games without the need for handcrafted heuristics or expert knowledge.
本論文では、強化学習を用いてゲームの状態評価関数を学習するための一連の手法を提案します。第一に、我々は「ツリーブートストラップ法」(対局中に観測された状態だけでなく、探索中に遭遇した状態の価値も学習する手法)を、非線形関数近似を用いた強化学習の枠組みへと一般化します。第二に、最善手列を終端状態まで拡張することで、「Unbounded Best-First Minimax」アルゴリズムを改良します。第三に、従来の「+1/-1」という二値的な勝敗結果の代わりに、早期勝利、遅延敗北、スコア獲得など、より情報量の多い強化学習シグナルを採用します。第四に、状態の解像度(state resolution)を活用した補完メカニズムを提案します。最後に、「順序分布(ordinal distribution)」と呼ぶ新しい行動選択分布を導入します。実験の結果、これらの各手法がプレイ強度の著しい向上に寄与することが示されました。我々はこれらを統合したアルゴリズム「Athénan」を構築し、事前知識を必要としない自己対戦型強化学習の代表的な手法である「ExIt」と比較した。その結果、Athénanは一貫してExItを上回る性能を示した。さらに、Hex、Othello、ArimaaといったゲームでAthénanを評価したところ、ドメイン固有の知識に頼ることなく、最先端の性能を上回る結果が得られた。加えて、一人用ゲームであるMorpion Solitaireにおいても、同様の制約下で最先端レベルの成果を達成した。総じて、これらの結果は、提案手法を組み合わせることで、強化学習が、手作業によるヒューリスティクスや専門家の知識を必要とせずに、多種多様なゲームにおいて最先端の性能を達成できることを示しています。
Graph-based Clustering Revisited: A Relaxation of Kernel k-Means Perspective
Graph-based Clustering Revisited: A Relaxation of Kernel k-Means Perspective / グラフベース・クラスタリングの再考:カーネルk平均法の緩和という視点から
The well-known graph-based clustering methods, including spectral clustering, symmetric non-negative matrix factorization, and doubly stochastic normalization, can be viewed as relaxations of the kernel k-means approach. However, we posit that these methods excessively relax their inherent low-rank, nonnegative, doubly stochastic, and orthonormal constraints to ensure numerical feasibility, potentially limiting their clustering efficacy. In this paper, guided by our systematic theoretical analyses, we propose Low-Rank Doubly stochastic clustering (LoRD), a model that only relaxes the orthonormal constraint to derive a probabilistic clustering results. Furthermore, by theoretically establishing the equivalence between orthogonality and Block diagonality under the doubly stochastic constraint, we propose B-LoRD. By integrating block diagonal regularization into LoRD, expressed as the maximization of the Frobenius norm, we enhance clustering performance. To ensure numerical solvability, we transform the non-convex doubly stochastic constraint into a linear convex constraint through the introduction of a class probability parameter. The theoretical demonstration of the gradient Lipschitz continuity of our LoRD and B-LoRD enables the proposal of a projected gradient algorithm whose exact iteration admits a sublinear convergence-rate bound and ensures first-order stationarity of every accumulation point for the exact projected gradient iteration. Extensive experiments underscore the effectiveness of our approaches. The code is publicly available at https://github.com/lwl-learning/LoRD.
スペクトルクラスタリング、対称非負行列因子分解、二重確率正規化といった、よく知られたグラフベースのクラスタリング手法は、カーネルk平均法(kernel k-means)の緩和形とみなすことができます。しかし、我々は、これらの手法が数値的な実行可能性を確保するために、本来備わっている低ランク性、非負性、二重確率性、正規直交性といった制約を過度に緩和しており、その結果クラスタリングの有効性が制限されている可能性があると考えています。本論文では、体系的な理論分析に基づき、確率的なクラスタリング結果を導出するために正規直交性の制約のみを緩和したモデル、「Low-Rank Doubly stochastic clustering (LoRD)」を提案します。さらに、二重確率性の制約下における直交性とブロック対角性の等価性を理論的に確立し、「B-LoRD」を提案します。フロベニウスノルムの最大化として定式化されるLoRDにブロック対角正則化を組み込むことで、クラスタリング性能を向上させます。数値的な解法を可能にするため、クラス確率パラメータを導入し、非凸な二重確率制約を線形凸制約へと変換します。LoRDおよびB-LoRDの勾配リプシッツ連続性を理論的に示すことで、射影勾配アルゴリズムを提案可能にしました。このアルゴリズムの厳密な反復は、劣線形な収束率の限界を満たし、かつ厳密な射影勾配反復におけるすべての集積点が一次の定常性を満たすことを保証します。広範な実験により、我々のアプローチの有効性が実証されています。コードはhttps://github.com/lwl-learning/LoRDにて公開されています。
End-to-End Deep Learning for Predicting Metric Space-Valued Outputs
End-to-End Deep Learning for Predicting Metric Space-Valued Outputs / 距離空間値出力を予測するためのエンドツーエンド・ディープラーニング
Many modern applications involve predicting structured, non-Euclidean outputs such as probability distributions, networks, and symmetric positive-definite matrices. These outputs are naturally modeled as elements of general metric spaces, where classical regression techniques that rely on vector space structure no longer apply. We introduce E2M (End-to-End Metric regression), a deep learning framework for predicting metric space-valued outputs. E2M performs prediction via weighted Fréchet means over training outputs, where the weights are learned by a neural network conditioned on the input. This construction provides a principled mechanism for geometry-aware prediction that avoids surrogate embeddings and restrictive parametric assumptions, while fully preserving the intrinsic geometry of the output space. We establish theoretical guarantees, including a universal approximation theorem that characterizes the expressive capacity of the model and a convergence analysis of the entropy-regularized training objective. Through extensive simulations involving probability distributions, networks, and symmetric positive-definite matrices, we show that E2M consistently achieves state-of-the-art performance, with its advantages becoming more pronounced at larger sample sizes. Applications to human mortality distributions and New York City taxi networks further demonstrate the flexibility and practical utility of this framework.
現代の多くの応用において、確率分布、ネットワーク、対称正定値行列といった、構造化された非ユークリッドな出力の予測が求められています。これらの出力は一般の距離空間の要素として自然にモデル化されますが、ベクトル空間の構造に依存する従来の回帰手法は適用できません。本研究では、距離空間値の出力を予測するためのディープラーニング・フレームワークであるE2M(End-to-End Metric regression)を提案します。E2Mは、学習データの出力に対する重み付きフレシェ平均(Fréchet mean)を用いて予測を行います。ここで、重みは入力に条件付けられたニューラルネットワークによって学習されます。この構成により、代替的な埋め込みや制約の多いパラメータ仮定を避けつつ、出力空間の本来の幾何学的構造を完全に保持した、幾何学的構造を考慮した予測が可能になります。我々は、モデルの表現能力を特徴付ける普遍近似定理や、エントロピー正則化された学習目的関数の収束解析など、理論的な保証を確立しました。確率分布、ネットワーク、対称正定値行列を用いた広範なシミュレーションを通じて、E2Mが一貫して最先端の性能を達成すること、そしてサンプルサイズが大きくなるにつれてその利点がより顕著になることを示します。さらに、人間の死亡率分布やニューヨーク市のタクシーネットワークへの適用例により、本フレームワークの柔軟性と実用的な有用性を示します。
The Sample Complexity of Parameter-Free Stochastic Convex Optimization
The Sample Complexity of Parameter-Free Stochastic Convex Optimization / パラメータフリーな確率的凸最適化におけるサンプル複雑度
We study the sample complexity of stochastic convex optimization when problem parameters such as the distance to optimality and the Lipschitz constant are unknown. We pursue two strategies. First, we develop a reliable model selection method that avoids overfitting to the validation set. This method allows us to generically tune the learning rate of stochastic optimization methods to match the optimal known-parameter sample complexity up to $\log\log$ factors. Second, we develop a regularization-based method that is specialized to the case that only the distance to optimality is unknown. More specifically, it uses norm-regularized empirical risk minimization to estimate the distance to optimality to within a constant factor, allowing known-parameter stochastic optimization methods to achieve optimal sample complexity. This method provides perfect adaptability to unknown distance to optimality, demonstrating a separation between the sample and computational complexity of parameter-free stochastic convex optimization. Combining these two methods allows us to simultaneously adapt to multiple problem structures. Experiments performing few-shot learning on CIFAR-10 by fine-tuning CLIP models and prompt engineering Gemini to count shapes indicate that our reliable model selection method can help mitigate overfitting to small validation sets.
本研究では、最適解までの距離やリプシッツ定数といった問題パラメータが未知である状況下での、確率的凸最適化のサンプル複雑性を考察します。我々は2つのアプローチをとります。第一に、検証セットへの過学習を回避する、信頼性の高いモデル選択手法を開発します。この手法を用いることで、確率的最適化法の学習率を汎用的に調整し、パラメータ既知の場合の最適サンプル複雑性を($\log\log$因子の範囲内で)達成することが可能になります。第二に、最適解までの距離のみが未知である場合に特化した、正則化に基づく手法を開発します。具体的には、ノルム正則化付き経験リスク最小化を用いて最適解までの距離を定数倍の精度で推定し、それによって、パラメータ既知の場合の確率的最適化法でも最適サンプル複雑性を達成できるようにします。この手法は未知の最適解までの距離に対して完全な適応性を提供し、パラメータフリーな確率的凸最適化におけるサンプル複雑性と計算複雑性の間の乖離(セパレーション)を実証するものです。これら2つの手法を組み合わせることで、複数の問題構造に同時に適応することが可能になります。CIFAR-10を用いたCLIPモデルのファインチューニングによるフューショット学習や、Geminiに対する形状カウントのプロンプトエンジニアリングといった実験において、我々の提案する信頼性の高いモデル選択手法が、小規模な検証セットへの過学習を抑制するのに有効であることが示されました。
Near-optimal Delta-convex Estimation of Lipschitz Functions
Near-optimal Delta-convex Estimation of Lipschitz Functions / リプシッツ関数の準最適に近いデルタ凸推定
This paper presents a tractable algorithm for estimating an unknown Lipschitz function from noisy observations and establishes an upper bound on its convergence rate. The approach extends max-affine methods from convex shape-restricted regression to the more general Lipschitz setting. A key component is a nonlinear feature expansion that maps max-affine functions into a subclass of delta-convex functions, which act as universal approximators of Lipschitz functions while preserving their Lipschitz constants. Leveraging this property, the estimator attains the minimax convergence rate (up to logarithmic factors) with respect to the intrinsic dimension of the data under squared loss and subgaussian distributions in the random design setting. The algorithm integrates adaptive partitioning to capture intrinsic dimension, a penalty-based regularization mechanism that removes the need to know the true Lipschitz constant, and a two-stage optimization procedure combining a convex initialization with local refinement. The framework is also straightforward to adapt to convex shape-restricted regression. Experiments demonstrate competitive performance relative to other theoretically justified methods, including nearest-neighbor and kernel-based regressors.
本論文では、ノイズを含む観測値から未知のLipschitz関数を推定するための扱いやすいアルゴリズムを提案し、その収束率の上界を確立します。この手法は、凸形状制約付き回帰におけるmax-affine(区分線形凸)法を、より一般的なLipschitz関数の設定へと拡張するものです。その鍵となるのは、max-affine関数をdelta-convex関数(2つの凸関数の差として表される関数)の特定のクラスへと写像する非線形特徴量展開です。このクラスの関数は、Lipschitz定数を維持しつつLipschitz関数の普遍近似器として機能します。この性質を活用することで、本推定器は、ランダムデザイン設定において、二乗損失およびサブガウス分布の下、データの真の次元(intrinsic dimension)に関して(対数因子を除き)ミニマックス収束率を達成します。本アルゴリズムは、データの真の次元を捉えるための適応的分割、真のLipschitz定数を事前に知る必要をなくすペナルティベースの正則化機構、そして凸最適化による初期化と局所的な精密化を組み合わせた2段階の最適化手順を統合しています。また、この枠組みは凸形状制約付き回帰にも容易に適用可能です。実験により、近傍法やカーネル回帰など、理論的裏付けのある他の手法と比較しても遜色のない性能を示すことが実証されています。
Error Analyses of Auto-Regressive Video Diffusion Models
Error Analyses of Auto-Regressive Video Diffusion Models / 自己回帰型動画拡散モデルの誤差解析
Auto-Regressive Video Diffusion Models (AR-VDMs) have shown strong capabilities in generating long, photorealistic videos, but suffer from two key limitations: (i) history forgetting, where the model loses track of previously generated content, and (ii) temporal degradation, where frame quality deteriorates over time. Yet a rigorous theoretical analysis of these phenomena is lacking, and existing empirical understanding remains insufficiently grounded. In this paper, we introduce Meta-ARVDM, a unified analytical framework that studies both errors through the shared autoregressive structure of AR-VDMs. We show that history forgetting is characterized by the conditional mutual information between the generated output and preceding frames, conditioned on inputs, and prove that incorporating more past frames monotonically alleviates history forgetting, thereby theoretically justifying a common belief in existing works. Moreover, our theory reveals that standard metrics fail to capture this effect, motivating a new evaluation protocol based on a “needle-in-a-haystack” task in closed-ended environments (DMLab and Minecraft). We further show that temporal degradation can be quantified by the cumulative sum of per-step errors, enabling prediction of degradation for different schedulers without video rollout. Finally, our evaluation uncovers a strong empirical correlation between history forgetting and temporal degradation, a connection not previously reported.
自己回帰型動画拡散モデル(AR-VDM)は、写実的で長い動画を生成する高い能力を示していますが、主に2つの限界を抱えています。(i)過去に生成された内容をモデルが追跡できなくなる「履歴の忘却(history forgetting)」と、(ii)時間の経過とともにフレームの品質が低下する「時間的劣化(temporal degradation)」です。しかし、これらの現象に関する厳密な理論的分析は不足しており、既存の経験的な理解も十分な根拠に基づいているとは言えません。本論文では、AR-VDMの共通の自己回帰構造を通じてこれら双方の誤差を研究する、統一的な分析フレームワーク「Meta-ARVDM」を提案します。我々は、履歴の忘却が、入力を条件とした「生成出力」と「先行フレーム」との間の条件付き相互情報量によって特徴付けられることを示します。また、より多くの過去フレームを組み込むことが履歴の忘却を単調に緩和することを証明し、既存研究における一般的な通念を理論的に裏付けます。さらに、標準的な指標ではこの効果を捉えられないことを理論的に明らかにし、閉じた環境(DMLabやMinecraft)における「干し草の山から針を探す(needle-in-a-haystack)」タスクに基づいた新しい評価プロトコルを提案します。加えて、時間的劣化はステップごとの誤差の累積和によって定量化できることを示し、動画のロールアウト(実際の生成実行)を行わずに、異なるスケジューラにおける劣化を予測可能にします。最後に、評価を通じて、履歴の忘却と時間的劣化の間に強い経験的相関があることを明らかにします。これはこれまで報告されていなかった関連性です。
High-Dimensional Analysis of Gradient Flow for Extensive-Width Quadratic Neural Networks
High-Dimensional Analysis of Gradient Flow for Extensive-Width Quadratic Neural Networks / 超広幅2次ニューラルネットワークにおける勾配流の高次元解析
We study the high-dimensional training dynamics of a shallow neural network with quadratic activation in a teacher–student setup. We focus on the extensive-width regime, where the teacher and student network widths scale proportionally with the input dimension, and the sample size grows quadratically. This scaling aims to describe overparameterized neural networks in which feature learning still plays a central role. In the high-dimensional limit, we derive a dynamical characterization of the gradient flow, in the spirit of dynamical mean-field theory (DMFT). Under $\ell_2$-regularization, we analyze these equations at long times and characterize the performance and spectral properties of the resulting estimator. This result provides a quantitative understanding of the effect of overparameterization on learning and generalization, and reveals a double descent phenomenon in the presence of label noise, where generalization improves beyond interpolation. In the small regularization limit, we obtain an exact expression for the perfect recovery threshold as a function of the network widths, providing a precise characterization of how overparameterization influences recovery.
本研究では、教師・生徒設定における、二次関数的な活性化関数を持つ浅いニューラルネットワークの、高次元での学習ダイナミクスを研究します。特に、教師および生徒ネットワークの幅が入力次元に比例して拡大し、サンプルサイズがその二乗のオーダーで増大する「広範な幅(extensive-width)」の領域に焦点を当てます。このスケーリングは、特徴学習が依然として中心的な役割を果たす過剰パラメータ化されたニューラルネットワークを記述することを意図しています。高次元極限において、我々は動的平均場理論(DMFT)の精神に基づき、勾配流の動的な特性を導出します。$\ell_2$正則化の下で、長時間極限におけるこれらの方程式を解析し、得られる推定量の性能およびスペクトル特性を明らかにします。この結果は、過剰パラメータ化が学習と汎化に及ぼす影響についての定量的な理解を提供し、ラベルノイズが存在する場合に、補間(interpolation)を超えて汎化性能が向上する「ダブルディセント(二重降下)」現象を明らかにします。正則化が小さい極限では、ネットワーク幅の関数として完全復元(perfect recovery)の閾値を表す厳密な式を導出し、過剰パラメータ化が復元にどのように影響するかを正確に特徴付けます。
Bridging Domain Invariance and Diversity: A Fine-Grained Risk Bound for Domain Generalization
Bridging Domain Invariance and Diversity: A Fine-Grained Risk Bound for Domain Generalization / ドメイン不変性と多様性の架け橋:ドメイン汎化のための詳細なリスク境界
Domain-invariant representation learning and domain augmentation algorithms are two principal methodological paradigms for addressing domain generalization. They are widely employed in the machine learning literature to enhance domain invariance and domain diversity, respectively. However, existing risk bounds for domain generalization do not simultaneously capture the contributions of both approaches. This limitation arises because bounds derived directly in the original latent space are typically too coarse-grained and ambiguous to characterize how invariance and diversity jointly influence generalization. Since these two properties are often regarded as being inherently contradictory, it becomes difficult to disentangle and rigorously characterize their individual effects. To address this issue, we first observe that the latent representation space can be decomposed into several distinct subspaces, each exhibiting different characteristics and therefore being better suited for analyzing the respective roles of domain invariance and domain diversity. Building on this observation, we propose a unified analytical framework for domain generalization. Specifically, we introduce a Tri-Space Latent Representation and establish its unique decomposability via a direct-sum decomposition. Under this decomposition, each data representation can be uniquely partitioned into three components: domain-invariant features, spurious invariant features, and domain-variant features. Within this framework, we derive a finer-grained bound on the target-domain risk, which consists of two principal terms corresponding to domain diversity and invariant factors. By theoretically analyzing these two terms, we show that domain-invariant representation learning and domain augmentation are both effective and, crucially, compatible strategies for addressing domain generalization. Finally, we design two sets of experiments to empirically validate the relationship between domain invariance and domain diversity, and to examine their respective effects on domain generalization performance.
ドメイン不変表現学習とドメイン拡張アルゴリズムは、ドメイン汎化(domain generalization)に対処するための主要な2つの方法論的パラダイムです。これらは、それぞれドメイン不変性とドメイン多様性を高める手法として、機械学習の文献で広く採用されています。しかし、ドメイン汎化に関する既存の「リスク境界(risk bounds)」は、これら両方のアプローチによる寄与を同時に捉えることができていません。このような限界が生じるのは、元の潜在空間で直接導出された境界が、不変性と多様性が汎化にどのように複合的に影響するかを特徴付けるには、一般にあまりに粗く曖昧だからです。これら2つの特性は本質的に相反するものと見なされることが多いため、それぞれの効果を切り分け、厳密に特徴付けることは困難です。この問題に対処するため、我々はまず、潜在表現空間がいくつかの異なる部分空間に分解可能であることに着目しました。各部分空間は異なる特性を示すため、ドメイン不変性とドメイン多様性のそれぞれの役割を分析するのに適しています。この知見に基づき、我々はドメイン汎化のための統一的な分析フレームワークを提案します。具体的には、「3空間潜在表現(Tri-Space Latent Representation)」を導入し、直和分解によってその一意な分解可能性を確立します。この分解の下では、各データ表現を「ドメイン不変な特徴」、「見かけ上の不変特徴(spurious invariant features)」、および「ドメイン可変な特徴」という3つの成分に一意に分割できます。このフレームワーク内で、我々はターゲットドメインにおけるリスクについて、より詳細な(きめ細かい)境界を導出します。この境界は、ドメイン多様性と不変要因に対応する2つの主要な項で構成されます。これら2つの項を理論的に分析することで、ドメイン不変表現学習とドメイン拡張が、ドメイン汎化に対処する上で有効であり、かつ重要なことに、互いに両立し得る戦略であることを示します。最後に、ドメイン不変性とドメイン多様性の関係を実証的に検証し、ドメイン汎化性能に対するそれぞれの効果を調べるために、2種類の実験を設計・実施します。