Journal of Statistical Software Volume 117に記載されている内容を一覧にまとめ、機械翻訳を交えて日本語化し掲載します。
目次
- 1 記事
- 1.1 Making, Updating, and Querying Causal Models with CausalQueries
- 1.2 nonprobsvy: An R Package for Modern Methods for Non-Probability Surveys
- 1.3 asgl: A Python Package for Penalized Quantile, Linear, and Logistic Regression
- 1.4 BACE: A gretl Package for Model Averaging in Bounded Dependent Variable Models
- 1.5 OGBoost: A Python Package for Ordinal Regression Gradient Boosting
- 2 参考文献
- 3 関連情報
記事
Making, Updating, and Querying Causal Models with CausalQueries
Making, Updating, and Querying Causal Models with CausalQueries / CausalQueriesによる因果モデルの作成・更新・クエリ実行
The R package CausalQueries can be used to make, update, and query causal models defined on binary nodes. Users provide a causal statement of the form X -> M <- Y; M <-> Y which is interpreted as a structural causal model over a collection of binary nodes. Then CausalQueries allows users to (1) identify the set of principal strata – causal types – required to characterize all possible causal relations between nodes that are consistent with the causal statement (2) determine a set of parameters needed to characterize distributions over these causal types (3) update beliefs over distributions of causal types, using a Stan model plus data, and (4) pose a wide range of causal queries of the model, using either the prior distribution, the posterior distribution, or a user-specified candidate vector of parameters.
Rパッケージ「CausalQueries」は、二値ノード上で定義された因果モデルの作成、更新、およびクエリ(問い合わせ)を行うために使用できます。ユーザーは「X -> M <- Y; M <-> Y」といった形式で因果関係を記述し、これは二値ノードの集合上における構造的因果モデルとして解釈されます。その上で、CausalQueriesを用いると、ユーザーは以下のことが可能になります。(1)その因果関係の記述と整合するノード間のあらゆる可能な因果関係を特徴づけるために必要な「主要層(principal strata)」(因果タイプ)の集合を特定します。(2)それらの因果タイプ上の分布を特徴づけるために必要なパラメータの集合を決定します。(3) Stanモデルとデータを用いて、因果タイプの分布に関する信念(確信度)を更新します。(4)事前分布、事後分布、あるいはユーザーが指定したパラメータの候補ベクトルを用いて、モデルに対して多岐にわたる因果クエリを行う。
nonprobsvy: An R Package for Modern Methods for Non-Probability Surveys
nonprobsvy: An R Package for Modern Methods for Non-Probability Surveys / nonprobsvy: 非確率サンプリング調査の現代的手法に対応したRパッケージ
The paper presents nonprobsvy – an R package for inference based on non-probability samples. The package implements various approaches that can be categorized into three groups: model-based prediction approach, inverse probability weighting and doubly robust estimation. We assume the existence of either population-level data or probability-based population information and leverage the survey package for inference. The package implements both analytical and bootstrap variance estimation for the proposed estimators. We present the theory behind the package, its functionalities and a case study that showcases the usage of the package. The package is aimed at scientists and researchers who would like to use non-probability samples (e.g., big data, opt-in web panels, social media) to accurately estimate population characteristics.
本論文では、非確率標本に基づく推論を行うためのRパッケージ「nonprobsvy」を紹介します。このパッケージは、モデルベース予測、逆確率重み付け、および二重ロバスト推定という3つのグループに分類される様々な手法を実装しています。母集団レベルのデータまたは確率標本に基づく母集団情報の存在を前提とし、推論には既存の「survey」パッケージを活用します。また、提案する推論手法(推定量)に対して、解析的な方法とブートストラップ法の両方を用いた分散推定機能も備えています。本論文では、パッケージの理論的背景や機能に加え、実際の使用例を示すケーススタディも提示します。本パッケージは、非確率標本(ビッグデータ、オプトイン型Webパネル、ソーシャルメディアなど)を用いて母集団の特性を正確に推定したいと考える科学者や研究者を対象としています。
asgl: A Python Package for Penalized Quantile, Linear, and Logistic Regression
asgl: A Python Package for Penalized Quantile, Linear, and Logistic Regression / asgl: ペナルティ付き分位点回帰・線形回帰・ロジスティック回帰のためのPythonパッケージ
asgl is an open-source Python package that offers a robust and versatile framework for fitting a variety of regression models including linear, logistic, and, notably, quantile regression. It implements a comprehensive suite of penalization techniques such as LASSO, ridge, group LASSO, sparse group LASSO, elastic net, and their adaptive variants. A key contribution of asgl is its extensive support for adaptive penalizations, critically offering a range of built-in methodologies for estimating the necessary adaptive weights as proposed by Mendez-Civieta, Aguilera-Morillo, and Lillo (2021). This feature addresses a significant practical challenge – the weight estimation process – in applying advanced adaptive methods, especially in high-dimensional settings, and is largely absent from other packages. Furthermore, asgl offers penalized quantile regression, a less commonly available feature in statistical software. The primary class, Regressor, ensures seamless integration with the scikit-learn ecosystem, facilitating straightforward model evaluation and hyperparameter optimization. asgl has demonstrated utility in variable selection and prediction tasks across both low- and high-dimensional data, positioning it as a comprehensive tool for modern statistical modeling.
「asgl」は、線形回帰、ロジスティック回帰、そして特に分位点回帰を含む様々な回帰モデルを適合させるための、堅牢かつ汎用性の高いフレームワークを提供するオープンソースのPythonパッケージです。LASSO、リッジ、グループLASSO、スパースグループLASSO、エラスティックネット、およびそれらの適応型(adaptive)バリエーションといった、包括的なペナルティ化手法を実装しています。asglの重要な貢献の一つは、適応型ペナルティ化への広範な対応です。特に、Mendez-Civieta, Aguilera-Morillo, and Lillo (2021)が提案した手法に基づき、必要な適応型重みを推定するための組み込みメソッドを豊富に提供している点が特筆されます。この機能は、高度な適応型手法(特に高次元データへの適用時)において大きな実務的課題となっていた「重み推定プロセス」に対処するものであり、他のパッケージではほとんど見られないものです。さらに、asglは統計ソフトウェアではあまり一般的ではない「ペナルティ付き分位点回帰」も提供しています。主要なクラスである「Regressor」は、scikit-learnのエコシステムとシームレスに統合されており、モデルの評価やハイパーパラメータの最適化を容易に行うことができます。asglは、低次元および高次元データの双方における変数選択や予測タスクにおいて有用性が実証されており、現代の統計モデリングのための包括的なツールとしての地位を確立しています。
BACE: A gretl Package for Model Averaging in Bounded Dependent Variable Models
BACE: A gretl Package for Model Averaging in Bounded Dependent Variable Models / BACE: 有界従属変数モデルにおけるモデル平均化のためのgretlパッケージ
This study presents a fast and consistent software package called BACE (i.e., Bayesian averaging of classical estimates) that offers a model-building strategy for various bounded dependent variable models, including – logit and probit models, ordered logit and probit models, multinomial logistic regression, Poisson regression, Tobit model, interval regression – as well as linear regression. The BACE approach, originally associated with normal linear regression, is a model selection method that combines classical estimation and Bayesian techniques. It solves the problem of computational speed and model uncertainty that arise when dealing with numerous competing advanced statistical models. Our package also provides an implementation of the well-established Bayesian information criterion variants and the latest measures of jointness. We use gretl, a popular, free, and open-source software for econometric analysis that features an easy-to-use graphical user interface.
本研究では、BACE(Bayesian averaging of classical estimates:古典的推定値のベイズ平均化)と呼ばれる、高速かつ一貫性のあるソフトウェアパッケージを提案します。これは、線形回帰に加え、ロジット/プロビットモデル、順序ロジット/プロビットモデル、多項ロジスティック回帰、ポアソン回帰、トービットモデル、区間回帰など、様々な「値の範囲が制限された従属変数」を扱うモデルの構築戦略を提供するものです。元来、正規線形回帰に関連して開発されたBACEアプローチは、古典的推定とベイズ的手法を組み合わせたモデル選択手法であり、多数の競合する高度な統計モデルを扱う際に生じる計算速度やモデルの不確実性といった課題を解決します。また、本パッケージは、確立されたベイズ情報量規準(BIC)の派生手法や、最新の「結合性(jointness)」指標の実装も提供します。実装には、使いやすいGUIを備えた、計量経済分析用の人気ある無料オープンソースソフトウェア「gretl」を使用しています。
OGBoost: A Python Package for Ordinal Regression Gradient Boosting
OGBoost: A Python Package for Ordinal Regression Gradient Boosting / OGBoost: 順序回帰勾配ブースティングのためのPythonパッケージ
This paper introduces OGBoost, a scikit-learn-compatible Python package for ordinal regression gradient boosting. Ordinal variables (e.g., rating scales, quality assessments) lie between nominal and continuous data, requiring specialized methods that respect their inherent ordering. OGBoost employs a novel coordinate-descent optimization approach within the cumulative link framework, jointly optimizing a continuous regression function via functional gradient descent and a threshold vector via standard gradient descent. Key features include: (i) latent score predictions via decision_function for enhanced ranking performance; (ii) heterogeneous boosting ensembles by enabling any combination of scikit-learn regressors as base learners; and (iii) cross-validation-based early stopping for improved robustness. We provide a comprehensive review of existing ordinal regression software and present empirical comparisons across 17 public datasets, demonstrating that OGBoost significantly outperforms representative methods including OrdinalGBT, scikit-learn’s gradient boosting classifier, and scikit-lego’s ordinal classifier. Formal statistical tests confirm OGBoost’s superiority on concordance index (achieving best performance on all 17 datasets), with consistent performance advantages across classification accuracy, mean absolute error, and Spearman correlation. The package integrates seamlessly with scikit-learn machine learning workflows and is available via pip install ogboost.
本論文では、順序回帰のための勾配ブースティング手法を実装した、scikit-learn互換のPythonパッケージ「OGBoost」を紹介します。順序変数(例:評価尺度や品質評価)は名義尺度と連続データの中間に位置する性質を持ち、その内在する順序性を考慮した専門的な手法を必要とします。OGBoostは、累積リンク(cumulative link)の枠組みの中で、新規の座標降下最適化アプローチを採用しています。具体的には、関数勾配降下法による連続回帰関数の最適化と、標準的な勾配降下法による閾値ベクトルの最適化を同時に行います。主な特徴は以下の通りです。(i)ランキング性能を向上させるための、`decision_function`による潜在スコア予測。(ii) scikit-learnの回帰器をベース学習器として任意に組み合わせることを可能にした、異種混合ブースティング・アンサンブル。(iii)ロバスト性を高めるための、交差検証に基づく早期終了(アーリーストッピング)。既存の順序回帰ソフトウェアの包括的なレビューを行うとともに、17の公開データセットを用いた実証的な比較を行い、OGBoostがOrdinalGBT、scikit-learnの勾配ブースティング分類器、scikit-legoの順序分類器といった代表的な手法を大幅に上回る性能を示すことを実証します。統計的検定により、コンコーダンス指数(C-index)においてOGBoostの優位性が確認され(全17データセットで最高性能を達成)、分類精度、平均絶対誤差、スピアマンの順位相関係数においても一貫した性能上の利点が示されました。本パッケージはscikit-learnの機械学習ワークフローとシームレスに統合されており、`pip install ogboost`で利用可能です。


