Kalman-filterをはじめとする、逐次的な2乗誤差最小化に基づくリアルタイム信号処理では、損失関数を構成する2次形式の基になる行列の逆行列を高速で求める必要がある。そのような状況で頼りになるSherman-Morrisonの公式を導出する。Woodburyの公式の特別な場合として片付けられるのが一般的だが、本記事ではJordan分解を用いて導出してみる。
\[
% to cope with definition mismatch between MathJax and LaTeX
\newcommand{\coloneq}{\mathrel{:=}}
\newcommand{\eqcolon}{\mathrel{=:}}
% general purpose
\newcommand{\ctext}[1]{\raise0.2ex\hbox{\textcircled{\scriptsize{#1}}}}
% mathematics
% general purpose
\DeclarePairedDelimiterX{\parens}[1]{\lparen}{\rparen}{#1}
\DeclarePairedDelimiterX{\braces}[1]{\lbrace}{\rbrace}{#1}
\DeclarePairedDelimiterX{\bracks}[1]{\lbrack}{\rbrack}{#1}
\DeclarePairedDelimiterX{\verts}[1]{|}{|}{#1}
\DeclarePairedDelimiterX{\Verts}[1]{\|}{\|}{#1}
\DeclarePairedDelimiterX{\setComprehension}[2]{\lbrace}{\rbrace}{#1\,\delimsize\vert\,#2}
\newcommand{\as}{{\quad\textrm{as}\quad}}
\newcommand{\st}{{\textrm{ s.t. }}}
\newcommand{\naturalNumbers}{\mathbb{N}}
\newcommand{\integers}{\mathbb{Z}}
\newcommand{\rationalNumbers}{\mathbb{Q}}
\newcommand{\realNumbers}{\mathbb{R}}
\newcommand{\nonNegRealNumbers}{\mathbb{R}_{\geq 0}}
\newcommand{\posRealNumbers}{\mathbb{R}_{> 0}}
\newcommand{\complexNumbers}{\mathbb{C}}
\newcommand{\field}{\mathbb{F}}
\newcommand{\EuclideanSpace}{\mathbb{E}}
\newcommand{\argmax}{\operatorname*{arg~max}}
\newcommand{\argmin}{\operatorname*{arg~min}}
% set theory
\newcommand{\range}[2]{\braces*{#1,\dotsc,#2}}
\renewcommand{\complement}{\mathrm{c}}
\newcommand{\ind}[2]{\mathbbm{1}_{#1}\parens*{#2}}
\newcommand{\indII}[1]{\mathbbm{1}\braces*{#1}}
% number theory
\newcommand{\abs}[1]{\verts*{#1}}
\newcommand{\combi}[2]{{_{#1}\mathrm{C}_{#2}}}
\newcommand{\perm}[2]{{_{#1}\mathrm{P}_{#2}}}
\newcommand{\GaloisField}{\mathrm{GF}}
% real analysis
\newcommand{\NapierE}{\mathrm{e}}
\newcommand{\sgn}{\operatorname{sgn}} % sign function
\newcommand{\rect}{\operatorname{rect}} % rectangular function
\newcommand{\cl}{\operatorname{cl}} % closure of a set
\newcommand{\img}{\operatorname{im}} % image of a function
\newcommand{\dom}{\operatorname{dom}} % domain of a function
\newcommand{\LittleO}[2]{\underset{#1}{o}\parens*{#2}} % little-o notation
\newcommand{\norm}[1]{\Verts*{#1}} % norm of a vector or function
\newcommand{\floor}[1]{\left\lfloor #1\right\rfloor}
\newcommand{\ceil}[1]{\left\lceil#1\right\rceil}
\newcommand{\sinc}{\operatorname{sinc}}
\newcommand{\nrmSinc}{\operatorname{nsinc}} % normalized sinc function
\newcommand{\erf}{\operatorname{erf}}
% inverse trigonometric functions
\newcommand{\asin}{\operatorname{Sin}^{-1}}
\newcommand{\acos}{\operatorname{Cos}^{-1}}
\newcommand{\atan}{\operatorname{Tan}^{-1}}
% derivative
\newcommand{\deriv}[3]{\frac{\mathrm{d}^{#3}#1}{\mathrm{d}{#2}^{#3}}}
\newcommand{\derivLong}[3]{\frac{\mathrm{d}^{#3}}{\mathrm{d}{#2}^{#3}}#1}
\newcommand{\partDeriv}[3]{\frac{\mathrm{\partial}^{#3}#1}{\mathrm{\partial}{#2}^{#3}}}
\newcommand{\partDerivLong}[3]{\frac{\mathrm{\partial}^{#3}}{\mathrm{\partial}{#2}^{#3}}#1}
\newcommand{\partDerivIIHetero}[3]{\frac{\mathrm{\partial}^2#1}{\partial#2\mathrm{\partial}#3}}
\newcommand{\partDerivIIHeteroLong}[3]{{\frac{\mathrm{\partial}^2}{\partial#2\mathrm{\partial}#3}#1}}
% integral
\newcommand{\pv}{{\textrm{ p.v. }}}
\newcommand{\integrate}[5]{\int_{#1}^{#2}{#3}{\;\mathrm{d}^{#4}}#5}
\newcommand{\LebInteg}[4]{\int_{#1} {#2} {#3}\parens*{\;\mathrm{d}#4}}
% complex analysis
\newcommand{\conj}[1]{\overline{#1}}
\renewcommand{\Re}{\operatorname{Re}}
\renewcommand{\Im}{\operatorname{Im}}
\newcommand{\Arg}{\operatorname{Arg}}
\newcommand{\Log}{\operatorname{Log}}
% Laplace transform
\newcommand{\LPLC}{\operatorname{\mathcal{L}}}
\newcommand{\ILPLC}{\operatorname{\mathcal{L}}^{-1}}
% discrete fourier transform
\newcommand{\DFT}{\operatorname{DFT}}
\newcommand{\IDFT}{\operatorname{IDFT}}
% Z-transform
\newcommand{\ZTrans}{\operatorname{\mathcal{Z}}}
\newcommand{\IZTrans}{\operatorname{\mathcal{Z}}^{-1}}
\newcommand{\SSZTrans}{\underset{\text{s.s.}}{\mathcal{Z}}} % single-sided Z-transform
% linear algebra
\newcommand{\bm}[1]{{\boldsymbol{#1}}}
\newcommand{\vecEntry}[2]{\bm{#1}\bracks*{#2}}
\newcommand{\matEntry}[3]{#1\bracks*{#2}\bracks*{#3}}
\newcommand{\matPart}[5]{\matEntry{#1}{#2:#3}{#4:#5}}
\newcommand{\minor}[3]{{#1}\bracks*{\setminus #2}\bracks*{\setminus #3}} % the minor of a matrix: row #2 and column #3 deleted
\newcommand{\diag}{\operatorname{diag}}
\newcommand{\transpose}[1]{{#1}^\top}
\newcommand{\HerConj}[1]{{#1}^*}
\newcommand{\tr}{\operatorname{tr}}
\newcommand{\inProd}[2]{\left\langle#1,#2\right\rangle}
\newcommand{\dotProd}[2]{#1 \cdot #2}
\newcommand{\HadamardProd}{\odot}
\newcommand{\HadamardDiv}{\oslash}
\newcommand{\vecSpan}{\operatorname{span}}
\newcommand{\rank}{\operatorname{rank}}
% vector
% unit vector
\newcommand{\vix}{\bm{i}_x}
\newcommand{\viy}{\bm{i}_y}
\newcommand{\viz}{\bm{i}_z}
% graph theory
\newcommand{\neighborhood}{\mathcal{N}}
% probability theory
\newcommand{\PDF}{\operatorname{PDF}}
\newcommand{\Ber}{\operatorname{Ber}}
\newcommand{\Beta}{\operatorname{Beta}}
\newcommand{\ExpDist}{\operatorname{ExpDist}}
\newcommand{\ErlangDist}{\operatorname{ErlangDist}}
\newcommand{\PoissonDist}{\operatorname{PoissonDist}}
\newcommand{\GammaDist}{\operatorname{Gamma}}
\newcommand{\cind}[2]{\ind{#1\left| #2\right.}} % conditional indicator function
\renewcommand{\Pr}{\operatorname{Pr}}
\DeclarePairedDelimiterX{\cPrParens}[2]{(}{)}{#1\,\delimsize\vert\,#2}
\newcommand{\Ev}{\operatorname{E}} % expected value
\newcommand{\Var}{\operatorname{Var}}
\newcommand{\Cov}{\operatorname{Cov}}
% physics
% unit of measurement
\newcommand{\second}{\text{s}}
\newcommand{\hertz}{\text{Hz}}
\newcommand{\decibel}{\text{dB}}
% signal processing
% Discrete Time Fourier Transform
\newcommand{\DTFT}{\operatorname{DTFT}}
\newcommand{\IDTFT}{\operatorname{IDTFT}}
% computer science
% fixed-point arithmetic
\newcommand{\IntPartBW}[1]{\underbracket[0.140ex]{#1}_\mathrm{i}} % 0.140ex is half of the default thickness. See: [How to make underbracket thinner](https://tex.stackexchange.com/questions/559078/how-to-make-underbracket-thinner)
\newcommand{\DecPartBW}[1]{\underbracket[0.140ex]{#1}_\mathrm{d}}
\newcommand{\TotalBW}[1]{\underbracket[0.140ex]{#1}}
\newcommand{\Rat}{\operatorname{Rat}} % Maps a fixed-point number to a corresponding rational number.
% programming
\newcommand{\plpl}{\mathrel{++}}
\newcommand{\pleq}{\mathrel{+}=}
\newcommand{\asteq}{\mathrel{*}=}
\]
補題1(行列式の中身のサイズ変更)
主張
$A\in\realNumbers^{n\times m}, B\in\realNumbers^{m\times n}$に対して$|I_m-AB| = |I_n-BA|$
導出
\[ M \coloneqq \begin{bmatrix} I_n & A \\ B & I_m \end{bmatrix} \]
とすると
\[
\det{M} = \det{
\left(M
\begin{bmatrix}
I_n & -A \\
O & I_m
\end{bmatrix}
\right)
} = \det{
\begin{bmatrix}
I_n & O \\
B & I_m-BA
\end{bmatrix}
} = |I_m-BA|
\]
一方で
\[
\det{M} = \det{
\left(M
\begin{bmatrix}
I_n & O \\
-B & I_m
\end{bmatrix}
\right)
} = \det{
\begin{bmatrix}
I_n-AB & A \\
~ & I_m
\end{bmatrix}
} = |I_n-AB|
\]
$\square$
補題2(Sherman-Morrisonの公式の特別な場合)
主張
$\bm{u},\bm{v} \in \complexNumbers^m,\; 1+\bm{v}^*\bm{u} \neq 0$とするとき次式が成り立つ。
\[ (I+\bm{u}\bm{v}^*)^{-1} = I – \frac{1}{1+\bm{v}^*\bm{u}}\bm{u}\bm{v}^* \]
導出
$\bm{u}\bm{v}^* = O$のときは定理の主張は明らかに成り立つ。以下では$\bm{u}\bm{v}^* \neq O$とする。仮定$1+\bm{v}^*\bm{u} \neq 0$と前掲の補題1より$I+\bm{u}\bm{v}^*$には逆行列が存在する。$\bm{u}\bm{v}^*$の階数は1であり($\Img{\bm{u}\bm{v}^*} = \Span{\bm{u}}$)、唯一の固有値は$\bm{v}^*\bm{u}$である(対応する固有ベクトルは$\bm{u}$)。よって$\bm{u}\bm{v}^*$をJordan分解すると、適当な正則行列$J$を用いて次のように表せる。
\[
\bm{u}\bm{v}^* = J
\begin{bmatrix}
\bm{v}^*\bm{u} & \bm{0}^\top \\
\bm{0} & O
\end{bmatrix}
J^{-1} \eqqcolon J\Lambda J^{-1}
\]
よって
\begin{align*}
(I+\bm{u}\bm{v}^*)^{-1} &= J(I + \Lambda)^{-1}J^{-1} = J
\begin{bmatrix}
1+\bm{v}^*\bm{u} & \bm{0}^\top \\
\bm{0} & I
\end{bmatrix}
^{-1}J^{-1} = J
\begin{bmatrix}
\frac{1}{1+\bm{v}^*\bm{u}} & \bm{0}^\top \\
\bm{0} & I
\end{bmatrix}
J^{-1} \\
&= J\frac{1}{1+\bm{v}^*\bm{u}}
\begin{bmatrix}
1 & \bm{0}^\top \\
\bm{0} & (1+\bm{v}^*\bm{u})I
\end{bmatrix}
J^{-1} = J\frac{1}{1+\bm{v}^*\bm{u}}((1+\bm{v}^*\bm{u})I – \Lambda)J^{-1} \\
&= I – \frac{1}{1+\bm{v}^*\bm{u}}\bm{u}\bm{v}^*
\end{align*}
$\square$
Sherman-Morrisonの公式
主張
$A \in \complexNumbers^{m\times m}, \bm{u},\bm{v} \in \complexNumbers^m$とする。$A$は可逆であるとし、$1+\bm{v}^*A^{-1}\bm{u} \neq 0$とする。このとき次式が成り立つ。
\[ (A+\bm{u}\bm{v}^*)^{-1} = \left( I – \frac{1}{1+\bm{v}^*A^{-1}\bm{u}}A^{-1}\bm{u}\bm{v}^* \right)A^{-1} \]
導出
$A+\bm{u}\bm{v}^* = A(I + A^{-1}\bm{u}\bm{v}^*)$だから$(A+\bm{u}\bm{v}^*)^{-1} = (I + A^{-1}\bm{u}\bm{v}^*)^{-1}A^{-1}$であり、1つ目の因子に前掲の補題2を適用すればよい。
$\square$
コメントを残す
コメントを投稿するにはログインしてください。