Sherman-Morrisonの公式の導出

Kalman-filterをはじめとする、逐次的な2乗誤差最小化に基づくリアルタイム信号処理では、損失関数を構成する2次形式の基になる行列の逆行列を高速で求める必要がある。そのような状況で頼りになるSherman-Morrisonの公式を導出する。Woodburyの公式の特別な場合として片付けられるのが一般的だが、本記事ではJordan分解を用いて導出してみる。

\[ % to cope with definition mismatch between MathJax and LaTeX \newcommand{\coloneq}{\mathrel{:=}} \newcommand{\eqcolon}{\mathrel{=:}} % general purpose \newcommand{\ctext}[1]{\raise0.2ex\hbox{\textcircled{\scriptsize{#1}}}} % mathematics % general purpose \DeclarePairedDelimiterX{\parens}[1]{\lparen}{\rparen}{#1} \DeclarePairedDelimiterX{\braces}[1]{\lbrace}{\rbrace}{#1} \DeclarePairedDelimiterX{\bracks}[1]{\lbrack}{\rbrack}{#1} \DeclarePairedDelimiterX{\verts}[1]{|}{|}{#1} \DeclarePairedDelimiterX{\Verts}[1]{\|}{\|}{#1} \DeclarePairedDelimiterX{\setComprehension}[2]{\lbrace}{\rbrace}{#1\,\delimsize\vert\,#2} \newcommand{\as}{{\quad\textrm{as}\quad}} \newcommand{\st}{{\textrm{ s.t. }}} \newcommand{\naturalNumbers}{\mathbb{N}} \newcommand{\integers}{\mathbb{Z}} \newcommand{\rationalNumbers}{\mathbb{Q}} \newcommand{\realNumbers}{\mathbb{R}} \newcommand{\nonNegRealNumbers}{\mathbb{R}_{\geq 0}} \newcommand{\posRealNumbers}{\mathbb{R}_{> 0}} \newcommand{\complexNumbers}{\mathbb{C}} \newcommand{\field}{\mathbb{F}} \newcommand{\EuclideanSpace}{\mathbb{E}} \newcommand{\argmax}{\operatorname*{arg~max}} \newcommand{\argmin}{\operatorname*{arg~min}} % set theory \newcommand{\range}[2]{\braces*{#1,\dotsc,#2}} \renewcommand{\complement}{\mathrm{c}} \newcommand{\ind}[2]{\mathbbm{1}_{#1}\parens*{#2}} \newcommand{\indII}[1]{\mathbbm{1}\braces*{#1}} % number theory \newcommand{\abs}[1]{\verts*{#1}} \newcommand{\combi}[2]{{_{#1}\mathrm{C}_{#2}}} \newcommand{\perm}[2]{{_{#1}\mathrm{P}_{#2}}} \newcommand{\GaloisField}{\mathrm{GF}} % real analysis \newcommand{\NapierE}{\mathrm{e}} \newcommand{\sgn}{\operatorname{sgn}} % sign function \newcommand{\rect}{\operatorname{rect}} % rectangular function \newcommand{\cl}{\operatorname{cl}} % closure of a set \newcommand{\img}{\operatorname{im}} % image of a function \newcommand{\dom}{\operatorname{dom}} % domain of a function \newcommand{\LittleO}[2]{\underset{#1}{o}\parens*{#2}} % little-o notation \newcommand{\norm}[1]{\Verts*{#1}} % norm of a vector or function \newcommand{\floor}[1]{\left\lfloor #1\right\rfloor} \newcommand{\ceil}[1]{\left\lceil#1\right\rceil} \newcommand{\sinc}{\operatorname{sinc}} \newcommand{\nrmSinc}{\operatorname{nsinc}} % normalized sinc function \newcommand{\erf}{\operatorname{erf}} % inverse trigonometric functions \newcommand{\asin}{\operatorname{Sin}^{-1}} \newcommand{\acos}{\operatorname{Cos}^{-1}} \newcommand{\atan}{\operatorname{Tan}^{-1}} % derivative \newcommand{\deriv}[3]{\frac{\mathrm{d}^{#3}#1}{\mathrm{d}{#2}^{#3}}} \newcommand{\derivLong}[3]{\frac{\mathrm{d}^{#3}}{\mathrm{d}{#2}^{#3}}#1} \newcommand{\partDeriv}[3]{\frac{\mathrm{\partial}^{#3}#1}{\mathrm{\partial}{#2}^{#3}}} \newcommand{\partDerivLong}[3]{\frac{\mathrm{\partial}^{#3}}{\mathrm{\partial}{#2}^{#3}}#1} \newcommand{\partDerivIIHetero}[3]{\frac{\mathrm{\partial}^2#1}{\partial#2\mathrm{\partial}#3}} \newcommand{\partDerivIIHeteroLong}[3]{{\frac{\mathrm{\partial}^2}{\partial#2\mathrm{\partial}#3}#1}} % integral \newcommand{\pv}{{\textrm{ p.v. }}} \newcommand{\integrate}[5]{\int_{#1}^{#2}{#3}{\;\mathrm{d}^{#4}}#5} \newcommand{\LebInteg}[4]{\int_{#1} {#2} {#3}\parens*{\;\mathrm{d}#4}} % complex analysis \newcommand{\conj}[1]{\overline{#1}} \renewcommand{\Re}{\operatorname{Re}} \renewcommand{\Im}{\operatorname{Im}} \newcommand{\Arg}{\operatorname{Arg}} \newcommand{\Log}{\operatorname{Log}} % Laplace transform \newcommand{\LPLC}{\operatorname{\mathcal{L}}} \newcommand{\ILPLC}{\operatorname{\mathcal{L}}^{-1}} % discrete fourier transform \newcommand{\DFT}{\operatorname{DFT}} \newcommand{\IDFT}{\operatorname{IDFT}} % Z-transform \newcommand{\ZTrans}{\operatorname{\mathcal{Z}}} \newcommand{\IZTrans}{\operatorname{\mathcal{Z}}^{-1}} \newcommand{\SSZTrans}{\underset{\text{s.s.}}{\mathcal{Z}}} % single-sided Z-transform % linear algebra \newcommand{\bm}[1]{{\boldsymbol{#1}}} \newcommand{\vecEntry}[2]{\bm{#1}\bracks*{#2}} \newcommand{\matEntry}[3]{#1\bracks*{#2}\bracks*{#3}} \newcommand{\matPart}[5]{\matEntry{#1}{#2:#3}{#4:#5}} \newcommand{\minor}[3]{{#1}\bracks*{\setminus #2}\bracks*{\setminus #3}} % the minor of a matrix: row #2 and column #3 deleted \newcommand{\diag}{\operatorname{diag}} \newcommand{\transpose}[1]{{#1}^\top} \newcommand{\HerConj}[1]{{#1}^*} \newcommand{\tr}{\operatorname{tr}} \newcommand{\inProd}[2]{\left\langle#1,#2\right\rangle} \newcommand{\dotProd}[2]{#1 \cdot #2} \newcommand{\HadamardProd}{\odot} \newcommand{\HadamardDiv}{\oslash} \newcommand{\vecSpan}{\operatorname{span}} \newcommand{\rank}{\operatorname{rank}} % vector % unit vector \newcommand{\vix}{\bm{i}_x} \newcommand{\viy}{\bm{i}_y} \newcommand{\viz}{\bm{i}_z} % graph theory \newcommand{\neighborhood}{\mathcal{N}} % probability theory \newcommand{\PDF}{\operatorname{PDF}} \newcommand{\Ber}{\operatorname{Ber}} \newcommand{\Beta}{\operatorname{Beta}} \newcommand{\ExpDist}{\operatorname{ExpDist}} \newcommand{\ErlangDist}{\operatorname{ErlangDist}} \newcommand{\PoissonDist}{\operatorname{PoissonDist}} \newcommand{\GammaDist}{\operatorname{Gamma}} \newcommand{\cind}[2]{\ind{#1\left| #2\right.}} % conditional indicator function \renewcommand{\Pr}{\operatorname{Pr}} \DeclarePairedDelimiterX{\cPrParens}[2]{(}{)}{#1\,\delimsize\vert\,#2} \newcommand{\Ev}{\operatorname{E}} % expected value \newcommand{\Var}{\operatorname{Var}} \newcommand{\Cov}{\operatorname{Cov}} % physics % unit of measurement \newcommand{\second}{\text{s}} \newcommand{\hertz}{\text{Hz}} \newcommand{\decibel}{\text{dB}} % signal processing % Discrete Time Fourier Transform \newcommand{\DTFT}{\operatorname{DTFT}} \newcommand{\IDTFT}{\operatorname{IDTFT}} % computer science % fixed-point arithmetic \newcommand{\IntPartBW}[1]{\underbracket[0.140ex]{#1}_\mathrm{i}} % 0.140ex is half of the default thickness. See: [How to make underbracket thinner](https://tex.stackexchange.com/questions/559078/how-to-make-underbracket-thinner) \newcommand{\DecPartBW}[1]{\underbracket[0.140ex]{#1}_\mathrm{d}} \newcommand{\TotalBW}[1]{\underbracket[0.140ex]{#1}} \newcommand{\Rat}{\operatorname{Rat}} % Maps a fixed-point number to a corresponding rational number. % programming \newcommand{\plpl}{\mathrel{++}} \newcommand{\pleq}{\mathrel{+}=} \newcommand{\asteq}{\mathrel{*}=} \]

補題1(行列式の中身のサイズ変更)

主張

$A\in\realNumbers^{n\times m}, B\in\realNumbers^{m\times n}$に対して$|I_m-AB| = |I_n-BA|$

導出

\[ M \coloneqq \begin{bmatrix} I_n & A \\ B & I_m \end{bmatrix} \]

とすると

\[ \det{M} = \det{ \left(M \begin{bmatrix} I_n & -A \\ O & I_m \end{bmatrix} \right) } = \det{ \begin{bmatrix} I_n & O \\ B & I_m-BA \end{bmatrix} } = |I_m-BA| \]

一方で

\[ \det{M} = \det{ \left(M \begin{bmatrix} I_n & O \\ -B & I_m \end{bmatrix} \right) } = \det{ \begin{bmatrix} I_n-AB & A \\ ~ & I_m \end{bmatrix} } = |I_n-AB| \]

$\square$

補題2(Sherman-Morrisonの公式の特別な場合)

主張

$\bm{u},\bm{v} \in \complexNumbers^m,\; 1+\bm{v}^*\bm{u} \neq 0$とするとき次式が成り立つ。

\[ (I+\bm{u}\bm{v}^*)^{-1} = I – \frac{1}{1+\bm{v}^*\bm{u}}\bm{u}\bm{v}^* \]

導出

$\bm{u}\bm{v}^* = O$のときは定理の主張は明らかに成り立つ。以下では$\bm{u}\bm{v}^* \neq O$とする。仮定$1+\bm{v}^*\bm{u} \neq 0$と前掲の補題1より$I+\bm{u}\bm{v}^*$には逆行列が存在する。$\bm{u}\bm{v}^*$の階数は1であり($\Img{\bm{u}\bm{v}^*} = \Span{\bm{u}}$)、唯一の固有値は$\bm{v}^*\bm{u}$である(対応する固有ベクトルは$\bm{u}$)。よって$\bm{u}\bm{v}^*$をJordan分解すると、適当な正則行列$J$を用いて次のように表せる。

\[ \bm{u}\bm{v}^* = J \begin{bmatrix} \bm{v}^*\bm{u} & \bm{0}^\top \\ \bm{0} & O \end{bmatrix} J^{-1} \eqqcolon J\Lambda J^{-1} \]

よって

\begin{align*} (I+\bm{u}\bm{v}^*)^{-1} &= J(I + \Lambda)^{-1}J^{-1} = J \begin{bmatrix} 1+\bm{v}^*\bm{u} & \bm{0}^\top \\ \bm{0} & I \end{bmatrix} ^{-1}J^{-1} = J \begin{bmatrix} \frac{1}{1+\bm{v}^*\bm{u}} & \bm{0}^\top \\ \bm{0} & I \end{bmatrix} J^{-1} \\ &= J\frac{1}{1+\bm{v}^*\bm{u}} \begin{bmatrix} 1 & \bm{0}^\top \\ \bm{0} & (1+\bm{v}^*\bm{u})I \end{bmatrix} J^{-1} = J\frac{1}{1+\bm{v}^*\bm{u}}((1+\bm{v}^*\bm{u})I – \Lambda)J^{-1} \\ &= I – \frac{1}{1+\bm{v}^*\bm{u}}\bm{u}\bm{v}^* \end{align*}

$\square$

Sherman-Morrisonの公式

主張

$A \in \complexNumbers^{m\times m}, \bm{u},\bm{v} \in \complexNumbers^m$とする。$A$は可逆であるとし、$1+\bm{v}^*A^{-1}\bm{u} \neq 0$とする。このとき次式が成り立つ。

\[ (A+\bm{u}\bm{v}^*)^{-1} = \left( I – \frac{1}{1+\bm{v}^*A^{-1}\bm{u}}A^{-1}\bm{u}\bm{v}^* \right)A^{-1} \]

導出

$A+\bm{u}\bm{v}^* = A(I + A^{-1}\bm{u}\bm{v}^*)$だから$(A+\bm{u}\bm{v}^*)^{-1} = (I + A^{-1}\bm{u}\bm{v}^*)^{-1}A^{-1}$であり、1つ目の因子に前掲の補題2を適用すればよい。

$\square$

コメントを残す