Gauss-Seidel法の収束の十分条件

はじめに

Gauss-Seidel法が厳密解へ収束する主要な十分条件の一つに、係数行列が狭義優対角であることが知られている。腕試しに証明を試みたらできたので書き残しておく。

\[ % to cope with definition mismatch between MathJax and LaTeX \newcommand{\coloneq}{\mathrel{:=}} \newcommand{\eqcolon}{\mathrel{=:}} % general purpose \newcommand{\ctext}[1]{\raise0.2ex\hbox{\textcircled{\scriptsize{#1}}}} % mathematics % general purpose \DeclarePairedDelimiterX{\parens}[1]{\lparen}{\rparen}{#1} \DeclarePairedDelimiterX{\braces}[1]{\lbrace}{\rbrace}{#1} \DeclarePairedDelimiterX{\bracks}[1]{\lbrack}{\rbrack}{#1} \DeclarePairedDelimiterX{\verts}[1]{|}{|}{#1} \DeclarePairedDelimiterX{\Verts}[1]{\|}{\|}{#1} \DeclarePairedDelimiterX{\setComprehension}[2]{\lbrace}{\rbrace}{#1\,\delimsize\vert\,#2} \newcommand{\as}{{\quad\textrm{as}\quad}} \newcommand{\st}{{\textrm{ s.t. }}} \newcommand{\naturalNumbers}{\mathbb{N}} \newcommand{\integers}{\mathbb{Z}} \newcommand{\rationalNumbers}{\mathbb{Q}} \newcommand{\realNumbers}{\mathbb{R}} \newcommand{\nonNegRealNumbers}{\mathbb{R}_{\geq 0}} \newcommand{\posRealNumbers}{\mathbb{R}_{> 0}} \newcommand{\complexNumbers}{\mathbb{C}} \newcommand{\field}{\mathbb{F}} \newcommand{\EuclideanSpace}{\mathbb{E}} \newcommand{\argmax}{\operatorname*{arg~max}} \newcommand{\argmin}{\operatorname*{arg~min}} % set theory \newcommand{\range}[2]{\braces*{#1,\dotsc,#2}} \renewcommand{\complement}{\mathrm{c}} \newcommand{\ind}[2]{\mathbbm{1}_{#1}\parens*{#2}} \newcommand{\indII}[1]{\mathbbm{1}\braces*{#1}} % number theory \newcommand{\abs}[1]{\verts*{#1}} \newcommand{\combi}[2]{{_{#1}\mathrm{C}_{#2}}} \newcommand{\perm}[2]{{_{#1}\mathrm{P}_{#2}}} \newcommand{\GaloisField}{\mathrm{GF}} % real analysis \newcommand{\NapierE}{\mathrm{e}} \newcommand{\sgn}{\operatorname{sgn}} % sign function \newcommand{\rect}{\operatorname{rect}} % rectangular function \newcommand{\cl}{\operatorname{cl}} % closure of a set \newcommand{\img}{\operatorname{im}} % image of a function \newcommand{\dom}{\operatorname{dom}} % domain of a function \newcommand{\LittleO}[2]{\underset{#1}{o}\parens*{#2}} % little-o notation \newcommand{\norm}[1]{\Verts*{#1}} % norm of a vector or function \newcommand{\floor}[1]{\left\lfloor #1\right\rfloor} \newcommand{\ceil}[1]{\left\lceil#1\right\rceil} \newcommand{\sinc}{\operatorname{sinc}} \newcommand{\nrmSinc}{\operatorname{nsinc}} % normalized sinc function \newcommand{\erf}{\operatorname{erf}} % inverse trigonometric functions \newcommand{\asin}{\operatorname{Sin}^{-1}} \newcommand{\acos}{\operatorname{Cos}^{-1}} \newcommand{\atan}{\operatorname{Tan}^{-1}} % derivative \newcommand{\deriv}[3]{\frac{\mathrm{d}^{#3}#1}{\mathrm{d}{#2}^{#3}}} \newcommand{\derivLong}[3]{\frac{\mathrm{d}^{#3}}{\mathrm{d}{#2}^{#3}}#1} \newcommand{\partDeriv}[3]{\frac{\mathrm{\partial}^{#3}#1}{\mathrm{\partial}{#2}^{#3}}} \newcommand{\partDerivLong}[3]{\frac{\mathrm{\partial}^{#3}}{\mathrm{\partial}{#2}^{#3}}#1} \newcommand{\partDerivIIHetero}[3]{\frac{\mathrm{\partial}^2#1}{\partial#2\mathrm{\partial}#3}} \newcommand{\partDerivIIHeteroLong}[3]{{\frac{\mathrm{\partial}^2}{\partial#2\mathrm{\partial}#3}#1}} % integral \newcommand{\pv}{{\textrm{ p.v. }}} \newcommand{\integrate}[5]{\int_{#1}^{#2}{#3}{\;\mathrm{d}^{#4}}#5} \newcommand{\LebInteg}[4]{\int_{#1} {#2} {#3}\parens*{\;\mathrm{d}#4}} % complex analysis \newcommand{\conj}[1]{\overline{#1}} \renewcommand{\Re}{\operatorname{Re}} \renewcommand{\Im}{\operatorname{Im}} \newcommand{\Arg}{\operatorname{Arg}} \newcommand{\Log}{\operatorname{Log}} % Laplace transform \newcommand{\LPLC}{\operatorname{\mathcal{L}}} \newcommand{\ILPLC}{\operatorname{\mathcal{L}}^{-1}} % discrete fourier transform \newcommand{\DFT}{\operatorname{DFT}} \newcommand{\IDFT}{\operatorname{IDFT}} % Z-transform \newcommand{\ZTrans}{\operatorname{\mathcal{Z}}} \newcommand{\IZTrans}{\operatorname{\mathcal{Z}}^{-1}} \newcommand{\SSZTrans}{\underset{\text{s.s.}}{\mathcal{Z}}} % single-sided Z-transform % linear algebra \newcommand{\bm}[1]{{\boldsymbol{#1}}} \newcommand{\vecEntry}[2]{\bm{#1}\bracks*{#2}} \newcommand{\matEntry}[3]{#1\bracks*{#2}\bracks*{#3}} \newcommand{\matPart}[5]{\matEntry{#1}{#2:#3}{#4:#5}} \newcommand{\minor}[3]{{#1}\bracks*{\setminus #2}\bracks*{\setminus #3}} % the minor of a matrix: row #2 and column #3 deleted \newcommand{\diag}{\operatorname{diag}} \newcommand{\transpose}[1]{{#1}^\top} \newcommand{\HerConj}[1]{{#1}^*} \newcommand{\tr}{\operatorname{tr}} \newcommand{\inProd}[2]{\left\langle#1,#2\right\rangle} \newcommand{\dotProd}[2]{#1 \cdot #2} \newcommand{\HadamardProd}{\odot} \newcommand{\HadamardDiv}{\oslash} \newcommand{\vecSpan}{\operatorname{span}} \newcommand{\rank}{\operatorname{rank}} % vector % unit vector \newcommand{\vix}{\bm{i}_x} \newcommand{\viy}{\bm{i}_y} \newcommand{\viz}{\bm{i}_z} % graph theory \newcommand{\neighborhood}{\mathcal{N}} % probability theory \newcommand{\PDF}{\operatorname{PDF}} \newcommand{\Ber}{\operatorname{Ber}} \newcommand{\Beta}{\operatorname{Beta}} \newcommand{\ExpDist}{\operatorname{ExpDist}} \newcommand{\ErlangDist}{\operatorname{ErlangDist}} \newcommand{\PoissonDist}{\operatorname{PoissonDist}} \newcommand{\GammaDist}{\operatorname{Gamma}} \newcommand{\cind}[2]{\ind{#1\left| #2\right.}} % conditional indicator function \renewcommand{\Pr}{\operatorname{Pr}} \DeclarePairedDelimiterX{\cPrParens}[2]{(}{)}{#1\,\delimsize\vert\,#2} \newcommand{\Ev}{\operatorname{E}} % expected value \newcommand{\Var}{\operatorname{Var}} \newcommand{\Cov}{\operatorname{Cov}} % physics % unit of measurement \newcommand{\second}{\text{s}} \newcommand{\hertz}{\text{Hz}} \newcommand{\decibel}{\text{dB}} % signal processing % Discrete Time Fourier Transform \newcommand{\DTFT}{\operatorname{DTFT}} \newcommand{\IDTFT}{\operatorname{IDTFT}} % computer science % fixed-point arithmetic \newcommand{\IntPartBW}[1]{\underbracket[0.140ex]{#1}_\mathrm{i}} % 0.140ex is half of the default thickness. See: [How to make underbracket thinner](https://tex.stackexchange.com/questions/559078/how-to-make-underbracket-thinner) \newcommand{\DecPartBW}[1]{\underbracket[0.140ex]{#1}_\mathrm{d}} \newcommand{\TotalBW}[1]{\underbracket[0.140ex]{#1}} \newcommand{\Rat}{\operatorname{Rat}} % Maps a fixed-point number to a corresponding rational number. % programming \newcommand{\plpl}{\mathrel{++}} \newcommand{\pleq}{\mathrel{+}=} \newcommand{\asteq}{\mathrel{*}=} \]

Gauss-Seidel法

以下に述べる定義はWikipediaの英語記事“Gauss Seidel method”からの引用である。

$n\in\naturalNumbers,\;A\in\complexNumbers^{n\times n},\;\bm{b}\in\complexNumbers^n$ とする。$A$ は正定値対称、または狭義優対角であるとする。Gauss-Seidel法とは、線型方程式 $A\bm{x}=\bm{b}$ の解を求める反復法である。$\bm{x}_1\in\complexNumbers^n$ を任意の初期解とし、次の漸化式で解候補を更新してゆく。

\[ L_* \bm{x}_{k+1} = -U\bm{x}_k \quad (k=1,2,\dots)\]

ここに $L_*$ は $A$ の対角成分およびその下側の要素からなる下三角行列であり、 $U$ は $A$ の対角成分の上側の要素からなる上三角行列である。

係数行列が狭義優対角ならば厳密解に収束すること

$\textit{Proof}$

$A$ の次数を $n$ とする。$\mathring{\bm{x}}$ を厳密解とすると $L_* \mathring{\bm{x}} = \bm{b} – U\mathring{\bm{x}}$ である。これを解の更新式から減じると次式を得る。

\[ L_* (\bm{x}_{k+1} – \mathring{\bm{x}}) = -U(\bm{x}_k – \mathring{\bm{x}}) \tag{1} \]

$\bm{v}_k \coloneqq \bm{x}_k – \mathring{\bm{x}}$ とおくと、式 (1) より次式が成り立つ。

\[ L_* \bm{v}_{k+1} = -U\bm{v}_k \tag{2} \]

$M_k \coloneqq \max_{i=1,\dots,n} |v_{k,i}|\;(v_{k,i}$ は $\bm{v}_k$ の第 $i$ 要素)とする。次の2つが同時に成り立つことが、$\bm{v}_k$ が $\bm{0}_n$ に収束するための十分条件である。

  1. ある $k \in \mathbb{N}$ に対して $M_k = 0$ ならば $M_l = 0\;(l=k+1,k+2,\dots)$
  2. 適当な $0 < \alpha < 1$ が存在して $M_k > 0 \Rightarrow M_{k+1} < \alpha M_k$

$L_*$ が正則であることと式 (2) より直ちに 1. が成り立つ。次に 2. を数学的帰納法で示す。$\tilde{\alpha}$ を次式で定義する。

\[ \tilde{\alpha} \coloneqq \min_{i=1,2,\dots,n} \frac{1}{|a_{i,i}|} \sum_{j=1,\dots,n \wedge j\neq i} |a_{i,j}| \]

$A$ は優対角だから $0 < \tilde{\alpha} < 1$ である。

\[ \begin{align*} a_{1,1} v_{k+1, 1} &= -\sum_{j=2}^n a_{1,j}v_{k,j} \\ |a_{1,1}| |v_{k+1, 1}| &= \left|\sum_{j=2}^n a_{1,j}v_{k,j}\right| \leq \sum_{j=2}^n |a_{1,j}||v_{k,j}| \leq M_k \sum_{j=2}^n |a_{1,j}| \\ |v_{k+1, 1}| &\leq \frac{M_k}{|a_{1,1}|} \sum_{j=2}^n |a_{1,j}| \leq \tilde{\alpha}M_k \end{align*} \]

$|v_{k+1, j}| \leq \tilde{\alpha} M_k\;(j=1,2,\dots,l)\;(l\in\{1,2,\dots,n-1\})$ ならば $|v_{k+1, l+1}| \leq \tilde{\alpha} M_k$ であることを示す。式 (2) の $l+1$ 行目を展開すると次式を得る。

\[ \begin{align*} \sum_{j=1}^{l+1} a_{l+1,j} v_{k+1, j} &= -\sum_{j=l+2}^n a_{l+1,j}v_{k,j} \\ a_{l+1,l+1} v_{k+1, l+1} &= -\sum_{j=1}^l a_{l+1,j} v_{k+1, j} – \sum_{j=l+2}^n a_{l+1,j}v_{k,j} \\ |a_{l+1,l+1}| |v_{k+1, l+1}| &= \left|-\sum_{j=1}^l a_{l+1,j} v_{k+1, j} – \sum_{j=l+2}^n a_{l+1,j}v_{k,j}\right| \leq \sum_{j=1}^l |a_{l+1,j}||v_{k+1, j}| + \sum_{j=l+2}^n |a_{l+1,j}||v_{k,j}| \\ &\leq M_k \sum_{j=1,\dots,n \wedge j\neq l+1} |a_{l+1,j}| \\ |v_{k+1, l+1}| &\leq \frac{M_k}{|a_{l+1,l+1}|} \sum_{j=1,\dots,n \wedge j\neq l+1} |a_{l+1,j}| \leq \tilde{\alpha} M_k \end{align*} \]

以上より帰納的に $|v_{k+1,j}| \leq \tilde{\alpha} M_k\;(j=1,2,\dots,n)$ が成り立つ。すなわち $M_{k+1} \leq \tilde{\alpha} M_k$ が成り立つ。$\tilde{\alpha} < \alpha < 1$ となるように $\alpha $ を定めることで 2. が示される。

$\square$

コメントを残す