Hình thức ma trận của backpropagation với chuẩn hóa hàng loạt


12

Chuẩn hóa hàng loạt đã được ghi nhận với những cải tiến hiệu suất đáng kể trong mạng lưới thần kinh sâu. Rất nhiều tài liệu trên internet cho thấy cách triển khai nó trên cơ sở kích hoạt bằng cách kích hoạt. Tôi đã triển khai backprop bằng cách sử dụng đại số ma trận và cho rằng tôi đang làm việc với các ngôn ngữ cấp cao (trong khi dựa vào Rcpp(và cuối cùng là GPU) để nhân ma trận dày đặc), trích xuất mọi thứ và sử dụng formã nguồn có thể sẽ làm chậm mã của tôi thực chất, ngoài việc là một nỗi đau rất lớn

Chức năng batch bình thường là

b(xp)=γ(xp−μxp)σxp−1+β
nơi
  • xp lànút thứp , trước khi nó được kích hoạt
  • γ vàβ là các thông số vô hướng
  • μxp vàσxp là giá trị trung bình và SD củaxp . (Lưu ý rằng căn bậc hai của phương sai cộng với hệ số mờ thường được sử dụng - giả sử các phần tử khác không cho độ gọn)

Ở dạng ma trận, hàng loạt bình thường cho toàn bộ một lớp sẽ là nơi

b(X)=(γ⊗1p)⊙(X−μX)⊙σX−1+(β⊗1p)
  • là N × pXN×p
  • là một vectơ cột của những cái1N
  • và β hiện nay có hàng p -vectors các thông số bình thường mỗi lớpγβp
  • và σ X là N × p ma trận, trong đó mỗi cột là một N -vector phương tiện theo cột và độ lệch chuẩnμXσXN×pN
  • là sản phẩm Kronecker và ⊙ là elementwise (Hadamard) Sản phẩm⊗⊙

Một một lớp lưới thần kinh rất đơn giản với không bình thường hàng loạt và một kết quả liên tục là

y=a(XΓ1)Γ2+ϵ

Ở đâu

  • là p 1 × p 2Γ1p1×p2
  • là p 2 × 1Γ2p2×1
  • là hàm kích hoạta(.)

Nếu sự mất mát là , sau đó các gradient là ∂ RR=N−1∑(y−y^)2

∂R∂Γ1=−2VTϵ^∂R∂Γ2=XT(a′(XΓ1)⊙−2ϵ^Γ2T)

Ở đâu

  • V=a(XΓ1)
  • ϵ^=y−y^

Dưới bình thường hàng loạt, net trở thành hoặc y = một ( ( gamma ⊗ 1 N ) ⊙ ( X Γ 1 - μ X Γ 1 ) ⊙ σ - 1 X Γ 1 + ( beta ⊗ 1 N ) ) Γ 2

y=a(b(XΓ1))Γ2
y=a((γ⊗1N)⊙(XΓ1−μXΓ1)⊙σXΓ1−1+(β⊗1N))Γ2
Tôi không biết làm thế nào để tính toán các dẫn xuất của các sản phẩm Hadamard và Kronecker. Về chủ đề của các sản phẩm Kronecker, tài liệu trở nên khá phức tạp.

Có cách nào thực tế của máy tính , ∂ R / ∂ beta , và ∂ R / ∂ gamma 1 trong khuôn khổ ma trận? Một biểu thức đơn giản, không dùng đến tính toán theo nút?∂R/∂γ∂R/∂β∂R/∂Γ1

Cập nhật 1:

∂R/∂β

1NT(a′(XΓ1)⊙−2ϵ^Γ2T)
set.seed(1)
library(dplyr)
library(foreach)

#numbers of obs, variables, and hidden layers
N <- 10
p1 <- 7
p2 <- 4
a <- function (v) {
  v[v < 0] <- 0
  v
}
ap <- function (v) {
  v[v < 0] <- 0
  v[v >= 0] <- 1
  v
}

# parameters
G1 <- matrix(rnorm(p1*p2), nrow = p1)
G2 <- rnorm(p2)
gamma <- 1:p2+1
beta <- (1:p2+1)*-1
# error
u <- rnorm(10)

# matrix batch norm function
b <- function(x, bet = beta, gam = gamma){
  xs <- scale(x)
  gk <- t(matrix(gam)) %x% matrix(rep(1, N))
  bk <- t(matrix(bet)) %x% matrix(rep(1, N))
  gk*xs+bk
}
# activation-wise batch norm function
bi <- function(x, i){
  xs <- scale(x)
  gk <- t(matrix(gamma[i]))
  bk <- t(matrix(beta[i]))
  suppressWarnings(gk*xs[,i]+bk)
}

X <- round(runif(N*p1, -5, 5)) %>% matrix(nrow = N)
# the neural net
y <- a(b(X %*% G1)) %*% G2 + u

Sau đó tính đạo hàm:

# drdbeta -- the matrix way
drdb <- matrix(rep(1, N*1), nrow = 1) %*% (-2*u %*% t(G2) * ap(b(X%*%G1)))
drdb
           [,1]      [,2]    [,3]        [,4]
[1,] -0.4460901 0.3899186 1.26758 -0.09589582
# the looping way
foreach(i = 1:4, .combine = c) %do%{
  sum(-2*u*matrix(ap(bi(X[,i, drop = FALSE]%*%G1[i,], i)))*G2[i])
}
[1] -0.44609015  0.38991862  1.26758024 -0.09589582

β⊗1N

∂A⊗B∂A=(Inq⊗Tmp)(In⊗vec(B)⊗Im)
mnpqABT
# playing with the kroneker derivative rule
A <- t(matrix(beta)) 
B <- matrix(rep(1, N))
diag(rep(1, ncol(A) *ncol(B))) %*% diag(rep(1, ncol(A))) %x% (B) %x% diag(nrow(A))
     [,1] [,2] [,3] [,4]
 [1,]    1    0    0    0
 [2,]    1    0    0    0
 snip
[13,]    0    1    0    0
[14,]    0    1    0    0
snip
[28,]    0    0    1    0
[29,]    0    0    1    0
[snip
[39,]    0    0    0    1
[40,]    0    0    0    1

γΓ1β⊗1

Cập nhật 2

∂R/∂Γ1∂R/∂γvec()∂R/∂Γ1w⊙XΓ1Γ1w≡(γ⊗1)⊙σXΓ1−1

w⊙XwX

∂(A⊙B)=∂A⊙B+A⊙∂B

và từ cái này , cái kia

∂vec(w⊙XΓ1)∂vec(Γ1)T=vec(XΓ1)I∂vec(w)∂vec(Γ1)T+vec(w)I∂vec(XΓ1)∂vec(Γ1)T

Cập nhật 3

Tiến bộ ở đây. Tôi thức dậy lúc 2 giờ tối qua với ý tưởng này. Toán học không tốt cho giấc ngủ.

∂R/∂Γ1

  • w≡(γ⊗1)⊙σXΓ1−1
  • "stub"≡a′(b(XΓ1))⊙−2ϵ^Γ2T

∂R∂Γ1=∂w⊙XΓ1∂Γ1("stub")
ijI
∂R∂Γij=(wi⊙Xi)T("stub"j)
∂R∂Γij=(IwiXi)T("stub"j)
∂R∂Γij=XiTIwi("stub"j)
∂R∂Γ=XT("stub"⊙w)

Và, trên thực tế, đó là:

stub <- (-2*u %*% t(G2) * ap(b(X%*%G1)))
w <- t(matrix(gamma)) %x% matrix(rep(1, N)) * (apply(X%*%G1, 2, sd) %>% t %x% matrix(rep(1, N)))
drdG1 <- t(X) %*% (stub*w)

loop_drdG1 <- drdG1*NA
for (i in 1:7){
  for (j in 1:4){
    loop_drdG1[i,j] <- t(X[,i]) %*% diag(w[,j]) %*% (stub[,j])
  }
}

> loop_drdG1
           [,1]       [,2]       [,3]       [,4]
[1,] -61.531877  122.66157  360.08132 -51.666215
[2,]   7.047767  -14.04947  -41.24316   5.917769
[3,] 124.157678 -247.50384 -726.56422 104.250961
[4,]  44.151682  -88.01478 -258.37333  37.072659
[5,]  22.478082  -44.80924 -131.54056  18.874078
[6,]  22.098857  -44.05327 -129.32135  18.555655
[7,]  79.617345 -158.71430 -465.91653  66.851965
> drdG1
           [,1]       [,2]       [,3]       [,4]
[1,] -61.531877  122.66157  360.08132 -51.666215
[2,]   7.047767  -14.04947  -41.24316   5.917769
[3,] 124.157678 -247.50384 -726.56422 104.250961
[4,]  44.151682  -88.01478 -258.37333  37.072659
[5,]  22.478082  -44.80924 -131.54056  18.874078
[6,]  22.098857  -44.05327 -129.32135  18.555655
[7,]  79.617345 -158.71430 -465.91653  66.851965

Cập nhật 4

∂R/∂γ

  • XΓ~≡(XΓ−μXΓ)⊙σXΓ−1
  • γ~≡γ⊗1N

∂R∂γ~=∂γ~⊙XΓ~∂γ~("stub")
Looping gives you
∂R∂γ~i=(XΓ~)iTIγ~i("stub"i)
Which, like before, is basically pre-multiplying the stub. It should therefore be equivalent to:
∂R∂γ~=(XΓ~)T("stub"⊙γ~)

It sort of matches:

drdg <- t(scale(X %*% G1)) %*% (stub * t(matrix(gamma)) %x% matrix(rep(1, N)))

loop_drdg <- foreach(i = 1:4, .combine = c) %do% {
  t(scale(X %*% G1)[,i]) %*% (stub[,i, drop = F] * gamma[i])  
}

> drdg
           [,1]      [,2]       [,3]       [,4]
[1,]  0.8580574 -1.125017  -4.876398  0.4611406
[2,] -4.5463304  5.960787  25.837103 -2.4433071
[3,]  2.0706860 -2.714919 -11.767849  1.1128364
[4,] -8.5641868 11.228681  48.670853 -4.6025996
> loop_drdg
[1]   0.8580574   5.9607870 -11.7678486  -4.6025996

The diagonal on the first is the same as the vector on the second. But really since the derivative is with respect to a matrix -- albeit one with a certain structure, the output should be a similar matrix with the same structure. Should I take the diagonal of the matrix approach and simply take it to be γ? I'm not sure.

It seems that I have answered my own question but I am unsure whether I am correct. At this point I will accept an answer that rigorously proves (or disproves) what I've sort of hacked together.

while(not_answered){
  print("Bueller?")
  Sys.sleep(1)
}

2
Chapter 9 section 14 of "Matrix Differential Calculus with Applications in Statistics and Econometrics" by Magnus and Neudecker, 3rd edition janmagnus.nl/misc/mdc2007-3rdedition covers differentials of Kronecker products and concludes with an exercise on differential of Hadamard product. "Notes on Matrix Calculus" by Paul L. Fackler www4.ncsu.edu/~pfackler/MatCalc.pdf has a lot of material on differentiating Kronceker products
— Mark L. Stone

Thanks for the references. I've found those MatCalc notes before, but it doesn't cover Hadamard, and anyway I'm never certain whether a rule from non-matrix calculus applies or doesn't apply to to matrix case. Product rules, chain rules, etc. I'll look into the book. I'd accept an answer that points me to all of the ingredients I need to pencil it out myself...
— generic_user

why are you doing this? why not use framewroks such as Keras/TensorFlow? It's a waste of productive time to implement these low level algorithms, that you could use on solving actual problems
— Aksakal

1
More precisely, I'm fitting networks that exploit known parametric structure -- both in terms of linear-in-parameters representations of input data, as well as longitudinal/panel structure. Established frameworks are so heavily optimized as to be beyond my ability to hack/modify. Plus math is helpful generally. Plenty of codemonkeys have no idea what they're doing. Likewise learning enough Rcpp to implement it efficiently is useful.
— generic_user

1
@MarkL.Stone not only is it theoretically sound, it's practically easy! A more or less mechanical process! &%#$!
— generic_user

Câu trả lời:


1

Not a complete answer, but to demonstrate what I suggested in my comment if

b(X)=(X−eNμXT)ΓΣX−1/2+eNβT
where Γ=diag⁡(γ), ΣX−1/2=diag⁡(σX1−1,σX2−1,…) and eN is a vector of ones, then by the chain rule
∇βR=[−2ϵ^(Γ2T⊗I)JX(a)(I⊗eN)]T
Noting that −2ϵ^(Γ2T⊗I)=vec⁡(−2ϵ^Γ2T)T and JX(a)=diag⁡(vec⁡(a′(b(XΓ1)))), we see that
∇βR=(I⊗eNT)vec⁡(a′(b(XΓ1))⊙−2ϵ^Γ2T)=eNT(a′(b(XΓ1))⊙−2ϵ^Γ2T)
via the identity vec⁡(AXB)=(BT⊗A)vec⁡(X). Similarly,
∇γR=[−2ϵ^(Γ2T⊗I)JX(a)(ΣXΓ1−1/2⊗(XΓ1−eNμXΓ1T))K]T=KTvec⁡((XΓ1−eNμXΓ1T)TWΣXΓ1−1/2)=diag⁡((XΓ1−eNμXΓ1T)TWΣXΓ1−1/2)
where W=a′(b(XΓ1))⊙−2ϵ^Γ2T (the "stub") and K is an Np×p binary matrix that selects the columns of the Kronecker product corresponding to the diagonal elements of a square matrix. This follows from the fact that dΓi≠j=0. Unlike the first gradient, this expression is not equivalent to the expression you derived. Considering that b is a linear function w.r.t γi, there should not be a factor of γi in the gradient. I leave the gradient of Γ1 to the OP, but I will say for derivation with fixed w creates the "explosion" the writers of the article seek to avoid. In practice, you will also need to find the Jacobians of ΣX and μX w.r.t X and use product rule.
Khi sử dụng trang web của chúng tôi, bạn xác nhận rằng bạn đã đọc và hiểu Chính sách cookie và Chính sách bảo mật của chúng tôi.
Licensed under cc by-sa 3.0 with attribution required.