定義
對於 $n \times n$ 實矩陣 $Q$,如果滿足:
$$Q^T Q = Q Q^T = I$$
其中 $I$ 是 $n \times n$ 單位矩陣,則稱 $Q$ 為正交矩陣
例子
$$ Q_1 = \begin{bmatrix} \cos\theta & -\sin\theta \\ \sin\theta & \cos\theta \end{bmatrix} $$ $$ Q_2 = \begin{bmatrix} 1 & 0 \\ 0 & -1 \end{bmatrix} $$ $$ Q_3 = \begin{bmatrix} \frac{1}{\sqrt{2}} & \frac{1}{\sqrt{2}} \\ -\frac{1}{\sqrt{2}} & \frac{1}{\sqrt{2}} \end{bmatrix} $$定義
對於矩陣 $A$,其轉置矩陣表示為 $A^T$,其中:
$$A_{ij}^T = A_{ji}$$
換句話說,原矩陣的第 $i$ 行第 $j$ 列元素,在轉置矩陣中變成第 $j$ 行第 $i$ 列元素
例子
$$\begin{bmatrix} 1 & 2 \\ 3 & 4 \end{bmatrix}^T = \begin{bmatrix} 1 & 3 \\ 2 & 4 \end{bmatrix}$$ $$\begin{bmatrix} 1 & 2 \\ 3 & 4 \\ 5 & 6 \end{bmatrix}^T = \begin{bmatrix} 1 & 3 & 5 \\ 2 & 4 & 6 \end{bmatrix}$$Introduction
Canvas fingerprinting is a powerful technique for detecting fraudsters and bots. This presentation explores how attackers modify canvas fingerprints to evade detection and the methods we can use to identify these manipulations.
What is Canvas Fingerprinting?
Canvas fingerprinting uses the HTML canvas API to draw invisible shapes and text that create unique identifiers based on your:
Browser
Operating system
GPU
Installed fonts
These fingerprints are stable and unique, making them valuable for tracking fraudsters even if they delete cookies.
For more information, Please redirect to Powerpoint - Detecting Noise in Canvas Fingerprinting
此為 機器學習中的優化器 系列文章 - 第 3 篇:
- 機器學習中的優化器 (1) - Gradient descent 梯度下降與其變體
- 機器學習中的優化器 (2) - 從梯度下降問題到動量優化
- 機器學習中的優化器 (3) - AdaGrad、RMSProp 與 Adam
前言
上篇文章提到 動量(Momentum) 的引入可以有效逃離局部最小值進而幫助找到全局的最佳解,但 Momentum 對所有的參數都套用同樣的學習率,導致在某些狀況下可能過度衝刺而不易找到全局最佳解。因此後來出現了一些優化算法 - AdaGrad, RMSProp, Adam 等讓學習率可以隨著訓練迭代的過程中自適應調整以更快的找到全局最佳解
此為 機器學習中的優化器 系列文章 - 第 2 篇:
- 機器學習中的優化器 (1) - Gradient descent 梯度下降與其變體
- 機器學習中的優化器 (2) - 從梯度下降問題到動量優化
- 機器學習中的優化器 (3) - AdaGrad、RMSProp 與 Adam
梯度下降面臨的問題
在上一篇 Gradient descent 梯度下降與其變體 文章中,我們介紹了梯度下降的基本原理和三種主要變形:Batch Gradient Descent、Stochastic Gradient Descent (SGD) 和 Mini-batch Gradient Descent。雖然這些方法在機器學習中發揮了重要作用,但在實際應用中仍然面臨一些挑戰:
1. 學習率選擇困難
問題描述:
- 學習率太小:收斂速度慢,需要大量迭代
- 學習率太大:可能導致震盪或發散
- 不同參數可能需要不同的學習率
數學表達:
$$w_{t+1} = w_t - \eta \nabla L(w_t)$$
其中 $\eta$ 是固定的學習率,但實際上:
- 在平坦區域需要較大學習率加速
- 在陡峭區域需要較小學習率避免震盪