Viva 2 colour cheat sheet · 6 modules, blocks stacked one below the other · A4. Print with "Background graphics" enabled.
1

Analysis and Computing foundations

MATLAB for signals and CT images · polynomials · linear algebra · Fourier and sampling
formulanumbers to quotetrap / critiquesay this in the exam

MATLAB essentials

Names: letter first, no specials, < 32 chars, case-sensitive. A(row,col) rows split by ; · A(1,:) row · A(1,1:2:5) cols 1,3,5 · .* ./ .^ element-wise, * matrix (inner dims agree) · zeros ones eye rand · for i=1:n … end, if … elseif … else … end, odd test fix(x/2)~=x/2 · index starts at 1 → a(n+1).

Course matrix A = [12 10 13 15 16; 22 45 65 1 0; 22 33 41 23 45; 21 30 12 6 2; 1 0 0 1 7]: sum(A) = [78 118 131 46 70], sum(sum(A)) = 443, diag = [12 45 41 6 7], trace 111. [3,5].*[4,8] = [12,40].

Rectangle rule & loops

$$\int_a^b f\,dt\approx\sum_k f(t_k)\,\Delta t\qquad \int_0^\pi\sin t\,dt:\ \Delta t=0.5\to1.98,\ \Delta t=0.1\to1.9995,\ \text{exact }2$$

[0 2 10 20 3 15] with +1 if > 10 else −1 → [−1 1 9 21 2 16]. Plot: t=0:0.01:2*pi (0:1:2π gives only 7 points).

Images in MATLAB

sprintf names CT_001…012→imread → uint8 → double→hist_im: find(Y==i), Hist(i+1)→volume histogram (sum of slices)→valley T = 95→Z = Y if Y<95 else 255→imwrite(Z./255)

uint8 0–255 (uint16 65,535); toy 4×4 histogram [5 4 3 4]. medfilt2(X,[3 3]) odd window → middle of 9: {0,20,0,127,112,100,128,135,0} → 100; {0,5,8,60,99,99,109,125,155} → 99; ECG 1-D [1 3]. 3-D: V(:,:,k), isosurface(…,15), daspect([1 1 .4]), plot3, trisurf.

Polynomials & fitting

Coefficients highest power first: x³+4x²+9x+16 → [1 4 9 16]. roots ↔ poly; conv = multiply ([1 2 3 4]⋆[1 4 9 16] = [1 6 20 50 75 84 64]); deconv; polyder → [3 8 9]; polyint → [0.25 1.333 4.5 16 0].

$$\text{polyfit: }\min_p\sum_i(y_i-P(x_i))^2\ \Rightarrow\ (V^TV)p=V^Ty\quad(\text{41 pts}\to p\approx[0.965,\,0.140,\,4.969]);\qquad\text{spline passes exactly through the data}$$

Linear algebra

$$x=A\backslash b\ (\text{LU, not inv});\quad A^{-1}=\frac1{ad-bc}\begin{bmatrix}d&-b\\-c&a\end{bmatrix};\quad A^{-1}=\frac{\mathrm{adj}A}{|A|},\ C_{ij}=(-1)^{i+j}M_{ij};\quad x_{LS}=(A^TA)^{-1}A^Tb;\quad Av=\lambda v$$

Solution exists iff rank(A) = rank([A b]) = r; unique if r = n (det ≠ 0); infinite if r < n (pinv, rref); A x = 0 non-trivial iff rank < n. Over-determined consistent → exact; inconsistent → least squares. Slide system A = [1 2 3; 4 5 6; 7 8 0], y = [366; 804; 351] → x = [25; 22; 99] (y/A is a bug). [3 −4; 6 −8] singular. Circuit A = [1 −1 1; −1 1 −1; 4 2 0; 0 2 5], b = [0;0;8;9] → i = [1, 2, 1] A. Chemical CO₂ + H₂O → O₂ + C₆H₁₂O₆: null space t·[6 6 6 1]. dot = projection, cross = moment.

Fourier series, transform, modulation, sampling

$$x(t)=a_0+\sum_n\big[a_n\cos(2\pi nf_0t)+b_n\sin(2\pi nf_0t)\big],\ a_0=\tfrac1T\!\int_T\!x,\ a_n=\tfrac2T\!\int_T\!x\cos,\ b_n=\tfrac2T\!\int_T\!x\sin;\quad \pm1\text{ square: }a_0=0,\ b_n=\tfrac{4}{\pi n}\ (n\text{ odd})$$ $$X(f)=\!\int\! x(t)e^{-j2\pi ft}dt;\ \ \mathrm{rect}_\tau\leftrightarrow\tau\,\mathrm{sinc}(f\tau)\ (\text{zeros }n/\tau);\ \ e^{-at}u(t)\leftrightarrow\tfrac1{a+j2\pi f};\ \ e^{-a|t|}\leftrightarrow\tfrac{2a}{a^2+(2\pi f)^2};\ \ x\cos(2\pi f_0t)\leftrightarrow\tfrac12[X(f\!-\!f_0)+X(f\!+\!f_0)]$$ $$x_s=x\cdot\textstyle\sum_k\delta(t-kT_s)\ \leftrightarrow\ f_s\sum_kX(f-kf_s):\ \text{copies at }kf_s\pm f_m\ \Rightarrow\ \boxed{f_s\ge2f_m}\ \text{(200 Hz → 400 Hz; 150 → 300)};\ \text{RC: }\tfrac{V_o}{V_{in}}=\tfrac1{1+sRC}$$

Series: periodic only ("main limitation"); transform: periodic and aperiodic ("main advantage"). syms t; fourier(exp(-t^2)) = √π e^{−w²/4}; fourier(exp(-abs(t))) = 2/(1+w²).

Numbers to quote

sum(sum(A))443 · trace 111∫sin, Δt .5 / .11.98 / 1.9995 (exact 2)Lung thresholdT = 95 (valley of 12-slice histogram)Medians100 and 99polyfit[0.965 0.140 4.969] ≈ x²+5A\y[25 22 99]Circuiti = [1 2 1] ASquare waveb₁ 1.273, b₃ 0.424, b₅ 0.255; even 0Nyquistfs ≥ 2fmvar([2 4 6])4 (N−1 denominator)

Say this in the exam

"Backslash is Gaussian elimination (LU); the inverse is only for theory." "A solution exists when b lies in the column space: rank(A) = rank([A b])." "Fitting minimises squared error and need not touch the points; interpolation must." "A square wave has only odd harmonics falling as 1/k, so its edges need infinite bandwidth." "Sampling replicates the spectrum every fs; keep the copies apart with fs ≥ 2fm, then low-pass to recover."

Traps

MATLAB indices start at 1 (Hist(i+1), a(n+1)). uint8 saturates at 255: convert to double. medfilt2 needs an odd window. y/A ≠ A\y. det ≈ 0 means numerically singular. Variable names are case-sensitive (items vs Items bug). Fourier series only for periodic signals.

Viva 2 cheat sheet · Analysis and Computingpage 1 of 6
2

Computer Vision foundations

MAI621/ECE621 · histograms, Otsu, morphology, Fourier filters, spatial filters, Canny, restoration, tracking

Images, neighbours, histograms

f = i·r (illumination × reflectance); bits = M·N·k (1024² × 8 = 8,388,608). N4 / ND / N8; m-connectivity removes multiple paths. D4 = |Δx|+|Δy|, D8 = max, De = √: (1,2)→(4,6) = 7, 4, 5. Two-pass labelling with equivalence table. Point ops: negative (L−1)−r, log c·log(1+r), power s = c·rᵞ (γ<1 brightens), slicing.

$$p(g)=\frac{h(g)}{RC},\quad s_k=\mathrm{round}\big((L-1)\textstyle\sum_{j\le k}p_j\big):\ h=[5,4,0,0,2,1,3,0,4,1],\ L=10\ \to\ s=2,4,4,4,5,5,7,7,9,9\ (\text{"almost, not completely, flat"})$$

Otsu & segmentation

$$P_1=\sum_{i\le k}p_i,\ m(k)=\sum_{i\le k}ip_i,\ m_G=\sum ip_i,\quad \sigma_B^2(k)=\frac{(m_GP_1-m)^2}{P_1(1-P_1)}=P_1P_2(m_1-m_2)^2,\quad k^*=\arg\max,\ \eta=\sigma_B^2/\sigma_G^2$$

Freq 8,7,2,6,9,4 (N 36): m_G 2.3611; σ²_B = 1.593, 2.564, 2.629, 2.142, 0.870 → k* = 2, η ≈ 0.84. Limitation: histogram only, no spatial info. Iterative T = ½(m₁+m₂); Niblack T = m + k·s. Split/merge (std < 5): 4 quadrants pass, R1∪R3 and R2∪R4 merge → 2 regions. K-means 7 points: C₁(1,1), C₂(5,7) → it.1 {1,2,3}/{4–7} → it.2 point 3 moves → (1.25,1.5), (3.9,5.1) → it.3 stable.

Morphology

$$A\ominus B=\{z:(B)_z\subseteq A\}\ \text{(fit, shrink)}\qquad A\oplus B=\{z:(\hat B)_z\cap A\ne\emptyset\}\ \text{(hit, grow)}\qquad (A\ominus B)^c=A^c\oplus\hat B$$

Opening = erode→dilate (salt, thin joints); closing = dilate→erode (pepper, holes). 3×3 block: erode by 3×3 ones → centre only; by vertical line → middle row; dilate by 3×3 ones → full 5×5; by cross → plus with zero corners. Boundary β = A − (A⊖B); fill X_k = (X_{k−1}⊕B) ∩ Aᶜ; coins: binarise → erode → label → count.

Fourier domain & frequency filters

$$F(u,v)=\sum_x\sum_yf\,e^{-j2\pi(ux/M+vy/N)},\ f=\tfrac1{MN}\sum_u\sum_vF\,e^{+j2\pi(\cdot)};\quad f\star h\leftrightarrow FH\ (\text{pad }2M\times2N,\ (-1)^{x+y}\text{ centres})$$ $$H_{ILPF}=\mathbb 1[D\le D_0]\ (\text{rings}),\ H_{BLPF}=\frac1{1+(D/D_0)^{2n}},\ H_{GLPF}=e^{-D^2/2D_0^2}\ (\text{no ringing}),\ H_{HP}=1-H_{LP},\ H_{Lap}=-4\pi^2D^2,\ g=f-\nabla^2f$$

Unsharp g = f + k(f − f_LP). Homomorphic: ln f = ln i + ln r, H = (γ_H−γ_L)[1−e^{−cD²/D₀²}]+γ_L (0.25, 2, 1, 80). Notch pairs ±(u_k,v_k) kill stripes. Sampling: fs > 2·f_max else aliasing. Mask weights sum 1 → DC passes (low-pass); sum 0 → edge detector; [0 −1 0; −1 5 −1; 0 −1 0] emphasises edges.

Spatial filters & edges

$$g(x,y)=\sum_{s=-a}^{a}\sum_{t=-b}^{b}w(s,t)f(x+s,y+t);\quad \nabla^2f=f(x\!+\!1,y)+f(x\!-\!1,y)+f(x,y\!+\!1)+f(x,y\!-\!1)-4f;\quad g_x=(z_7\!+\!2z_8\!+\!z_9)-(z_1\!+\!2z_2\!+\!z_3)$$

Box 3×3 of 106,104,99,95,100,108,98,90,85 → 98.33; median = 5th of the 9 sorted values (slide: 0,0,1,1,1,2,2,2,4 → 1). Borders: discard (512→510), zero-pad (false edges), replicate. Second derivative: thin double edges, fine detail. Canny: Gaussian 5×5/159 (→ 41) → Sobel G_x −191, G_y −181, |G| 263, θ 134° → NMS (263 vs 7, 255 → kept) → hysteresis T_H 200 / T_L 50 (weak kept if 8-connected to strong).

Restoration

$$g=h\star f+\eta;\quad \hat f=g-\frac{\sigma_\eta^2}{\sigma_L^2}(g-m_L);\quad \hat F=\Big[\frac1H\frac{|H|^2}{|H|^2+K}\Big]G;\quad I^{t+1}=I^t+\lambda\!\sum_{N,S,E,W}\!c_d\nabla_dI,\ c=e^{-(|\nabla I|/K)^2}$$

Contraharmonic Q>0 pepper, Q<0 salt; median for salt-and-pepper. Local filter μ 75.14, σ² 6259.5, v² 400: 186→178.9, 95→93.7, 36→38.5. Diffusion 186 with N255 S157 E212 W208, K 30, λ 1/7: c = .005, .392, .472, .584 → 188.0 (edge kept).

Video: background, GMM/EM, tracking

$$\gamma(z_{nk})=\frac{\pi_k\mathcal N(x_n|\mu_k,\Sigma_k)}{\sum_j\pi_j\mathcal N(x_n|\mu_j,\Sigma_j)}\ (\text{E}),\quad \mu_k=\tfrac1{N_k}\sum_n\gamma_{nk}x_n,\ \pi_k=\tfrac{N_k}N\ (\text{M});\qquad P(X_t|y_{0:t})\propto P(y_t|X_t)\!\int\!P(X_t|X_{t-1})P(X_{t-1}|y_{0:t-1})$$

Frame differencing = boundaries + ghosts; temporal median or Stauffer–Grimson GMM background. E-step example x=1, N(0,1)/N(4,1): γ₁ = 0.982. Tracking = detection + prediction; Markov on state and observation. Kalman: linear-Gaussian, mean + covariance, one object. Particle filter: weighted samples, clutter, multimodal. Too strong dynamics → ignores data; too strong observation → repeated detection; drift.

Numbers to quote

Otsuk* = 2, σ²_B 2.629, η .84Equalisation0→2, 1→4, 4→5, 6→7, 8→9K-means(1.25,1.5), (3.9,5.1) after 3 it.Sobel/Canny−191, −181, 263; T 200/50Diffusion188.0Wiener-local178.9 / 93.7 / 38.5Laplacian H−4π²D² (−39.48 at D=1)Homomorphicγ_L .25, γ_H 2, c 1, D₀ 80PadP ≥ A+C−1, 2M×2N

Say this in the exam

"Otsu maximises between-class variance from the histogram alone." "Erosion keeps a pixel only where the structuring element fits; dilation wherever it hits; they are duals." "Filtering in space is multiplication in frequency, but DFT convolution is circular, so we pad." "Canny: smooth, gradient, thin by non-maximum suppression, link by hysteresis." "EM alternates responsibilities and parameter updates; a Kalman filter predicts then corrects."

Traps

Ideal low-pass rings (sinc kernel). Zero padding of borders creates false edges. Otsu fails on unimodal or unevenly lit images. Second derivative amplifies noise: smooth first. 8-connectivity merges diagonal blobs. Mask sums: 1 passes DC, 0 rejects it.

Viva 2 cheat sheet · Computer Visionpage 2 of 6
3

Intelligent Robots foundations

MAI675/DEN775 · PID, Hough, RANSAC, camera, YOLO/MIO, Kalman, ROS 2

The loop and the pipeline

Sense→Compute→Actuate⟳image→grey→edges (Sobel/Canny)→ROI→Hough→filter L/R→centre, error→PID→servo / wheels

PID

$$u=K_pe+K_i\!\int\! e\,dt+K_d\frac{de}{dt}\qquad u[n]=K_pe[n]+K_i\sum e[i]\Delta t+K_d\frac{e[n]-e[n-1]}{\Delta t}$$

$e=SP-PV$. P only → steady-state error (0 output at $e=0$); I removes it; D damps, amplifies noise. Tune $K_p$, then $K_d$, then $K_i$ (0.0001). $K_p=125/3500=0.0357$. Example (2.0,0.5,0.1), $\Delta t$ 0.1, $e=-0.48$, $e_{prev}=-0.30$, $\sum=-0.12$: $-0.96-0.06-0.18=\mathbf{-1.20}$.

Hough & RANSAC

$$\rho=x\cos\theta+y\sin\theta,\quad m=-\cot\theta,\ b=\rho/\sin\theta\qquad S=\frac{\log(1-P)}{\log(1-p^k)}$$

Normal form: vertical lines finite. (3,3),(4,3),(5,3) → $A(3,90°)=3$. RANSAC $P$ 0.99, $p$ 0.5: $k$=2→17, 3→35, 4→72. Reject lane parabola $|a|\ge0.003$.

Camera & stereo

$$\lambda\begin{bmatrix}u\\v\\1\end{bmatrix}=K[R|t]\begin{bmatrix}X\\Y\\Z\\1\end{bmatrix},\ K=\begin{bmatrix}f_x&s&c_x\\0&f_y&c_y\\0&0&1\end{bmatrix},\ X=\frac{(u-c_x)Z}{f_x},\ Z=\frac{f_xB}{d}$$

(700,400), $Z$ 10, $f$ 800, $c$ (640,360) → (0.75, 0.5, 10). Stereo $d=80$, $f_x$ 795, $B$ 0.2 → $Z=1.9875$ m. One image: 2 eq, 3 unknowns. Quality = reprojection error.

YOLO & MIO

Output $S\times S\times(5B+C)$: 7×7×30. Loss: $\sqrt w,\sqrt h$; $\lambda_{noobj}=0.5$; NMS IoU > 0.5. MIO: in lane if $x_L(y)\le x\le x_R(y)$, $x(y)=(y-b)/m$; MIO = argmax $y_{bottom}$. Three-car example: Car 3 (x 500) out of lane → Car 2 (290 > 260). FCW: tracks (confirm [2 3], delete 5), closest in lane; $d=1.2v+v^2/(2\cdot0.4\cdot9.8)$ → 24.8 m at 10 m/s.

Kalman filter · five scalar equations and their origin

$$\underbrace{\mu_p=\mu+v\Delta t+\tfrac12a\Delta t^2}_{\text{mechanics}}\quad \underbrace{p_p=p+q}_{\text{variances add}}\quad \underbrace{K=\frac{p_p}{p_p+r}}_{\text{confidence}}\quad \underbrace{\mu=\mu_p+K(z-\mu_p)}_{\text{Bayes mean}}\quad \underbrace{p=(1-K)p_p}_{\text{Bayes variance}}$$

Derivation: $N(\mu_p,p)\times N(z,r)$ → $\frac1{\sigma^2}=\frac1p+\frac1r$, $\mu=\frac{r\mu_p+pz}{p+r}$; set $K=\frac{p}{p+r}$. $K\to1$ trust sensor, $K\to0$ trust prediction; $0

Fusion: $z_f=\frac{\sum z_i/r_i}{\sum 1/r_i},\ r_f=\frac1{\sum1/r_i}$; (0.9,1.1),(1,4) → 0.94, 0.8; prior 10.1 → $K=0.927$, $x=0.87$, $p=0.74$. $r_i$ never changes during filtering.

Numbers to quote

Ultrasonicd = t·0.034/2 cm (2000 µs → 34 cm)L293DIN 10 fwd, 01 rev, 00 stop; EN = PWMLine bar I²Cbar 0x3E, robot 81, request 240Grey0.299R+0.587G+0.114B; (120,200,80) → 162Sobel ex.Gx 275, Gy 145 → M 311, θ 27.8°HoughLinesPrho 2, θ π/180, thr 50, minLen 10, gap 30BC net24/36/48 (5×5 s2), 64/64, FC 1254/1254/256/1, lr 1e-4ROS 2Jazzy · colcon · /thing_on Bool q10 · frame "map"

Say this in the exam

"Bang-bang chooses a direction; PID chooses how much." "Hough votes in $(\rho,\theta)$ because slope is infinite for vertical lines." "RANSAC keeps the model most points agree with; least squares is pulled by outliers." "The Kalman filter is a recursive Bayesian estimator: predict widens, update shrinks." "MIO comes from confirmed tracks because detections flicker."

Traps

Pull-up button pressed = LOW. Never delay(), use millis(). Pooling has no parameters; PID I-term needs anti-windup in practice. Behaviour cloning is regression (ELU + regression layer), not softmax. Gazebo = world, RViz = belief.

Viva 2 cheat sheet · Intelligent Robotspage 3 of 6
4

Paper 1 · Solar power forecasting with FFNN, LSTM, GRU

du Plessis, Strauss & Rix 2021, Applied Energy · 75 MW plant, macro vs aggregated inverter-level forecasts

Pipeline

4 yr 1-min data, 84 inverters→clean: eliminate / interpolate ≤1 h / impute→5-min avg, sun angles, one-hot month & wind dir→split 2/1/1 yr→3-phase tuning→FFNN · LSTM · GRU→macro vs Σ inverters (+ loss model)→NRMSE / MAE / MAPE + bootstrap CI

Metrics (capacity-normalised)

$$\mathrm{NRMSE}=\sqrt{\tfrac1N\textstyle\sum\big(\tfrac{\hat P_i-P_i}{P_{cap}}\big)^2}\cdot100\%,\quad \mathrm{MAE}=\tfrac1N\textstyle\sum|\hat P_i-P_i|\ (\text{kW}),\quad \mathrm{MAPE}=\tfrac1N\textstyle\sum\tfrac{|\hat P_i-P_i|}{P_{cap}}\cdot100\%,\quad P_{cap}=75\ \text{MW}$$

NRMSE punishes large errors quadratically (selection criterion); MAE is what an operator reads; capacity normalisation avoids dividing by night-time zeros. Bootstrap: resample MAPEs m = 10,000 times → 95% CI of the mean; width grows with horizon.

Models & windows

$$\text{GRU: }z_t=\sigma(W_z[h_{t-1},x_t]),\ r_t=\sigma(W_r[h_{t-1},x_t]),\ \tilde h_t=\tanh(W[r_t\odot h_{t-1},x_t]),\ h_t=(1-z_t)\odot h_{t-1}+z_t\odot\tilde h_t$$

LSTM adds a cell state with forget/input/output gates (more parameters, similar accuracy here). Sliding window of past samples (1, 2, 3, 6, 24 h, HISIMI+x) → 21 outputs: 1–6 h ahead at 15-min steps (multi-target regression). FFNN sees the flattened window; recurrent models see the sequence.

Tuning protocol & clustering

Phase 1 feature/window combinations + extensive grid HL ∈ {2,3}, HU powers of 2, mb ∈ {32,64} (≥ 240 runs/model). Phase 2 guided grid search: expand any hyper-parameter sitting on the grid edge until accuracy saturates. Phase 3 mb ∈ {16,128,256}. Fixed: MSE, Adam 1e-4, early stopping patience 20. Clustering: Euclidean distance between inverter series → 84×84 matrix → K-means K = 10 (elbow, CH, gap) → tune one representative per cluster, assign its top-2 settings to the rest.

$$d(a,b)=\sqrt{\textstyle\sum_t(a_t-b_t)^2},\qquad J=\sum_{k}\sum_{a\in C_k}\|a-\mu_k\|^2$$

Numbers to quote

Plant75 MW, 84 inverters, ~100 ha, South AfricaData4 years 1-min → 5-min; 2/1/1 splitHorizon1–6 h, 15-min steps = 21 outputsBest macro (GRU)NRMSE 8.12 %, MAPE 3.42 %Best inverter-levelNRMSE 8.02 %, MAPE 3.39 %All-84 exhaustive8.15 / 3.45 vs cluster 8.15 / 3.45Bootstrapm = 10,000, 95 % CIClustersK = 10Inter-inverter variationup to 3 % (wind)OptimiserAdam 1e-4, patience 20

Theory in one breath

Macro model: one network on the plant total. Aggregated inverter-level: 84 networks (one per inverter), summed, plus a loss-correction FFNN for the difference between Σ inverters and the plant meter. Hypothesis: local effects (wind, dust, shading) are visible per inverter and lost in the total. Weather types (clear, overcast, variable) are compared separately.

Results in one breath

GRU ≥ LSTM > FFNN at all horizons; inverter-level beats macro by a small margin (8.02 vs 8.12 NRMSE), mainly on variable-cloud days and short horizons; gains vanish on clear/overcast days and at 6 h. Cluster-based tuning matches exhaustive tuning. Error grows roughly linearly with horizon; CIs widen with horizon.

Critique (prepare calmly)

  • Headline gain ≈ 0.1 NRMSE point with no seed variance or significance test.
  • 84 models + loss model for ~1 %: cost not reported.
  • No persistence / ARIMA baseline in main tables; MAPE by capacity flatters.
  • One flat site; hilly-site claim untested; raw Euclidean clustering may just follow capacity (try normalised or DTW).
  • Strengths: real utility data, full-year test, transparent equal-effort tuning, robustness to bad data.

Two-minute summary

"Operators need intra-day solar forecasts. Almost all models forecast the plant total. On a 75 MW plant with 84 inverters, local wind, dust and shading differ across the site, so the authors forecast every inverter and sum them. Four years of data, equal-effort three-phase tuning, FFNN/LSTM/GRU, 21 outputs over 1–6 h. GRU is best; inverter-level forecasts are slightly better (8.02 vs 8.12 % NRMSE), mostly on variable days. Clustering inverters with K-means lets you tune ten instead of 84 at no loss. The gain is real but small, and its cost and significance are not quantified."

Viva 2 cheat sheet · Paper 1page 4 of 6
5

Paper 2 · Robots that simulate themselves from video

Hu, Lin & Lipson 2025, Nature Machine Intelligence · Free-Form Kinematic Self-Model (FFKSM)

Pipeline

motor babbling (9° steps)→one RGB camera, 100×100→colour segmentation + median background→silhouette GT→train FFKSM (3-D point + joints → σ, α)→render Σσα along 64 ray points→MSE vs silhouette→IK · RRT planning · damage recovery

Model equations

$$X'=T^{-1}X,\ T=R_{pitch}(A_1)R_{yaw}(A_0);\quad H_1=C(\gamma(X'))\ (3\to33),\ H_2=K(\gamma(A_2,A_3))\ (2\to22);\quad (\sigma_{ijk},\alpha_{ijk})=P(H_1,H_2)$$ $$\sigma=1-e^{-\mathrm{ReLU}(B)},\qquad \mathrm{Pred}_{ij}=\sum_{k=1}^{64}\sigma_{ijk}\alpha_{ijk},\qquad \mathcal L=\frac1{WH}\sum_{ij}(\mathrm{Pred}_{ij}-\mathrm{GT}_{ij})^2,\qquad \gamma(p)=\big(p,\sin2^l\pi p,\cos2^l\pi p\big)_{l=0..4}$$

σ = density (is the point inside the robot?), α = visibility (does the camera see it?). Gradient to a point ∝ pixel error × α, so hidden points are not trained. Plain NeRF volume rendering failed (no depth ordering from one view).

Using the self-model

$$p_{ee}(\theta)=\frac1{|Q_{occ}|}\sum_{q\in Q_{occ}}q,\qquad \theta^*=\arg\min_\theta\|p_{ee}(\theta)-p^*\|^2\ \ (\text{Adam, lr }0.04,\ \text{stop at }10^{-4}\text{ or 1,000 it.})$$

IK by gradient descent through the differentiable model, warm-started from the previous pose (1,000-point 3-D spiral). Planning: whole-body FFKSM flags configurations whose occupied points hit an obstacle; end-effector FFKSM gives the heuristic; RRT grows the tree. Damage: bent link → re-babble, re-train, re-plan; mirror self-recognition by comparing predicted and observed silhouettes.

Numbers to quote

Data12,000 samples: 10k train/val (8:2), 2k testImages100 × 100, one fixed RGB cameraRobots4-DOF arms, sim + real (two robots)RayM = 64 pointsEncoding3→33, 2→22 (L = 5)Error (px²)OM 0.004 vs NN 0.010 vs RS 0.029IKAdam lr 0.04, < 1e-4, ≤ 1,000 it.Model size333 kBBabbling step9°

Theory in one breath

A self-model is a learned simulator of the robot's own morphology and kinematics. Instead of CAD and forward-kinematics equations, a neural field maps (3-D query point, joint angles) → occupancy and visibility; the first two joints are handled analytically as rotations (virtual coordinates), the last two by the network. Self-supervised: the only labels are the robot's own silhouettes.

Results in one breath

Morphology predicted in simulation and reality for two arms; end-effector isolated; spiral tracked by gradient-descent IK; collision-free RRT paths executed; after a bent link the re-learned model restores task performance. Baselines (random search RS, nearest-neighbour NN) are 2.5–7× worse in silhouette MSE.

Critique (prepare calmly)

  • Evaluation is 2-D silhouette MSE from one view; no 3-D ground truth or CAD comparison.
  • Baselines trivial; no comparison with the depth-camera predecessor or a CAD simulator; no numeric tracking error or RRT success rate.
  • Needs painted robot, clean background, calibrated camera, joint encoders: "no kinematic priors" is overstated.
  • 4-DOF, quasi-static, no dynamics, loads or compliance; single-view depth ambiguity.
  • Strengths: elegant self-supervised formulation, real hardware, damage recovery, 333 kB model.

Two-minute summary

"Robots usually need a hand-built simulator. Here a 4-DOF arm builds its own by watching itself with one camera. It babbles, segments its silhouette, and trains a small neural field that says for any 3-D point and joint angles whether the point is occupied and visible. Rendering sums density × visibility along each pixel ray and is compared with the silhouette by MSE. With 12,000 images the model predicts the robot's shape in simulation and reality, isolates the end effector, tracks a spiral by gradient-descent inverse kinematics, plans collision-free RRT paths, and recovers from a bent link. The weak point is that everything is judged by 2-D silhouettes from one view."

Viva 2 cheat sheet · Paper 2page 5 of 6
6

Paper 3 · AI cybersecurity framework for IoT

Saeed 2025, IEEE Access · LSTM anomaly detection, salted hashing, Q-learning response, LWE post-quantum encryption

Architecture

IoT traffic→LSTM predicts next value (30-s lookback)→E_t > τ (95th pct)?→anomaly→Q-learning: block / countersalted SHA-256 integrity+LWE lattice encryption→gateway (LSTM/RL) · device (light hash) · precomputed keys

Core equations

$$x_{t+1}=f(x_t,\dots,x_{t-k});\quad L=\tfrac1N\textstyle\sum(x-\hat x)^2;\quad E_t=|x_t-\hat x_t|,\ \text{anomaly if }E_t>\tau\ (\tau=95^{th}\text{ percentile of MSE})$$ $$H(m,s)=\mathrm{SHA256}(m\oplus s),\ \text{verify }H(m',s)\overset{?}{=}H(m,s);\qquad Q(s,a)\leftarrow Q(s,a)+\alpha\big[R+\gamma\max_{a'}Q(s',a')-Q(s,a)\big];\qquad y=Ax+e\ (\mathrm{mod}\ q)$$

α learning rate, γ discount, R reward; A public random matrix, x secret, e small noise, q modulus. The noise e turns an easy linear system (Gaussian elimination) into a hard lattice problem: that is the LWE assumption.

Q-table and policy

stateblockcounter-attacklearned action
no attack––monitor
mild attack105block
severe attack515counter-attack

States {no, mild, severe}, actions {block, counter}: a 3×2 table, so the "learned" policy is essentially the hand-set rewards. Response effectiveness score ranks actions.

Numbers to quote

Dataset150,000 simulated events, 60/40 splitLookback30 sThreshold95th percentile of MSEExamplet = 810: error 1.25 vs τ 0.72 → anomalyHashSHA-256 with secret salt (XOR)Q-valuesmild 10/5 · severe 5/15CryptoLWE y = Ax + e mod q; Ring-LWE, Foxtail+ citedQuantum threatShor breaks RSA/ECC (factoring, discrete log)Reported metricsnone (no precision/recall/F1/latency)

Theory in one breath

Signature-based defences miss zero-day and disguised attacks, so detect surprise: train on normal traffic, flag large prediction errors. Integrity by keyed hashing (an HMAC in spirit). Adaptive response by tabular reinforcement learning. Confidentiality that survives quantum computers via lattices (LWE): keys are matrices, ciphertexts expand, arithmetic mod q is heavier than AES, hence precomputed keys and offloading to gateways.

Results in one breath

Qualitative only: the LSTM flags injected anomalies in a simulated network, tampered packets fail hash verification, the Q-table converges to block-for-mild and counter-for-severe, and LWE encryption is described as deployable with precomputation. No accuracy, false-positive rate, latency or energy figures.

Critique (prepare calmly)

  • "Homomorphic hashing" is a misnomer: a salted SHA-256 is not homomorphic (rename to HMAC or implement a real homomorphic scheme).
  • No quantitative evaluation or baselines (threshold statistics, autoencoder, isolation forest) on public data (TON_IoT, Bot-IoT); no ablation.
  • RL with 3 states and hand-set rewards is trivial; show learning curves and compare with a fixed rule.
  • LWE without parameters, key sizes, latency or energy on real devices; Eq. 11 and XOR/concatenation inconsistency.
  • Credit: clear modular architecture and sensible deployment plan.

Two-minute summary

"IoT devices are numerous and weak, and signature defences miss new attacks. The paper combines four parts: an LSTM trained on normal traffic flags prediction errors above the 95th percentile; a salted SHA-256 detects tampering; Q-learning over three states learns to block mild and counter severe attacks; and LWE lattice encryption gives post-quantum confidentiality. Heavy parts run on gateways, light hashing on devices. The design is sensible, but the evaluation is qualitative on 150,000 simulated events, the hashing is wrongly called homomorphic, and no accuracy, latency or key-size numbers are given."

Viva 2 cheat sheet · Paper 3page 6 of 6