· poster, CHES 2026 · interactive version: press + on a block

Gradient-free Profiling for Deep Learning based Side-Channel Analysis against Masking: the Solution for the Plateau Effect?

Nathan Rousselot1,2 · Karine Heydemann1 · Loïc Masure2 · Vincent Migairou1

1Thales, Meyreuil, France · 2LIRMM, Univ. Montpellier, CNRS, France

TL;DR: no.

A Bayesian network trained without any gradient (TAGI, Tractable Approximate Gaussian Inference) still plateaus, and its plateau grows with the masking order like Adam’s. So do all five profilers we tested: the plateau has deeper origins than gradient-based optimization.

Masking and the plateau effect

Profiled attacks

  • Learn P[Y∣X]\mathbb{P}[Y \mid X] from labelled traces
  • Attack unseen traces

Masking of order d\sym{d}{d}

  • Y=S0⊕⋯⊕SdY = S_0 \oplus \cdots \oplus S_{\sym{d}{d}} (d+1\sym{d}{d}+1 random shares)
  • No leakage moment of order ≤d\leq \sym{d}{d} depends on YY
Schematic: the training loss sits at H(Y), the loss of a blind guess, then drops sharply. The drop comes later for masking order d = 2 than d = 1, and later still for d = 3.H(Y)training time (log scale)lossschematicd=1d=2d=3plateau (d=3)

Plateau effect [MCLS23, RHMM25]

  • Losses stall at H(Y)\sym{entropy}{H(Y)}, then drop sharply
  • Plateau length grows exponentially with d\sym{d}{d}
  • Persists when shares are unknown during profiling
  • Still occurs with noiseless leakages
$ man block-1 : Masking and the plateau effect

Neural networks trained against masked implementations stall before they learn: the loss sits at the value of a blind guess, then drops sharply, and the stall grows exponentially with the masking order d\sym{d}{d}. The figure in the block is a sketch. The experiment below is real, but small: it trains a tiny network in your browser, on the same kind of task as block 5.

● live · toy re-run in your browsernot the poster's numbers
Training loss against samples seen, log scale, for the masking orders you run. Each curve starts at H(Y) = 4 bits and falls toward its dashed floor, the best any model can do from Hamming weights; a dot marks the escape.10²10³10⁴10⁵234H(Y) = 4 bits, a blind guesssamples seen (log)loss (bits)

Pick an order and press run. d = 3 takes a few seconds; d = 4 rarely escapes within the demo budget.

Each order adds a random share: Y = S_0 ⊕ … ⊕ S_d, each share leaking its Hamming weight, no noise. The network is an MLP (d+1) → 32 → 32 → 16 trained with Adam, much smaller than the poster's. Hamming weights can't pin the shares down, so even a perfect model keeps some loss: the dashed floor, H(Y | leakage). Expect the same shape (a flat loss at H(Y), then a fall to the floor, later for every extra share), not the same numbers.

Run d=0\sym{d}{d} = 0, then 11, 22, 33. Each extra share stretches the flat part, by a lot more each time. During the plateau the network isn’t slowly improving: its loss sits at H(Y)\sym{entropy}{H(Y)}. Then it leaves and falls to its floor, the dashed line: with noise-free Hamming weights, even a perfect model can’t tell every value of the secret apart.

Is the gradient to blame?

Scoop [RHMM25]: the plateau is caused by gradient training

  • Gradient magnitudes decay exponentially with d\sym{d}{d}
  • Saddle points trap the optimizer
  • Theory: on XOR-like targets, gradients carry almost no label information [SSS17]

Research question. If gradients cause the plateau, a learner that never computes one should avoid it. Is gradient-free profiling the solution?

$ man block-2 : Is the gradient to blame?

In Scoop we blamed gradient-based training for the plateau. Gradients shrink exponentially with d\sym{d}{d}, a saddle point sits near where training starts, and on XOR-like targets a gradient carries almost no information about the label. Scoop changes the optimizer accordingly, and it does shorten the plateau.

That makes a testable prediction. If the gradient causes the plateau, a learner that never computes one shouldn’t have it. So we need a learner that trains a neural network with no gradient at all (block 3), and a fair way to compare it with the others (block 4).

Bayesian learning, no gradient

In TAGI [GNA21], every weight is a Gaussian random variable.

Illustration of TAGI: a trace X enters a small network whose weights are each a Gaussian distribution, w ∼ N(μ_w, σ_w²). Step 1, forward: means and variances propagate in closed form to a predicted distribution P[Y|X]. Step 2, update: Bayes' rule, layer by layer, backpropagation style.trace Xw ∼ N(μ_w, σ_w²)P[Y | X]1 forward: propagate (μ, σ²) in closed form2 update: Bayes' rule, layer by layer

The updates use Bayesian statistics:

Δμθ=Cov⁡(θ,z) y−μzσz2+σv2\sym{dmu}{\Delta\mu_\theta} = \sym{cov}{\operatorname{Cov}}(\sym{wt}{\theta}, \sym{out}{z})\,\frac{\sym{err}{\sym{y}{y} - \sym{muz}{\mu_z}}}{\sym{unc}{\sym{sz}{\sigma_z^2} + \sym{sv}{\sigma_v^2}}}
  • No gradient is ever computed, and no learning rate is used.
  • Bayesian layers run on GPU (CUDA, built on cuTAGI [NG25]).
$ man block-3 : Bayesian learning, no gradient

TAGI (tractable approximate Gaussian inference, Goulet et al.) trains a neural network without backpropagating a gradient. Every weight is a Gaussian random variable with a mean and a variance. The forward pass pushes means and variances through the network in closed form. Learning is Bayes’ rule: condition on the label, then pass the correction back layer by layer, the way backpropagation would, but with covariances instead of derivatives.

Δμθ=Cov⁡(θ,z) y−μzσz2+σv2\sym{dmu}{\Delta\mu_\theta} = \sym{cov}{\operatorname{Cov}}(\sym{wt}{\theta}, \sym{out}{z})\,\frac{\sym{err}{\sym{y}{y} - \sym{muz}{\mu_z}}}{\sym{unc}{\sym{sz}{\sigma_z^2} + \sym{sv}{\sigma_v^2}}}

The prediction error y−μz\sym{err}{\sym{y}{y} - \sym{muz}{\mu_z}} is scaled by how much the weight θ\sym{wt}{\theta} covaries with the output z\sym{out}{z}, relative to the output’s total uncertainty. Because every weight carries a variance, the network doesn’t output one function but a distribution over functions, and it knows where it is unsure. Below, a tiny network of this kind (one input, 32 hidden units, one output) fits a curve. The band is its prediction ± 2σ; the thin lines are five networks drawn from its weight distributions.

● live · a Bayesian network, 1 → 32 → 1, in your browserillustration, not the poster's networks
A small Bayesian network's prediction against its input x. A shaded band shows its uncertainty, ±2σ, and thin lines show networks drawn from its weight distributions. Near the labels the band is narrow and the drawn networks agree; for x above 1.5, where no label was drawn, the band stays wide.no random labels−202−3−2−10123input x
  • mean ± 2σ
  • 5 networks drawn from the weights
  • labels
  • the function they come from

30 labels, three passes. ±2σ at x = 0: 0.37; at x = 2.5, where no label fell: 1.89.

Before any label the band is the prior: the drawn networks disagree everywhere. Click the plot to give a label where you like. Each one is a single Bayesian update of every weight's mean and variance: the band narrows around it and the drawn networks fall into line. Nothing changes where no label has been, and the band there stays wide: the network says it doesn't know. A small σ_v trusts each label more, a large one barely moves.

Three things to notice. Each label is one closed-form update, and the step size comes out of the variances: nothing plays the role of a learning rate. Every update also shrinks variances, so the band closes where labels fall and stays open where none did. And the prediction is a distribution, so “I don’t know” is part of the answer.

Protocol: five profiling techniques

gradient-basedgradient-freepoint estimate
back-propagation
evolution strategy
Bayesian
variational inference
closed-form updatesour focus
click a learner to follow it through blocks 5 and 6
  • Every baseline is tuned with the same budget, for a fair comparison.
  • Plateau: samples seen until the loss <H(Y)−ε< \sym{entropy}{H(Y)} - \sym{tol}{\varepsilon}.
$ man block-4 : Protocol: five profiling techniques

TAGI is one corner of a two-by-two grid: gradient-based or not, point estimate or Bayesian. The other corners are filled by Adam and Scoop (backpropagation), an evolution strategy, and variational inference. If the plateau came from the gradient, the two gradient-free learners would escape it.

Gradient-basedGradient-free
Point estimateAdam, Scoop (backpropagation)ES (evolution strategy)
BayesianVI (variational inference)TAGI (closed-form updates)

Each one gets the same tuning budget. The plateau is measured in training samples seen until the loss falls below H(Y)−ε\sym{entropy}{H(Y)} - \sym{tol}{\varepsilon}, with ε\sym{tol}{\varepsilon} = [placeholder: value of ε]. Click a learner in the grid, or in any legend, to follow it through blocks 5 and 6.

TAGI plateaus too

Samples until escape against masking order d, log scale. Adam and TAGI rise exponentially from 32 at d = 0 to about a million and more at d = 4; Scoop matches them at d = 0 and 1.10¹10³10⁵10⁷01234masking order dplateau size4.2·10⁶1.0·10⁶
  • one epoch

Setup: exhaustive dataset · 4-bit secret masked at order d (16d+1 traces) · noise-free Hamming-weight leakage · MLP (d+1) → 200 → 200 → 16 · 20 seeds

TAGI suffers from the same exponential increase (at the same rate) as gradient-based DL-SCA.

$ man block-5 : TAGI plateaus too

The cleanest setting: a 4-bit secret masked at order d\sym{d}{d}, every combination of share values in the dataset (16d+116^{\sym{d}{d}+1} traces), noise-free Hamming-weight leakage, the same small MLP for every learner, 20 seeds. Median samples seen before the loss breaks:

Order d\sym{d}{d}01234
Adam322,30436,864196,6084,194,304
TAGI321,28053,248327,6801,048,576
Scoop322,560still running

TAGI never computes a gradient, and its plateau still grows exponentially with d\sym{d}{d}, at the same rate as Adam’s. At some orders it escapes sooner, at others later; the trend is the same. The dotted line is one epoch, 16d+116^{\sym{d}{d}+1} samples: every learner needs to see the whole dataset many times over before it escapes. A toy version of this experiment runs live in block 1.

Every profiling method plateaus

We consider a noisy leakage of a single bit with dummy samples. The noise is additive Gaussian.

noise
Median samples until escape against masking order d = 0, 1, 2, for Adam, Scoop, VI, ES and TAGI, log scale. Every learner rises with d. At noise σ = 1.5 and d = 2, Scoop, VI and ES do not escape within the 2·10⁶ sample budget, and 6 of 10 TAGI runs do.10²10⁴10⁶012masking order dplateau sizebudget: 2·10⁶ samples0/10 escaped6/10
  • not escaped
  • budget
  • Every method suffers from the plateau
  • Some sort of masking hardness
$ man block-6 : Every profiling method plateaus

Noisy leakage: a single masked bit plus dummy samples, with additive Gaussian noise, 10 seeds, and a budget of two million samples. Median samples until escape, at d=0,1,2\sym{d}{d} = 0, 1, 2:

LearnerNoise σ=1.0\sym{sigma}{\sigma} = 1.0Noise σ=1.5\sym{sigma}{\sigma} = 1.5
Adam128 / 576 / 3,776128 / 1,472 / 132,480
Scoop448 / 1,408 / 6,912512 / 2,816 / not escaped
VI128 / 576 / 4,160128 / 1,408 / not escaped
ES1,152 / 3,840 / 41,9841,472 / 9,088 / not escaped
TAGI512 / 640 / 3,776512 / 2,048 / 1,027,840

All five plateau, and all five plateau longer as d\sym{d}{d} grows: gradient or not, Bayesian or not. At the higher noise level and d=2\sym{d}{d} = 2, only Adam and TAGI escape within the budget, and only six of the ten TAGI runs do. Flip the noise switch on the figure to watch every learner climb.

Every method seems to suffer from masking in the same way. This lets us conjecture that the plateau effect comes from some sort of hardness inherent to masking, not from the way a network is trained.

Why does TAGI plateau?

σ = 1.0 · last-layer weights · median over 10 seeds · relative to d = 0

TAGI, relative to d = 0. Size of the update: 1, 0.93, 1.01 at d = 0, 1, 2. Output–label covariance: 1, 0.53, 0.20.Size of the update00.51d=0d=1d=2≈1Output–label covarianced=0d=1d=2≈0.2relative to d = 0
  • Training does not stall: updates do not vanish with d\sym{d}{d}
  • Output–label covariance ≈ 5× lower at d=2\sym{d}{d} = 2
  • Updates decorrelate from the secret as d\sym{d}{d} grows, similar to [SSS17, RHMM25]
$ man block-7 : Why does TAGI plateau?

TAGI’s updates are easy to instrument, so we looked at what changes with d\sym{d}{d}. Two quantities, measured on the last layer at σ=1.0\sym{sigma}{\sigma} = 1.0 and expressed relative to d=0\sym{d}{d} = 0 (median over 10 seeds):

Order d\sym{d}{d}012
Size of TAGI’s update10.931.01
Output–label covariance, TAGI10.530.20
Output–label covariance, Adam10.510.20

The updates don’t shrink. TAGI keeps moving its weights as much at d=2\sym{d}{d} = 2 as at d=0\sym{d}{d} = 0. What collapses is the covariance between the network’s output and the label, about five times smaller at d=2\sym{d}{d} = 2, and Adam shows the same drop. The learner isn’t stuck; its updates have stopped pointing at the secret. That’s the same symptom Shalev-Shwartz et al. describe for gradients, and that we saw in Scoop, now without any gradient involved.

In the update of block 3, that covariance is the factor Cov⁡(θ,z)\sym{cov}{\operatorname{Cov}}(\sym{wt}{\theta}, \sym{out}{z}) in front of the error: the step keeps its size, but it is spent in directions that say little about YY.

What's next? Towards more frugal DL-SCA

Synthetic 1st-order masked bytes (σ = 0.5) · MLP 16–200–200–256 · 5 seeds

Against the number of profiling traces N_p, from 500 to 100,000: validation perceived information and attack traces needed, TAGI and Adam. TAGI's PI is positive from 500 traces while Adam's stays negative up to 2,000. At 2,000 profiling traces TAGI needs 5 attack traces and Adam 3,070. From 10,000 on, both are close.validation PI (bits)02461k10k100kattack traces to GE ≤ 11101001k10k1k10k100kprofiling traces N_p (log)MI = 6.69attack set
  • no attack within 10k traces

N_p = 2,000 · PI: TAGI 1.76, Adam −0.54 bits · attack traces: TAGI 5, Adam 3,070

  • Profiling traces needed: TAGI 500 vs. Adam 5,000
  • At Np=2,000\sym{Np}{N_p} = 2{,}000: TAGI 5 vs. Adam 3,070 attack traces
  • Few traces: Adam is over-confident
  • The baseline has been tuned.
$ man block-8 : What's next? Towards more frugal DL-SCA

One thing did come out in TAGI’s favour. On synthetic first-order masked bytes (σ=0.5\sym{sigma}{\sigma} = 0.5, 5 seeds), TAGI needs far fewer profiling traces than a tuned Adam baseline:

Profiling traces Np\sym{Np}{N_p}5001,0002,0005,00010,000
PI, TAGI (bits)0.120.661.763.435.12
PI, Adam (bits)−1.10−0.86−0.541.714.72
Attack traces, TAGI3511532
Attack traces, Adam> 10,00010,0003,07042

With few traces, Adam is overconfident: its perceived information is negative, so its predictions are worse than a uniform guess. TAGI is positive from 500 traces. At 2,000 profiling traces, TAGI’s model recovers the key in 5 attack traces and Adam’s in 3,070. With 10,000 or more, the two meet. Drag the slider under the figure to read both panels at one Np\sym{Np}{N_p}.

This is synthetic data only; whether TAGI stays this frugal on real traces is the next thing to test.

References

  • [RHMM25] N. Rousselot, K. Heydemann, L. Masure, V. Migairou. Scoop: an optimization algorithm for profiling attacks against higher-order masking. TCHES 2025(3). ePrint 2025/498
  • [MCLS23] L. Masure, V. Cristiani, M. Lecomte, F.-X. Standaert. Don’t learn what you already know. TCHES 2023(1).
  • [GNA21] J.-A. Goulet, L. H. Nguyen, S. Amiri. Tractable approximate Gaussian inference for Bayesian neural networks. JMLR 2021.
  • [SSS17] S. Shalev-Shwartz, O. Shamir, S. Shammah. Failures of gradient-based deep learning. ICML 2017.
  • [BEG+22] B. Barak et al. Hidden progress in deep learning. NeurIPS 2022.
  • [NG25] L. H. Nguyen, J.-A. Goulet. cuTAGI / pytagi. github.com/lhnguyen102/cuTAGI

Acknowledgment

This work was partially funded by the France 2030 program, managed by the French National Research Agency under grant agreement No. ANR-22-PETQ-0008 PQ-TLS.

Take-home message: the plateau effect has long days ahead of it

  • Without any gradient, TAGI still plateaus.
  • Masking decorrelates the leakage from the labels.
  • The plateau effect seems intrinsic to masking and is not an optimization artifact.
  • TAGI might be more frugal than regular DL techniques.

Next: not the gradient… so where does the plateau really come from?

CHES 2026 · Antalya, Türkiye · October 11–15, 2026 Ongoing work, ideas and developments were made by humans and LLMs were not used on any research tasks.

$ cat rousselot2026poster_tagi.bib
@misc{rousselot2026poster_tagi,  author       = {Nathan Rousselot and Karine Heydemann and Lo{\"\i}c Masure and Vincent Migairou},  title        = {Gradient-free Profiling for Deep Learning based Side-Channel Analysis against Masking: the Solution for the Plateau Effect?},  howpublished = {Poster, CHES 2026},  address      = {Antalya, Turkey},  year         = {2026},  month        = {October},  type         = {poster},  url          = {https://nathan-rousselot.com/posters/ches2026-gradient-free-plateau.pdf}}
$ ./welcome

// poster · ches 2026 · antalya

This poster is interactive.

The figures respond: hover a mark for its numbers, click a learner to follow it from block to block. Every block opens into a longer explanation, some with an experiment that runs in your browser.

A short tour shows you where things are, and lets you try each one.

the tour stays one click away: ? tour, above the title