One model (qwen2.5-0.5b-instruct, q4). Choose N: how many transformer blocks run in this browser before the hidden state goes to the server. What leaves this device at N is the hidden state at boundary N (N = 0: the token ids themselves), and a server can recover input tokens from it at the rate the privacy band shows. Nothing here hides the prompt from the server; local mode sends nothing.
Planner inputs, measured on this device
Device
Microbench (this variant)
Network
Server /plan and rate
Split point
estimated ms / token (decode)
–
server share and cost per 1M tokens
–
privacy band at this boundary
–
every candidate (click to select)
mode
N
feasible
ms/token
share
cost/1M
linear-500k top-1
inversion top-1
GPU MiB
download MiB
why
Run
measuring inputs…
Measured
per-step breakdown (medians over decode steps)
client blocksexportnetwork (round trip − server busy)server busylm_headsampling